Method and apparatus for evaluating server performance level

By testing the server's accelerator card, communication, video memory, and model performance, and combining this with weighted values, the server's overall evaluation value and grade are automatically determined. This solves the error problem caused by manual grading and achieves a more accurate performance grade evaluation.

CN120973612BActive Publication Date: 2026-02-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511517527.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-02-03
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

In existing server performance rating methods, manual grading leads to ambiguity in the weighting of performance indicators and strong subjectivity, resulting in significant errors.

Method used

By conducting performance tests on the accelerator cards in the server, performing various communication tests, memory performance tests, and tests on the model's inference process and expert partition performance, and combining the various test indicators and weight values, the server's overall evaluation value and performance level are automatically determined.

Benefits of technology

It improves the accuracy of server performance rating, reduces the subjectivity of human intervention, and provides a more accurate performance rating assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973612B_ABST
    Figure CN120973612B_ABST
Patent Text Reader

Abstract

This application discloses a method and device for evaluating server performance levels, relating to the field of server technology. The method includes: performing performance testing on an accelerator card to obtain an accelerator card performance evaluation value; performing multi-type communication tests on multiple interconnect types corresponding to each device to obtain communication performance values; performing memory performance tests on the video memory unit to obtain memory performance values; testing the inference process and expert partitioning performance of the model deployed on the server to obtain a first separation fitness value and a second separation fitness value for the server; determining the server's total evaluation value based on the accelerator card performance evaluation value, communication performance value, memory performance value, first separation fitness value, second separation fitness value, and corresponding target weight values; and determining the server's target performance level from multiple performance levels based on the total evaluation value, thereby improving the accuracy of server performance level evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of servers, and in particular to a server performance level evaluation method and device. BACKGROUND

[0002] In related technologies, the current server performance level evaluation method is usually manually graded by a server performance evaluator according to server running indicators to obtain a corresponding performance level. However, in related technologies, the manual grading method performed by the server performance evaluator according to the server running data is prone to problems of ambiguous running indicator weights and strong subjectivity, resulting in a large error in server performance level evaluation. SUMMARY

[0003] The present application provides a server performance level evaluation method and device to at least solve the problem of a large error in server performance level evaluation due to the manual grading method performed by the server performance evaluator according to the server running data being prone to problems of ambiguous running indicator weights and strong subjectivity in related technologies.

[0004] The present application provides a server performance level evaluation method, comprising:

[0005] For any server, performance testing is performed on an acceleration card in the server according to a preset multi-class test precision to obtain a corresponding acceleration card performance evaluation value;

[0006] Multi-class communication testing is performed on a plurality of interconnection types corresponding to each device in the server to obtain a corresponding communication performance value;

[0007] Memory performance testing is performed on a memory unit in the server to obtain a corresponding memory performance value;

[0008] The inference process and expert partition performance of a model deployed on the server are tested to obtain a first separation adaptation value and a second separation adaptation value of the server;

[0009] According to the acceleration card performance evaluation value, the communication performance value, the memory performance value, the first separation adaptation value, the second separation adaptation value, and corresponding target weight values, a total evaluation value of the server is determined;

[0010] According to the total evaluation value, a target performance level of the server is determined from a plurality of preset performance levels.

[0011] The present application also provides an electronic device, comprising a memory for storing a computer program and a processor for executing the computer program to implement the steps of any of the above server performance level evaluation methods.

[0012] The server performance level evaluation method and device provided in this application embodiment perform the following steps: For any server, the performance of the accelerator card in the server is tested according to preset multi-type test precision to obtain the corresponding accelerator card performance evaluation value; multiple communication tests are performed on the multiple interconnection types corresponding to each device in the server to obtain the corresponding communication performance value; the video memory unit in the server is tested to obtain the corresponding video memory performance value; the inference process and expert partitioning performance of the model deployed on the server are tested to obtain the server's first separation fitness value and second separation fitness value; based on the accelerator card performance evaluation value, communication performance value, video memory performance value, first separation fitness value, second separation fitness value, and corresponding target weight values, the server's total evaluation value is determined; based on the total evaluation value, the server's target performance level is determined from preset multiple performance levels. The target performance level of the server is determined through each test indicator and its corresponding weight, without the need for manual intervention, thus improving the accuracy of server performance level evaluation. Attached Figure Description

[0013] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a schematic diagram illustrating an application scenario of the server performance level evaluation method provided in the embodiments of this application;

[0015] Figure 2 A flowchart illustrating the server performance level evaluation method provided in this application embodiment. Figure 1 ;

[0016] Figure 3 A flowchart illustrating the server performance level evaluation method provided in this application embodiment. Figure 2 ;

[0017] Figure 4 This is a schematic diagram of the server performance level evaluation device provided in the embodiments of this application;

[0018] Figure 5 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0020] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0021] In related technologies, current methods for evaluating server performance levels typically involve server performance evaluators manually classifying servers based on their operational metrics to obtain the corresponding performance level. However, this manual classification method, relying on server performance evaluators to manually classify servers based on operational data, is prone to issues such as ambiguous weighting of operational metrics and strong subjectivity, resulting in significant errors in server performance level evaluation.

[0022] To address the aforementioned technical problems, this application proposes the following technical concept: The inventors consider the performance evaluation values ​​of the accelerator card, communication performance, video memory performance, first separation adaptation value, and second separation adaptation value generated from server testing. Based on these values ​​and their corresponding target weights, the inventors determine the server's overall evaluation value. Based on this overall evaluation value, the inventors determine the server's target performance level from a set of preset performance levels. This target performance level is determined through various test indicators and their corresponding weights, eliminating the need for manual intervention and thus improving the accuracy of server performance level evaluation.

[0023] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0024] This section describes the specific application environment architecture or hardware architecture upon which the server performance evaluation method depends. (References) Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario for the server performance level evaluation method provided in this application embodiment.

[0025] likeFigure 1 As shown, the application scenarios of this server performance level evaluation method include: Server 10.

[0026] The server 10 includes at least a controller 101, an accelerator card 102, a video memory unit 103, and multiple devices 104.

[0027] The video memory unit 103 includes multiple video memory cards.

[0028] Multiple devices 104, including at least a central processing unit, a graphics processing unit, and bus devices.

[0029] The controller 101 performs performance tests on the accelerator card 102 in the server 10 according to preset multi-type test precision to obtain the corresponding accelerator card performance evaluation value; performs multi-type communication tests on the multiple interconnection types corresponding to each device 104 in the server 10 to obtain the corresponding communication performance value; performs video memory performance tests on the video memory unit 103 in the server 10 to obtain the corresponding video memory performance value; tests the inference process and expert partitioning performance of the model deployed on the server 10 to obtain the first separation fitness value and the second separation fitness value of the server 10; determines the total evaluation value of the server 10 based on the accelerator card performance evaluation value, communication performance value, video memory performance value, first separation fitness value, second separation fitness value and corresponding target weight values; and determines the target performance level of the server 10 from preset multiple performance levels based on the total evaluation value.

[0030] Figure 2 A flowchart illustrating the server performance level evaluation method provided in this application embodiment. Figure 1 ,like Figure 2 As shown, embodiments of this application provide a method for evaluating server performance levels. The method is described in detail below:

[0031] S201: For any server, perform performance tests on the accelerator cards in the server according to preset multi-type test precision, and obtain the corresponding accelerator card performance evaluation value.

[0032] In this embodiment, the server can be a supernode server or other servers.

[0033] Among them, a supernode server is a high-density computing unit formed by connecting multiple independent server nodes through high-speed interconnection technology.

[0034] In addition, the server nodes in the supernode server achieve high-bandwidth, low-latency communication interconnection between nodes through a unified communication protocol and memory addressing.

[0035] In addition, the accelerator card configuration parameters of the acquisition server, the multiple interconnection types corresponding to each device, the video memory configuration parameters, and the network topology are also included.

[0036] The accelerator card configuration parameters include the model and quantity of the accelerator cards; multiple interconnect types include CXL, PCIe and RoCE; the memory configuration parameters include the single memory card capacity, total memory capacity and nominal bandwidth; and the network topology includes the interconnection method of accelerator cards within a node and the interconnection topology between nodes.

[0037] Among them, CXL, short for Compute Express Link, is a high-speed interconnect protocol designed to solve the problem of efficient communication between processors, accelerators, and memory expansion devices in data centers; PCIe, short for Peripheral Component Interconnect Express, is a high-speed serial computer expansion bus standard that supports dedicated channel bandwidth, hot-plugging, and multi-channel configuration; RoCE, short for RDMA over Converged Ethernet, is a network protocol that allows RDMA technology to be implemented over Ethernet.

[0038] RDMA, short for Remote Direct Memory Access, is a technology that allows computers to directly access memory, enabling high-speed data transfer without processor intervention.

[0039] Specifically, step S201 includes steps a to e:

[0040] Step a: Determine the first weight value for the accuracy of each type of test.

[0041] In this embodiment, various test accuracies include FP32, FP16, BF16, INT8, FP8, and other test accuracies.

[0042] Among them, FP32, short for Single-Precision Floating-Point, is a 32-bit floating-point format for representing real numbers in computers; FP16, short for Half-Precision Floating-Point, is a 16-bit floating-point format; BF16, short for Brain Floating Point 16-bit, is a 16-bit floating-point format optimized for deep learning; INT8, short for 8-bit Integer, is a low-precision integer data type mainly used for model quantization, inference acceleration, and storage optimization in deep learning; FP8, short for 8-bit Floating Point, is a low-precision floating-point format optimized for deep learning, aiming to improve the training and inference efficiency of AI models by reducing storage and computational overhead.

[0043] AI (Artificial Intelligence) refers to artificial intelligence.

[0044] Step b: Perform performance tests on the accelerator card according to various test accuracies to obtain the throughput of the accelerator card under each test accuracies.

[0045] Specifically, the performance of the accelerator card is tested using testing tools, test cases, and various test accuracies to obtain the throughput of the accelerator card at different test accuracies.

[0046] For example, the testing tools are TensorFlow Benchmarks and PyTorch Lightning; the test cases are ResNet-50, BERT-base, and GPT-2.

[0047] TensorFlow Benchmarks is an open-source benchmark tool used to evaluate the performance of TensorFlow under different hardware and software environments.

[0048] TensorFlow is an open-source machine learning framework that is widely used in deep learning, neural network construction, and various AI tasks.

[0049] PyTorch Lightning is a lightweight deep learning framework designed to simplify model development processes and improve code maintainability and scalability.

[0050] ResNet-50, short for Residual Network 50, indicates that the neural network contains 50 convolutional layers.

[0051] Among them, BERT-base, which stands for Bidirectional Encoder Representations from Transformers, is a large-scale pre-trained bidirectional language model use case.

[0052] Among them, GPT-2, which stands for Generative Pre-trained Transformer 2, is used to achieve multi-task text generation and understanding through unsupervised learning and is widely used in the field of natural language processing.

[0053] Step c: Convert each throughput into the number of floating-point operations for a preset duration.

[0054] In this embodiment, the preset duration can be any of 1 second, 3 seconds, or 5 seconds, or other durations.

[0055] Step d: Compare the number of floating-point operations with the corresponding baseline number of floating-point operations to determine the first evaluation value of each number of floating-point operations.

[0056] In this embodiment, the calculation formula for the first evaluation value of each floating-point operation is determined by comparing the number of floating-point operations with the corresponding baseline number of floating-point operations, including:

[0057]

[0058] In the formula, The first evaluation value; The base number of floating-point operations; This represents the number of floating-point operations.

[0059] Step e: Calculate the accelerator card performance evaluation value of the server based on each first evaluation value and each first weight value.

[0060] In this embodiment, the calculation formula for the server's accelerator card performance evaluation value based on each first evaluation value and each first weight value includes:

[0061]

[0062] In the formula, This is a performance evaluation value for the accelerator card; This is the first evaluation value corresponding to a single-precision floating-point number. This corresponds to the first weight value; The first evaluation value corresponds to both half-precision floating-point numbers and 16-bit floating-point numbers. This corresponds to the first weight value; The first evaluation value corresponds to an 8-digit integer. This corresponds to the first weight value; The first evaluation value corresponding to the 8-bit floating-point number. This is the corresponding first weight value.

[0063] S202: Perform multiple communication tests on the various interconnection types corresponding to each device in the server to obtain the corresponding communication performance values.

[0064] S203: Perform memory performance testing on the memory units in the server to obtain the corresponding memory performance values.

[0065] Specifically, step S203 includes steps a~d:

[0066] Step a: Perform memory performance tests on the memory units to obtain various memory performance values.

[0067] In this embodiment, the video memory performance test includes video memory capacity testing, video memory bandwidth testing, and access latency testing; correspondingly, step a specifically includes steps a1 to a4:

[0068] Step a1: Perform a memory capacity test on the memory unit to obtain the corresponding memory capacity value.

[0069] Step a2: Perform a memory bandwidth test on the memory unit to obtain the corresponding memory bandwidth value.

[0070] Step a3: Perform an access latency test on the video memory unit to obtain the corresponding access latency value.

[0071] Step a4: Determine the video memory capacity value, video memory bandwidth value, and access latency value as multiple types of video memory performance values.

[0072] Step b: Determine the third weight value corresponding to each type of video memory performance value.

[0073] Step c: Compare the various memory performance values ​​with the corresponding benchmark memory performance values ​​to obtain the performance evaluation values ​​for each type of memory.

[0074] Step d: Calculate the server's memory performance value based on various memory performance evaluation values ​​and each third weight value.

[0075] In this embodiment, the calculation formula for the server's memory performance value based on various memory performance evaluation values ​​and each third weight value includes:

[0076]

[0077] In the formula, This represents the video memory performance value. This is the memory performance evaluation value corresponding to the memory capacity test, and 0.07 is the corresponding third weight value; This is the memory performance evaluation value corresponding to the memory bandwidth test, and 0.05 is the corresponding third weight value; This is the memory performance evaluation value corresponding to the access latency test, with 0.03 being the corresponding third weight value.

[0078] S204: Test the inference process and expert partitioning performance of the model deployed on the server to obtain the server's first separation fitness value and second separation fitness value.

[0079] Specifically, step S204 includes steps a~g:

[0080] Step a: Determine the corresponding pre-filling and decoding separation deployment process from the model's inference process.

[0081] In this embodiment, the model can be a pre-trained language model. A pre-trained language model generally refers to a language model training task designed based on a large-scale corpus (including language training materials such as sentences and paragraphs), training a large-scale neural network algorithm structure to learn and implement it. The final large-scale neural network algorithm structure and parameters constitute the pre-trained language model. Subsequent tasks can use this model for feature extraction or task fine-tuning to achieve specific task objectives. The idea of ​​pre-training is to first train a set of model parameters for one task, then use these parameters to initialize the network model parameters, and finally use the initialized network model to train other tasks, obtaining models adapted for those tasks. By pre-training on a large-scale corpus, the neural language representation model can learn powerful language representation capabilities, extracting rich syntactic and semantic information from text. The pre-trained language model can provide tokens containing rich semantic information and sentence-level features for downstream tasks. Fine-tuning can also be performed directly on the pre-trained model for downstream tasks, conveniently and quickly obtaining downstream-specific models. The neural network algorithm structure used to train the pre-trained language model can be CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), etc., or it can be a model built with attention networks, such as Transformer, BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), Clip (Contrastive Language-Image Pre-training), etc., which are not limited here. Among them, attention network refers to a network model trained using an attention mechanism. This model extracts more important feature information from the input sequence by assigning different weights to each part of the input sequence, resulting in a more accurate output.

[0082] Step b: Analyze the pre-filling and decoding separate deployment process to obtain the pre-filling process and the decoding process.

[0083] The pre-filling process is used to process the input sequence in parallel to generate the KV Cache; the decoding process is used to output delays per token based on the KV Cache.

[0084] KV Cache, short for Key-Value Cache, is used to avoid redundant calculations by caching key and value tensors in the attention mechanism.

[0085] Step c: Determine the pre-fill throughput corresponding to the pre-fill process.

[0086] The formula for calculating the pre-filled throughput is as follows:

[0087] Throughput = (batch_size × input length) / computation time

[0088] In the formula, batch_size is the batch size.

[0089] Step d: Determine the first word delay and subsequent word delay corresponding to the decoding process.

[0090] For example, a delay of less than 100ms for the first character is preferred; a delay of less than 10ms for subsequent characters is preferred.

[0091] Step e: Determine the data transmission efficiency corresponding to the pre-filling process and the decoding process.

[0092] In this embodiment, the formula for calculating the data transmission efficiency corresponding to the pre-filling process and the decoding process is as follows:

[0093] Transmission efficiency = Actual transmission rate / Theoretical bandwidth

[0094] Step f: Determine the fourth weight values ​​corresponding to pre-fill throughput, first word delay, subsequent word delay, and data transmission efficiency, respectively.

[0095] Step g: Determine the server's first separation fitness value based on pre-filled throughput, first word latency, subsequent word latency, data transmission efficiency, and each fourth weight value.

[0096] Specifically, step g includes steps g1 to g5:

[0097] Step g1: Compare the pre-filled throughput with the baseline throughput to obtain the pre-filled throughput evaluation value.

[0098] Step g2: Compare the first character delay with the baseline first character delay to obtain the first character delay evaluation value.

[0099] Step g3: Compare the subsequent word delay with the baseline subsequent word delay to obtain the subsequent word delay evaluation value.

[0100] Step g4: Compare the data transmission efficiency with the benchmark data transmission efficiency to obtain the data transmission efficiency evaluation value.

[0101] Step g5: Determine the server's first separation fitness value based on the pre-filled throughput evaluation value, the first word delay evaluation value, the subsequent word delay evaluation value, the data transmission efficiency evaluation value, and each fourth weight.

[0102] In this embodiment, the calculation formula for determining the server's first separation fitness value based on the pre-filled throughput evaluation value, the first-word latency evaluation value, the subsequent-word latency evaluation value, the data transmission efficiency evaluation value, and each fourth weight includes:

[0103]

[0104] In the formula, This is the first separation fitness value; This is the pre-filled throughput evaluation value, with 0.045 being the corresponding fourth weight value; The first character delay evaluation value is 0.06, which is the corresponding fourth weight value. This is the evaluation value for subsequent character delays, with 0.03 being the corresponding fourth weight value; This is the data transmission efficiency evaluation value, with 0.015 being the corresponding fourth weight value.

[0105] Specifically, step S204 further includes step h~m:

[0106] Step h: Test the expert partitioning performance of the hybrid expert model to obtain scheduling latency, communication ratio, and throughput.

[0107] In this embodiment, the expert partitioning performance can be the expert partitioning performance of a hybrid expert model.

[0108] For example, the hybrid expert model can also be a 16-expert-layer Switch Transformer, i.e., a sparse activation expert hybrid model, corresponding to an input token length of 512.

[0109] Among them, scheduling latency is the time from routing decision to expert influence, and less than 5us is considered optimal; communication ratio is the cross-node expert communication time / total time, and less than 20% is considered optimal; throughput is the number of tokens processed per second.

[0110] Step i: Determine the fifth weight values ​​corresponding to scheduling delay, communication ratio, and throughput.

[0111] Step j: Compare the scheduling delay with the baseline scheduling delay to obtain the scheduling delay evaluation value.

[0112] Step k: Compare the communication ratio with the benchmark communication ratio to obtain the communication ratio evaluation value.

[0113] Step 1: Compare the throughput with the benchmark throughput to obtain the throughput evaluation value.

[0114] Step m: Determine the second separation fitness value of the server based on the scheduling delay evaluation value, communication ratio evaluation value, throughput evaluation value, and each fifth weight value.

[0115] In this embodiment, the calculation formula for determining the server's second separation fitness value based on the scheduling delay evaluation value, communication ratio evaluation value, throughput evaluation value, and each fifth weight value includes:

[0116]

[0117] In the formula, This is the second separation fitness value; This is the scheduling delay evaluation value, with 0.04 being the corresponding fifth weight value; This is the communication ratio evaluation value, with 0.03 being the corresponding fifth weight value; This is the throughput evaluation value, and 0.03 is the corresponding fifth weight value.

[0118] S205: Determine the server's total evaluation value based on the accelerator card performance evaluation value, communication performance value, video memory performance value, first separation adaptation value, second separation adaptation value, and the corresponding target weight values.

[0119] In this embodiment, the formula for calculating the server's total evaluation value is determined based on the accelerator card performance evaluation value, communication performance value, video memory performance value, first separation adaptation value, second separation adaptation value, and corresponding target weight values, including:

[0120]

[0121] In the formula, The server's overall rating; n=5; This is the i-th value mentioned above; for The corresponding weights.

[0122] For example, ; ; ; ; .

[0123] S206: Determine the target performance level of the server from multiple preset performance levels based on the overall evaluation value.

[0124] For example, the multiple performance levels are: A score of 90 or higher is Grade A (Excellent). A score of 75 or higher and less than 90 is grade B (excellent). A score of 60 or higher and less than 75 is Grade C (Good). If the value is less than 60, it is classified as Grade D (to be optimized).

[0125] In addition, after step S206, the method further includes: determining the application scenario of the server based on preset scenario recommendation rules and the server's target performance level.

[0126] For example, the application scenarios of the server are real-time inference scenarios, model training scenarios, and batch inference scenarios.

[0127] In real-time inference scenarios: if the first character latency is greater than or equal to 90ms and the latency is less than or equal to 5µs, it is recommended for intelligent customer service and code generation; for model training scenarios: if the scheduling latency is greater than or equal to 85ms and the bandwidth is greater than or equal to 200GB / s, it is recommended for hybrid expert models with hundreds of billions of parameters; for batch inference scenarios: if the INT8 value is greater than or equal to 90ms and the bandwidth is greater than or equal to 2TB / s, it is recommended for document summarization and batch classification.

[0128] In summary, the server performance level evaluation method provided in this embodiment obtains the corresponding accelerator card performance evaluation value by performing performance tests on the accelerator cards of any server according to preset multi-type test precision; performs multi-type communication tests on multiple interconnection types corresponding to each device in the server to obtain the corresponding communication performance value; performs memory performance tests on the memory units in the server to obtain the corresponding memory performance value; tests the inference process and expert partitioning performance of the model deployed on the server to obtain the server's first separation fitness value and second separation fitness value; determines the server's total evaluation value based on the accelerator card performance evaluation value, communication performance value, memory performance value, first separation fitness value, second separation fitness value, and corresponding target weight values; and determines the server's target performance level from preset multiple performance levels based on the total evaluation value. This method, which determines the server's target performance level through various test indicators and corresponding weights, requires no manual intervention, thus improving the accuracy of server performance level evaluation.

[0129] In addition, the server performance level evaluation method provided in this embodiment tests the inference process and expert partitioning performance of the model deployed on the server to obtain the first separation fitness value and the second separation fitness value of the server, filling the gap in the evaluation of the adaptability of large model specialization architecture in related technologies, and making the server performance level evaluation more complete.

[0130] Furthermore, the server performance level evaluation method provided in this embodiment determines the corresponding pre-filling and decoding separation deployment process from the model's inference process; parses the pre-filling and decoding separation deployment process to obtain the pre-filling process and the decoding process; determines the pre-filling throughput corresponding to the pre-filling process; determines the first-word latency and subsequent-word latency corresponding to the decoding process; determines the data transmission efficiency corresponding to the pre-filling process and the decoding process; determines the fourth weight values ​​corresponding to the pre-filling throughput, first-word latency, subsequent-word latency, and data transmission efficiency, respectively; and determines the server's first separation fitness value based on the pre-filling throughput, first-word latency, subsequent-word latency, data transmission efficiency, and each fourth weight value. Determining the server's first separation fitness value through multi-dimensional data makes the first separation fitness value more accurate, which is beneficial for determining the subsequent server performance level. Moreover, by independently optimizing the first-word latency and subsequent-word latency, diverse service level objectives can be met.

[0131] Furthermore, the server performance level evaluation method provided in this embodiment tests the performance of expert partitions in a hybrid expert model to obtain scheduling latency, communication ratio, and throughput; determines the fifth weight values ​​corresponding to scheduling latency, communication ratio, and throughput; compares the scheduling latency with the baseline scheduling latency to obtain a scheduling latency evaluation value; compares the communication ratio with the baseline communication ratio to obtain a communication ratio evaluation value; compares the throughput with the baseline throughput to obtain a throughput evaluation value; and determines the server's second separation fitness value based on the scheduling latency evaluation value, communication ratio evaluation value, throughput evaluation value, and each fifth weight value. By measuring scheduling latency, communication ratio, and throughput, performance bottlenecks in expert partition performance during use can be identified. Moreover, determining the server's second separation fitness value through multi-dimensional evaluation values ​​makes the second separation fitness value more accurate, which is beneficial for subsequent determination of the server performance level.

[0132] Figure 3 A flowchart illustrating the server performance level evaluation method provided in this application embodiment. Figure 2 In the embodiments of this application, in Figure 2 Based on the provided embodiments, the specific implementation method for performing multiple communication tests on multiple interconnection types corresponding to each device in the server in step S202 to obtain the corresponding communication performance values ​​is described in detail, such as... Figure 3 As shown, the method includes:

[0133] S301: Perform multiple communication tests on each interconnection type to obtain the corresponding multiple communication test values ​​for each interconnection type.

[0134] In this embodiment, the discussion of each interconnection type has been explained in detail in step S201, and will not be repeated here.

[0135] In this embodiment, the various communication tests include aggregated communication tests, decentralized communication tests, one-sided communication tests, and collaborative communication tests; correspondingly, step S301 specifically includes steps a~f:

[0136] Step a: Perform a set communication test on any interconnection type to obtain the first communication test value.

[0137] Among them, the collection communication test is used for bandwidth and latency testing of the corresponding Allreduce and Alltoall.

[0138] Allreduce is the core communication primitive in distributed deep learning training, mainly used for gradient synchronization and parameter aggregation between multiple machines / cards; Alltoall is the core communication primitive in distributed computing, mainly used for data exchange and reassembly between multiple nodes.

[0139] Step b: Perform decentralized communication tests on the interconnection type to obtain the second communication test value.

[0140] Among them, the decentralized communication test is used to test the bandwidth and latency of communication between any two accelerator cards.

[0141] Step c: Perform a one-sided communication test on the interconnection type to obtain the third communication test value.

[0142] Among them, the one-sided communication test is the bandwidth and latency test of the corresponding write and read operations.

[0143] Step d: Determine the first communication test value, the second communication test value, and the third communication test value corresponding to each interconnection type.

[0144] Step e: Perform collaborative communication tests on each interconnection type to obtain the fourth communication test value.

[0145] Among them, the collaborative communication test is used to test the performance loss when multiple types of interconnected hybrid deployments are deployed.

[0146] Step f: Determine the first communication test value, the second communication test value, the third communication test value, and the fourth communication test value as the multi-type communication test values ​​corresponding to each interconnection type.

[0147] S302: Compare the multi-type communication test values ​​corresponding to each interconnection type with the corresponding benchmark communication values ​​to obtain the multi-type communication evaluation values ​​corresponding to each interconnection type.

[0148] S303: Determine the second weight value corresponding to each type of communication test.

[0149] S304: Calculate the server's communication performance value based on each second weight value and the multi-class communication evaluation value corresponding to each interconnection type.

[0150] In this embodiment, the calculation formula for the server's communication performance value based on each second weight value and the multi-class communication evaluation value corresponding to each interconnection type includes:

[0151]

[0152] In the formula, This represents the communication performance value. For any interconnect type; T is a set of multiple interconnect types; 0.4 is the communication evaluation value corresponding to any set communication test, and 0.4 is the corresponding second weight value; 0.25 is the communication evaluation value corresponding to any decentralized communication test, and 0.25 is the corresponding second weight value; 0.2 represents the communication evaluation value corresponding to any one-sided communication test, and 0.2 represents the corresponding second weight value. This refers to the communication evaluation value corresponding to the collaborative communication test. This is the corresponding second weight value.

[0153] In summary, the server performance evaluation method provided in this embodiment obtains multi-type communication test values ​​for each interconnection type by conducting multiple communication tests on each interconnection type; compares each multi-type communication test value with the corresponding benchmark communication value to obtain a multi-type communication evaluation value for each interconnection type; determines the second weight value corresponding to each type of communication test; and calculates the server's communication performance value based on each second weight value and the multi-type communication evaluation value corresponding to each interconnection type. By treating the four interconnection types between nodes—NVLINK, CXL, PCIe, and RoCE—as independent test objects to conduct aggregated communication, decentralized communication, unilateral communication, and collaborative communication tests, the method can accurately capture the impact of different interconnection technologies on the scalability of supernodes.

[0154] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0155] Figure 4 This is a schematic diagram of the server performance level evaluation device provided in an embodiment of this application. Figure 4 As shown, embodiments of this application also provide a server performance level evaluation device, including: a first test module 401, a second test module 402, a third test module 403, a fourth test module 404, a first determination module 405, and a second determination module 406.

[0156] The first test module 401 is used to perform performance tests on the accelerator cards in any server according to preset multiple test precisions, and obtain the corresponding accelerator card performance evaluation value.

[0157] The second test module 402 is used to perform multiple communication tests on multiple interconnection types corresponding to various devices in the server, and obtain the corresponding communication performance values.

[0158] The third test module 403 is used to perform memory performance tests on the memory units in the server and obtain the corresponding memory performance values.

[0159] The fourth test module 404 is used to test the inference process and expert partitioning performance of the model deployed on the server, and obtain the first separation fitness value and the second separation fitness value of the server.

[0160] The first determining module 405 is used to determine the total evaluation value of the server based on the accelerator card performance evaluation value, communication performance value, video memory performance value, first separation adaptation value, second separation adaptation value, and corresponding target weight values.

[0161] The second determining module 406 is used to determine the target performance level of the server from a set of preset performance levels based on the total evaluation value.

[0162] In one possible implementation, the first test module 401 specifically includes:

[0163] The determination unit is used to determine the first weight value for various test accuracies.

[0164] The test unit is used to perform performance tests on the accelerator card according to various test accuracies, and obtain the throughput of the accelerator card under each test accuracies.

[0165] The conversion unit is used to convert each throughput into a number of floating-point operations for a preset duration.

[0166] The comparison unit is used to compare the number of floating-point operations with the corresponding baseline number of floating-point operations to determine the first evaluation value of each number of floating-point operations.

[0167] The calculation unit is used to calculate the performance evaluation value of the server's accelerator card based on each first evaluation value and each first weight value.

[0168] In one possible implementation, the second test module 402 specifically includes:

[0169] The test unit is used to perform multiple communication tests on each interconnection type and obtain the corresponding multiple communication test values ​​for each interconnection type.

[0170] The comparison unit is used to compare the multi-type communication test values ​​corresponding to each interconnection type with the corresponding benchmark communication values ​​to obtain the multi-type communication evaluation values ​​corresponding to each interconnection type.

[0171] The determination unit is used to determine the second weight value corresponding to various communication tests.

[0172] The calculation unit is used to calculate the communication performance value of the server based on each second weight value and the multi-class communication evaluation value corresponding to each interconnection type.

[0173] In one possible implementation, the various communication tests include aggregated communication tests, decentralized communication tests, unilateral communication tests, and cooperative communication tests; correspondingly, the test unit specifically includes:

[0174] The first test unit is used to perform a collective communication test on any interconnection type and obtain the first communication test value.

[0175] The second test unit is used to perform decentralized communication tests on the interconnection type and obtain the second communication test value.

[0176] The third test unit is used to perform unilateral communication tests on the interconnection type and obtain the third communication test value.

[0177] The first determining unit is used to determine the first communication test value, the second communication test value, and the third communication test value corresponding to each interconnection type.

[0178] The fourth test unit is used to perform collaborative communication tests on each interconnection type and obtain the fourth communication test value.

[0179] The second determining unit is used to determine each first communication test value, each second communication test value, each third communication test value, and the fourth communication test value as multiple types of communication test values ​​corresponding to each interconnection type.

[0180] In one possible implementation, the third test module 403 specifically includes:

[0181] The test unit is used to test the performance of the video memory unit and obtain various video memory performance values.

[0182] The determination unit is used to determine the third weight value corresponding to various video memory performance values.

[0183] The comparison unit is used to compare various memory performance values ​​with their corresponding benchmark memory performance values ​​to obtain various memory performance evaluation values.

[0184] The calculation unit is used to calculate the server's memory performance value based on various memory performance evaluation values ​​and each third weight value.

[0185] In one possible implementation, the video memory performance test includes video memory capacity testing, video memory bandwidth testing, and access latency testing; correspondingly, the test unit specifically includes:

[0186] The first test unit is used to test the memory capacity of the memory unit and obtain the corresponding memory capacity value.

[0187] The second test unit is used to test the memory bandwidth of the memory unit and obtain the corresponding memory bandwidth value.

[0188] The third test unit is used to perform access latency tests on the video memory units and obtain the corresponding access latency values.

[0189] The determination unit is used to determine the video memory capacity value, video memory bandwidth value, and access latency value as various types of video memory performance values.

[0190] In one possible implementation, the fourth test module 404 specifically includes:

[0191] The first determining unit is used to determine the corresponding pre-filling and decoding separation deployment process from the inference process of the model.

[0192] The parsing unit is used to parse the pre-filling and decoding separate deployment processes to obtain the pre-filling process and the decoding process.

[0193] The second determining unit is used to determine the pre-filling throughput corresponding to the pre-filling process.

[0194] The third determining unit is used to determine the first word delay and subsequent word delay corresponding to the decoding process.

[0195] The fourth determining unit is used to determine the data transmission efficiency corresponding to the pre-filling process and the decoding process.

[0196] The fifth determining unit is used to determine the fourth weight values ​​corresponding to the pre-filling throughput, first word delay, subsequent word delay, and data transmission efficiency, respectively.

[0197] The sixth determining unit is used to determine the server's first separation fitness value based on the pre-filled throughput, first word delay, subsequent word delay, data transmission efficiency, and each fourth weight value.

[0198] In one possible implementation, the sixth determining unit specifically includes:

[0199] The first comparison unit is used to compare the pre-filled throughput with the baseline throughput to obtain the pre-filled throughput evaluation value.

[0200] The second comparison unit is used to compare the first character delay with the benchmark first character delay to obtain the first character delay evaluation value.

[0201] The third comparison unit is used to compare the subsequent word delay with the benchmark subsequent word delay to obtain the subsequent word delay evaluation value.

[0202] The fourth comparison unit is used to compare the data transmission efficiency with the benchmark data transmission efficiency to obtain a data transmission efficiency evaluation value.

[0203] The determination unit is used to determine the first separation fitness value of the server based on the pre-filled throughput evaluation value, the first word delay evaluation value, the subsequent word delay evaluation value, the data transmission efficiency evaluation value, and each fourth weight.

[0204] In one possible implementation, the model is a hybrid expert model, and the fourth test module 404 specifically includes:

[0205] The test unit is used to test the performance of expert partitioning in the hybrid expert model to obtain scheduling latency, communication ratio, and throughput.

[0206] The first determining unit is used to determine the fifth weight values ​​corresponding to scheduling delay, communication ratio, and throughput.

[0207] The first comparison unit is used to compare the scheduling delay with the baseline scheduling delay to obtain a scheduling delay evaluation value.

[0208] The second comparison unit is used to compare the communication ratio with the benchmark communication ratio to obtain the communication ratio evaluation value.

[0209] The third comparison unit is used to compare the throughput with the benchmark throughput to obtain a throughput evaluation value.

[0210] The second determining unit is used to determine the second separation adaptation value of the server based on the scheduling delay evaluation value, the communication ratio evaluation value, the throughput evaluation value, and each fifth weight value.

[0211] For a description of the features in the embodiment corresponding to the server performance level evaluation device, please refer to the relevant description in the embodiment corresponding to the server performance level evaluation method, which will not be repeated here.

[0212] Figure 5 A schematic diagram of the structure of the electronic device provided in this application. Figure 5 As shown, the electronic device provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the electronic device further includes a communication component 503. The processor 501, memory 502, and communication component 503 are connected via a bus.

[0213] In the specific implementation process, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to execute the above-described server performance level evaluation method embodiment.

[0214] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0215] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0216] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0217] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0218] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described server performance level evaluation method embodiments when it runs.

[0219] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0220] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described server performance level evaluation method embodiments.

[0221] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described server performance level evaluation method embodiments.

[0222] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0223] The above provides a detailed description of a server performance level evaluation method and device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for evaluating server performance levels, characterized in that, include: For any given server, the performance of the accelerator card in the server is tested according to preset multiple test precisions to obtain the corresponding accelerator card performance evaluation value; Multiple communication tests were performed on the various interconnection types corresponding to each device in the server to obtain the corresponding communication performance values; The video memory units in the server are subjected to video memory performance tests to obtain the corresponding video memory performance values. The pre-filling and decoding separation deployment process is determined from the inference process of the model deployed on the server; the pre-filling and decoding separation deployment process is parsed to obtain the pre-filling process and the decoding process; the pre-filling throughput corresponding to the pre-filling process is determined; the first word latency and subsequent word latency corresponding to the decoding process are determined; the data transmission efficiency corresponding to the pre-filling process and the decoding process is determined; the fourth weight values ​​corresponding to the pre-filling throughput, the first word latency, the subsequent word latency, and the data transmission efficiency are determined respectively; the first separation fitness value of the server is determined based on the pre-filling throughput, the first word latency, the subsequent word latency, the data transmission efficiency, and each fourth weight value. The model deployed on the server is a hybrid expert model. The performance of the expert partition of the hybrid expert model is tested to obtain the scheduling latency, communication ratio and throughput. The fifth weight values ​​corresponding to the scheduling delay, the communication ratio, and the throughput are determined; the scheduling delay is compared with the baseline scheduling delay to obtain a scheduling delay evaluation value; the communication ratio is compared with the baseline communication ratio to obtain a communication ratio evaluation value; the throughput is compared with the baseline throughput to obtain a throughput evaluation value; and a second separation fitness value for the server is determined based on the scheduling delay evaluation value, the communication ratio evaluation value, the throughput evaluation value, and the fifth weight values. The overall evaluation value of the server is determined based on the accelerator card performance evaluation value, the communication performance value, the video memory performance value, the first separation adaptation value, the second separation adaptation value, and the corresponding target weight values. Based on the overall evaluation value, the target performance level of the server is determined from a set of preset performance levels.

2. The server performance level evaluation method according to claim 1, characterized in that, The process of performing performance tests on the accelerator cards in the server according to preset multiple test precisions to obtain corresponding accelerator card performance evaluation values ​​includes: Determine the first weight value for the accuracy of each type of test; The accelerator card is tested for performance according to the various test accuracies to obtain the throughput of the accelerator card under each test accuracies. Convert each throughput into the number of floating-point operations for a preset duration; The number of floating-point operations is compared with the corresponding baseline number of floating-point operations to determine the first evaluation value of each number of floating-point operations; The performance evaluation value of the accelerator card of the server is calculated based on each first evaluation value and each first weight value.

3. The server performance level evaluation method according to claim 1, characterized in that, The step of performing multiple communication tests on various interconnection types corresponding to each device in the server to obtain corresponding communication performance values ​​includes: Multiple communication tests were performed on each interconnection type to obtain the corresponding multiple communication test values ​​for each interconnection type. The test values ​​of multiple communication types corresponding to each interconnection type are compared with the corresponding benchmark communication values ​​to obtain the evaluation values ​​of multiple communication types corresponding to each interconnection type. Determine the second weight value corresponding to each type of communication test; The communication performance value of the server is calculated based on each second weight value and the multi-class communication evaluation value corresponding to each interconnection type.

4. The server performance level evaluation method according to claim 3, characterized in that, The various communication tests mentioned above include aggregated communication tests, decentralized communication tests, unilateral communication tests, and collaborative communication tests. Accordingly, the step of performing multiple communication tests on each interconnection type to obtain multiple communication test values ​​corresponding to each interconnection type includes: Perform a collective communication test on any interconnection type to obtain the first communication test value; A decentralized communication test is performed on the aforementioned interconnection type to obtain a second communication test value; A unilateral communication test is performed on the aforementioned interconnection type to obtain a third communication test value; Determine the first communication test value, the second communication test value, and the third communication test value corresponding to each interconnection type; Cooperative communication tests were conducted on each interconnection type to obtain the fourth communication test value; Each first communication test value, each second communication test value, each third communication test value, and the fourth communication test value are determined as the multi-type communication test values ​​corresponding to each interconnection type.

5. The server performance level evaluation method according to claim 1, characterized in that, The step of performing video memory performance testing on the video memory units in the server to obtain the corresponding video memory performance values ​​includes: The memory performance of the memory unit was tested to obtain various memory performance values. Determine the third weight value corresponding to each type of video memory performance value; The various memory performance values ​​are compared with the corresponding benchmark memory performance values ​​to obtain various memory performance evaluation values. The memory performance value of the server is calculated based on various memory performance evaluation values ​​and each third weight value.

6. The server performance level evaluation method according to claim 5, characterized in that, The memory performance test includes memory capacity test, memory bandwidth test, and access latency test. Accordingly, the memory performance test is performed on the memory unit to obtain various memory performance values, including: The memory capacity of the memory unit is tested to obtain the corresponding memory capacity value; The memory unit is tested for memory bandwidth to obtain the corresponding memory bandwidth value; The access latency of the video memory unit is tested to obtain the corresponding access latency value; The video memory capacity value, the video memory bandwidth value, and the access latency value are determined as multiple types of video memory performance values.

7. The server performance level evaluation method according to claim 1, characterized in that, The step of determining the server's first separation fitness value based on the pre-filled throughput, the first-word latency, the subsequent-word latency, the data transmission efficiency, and each fourth weight value includes: The pre-filled throughput is compared with the baseline throughput to obtain the pre-filled throughput evaluation value; The first character delay is compared with the benchmark first character delay to obtain the first character delay evaluation value; The subsequent word delay is compared with the benchmark subsequent word delay to obtain the subsequent word delay evaluation value; The data transmission efficiency is compared with the benchmark data transmission efficiency to obtain a data transmission efficiency evaluation value; The first separation fitness value of the server is determined based on the pre-filled throughput evaluation value, the first word latency evaluation value, the subsequent word latency evaluation value, the data transmission efficiency evaluation value, and each fourth weight.

8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the server performance level evaluation method as described in any one of claims 1 to 7 when executing the computer program.

Citation Information

Patent Citations

  • Method and device for testing performance of accelerator card in server, electronic equipment and storage medium

    CN114238003A

  • Method for testing training and reasoning performance of artificial intelligence acceleration card product

    CN116090552A