Performance testing method, device and equipment of graphics processor, storage medium and program product
By employing a multi-level tree-structured data transmission mode in the graphics processor to generate performance test results, the problem of high performance testing difficulty in complex hardware topology design scenarios is solved, and efficient performance evaluation is achieved.
Patent Information
- Application Number
- CN202511181738.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-08-22
AI Technical Summary
In complex hardware topology design scenarios, the performance testing of graphics processors becomes more difficult, and existing technologies cannot effectively evaluate the performance of multiple graphics processors.
A multi-level tree-structured data transmission mode is adopted. Test data is transmitted from the root node graphics processor to the leaf node graphics processor and back from the leaf node graphics processor to the root node graphics processor through a command queue to generate performance test results.
It reduces the difficulty of graphics processor performance testing, improves testing efficiency and accuracy, and can effectively evaluate the performance of multiple graphics processors.
Smart Images

Figure CN120723609B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of testing, and particularly relates to a performance testing method and device of a graphics processor, equipment, a storage medium and a program product. BACKGROUND
[0002] With the development of deep learning models, in the training and inference process of the models, the computing and storage capacity of a single graphics processor is difficult to meet the needs of model training and model inference. Multiple graphics processors are configured in a server for parallel computing to adapt to the needs of model training and model inference.
[0003] At present, in the related art, the performance of the graphics processors is evaluated by testing the performance of the links between the graphics processors. However, in the related art, the performance testing of the graphics processors is mainly link testing, and the difficulty of performance testing is increased when facing complex hardware topology design scenarios. SUMMARY
[0004] The present application provides a performance testing method and device of a graphics processor, equipment, a storage medium and a program product to at least solve the problem of increased difficulty of performance testing in the related art.
[0005] The present application provides a performance testing method of a graphics processor, comprising:
[0006] According to the performance testing instruction, first test data and second test data are generated, wherein the graphics processor comprises a root node graphics processor and a plurality of leaf node graphics processors, the first test data is test data transmitted from the root node graphics processor to the plurality of leaf node graphics processors, and the second test data is test data transmitted from the plurality of leaf node graphics processors to the root node graphics processor.
[0007] The first test data is transmitted from the root node graphics processor to the plurality of leaf node graphics processors through a command queue to generate first test information.
[0008] The second test data is transmitted from the plurality of leaf node graphics processors to the root node graphics processor through a command queue to generate second test information.
[0009] According to the first test information and the second test information, a performance testing result of the graphics processor is generated.
[0010] The present application also provides a performance testing device of a graphics processor, comprising:
[0011] The first generation module is configured to generate first test data and second test data according to the performance test instruction, wherein the graphic processor comprises a root node graphic processor and a plurality of leaf node graphic processors, the first test data is test data transmitted from the root node graphic processor to the plurality of leaf node graphic processors, and the second test data is test data transmitted from the plurality of leaf node graphic processors to the root node graphic processor.
[0012] The first transmission module is configured to transmit the first test data from the root node graphic processor to the plurality of leaf node graphic processors through a command queue to generate first test information.
[0013] The second transmission module is configured to transmit the second test data from the plurality of leaf node graphic processors to the root node graphic processor through the command queue to generate second test information.
[0014] The second generation module is configured to generate performance test results of the graphic processor according to the first test information and the second test information.
[0015] The application further provides an electronic device, comprising a memory configured to store a computer program and a processor configured to execute the computer program to implement the steps of the performance test method of any one of the graphic processors.
[0016] The application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the performance test method of any one of the graphic processors.
[0017] The application further provides a computer program product comprising a computer program, and the computer program is executed by a processor to implement the steps of the performance test method of any one of the graphic processors.
[0018] According to the application, the graphic processor is set to a multi-level tree type data transmission mode, the first test information of the first test data transmitted from the root node to the leaf node and the second test information of the second test data transmitted from the leaf node to the root node are collected, and the performance test results of the graphic processor are generated according to the first test information and the second test information, so that the difficulty of the performance test of the graphic processor is reduced, and the technical problem of the difficulty of the performance test is solved, and the technical effect of reducing the difficulty of the performance test is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the application, the drawings required in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0020] Figure 1 The application scenario of the performance test method of the graphics processor provided by the embodiment of the application is shown in the figure.
[0021] Figure 2 The flow of the performance test method of the graphics processor provided by the embodiment of the application is shown in the figure. Figure One ;
[0022] Figure 3 The hardware topology of the graphics processor provided by the embodiment of the application is shown in the figure. Figure One ;
[0023] Figure 4 The hardware topology of the graphics processor provided by the embodiment of the application is shown in the figure. Figure Two ;
[0024] Figure 5 The flow of the performance test method of the graphics processor provided by the embodiment of the application is shown in the figure. Figure Two ;
[0025] Figure 6 The flow of the performance test method of the graphics processor provided by the embodiment of the application is shown in the figure. Figure Three ;
[0026] Figure 7 The flow of the performance test method of the graphics processor provided by the embodiment of the application is shown in the figure. Figure Four ;
[0027] Figure 8 The flow of the performance test method of the graphics processor provided by the embodiment of the application is shown in the figure. Figure Five ;
[0028] Figure 9 The flow of the performance test method of the graphics processor provided by the embodiment of the application is shown in the figure. Figure Six ;
[0029] Figure 10 The flow of the performance test method of the graphics processor provided by the embodiment of the application is shown in the figure. Figure Seven ;
[0030] Figure 11 The flow of the performance test method of the graphics processor provided by the embodiment of the application is shown in the figure. Figure Eight ;
[0031] Figure 12 The structure of the performance test device of the graphics processor provided by the embodiment of the application is shown in the figure.
[0032] Figure 13 The structure of the electronic device provided by the application is shown in the figure. DETAILED DESCRIPTION
[0033] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0034] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0035] In order to solve the problem of low test efficiency of the server in the related art, the present application embodiment proposes the following technical concept: the inventor considers setting the graphic processor to a multi-level tree type data transmission mode, generates first test data and second test data, transmits the first test data from the root node to the leaf node, transmits the second test data from the leaf node to the root node, and generates the performance test result of the graphic processor according to the first test information and the second test information, thereby reducing the performance test difficulty of the graphic processor.
[0036] In order to enable a person skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0037] In combination with the specific application environment architecture or specific hardware architecture on which the performance test method of the graphic processor is dependent, the specific application environment architecture or specific hardware architecture is described here.
[0038] Reference Figure 1 , Figure 1 The application scenario of the performance test method of the graphic processor provided by the present application embodiment is shown in the figure. The scenario includes a device under test 101 and a test device 102.
[0039] The test device 102 generates performance test instructions, and generates first test data and second test data according to the performance test instructions, transmits the first test data from the root node graphics processor of the device under test 101 to the plurality of leaf node graphics processors through a command queue, generates first test information, transmits the second test data from the plurality of leaf node graphics processors to the root node graphics processor through the command queue, generates second test information, and the test device 102 generates a performance test result of the graphics processor of the device under test 101 according to the first test information and the second test information.
[0040] Figure 2 Flowchart of a performance test method of a graphics processor provided by an embodiment of the application Figure One As shown in Figure 2 , an embodiment of the application provides a performance test method of a graphics processor, which is described in detail as follows.
[0041] S201: generating first test data and second test data according to performance test instructions, wherein the graphics processor includes a root node graphics processor and a plurality of leaf node graphics processors, the first test data is test data transmitted from the root node graphics processor to the plurality of leaf node graphics processors, and the second test data is test data transmitted from the plurality of leaf node graphics processors to the root node graphics processor.
[0042] Figure 3 Hardware topology of a graphics processor provided by an embodiment of the application Figure One .
[0043] Figure 4 Hardware topology of a graphics processor provided by an embodiment of the application Figure Two .
[0044] As shown in Figure 3 and Figure 4 , a hardware environment of the graphics processor is configured, for example, 2 CPUs and 16 GPUs (Graphics Processing Unit, graphics processor) are configured in a server, each CPU is connected to 8 GPUs through a switch card, and every 4 GPUs are physically connected through a bridge, and the network ports between the GPUs are connected between the bridges through a DOC cable (high-speed cable).
[0045] For example, the 16 GPUs are constructed as a tree structure, each CPU is connected to 8 GPUs as branches of the tree, and every 4 GPUs are connected as sub-branches through a bridge, and the link state is checked through a hardware monitoring tool.
[0046] S202: transmitting the first test data from the root node graphics processor to the plurality of leaf node graphics processors through a command queue, and generating first test information.
[0047] Specifically, a forward transmission buffer division strategy is generated according to the first test data, the buffer is divided to obtain a sending buffer of the root node, a receiving buffer of the branch node and a receiving buffer of the leaf node, a first transmission instruction and a second transmission instruction are generated according to the first test data, and the first test data is transmitted from the sending buffer of the root node to the receiving buffer of the branch node according to the first transmission instruction, the first test data is transmitted from the receiving buffer of the branch node to the receiving buffer of the leaf node according to the second transmission instruction, and first test information is generated.
[0048] S203: Transmit the second test data from the plurality of leaf node graphic processors to the root node graphic processor through the command queue to generate second test information.
[0049] Specifically, a reverse transmission buffer division strategy is generated according to the second test data, the buffer is divided to obtain a receiving buffer of the root node, a receiving buffer of the branch node and a sending buffer of the leaf node, a third transmission instruction and a fourth transmission instruction are generated according to the second test data, and the second test data is transmitted from the sending buffer of the leaf node to the receiving buffer of the branch node according to the third transmission instruction, the second test data is transmitted from the receiving buffer of the branch node to the receiving buffer of the root node according to the fourth transmission instruction, and second test information is generated.
[0050] S204: Generate a performance test result of the graphic processor according to the first test information and the second test information.
[0051] Specifically, the data integrity is checked according to the first test information and the second test information, a check result is generated, and the data transmission time information and the performance information of the graphic processor are determined according to the check result.
[0052] Specifically, the performance bottleneck of the graphic processor is analyzed according to the performance test result, the transmission delay difference of each node is analyzed, and if the delay of a certain branch node is too high, it is checked whether there is a PCIe (peripheral component interconnect express, high-speed serial computer expansion bus) link bottleneck or data processing overload.
[0053] Specifically, the load balancing rate of each branch node is calculated according to the performance test result, wherein the load balancing rate includes a minimum load and a maximum load, and if the deviation degree of the load balancing rate from 1 exceeds a preset deviation threshold, the data division strategy is optimized.
[0054] Specifically, the tree type test result is compared with the performance test results of the graphic processors in ring type and star type data transmission modes, and the graphic processor is optimized according to different data transmission modes.
[0055] From the above embodiment, by setting the graphics processor to a multi-level tree type data transmission mode, respectively collecting first test information of transmitting the first test data from the root node to the leaf node and second test information of transmitting the second test data from the leaf node to the root node, and generating the performance test result of the graphics processor according to the first test information and the second test information, the difficulty of performance test of the graphics processor is reduced.
[0056] In an embodiment of the present application, step S202 comprises:
[0057] S202a: generating a forward transmission buffer area division strategy of the graphics processor according to the first test data.
[0058] In the embodiment, the transmission path from the root node graphics processor to the branch node graphics processor to the leaf node graphics processor is defined as a forward transmission path.
[0059] S202b: dividing the forward transmission buffer area of the root node graphics processor, the branch node graphics processor and the leaf node graphics processor according to the forward transmission buffer area division strategy of the graphics processor, to obtain a sending buffer area of the root node, a receiving buffer area of the branch node and a receiving buffer area of the leaf node.
[0060] Specifically, a data buffer area is created on the graphics processor, the root node graphics processor initializes the sending buffer area, and the branch node graphics processor and the leaf node graphics processor initialize the receiving buffer area.
[0061] S202c: determining the forward transmission parameters of the root node graphics processor and the branch node graphics processor according to the first test data.
[0062] In the embodiment, the forward transmission parameters of the root node graphics processor and the branch node graphics processor include but are not limited to the graphics processor type, the maximum number of graphics processors and the object pointer of the root node and the branch node.
[0063] S202d: generating a first transmission instruction according to the forward transmission parameters of the root node graphics processor and the branch node graphics processor.
[0064] Specifically, the context attribute of the Open Computing Language is defined, the Open Computing Language queue attribute is set, the command queue is created according to the context attribute and the queue attribute, the forward transmission parameters of the root node graphics processor and the branch node graphics processor are written into the command queue, and the first transmission instruction is obtained.
[0065] In the embodiment, the first transmission instruction is used to instruct the graphics processor to transmit the first test data from the root node graphics processor to the branch node graphics processor.
[0066] S202e: transmitting the first test data from the sending buffer of the root node to the receiving buffer of the branch node according to the first transmission instruction through the command queue to obtain first transmission data.
[0067] In the embodiment, the first transmission data includes but is not limited to a transmission duration of the root node to the branch node, a bandwidth occupation of the root node to the branch node, and a node load of the root node to the branch node.
[0068] S202f: determining forward transmission parameters of the branch node graphic processor and the leaf node graphic processor according to the first test data.
[0069] In the embodiment, the forward transmission parameters of the branch node graphic processor and the leaf node graphic processor include but are not limited to a graphic processor type of the branch node and the leaf node, a maximum number of graphic processors, and an object pointer.
[0070] S202g: generating a second transmission instruction according to the forward transmission parameters of the branch node graphic processor and the leaf node graphic processor.
[0071] Specifically, a context attribute of an open computing language is defined, an open computing language queue attribute is set, a command queue is created according to the context attribute and the queue attribute, the forward transmission parameters of the branch node graphic processor and the leaf node graphic processor are written into the command queue, and the second transmission instruction is obtained.
[0072] In the embodiment, the second transmission instruction is used to instruct the graphic processor to transmit the first test data from the branch node graphic processor to the leaf node graphic processor.
[0073] S202h: transmitting the first test data from the receiving buffer of the branch node to the receiving buffer of the leaf node according to the second transmission instruction through the command queue to obtain second transmission data.
[0074] In the embodiment, the second transmission data includes but is not limited to a transmission duration of the branch node to the leaf node, a bandwidth occupation of the branch node to the leaf node, and a node load of the branch node to the leaf node.
[0075] S202i: generating first test information according to the first transmission data and the second transmission data.
[0076] Specifically, the first test information of the first test data from the root node to the leaf node is obtained according to the first transmission data and the second transmission data.
[0077] From the above embodiment, by dividing the forward transmission buffer of the graphics processor, the sending buffer of the root node, the receiving buffer of the branch node and the receiving buffer of the leaf node are obtained, the first test data is sent from the sending buffer of the root node to the receiving buffer of the branch node to obtain first transmission data, and then the first test data is sent from the receiving buffer of the branch node to the receiving buffer of the leaf node to obtain second transmission data, the first test information is generated according to the first transmission data and the second transmission data, the node is divided into multiple levels, the performance bottleneck in the data transmission process is avoided, and the data transmission efficiency is improved.
[0078] In an embodiment of the present application, step S203 comprises:
[0079] S203a: generating a reverse transmission buffer division strategy of the graphics processor according to the second test data.
[0080] In this embodiment, the transmission path from the leaf node graphics processor to the branch node graphics processor to the root node graphics processor is defined as a reverse transmission path.
[0081] S203b: dividing the reverse transmission buffer of the root node graphics processor, the branch node graphics processor and the leaf node graphics processor according to the reverse transmission buffer division strategy of the graphics processor, to obtain the receiving buffer of the root node, the receiving buffer of the branch node and the sending buffer of the leaf node.
[0082] Specifically, a data buffer is created on the graphics processor, the leaf node graphics processor initializes the sending buffer, and the branch node graphics processor and the root node graphics processor initialize the receiving buffer.
[0083] S203c: determining the reverse transmission parameters of the leaf node graphics processor and the branch node graphics processor according to the second test data.
[0084] In this embodiment, the reverse transmission parameters of the leaf node graphics processor and the branch node graphics processor include but are not limited to the graphics processor types of the leaf node and the branch node, the maximum number of graphics processors and the object pointer.
[0085] S203d: generating a third transmission instruction according to the reverse transmission parameters of the leaf node graphics processor and the branch node graphics processor.
[0086] Specifically, the context attribute of the Open Computing Language is defined, the Open Computing Language queue attribute is set, the command queue is created according to the context attribute and the queue attribute, the reverse transmission parameters of the leaf node graphics processor and the branch node graphics processor are written into the command queue, and the third transmission instruction is obtained.
[0087] In the embodiment, the third transmission instruction is used to instruct the graphic processor to transmit the second test data from the leaf node graphic processor to the branch node graphic processor.
[0088] S203e: transmitting the second test data from the sending buffer of the leaf node to the receiving buffer of the branch node according to the third transmission instruction through the command queue to obtain third transmission data.
[0089] In the embodiment, the third transmission data includes but is not limited to transmission duration of the leaf node to the branch node, bandwidth occupation of the leaf node to the branch node and node load of the leaf node to the branch node.
[0090] S203f: determining reverse transmission parameters of the branch node graphic processor and the root node graphic processor according to the second test data.
[0091] In the embodiment, the reverse transmission parameters of the branch node graphic processor and the root node graphic processor include but are not limited to graphic processor types of the branch node and the root node, maximum graphic processor number and object pointer.
[0092] S203g: generating fourth transmission instruction according to the reverse transmission parameters of the branch node graphic processor and the root node graphic processor.
[0093] Specifically, context attributes of the OpenCL are defined, OpenCL queue attributes are set, the command queue is created according to the context attributes and the queue attributes, the reverse transmission parameters of the branch node graphic processor and the root node graphic processor are written into the command queue to obtain the fourth transmission instruction.
[0094] In the embodiment, the fourth transmission instruction is used to instruct the graphic processor to transmit the second test data from the branch node graphic processor to the root node graphic processor.
[0095] S203h: transmitting the second test data from the receiving buffer of the branch node to the receiving buffer of the root node according to the fourth transmission instruction through the command queue to obtain fourth transmission data.
[0096] In the embodiment, the fourth transmission data includes but is not limited to transmission duration of the branch node to the root node, bandwidth occupation of the branch node to the root node and node load of the branch node to the root node.
[0097] S203i: generating second test information according to the third transmission data and the fourth transmission data.
[0098] Specifically, the second test information of the second test data from the leaf node to the root node is obtained according to the third transmission data and the fourth transmission data.
[0099] From the above embodiment, by dividing the reverse transmission buffer of the graphics processor, the receiving buffer of the root node, the receiving buffer of the branch node and the sending buffer of the leaf node are obtained, the second test data is sent from the sending buffer of the leaf node to the receiving buffer of the branch node to obtain third transmission data, and then the second test data is sent from the receiving buffer of the branch node to the sending buffer of the root node to obtain fourth transmission data, the second test information is generated according to the third transmission data and the fourth transmission data, the nodes are divided into multiple levels, the performance bottleneck in the data transmission process is avoided, and the data transmission efficiency is improved.
[0100] Figure 5 The performance test method of the graphics processor provided in the embodiment of the application is described in detail as follows. Figure Two As shown in Figure 5 The method for generating a plurality of groups of aggregated data to be transmitted provided in the embodiment of the application is described in detail as follows:
[0101] S301: Obtain the data size of the second test data.
[0102] In this embodiment, the data size of the second test data records the data size of a single transmission task and the number of tasks to be transmitted.
[0103] S302: If the data size of the second test data exceeds the preset data threshold of the root node, set the data aggregation frequency according to the data size of the second test data.
[0104] Specifically, if the data size of the second test data does not exceed the preset data threshold, the full-precision data of the second test data is retained by using the accurate aggregation method, and the second test data is sent to the root node.
[0105] S303: Divide the second test data according to the data aggregation frequency by using a compression algorithm to generate a plurality of groups of aggregated data to be transmitted.
[0106] Specifically, the gradient compression algorithm is used to divide the second test data according to the set data aggregation frequency to generate a plurality of groups of aggregated data to be transmitted.
[0107] Specifically, the division strategy is dynamically adjusted according to the generated plurality of groups of aggregated data to be transmitted.
[0108] From the above embodiment, by obtaining the data size of the second test data, if the data size of the second test data exceeds the preset data threshold of the root node, the data aggregation frequency is set according to the data size, and a plurality of groups of aggregated data to be transmitted are generated by dividing according to the data aggregation frequency, thereby reducing the resource pressure of the branch node on the root node in data transmission.
[0109] In an embodiment of the application, step S204 comprises:
[0110] S204a: Verify the data integrity of the leaf node graphics processor based on the first test information and generate the first verification result.
[0111] Specifically, the first test data is compared with the data received by the leaf nodes in the first test information to verify whether the data is complete, and a first verification result is generated.
[0112] S204b: Verify the data integrity of the root node graphics processor based on the second test information, and generate a second verification result.
[0113] Specifically, the second test data is compared with the data received by the root node in the second test information to verify whether the data is complete, and a second verification result is generated.
[0114] S204c: Determine the data transmission time and performance information of the graphics processor based on the first and second verification results.
[0115] In this embodiment, performance information includes, but is not limited to, transmission latency, node load balancing rate, and bandwidth usage.
[0116] In this embodiment, the data transmission time information includes the timestamp of when data transmission begins, the timestamp of when data transmission ends, and the transmission duration.
[0117] S204d: Generates graphics processor performance test results based on the graphics processor's data transfer time and performance information.
[0118] In this embodiment, the performance test results record, including but not limited to, the total data transmission time from the root node to the leaf node and from the leaf node to the root node, the actual bandwidth of data transmission, GPU utilization, video memory usage, data integration compression ratio, and data processing time.
[0119] As can be seen from the above embodiments, by verifying the data integrity of the leaf node graphics processor and the root node graphics processor according to the first test information and the second test information respectively, a first verification result and a second verification result are generated. The data transmission time information and performance information of the graphics processor are determined according to the first verification result and the second verification result, and the performance of the graphics processor is evaluated according to the data transmission time and performance.
[0120] Figure 6 A flowchart illustrating the performance testing method for a graphics processor provided in this application embodiment. Figure Three ,like Figure 6 The method for generating optimization strategies provided in the embodiments of this application is described in detail below:
[0121] S401: Obtain link quality information between a plurality of graphic processors.
[0122] In the embodiment, the link quality information includes, but is not limited to, communication delay, bandwidth utilization and error rate.
[0123] S402: Determine whether the link quality information exceeds a preset link quality threshold.
[0124] Specifically, when one or more of the link quality information exceeds the preset link quality threshold, a node optimization strategy is generated.
[0125] S403: If the link quality information exceeds the preset link quality threshold, an optimization strategy is generated according to the link quality information.
[0126] Specifically, a node optimization strategy is generated according to the node distribution information of the graphic processor, link connection information of the optimized graphic processor is generated according to the node optimization strategy, and the optimization strategy is determined according to the node optimization strategy and the link connection information of the optimized graphic processor.
[0127] From the above embodiment, by obtaining the link quality information between a plurality of graphic processors, it is determined whether the link quality information exceeds a preset link quality threshold, and if the link quality threshold is exceeded, an optimization strategy is generated to optimize the graphic processor, thereby improving the performance of the graphic processor.
[0128] In an embodiment of the present application, step S403 comprises:
[0129] S403a: If the link quality information exceeds the preset link quality threshold, obtain the node distribution information of the graphic processor.
[0130] In the embodiment, the node distribution information of the graphic processor includes leaf nodes, root nodes and branch nodes.
[0131] S403b: Generate a node optimization strategy according to the node distribution information of the graphic processor.
[0132] In the embodiment, the node optimization strategy is a node distribution optimization strategy.
[0133] For example, if the original node distribution information of the graphic processor is 1 root node, 3 branch nodes and 12 leaf nodes, the node distribution information of the optimized graphic processor is 2 root nodes, 2 branch nodes and 12 leaf nodes.
[0134] S403c: Generate link connection information of the optimized graphic processor according to the node optimization strategy.
[0135] Specifically, the link between the optimized graphics processors is created according to the node optimization strategy, and link connection information of the optimized graphics processors is generated.
[0136] S403d: determining the optimization strategy according to the node optimization strategy and the link connection information of the optimized graphics processors.
[0137] Specifically, the optimization strategy is determined according to the node optimization strategy and the link connection information of the optimized graphics processors, and test data is generated to test the generated optimization strategy.
[0138] From the above embodiments, if the link quality information exceeds the preset link quality threshold, the node optimization strategy is generated according to the node distribution information of the graphics processors, and the optimization strategy of the link connection information is generated, thereby improving the performance of the graphics processors.
[0139] Figure 7 The performance test method of the graphics processors provided in the embodiments of the present application is described in detail as follows. Figure Four As shown in Figure 7 The method for generating the optimized performance test data provided in the embodiments of the present application is described in detail as follows:
[0140] S501: If the performance test result is that the performance test fails, the target graphics processor is obtained according to the performance test result.
[0141] In this embodiment, the target graphics processor is the graphics processor whose performance test fails.
[0142] S502: The processor node type information is determined according to the target graphics processor.
[0143] In this embodiment, the processor node type information includes leaf nodes, root nodes and branch nodes.
[0144] S503: The node optimization strategy is generated according to the node type information of the processor.
[0145] In this embodiment, if the target node is a leaf node, the data transmission task is taken over by the branch node.
[0146] In this embodiment, if the target node is a branch node, the number of standby leaf nodes is increased.
[0147] In this embodiment, if the target node is a root node, the root node is re-screened to generate a new root node.
[0148] S504: The target graphics processor is optimized according to the node optimization strategy, and the data retransmission is performed on the optimized target graphics processor to generate the optimized performance test data.
[0149] Specifically, according to the node optimization strategy, new test data is generated, the new test data is retransmitted, the performance test data after node optimization is recorded, and if the performance test data after optimization still fails the test, the node optimization strategy is regenerated.
[0150] From the above embodiments, if the performance test result fails, the target graphics processor that fails the test is obtained, the node type of the target graphics processor is determined, the optimization strategy corresponding to the node type is generated according to the node type, and the data retransmission of the optimized target graphics processor is performed. The performance test data after optimization is verified, and the performance of the graphics processor is improved.
[0151] Figure 8 The flowchart of the performance test method of the graphics processor provided in the embodiments of the present application is shown in the figure Figure Five As shown in the figure Figure 8 The method for generating the node allocation strategy of the graphics processor provided in the embodiments of the present application is described in detail as follows:
[0152] S601: Obtain link connection information between a plurality of graphics processors.
[0153] In this embodiment, the link connection information includes the graphics processors at both ends of the link connection and the link connection line.
[0154] S602: Determine the link type according to the link connection information.
[0155] In this embodiment, the link type includes but is not limited to a single bridge connection link, a cross-bridge connection link under the same CPU, and a cross-CPU communication link.
[0156] S603: Calculate and generate corresponding data layer bandwidth according to the link type.
[0157] In this embodiment, the formula for calculating and generating the data layer bandwidth is:
[0158]
[0159] In the formula, represents the bandwidth of the data layer i; represents the total bandwidth; represents the exponential priority of the data layer i; represents the exponential priority of the total data layer j.
[0160] S604: Generate the node allocation strategy of the graphics processor according to the data layer bandwidth.
[0161] For example, the bridge is preferentially used for data transmission, the cross-bridge under the same CPU is secondly used for data transmission, and the cross-CPU communication is finally used for data transmission.
[0162] As can be seen from the above embodiments, by obtaining the link connection information between graphics processors, determining the link type, calculating the data layer bandwidth of each type of link, and generating a node allocation strategy for graphics processors based on the calculated data layer bandwidth, the performance of the allocated graphics processors is improved.
[0163] Figure 9 A flowchart illustrating the performance testing method for a graphics processor provided in this application embodiment. Figure Six ,like Figure 9 The method for generating resource allocation information provided in the embodiments of this application is described in detail below:
[0164] S701: Input the first test data and the second test data into the pre-trained task classifier to generate task classification information for the first test data and the second test data.
[0165] In this embodiment, task classification information includes, but is not limited to, gradient synchronization, parameter broadcasting, and checkpoint saving.
[0166] S702: Query the mapping table between transmission tasks and transmission modes based on the task classification information to obtain the transmission mode corresponding to the task classification information.
[0167] In this embodiment, the transmission modes include, but are not limited to, converged transmission from the leaf node graphics processor to the root node graphics processor, transmission from the root node graphics processor to the leaf node graphics processor, and cross-CPU compressed transmission.
[0168] S703: Allocate transmission resources for the first test data and the second test data according to the transmission mode, and generate resource allocation information.
[0169] In this embodiment, the resource allocation information includes, but is not limited to, bandwidth allocation information, video memory usage allocation information, and data volume allocation information.
[0170] As can be seen from the above embodiments, by inputting the first test data and the second test data into the pre-trained task classifier, task classification information is generated. By querying the mapping table between transmission tasks and transmission modes, the transmission mode corresponding to the task classification information is obtained. Based on the transmission mode, the transmission resources of the first test data and the second test data are divided to generate resource allocation information, thereby improving the efficiency of transmitting the first test data and the second test data.
[0171] Figure 10 A flowchart illustrating the performance testing method for a graphics processor provided in this application embodiment. Figure Seven ,like Figure 10 The method for optimizing the mapping table between transmission tasks and transmission modes provided in this application embodiment is described in detail below:
[0172] S704: Obtain resource occupation information of executing the first test data and the second test data transmission task through the resource monitoring probe.
[0173] Specifically, the resource monitoring probe is deployed on the test data transmission path, and the resource occupation information of executing the test data transmission is collected according to the resource monitoring probe.
[0174] S705: Determine whether the resource occupation information exceeds a preset resource threshold.
[0175] In this embodiment, the resource occupation information includes but is not limited to GPU memory occupation information, CPU occupation information, and memory occupation information.
[0176] S706: If the resource occupation information exceeds the preset resource threshold, a resource division optimization strategy is generated, and a mapping table of the transmission task and the transmission mode is optimized according to the resource division optimization strategy.
[0177] Specifically, the resource occupation item exceeding the threshold in the resource occupation information is obtained, the transmission task is optimized according to the resource occupation item exceeding the threshold, the resource division optimization strategy is generated, the resource optimization strategy is executed, and if the resource occupation item after executing the resource optimization strategy does not exceed the threshold, the mapping table of the transmission task and the transmission mode is optimized through the resource optimization strategy.
[0178] From the above embodiment, it can be seen that the resource occupation information of executing the data transmission task is obtained through the resource monitoring probe, if the resource occupation information exceeds the preset resource threshold, the resource division optimization strategy is generated, and the mapping table of the transmission task and the transmission mode is optimized according to the resource division optimization strategy, thereby improving the matching accuracy of the mapping table of the transmission task and the transmission mode.
[0179] Figure 11 The performance test method of the graphics processor provided in the embodiment of the present application is shown in the flowchart Figure Eight The method for generating a visual file provided in the embodiment of the present application is described in detail as follows: Figure 11
[0180] S801: Obtain node distribution information and resource occupation information of the graphics processor.
[0181] In this embodiment, the node distribution information includes but is not limited to the number of nodes, the type of nodes, and the link information between nodes.
[0182] In this embodiment, the resource occupation information includes but is not limited to GPU memory occupation information, CPU occupation information, and memory occupation information.
[0183] S802: Standardize the node distribution information and the resource occupation information of the graphics processor to generate standardized information.
[0184] Specifically, the node distribution information and the resource occupation information of the graphics processor are format-converted according to a predefined data format, a unified file format is generated, and standardized information is determined.
[0185] S803: A visualization file is generated according to the standardized information.
[0186] In this embodiment, the visualization file is an editable file.
[0187] In this embodiment, the content recorded in the visualization file includes, but is not limited to, the node distribution information of the graphics processor, the line connection information between the graphics processors, the type of the graphics processor, and the resource occupation information of each graphics processor.
[0188] From the above embodiment, it can be known that the node distribution information and the resource occupation information of the graphics processor are acquired, the node distribution information and the resource occupation information of the graphics processor are standardized, the visualization file is generated for the standardized information, and the developer can optimize the graphics processor according to the visualization file, thereby improving the convenience of optimizing the graphics processor.
[0189] Figure 12 A structural schematic diagram of a performance testing device of a graphics processor provided by an embodiment of the present application is shown in FIG. 1. Figure 12 As shown in FIG. 1, the embodiment of the present application further provides a performance testing device 120 of a graphics processor, which includes a first generation module 1201, a first transmission module 1202, a second transmission module 1203, and a second generation module 1204.
[0190] The first generation module 1201 is configured to generate first test data and second test data according to a performance testing instruction, wherein the graphics processor includes a root node graphics processor and a plurality of leaf node graphics processors, the first test data is test data transmitted from the root node graphics processor to the plurality of leaf node graphics processors, and the second test data is test data transmitted from the plurality of leaf node graphics processors to the root node graphics processor.
[0191] The first transmission module 1202 is configured to transmit the first test data from the root node graphics processor to the plurality of leaf node graphics processors through a command queue, and generate first test information.
[0192] The second transmission module 1203 is configured to transmit the second test data from the plurality of leaf node graphics processors to the root node graphics processor through the command queue, and generate second test information.
[0193] The second generation module 1204 is configured to generate a performance testing result of the graphics processor according to the first test information and the second test information.
[0194] In an embodiment of the present application, the first transmission module 1202 comprises:
[0195] The first generation unit is configured to generate a forward transmission buffer partition strategy of the graphic processor according to the first test data.
[0196] The first partition unit is configured to partition the forward transmission buffer of the root node graphic processor, the branch node graphic processor and the leaf node graphic processor according to the forward transmission buffer partition strategy of the graphic processor, to obtain the sending buffer of the root node, the receiving buffer of the branch node and the receiving buffer of the leaf node.
[0197] The first determination unit is configured to determine the forward transmission parameters of the root node graphic processor and the branch node graphic processor according to the first test data.
[0198] The second generation unit is configured to generate the first transmission instruction according to the forward transmission parameters of the root node graphic processor and the branch node graphic processor.
[0199] The first sending unit is configured to send the first test data from the sending buffer of the root node to the receiving buffer of the branch node according to the first transmission instruction through the command queue, to obtain the first transmission data.
[0200] The second determination unit is configured to determine the forward transmission parameters of the branch node graphic processor and the leaf node graphic processor according to the first test data.
[0201] The third generation unit is configured to generate the second transmission instruction according to the forward transmission parameters of the branch node graphic processor and the leaf node graphic processor.
[0202] The second sending unit is configured to send the first test data from the receiving buffer of the branch node to the receiving buffer of the leaf node according to the second transmission instruction through the command queue, to obtain the second transmission data.
[0203] The fourth generation unit is configured to generate the first test information according to the first transmission data and the second transmission data.
[0204] In an embodiment of the present application, the second transmission module 1203 comprises:
[0205] The fifth generation unit is configured to generate a reverse transmission buffer partition strategy of the graphic processor according to the second test data.
[0206] The second partition unit is configured to partition the reverse transmission buffer of the root node graphic processor, the branch node graphic processor and the leaf node graphic processor according to the reverse transmission buffer partition strategy of the graphic processor, to obtain the receiving buffer of the root node, the receiving buffer of the branch node and the sending buffer of the leaf node.
[0207] The third determining unit is configured to determine reverse transmission parameters of the leaf node graphic processor and the branch node graphic processor according to the second test data.
[0208] The sixth generating unit is configured to generate third transmission instructions according to the reverse transmission parameters of the leaf node graphic processor and the branch node graphic processor.
[0209] The third sending unit is configured to send the second test data from the sending buffer of the leaf node to the receiving buffer of the branch node according to the third transmission instructions through the command queue, to obtain third transmission data.
[0210] The fourth determining unit is configured to determine reverse transmission parameters of the branch node graphic processor and the root node graphic processor according to the second test data.
[0211] The seventh generating unit is configured to generate fourth transmission instructions according to the reverse transmission parameters of the branch node graphic processor and the root node graphic processor.
[0212] The fourth sending unit is configured to send the second test data from the receiving buffer of the branch node to the receiving buffer of the root node according to the fourth transmission instructions through the command queue, to obtain fourth transmission data.
[0213] The eighth generating unit is configured to generate second test information according to the third transmission data and the fourth transmission data.
[0214] In an embodiment of the present application, the performance testing device 120 of the graphic processor further comprises:
[0215] The first obtaining module is configured to obtain a data size of the second test data.
[0216] The setting module is configured to, if the data size of the second test data exceeds a preset data threshold of the root node, set a data aggregation frequency according to the data size of the second test data.
[0217] The dividing module is configured to divide the second test data according to the data aggregation frequency by using a compression algorithm, to generate a plurality of groups of aggregated data to be transmitted.
[0218] In an embodiment of the present application, the second generating module 1204 comprises:
[0219] The first checking unit is configured to check data integrity of the leaf node graphic processor according to the first test information, to generate a first checking result.
[0220] The second checking unit is configured to check data integrity of the root node graphic processor according to the second test information, to generate a second checking result.
[0221] The fifth determining unit is configured to determine data transmission time information and performance information of the graphic processor according to the first check result and the second check result.
[0222] The ninth generating unit is configured to generate a performance test result of the graphic processor according to the data transmission time information and the performance information of the graphic processor.
[0223] In an embodiment of the present application, the second generating module 1204 further comprises:
[0224] The obtaining unit is configured to obtain link quality information among the plurality of graphic processors.
[0225] The judging unit is configured to judge whether the link quality information exceeds a preset link quality threshold.
[0226] The tenth generating unit is configured to generate an optimization strategy according to the link quality information if the link quality information exceeds the preset link quality threshold.
[0227] In an embodiment of the present application, the tenth generating unit comprises:
[0228] The obtaining sub-unit is configured to obtain node distribution information of the graphic processor if the link quality information exceeds the preset link quality threshold.
[0229] The first generating sub-unit is configured to generate a node optimization strategy according to the node distribution information of the graphic processor.
[0230] The second generating sub-unit is configured to generate link connection information of the optimized graphic processor according to the node optimization strategy.
[0231] The determining sub-unit is configured to determine the optimization strategy according to the node optimization strategy and the link connection information of the optimized graphic processor.
[0232] In an embodiment of the present application, the performance testing device 120 of the graphic processor further comprises:
[0233] The second obtaining module is configured to obtain a target graphic processor according to the performance test result if the performance test result is that the performance test fails.
[0234] The first determining module is configured to determine processor node type information according to the target graphic processor.
[0235] The third generating module is configured to generate a node optimization strategy according to the node type information of the processor.
[0236] The optimization module is configured to optimize the target graphic processor according to the node optimization strategy, perform data retransmission on the optimized target graphic processor, and generate optimized performance test data.
[0237] In an embodiment of the present application, the performance testing device 120 of the graphics processor further comprises:
[0238] The third obtaining module is configured to obtain link connection information between the plurality of graphics processors.
[0239] The second determining module is configured to determine a link type according to the link connection information.
[0240] The calculating module is configured to calculate a corresponding data hierarchical bandwidth according to the link type.
[0241] The fourth generating module is configured to generate a node allocation strategy of the graphics processor according to the data hierarchical bandwidth.
[0242] In an embodiment of the present application, the performance testing device 120 of the graphics processor further comprises:
[0243] The fifth generating module is configured to input the first test data and the second test data into a pre-trained task classifier to generate task classification information of the first test data and the second test data.
[0244] The querying module is configured to query a mapping table of transmission tasks and transmission modes according to the task classification information to obtain a transmission mode corresponding to the task classification information.
[0245] The sixth generating module is configured to divide transmission resources of the first test data and the second test data according to the transmission mode to generate resource division information.
[0246] In an embodiment of the present application, the performance testing device 120 of the graphics processor further comprises:
[0247] The fourth obtaining module is configured to obtain resource occupation information of the transmission tasks of the first test data and the second test data by a resource monitoring probe.
[0248] The judging module is configured to judge whether the resource occupation information exceeds a preset resource threshold.
[0249] The seventh generating module is configured to generate a resource division optimization strategy if the resource occupation information exceeds the preset resource threshold, and optimize the mapping table of the transmission tasks and the transmission modes according to the resource division optimization strategy.
[0250] The features of the embodiments of the performance testing device of the graphics processor can be referred to the related descriptions of the embodiments of the performance testing method of the graphics processor, which will not be repeated here.
[0251] Figure 13 The structural schematic diagram of the electronic device provided in the present application is shown in FIG. 1. Figure 13As shown, the electronic device 130 provided by the embodiment includes at least one processor 1301 and a memory 1302. Optionally, the electronic device 130 further includes a communication component 1303. Wherein, the processor 1301, the memory 1302 and the communication component 1303 are connected through a bus.
[0252] In the process of implementation, the at least one processor 1301 executes the computer execution instructions stored in the memory 1302, so that the at least one processor 1301 executes the performance test method embodiment of the graphic processor described above.
[0253] The specific implementation process of the processor 1301 can refer to the method embodiments described above, which has similar implementation principles and technical effects, and will not be described here in detail.
[0254] In the above embodiment, it should be understood that the processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC) and the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the application can be directly embodied as hardware processor execution or executed by hardware and software module combination in the processor.
[0255] The memory can contain a random access memory (RAM), and can also include a non-volatile memory (NVM), such as at least one disk memory.
[0256] The bus can be an industry standard architecture (ISA) bus, a peripheral component (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of the present application does not limit to only one bus or one type of bus.
[0257] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, wherein the computer program is arranged to execute the steps in any of the above-mentioned performance testing methods of a graphic processor.
[0258] In an example embodiment, the above-mentioned computer readable storage medium can include, but is not limited to, a U disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0259] The embodiment of the present application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned performance testing methods of a graphic processor.
[0260] The embodiment of the present application further provides another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned performance testing methods of a graphic processor.
[0261] The skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general in the above description. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0262] The above provides a detailed introduction to the performance testing method, device, equipment, storage medium and program product of a graphic processor. The principle and implementation mode of the present application are described by applying specific examples in this paper, and the above-mentioned example description is only applicable to help understand the method and core idea of the present application. It should be pointed out that for the ordinary skilled in the art, some improvements and modifications can be made to the present application without departing from the principle of the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method of performance testing a graphics processor, characterized by, The method comprises the following steps: generating first test data and second test data according to performance test instructions, wherein the graphic processor comprises a root node graphic processor, a branch node graphic processor and a plurality of leaf node graphic processors, the first test data is test data transmitted from the root node graphic processor to the branch node graphic processor and then transmitted from the branch node graphic processor to the plurality of leaf node graphic processors, and the second test data is test data transmitted from the plurality of leaf node graphic processors to the branch node graphic processor and then transmitted from the branch node graphic processor to the root node graphic processor; generating a forward transmission buffer division strategy according to the first test data, dividing the buffer to obtain a sending buffer of the root node, a receiving buffer of the branch node and a receiving buffer of the leaf node, generating a first transmission instruction and a second transmission instruction according to the first test data, and sending the first test data from the sending buffer of the root node to the receiving buffer of the branch node according to the first transmission instruction, sending the first test data from the receiving buffer of the branch node to the receiving buffer of the leaf node according to the second transmission instruction, and generating first test information; generating a reverse transmission buffer division strategy according to the second test data, dividing the buffer to obtain a receiving buffer of the root node, a receiving buffer of the branch node and a sending buffer of the leaf node, generating a third transmission instruction and a fourth transmission instruction according to the second test data, and sending the second test data from the sending buffer of the leaf node to the receiving buffer of the branch node according to the third transmission instruction, sending the second test data from the receiving buffer of the branch node to the receiving buffer of the root node according to the fourth transmission instruction, and generating second test information; checking the data integrity of the leaf node graphic processor according to the first test information, and generating a first check result; checking the data integrity of the root node graphic processor according to the second test information, and generating a second check result; determining data transmission time information and performance information of the graphic processor according to the first check result and the second check result; generating a performance test result of the graphic processor according to the data transmission time information and the performance information of the graphic processor.
2. The method of claim 1, wherein, The method of generating the first test data and the second test data according to the performance test instructions, wherein the graphic processor comprises a root node graphic processor, a branch node graphic processor and a plurality of leaf node graphic processors, the first test data is test data transmitted from the root node graphic processor to the branch node graphic processor and then transmitted from the branch node graphic processor to the plurality of leaf node graphic processors, and the second test data is test data transmitted from the plurality of leaf node graphic processors to the branch node graphic processor and then transmitted from the branch node graphic processor to the root node graphic processor, comprises the following steps: generating a forward transmission buffer division strategy of the graphic processor according to the first test data; dividing the forward transmission buffer of the root node graphic processor, the branch node graphic processor and the leaf node graphic processor according to the forward transmission buffer division strategy of the graphic processor to obtain a sending buffer of the root node, a receiving buffer of the branch node and a receiving buffer of the leaf node; generating the first test information, comprises the following steps: sending the first test data from the sending buffer of the root node to the receiving buffer of the branch node according to the first transmission instruction, and sending the first test data from the receiving buffer of the branch node to the receiving buffer of the leaf node according to the second transmission instruction. determining forward transmission parameters of the root node graphic processor and the branch node graphic processor according to the first test data; generating first transmission instructions according to the forward transmission parameters of the root node graphic processor and the branch node graphic processor; sending the first test data from the sending buffer of the root node to the receiving buffer of the branch node according to the first transmission instructions through the command queue to obtain first transmission data; determining forward transmission parameters of the branch node graphic processor and the leaf node graphic processor according to the first test data; generating second transmission instructions according to the forward transmission parameters of the branch node graphic processor and the leaf node graphic processor; sending the first test data from the receiving buffer of the branch node to the receiving buffer of the leaf node according to the second transmission instructions through the command queue to obtain second transmission data; generating first test information according to the first transmission data and the second transmission data.
3. The method of claim 1, wherein the performance test of the graphic processor is performed by a plurality of graphic processors. The method further comprises: generating a reverse transmission buffer division strategy of the graphic processor according to the second test data; dividing reverse transmission buffers of the root node graphic processor, the branch node graphic processor and the leaf node graphic processor according to the reverse transmission buffer division strategy of the graphic processor to obtain the receiving buffer of the root node, the receiving buffer of the branch node and the sending buffer of the leaf node; determining reverse transmission parameters of the leaf node graphic processor and the branch node graphic processor according to the second test data; generating third transmission instructions according to the reverse transmission parameters of the leaf node graphic processor and the branch node graphic processor; sending the second test data from the sending buffer of the leaf node to the receiving buffer of the branch node according to the third transmission instructions through the command queue to obtain third transmission data; determining reverse transmission parameters of the branch node graphic processor and the root node graphic processor according to the second test data; generating fourth transmission instructions according to the reverse transmission parameters of the branch node graphic processor and the root node graphic processor; sending the second test data from the receiving buffer of the branch node to the receiving buffer of the root node according to the fourth transmission instructions through the command queue to obtain fourth transmission data; and generating second test information according to the third transmission data and the fourth transmission data. The method further comprises: generating a reverse transmission buffer division strategy of the graphic processor according to the second test data; dividing reverse transmission buffers of the root node graphic processor, the branch node graphic processor and the leaf node graphic processor according to the reverse transmission buffer division strategy of the graphic processor to obtain the receiving buffer of the root node, the receiving buffer of the branch node and the sending buffer of the leaf node; determining reverse transmission parameters of the leaf node graphic processor and the branch node graphic processor according to the second test data; generating third transmission instructions according to the reverse transmission parameters of the leaf node graphic processor and the branch node graphic processor; sending the second test data from the sending buffer of the leaf node to the receiving buffer of the branch node according to the third transmission instructions through the command queue to obtain third transmission data; determining reverse transmission parameters of the branch node graphic processor and the root node graphic processor according to the second test data; generating fourth transmission instructions according to the reverse transmission parameters of the branch node graphic processor and the root node graphic processor; sending the second test data from the receiving buffer of the branch node to the receiving buffer of the root node according to the fourth transmission instructions through the command queue to obtain fourth transmission data; and generating second test information according to the third transmission data and the fourth transmission data. The method further comprises: generating a reverse transmission buffer division strategy of the graphic processor according to the second test data; dividing reverse transmission buffers of the root node graphic processor, the branch node graphic processor and the leaf node graphic processor according to the reverse transmission buffer division strategy of the graphic processor to obtain the receiving buffer of the root node, the receiving buffer of the branch node and the sending buffer of the leaf node; determining reverse transmission parameters of the leaf node graphic processor and the branch node graphic processor according to the second test data; generating third transmission instructions according to the reverse transmission parameters of the leaf node graphic processor and the branch node graphic processor; sending the second test data from the sending buffer of the leaf node to the receiving buffer of the branch node according to the third transmission instructions through the command queue to obtain third transmission data; determining reverse transmission parameters of the branch node graphic processor and the root node graphic processor according to the second test data; generating fourth transmission instructions according to the reverse transmission parameters of the branch node graphic processor and the root node graphic processor; sending the second test data from the receiving buffer of the branch node to the receiving buffer of the root node according to the fourth transmission instructions through the command queue to obtain fourth transmission data; and generating second test information according to the third transmission data and the fourth transmission data. 4. The method of claim 3, wherein the performance test of the graphic processor is performed by using a graphic benchmark program. acquire a data size of the second test data; if the data size of the second test data exceeds a preset data threshold of a root node, set a data aggregation frequency according to the data size of the second test data; divide the second test data according to the data aggregation frequency by using a compression algorithm to generate a plurality of groups of aggregated data to be transmitted.
5. The method of claim 1, wherein the performance test of the graphic processor is performed by a graphic driver. after the performance test result of the graphics processor is generated according to the data transmission time information and the performance information of the graphics processor, the method further comprises: acquiring link quality information between the plurality of graphics processors; judging whether the link quality information exceeds a preset link quality threshold; if the link quality information exceeds the preset link quality threshold, generating an optimization strategy according to the link quality information.
6. The method of claim 5, wherein the performance test of the graphic processor is performed by a plurality of graphic processors. if the link quality information exceeds the preset link quality threshold, the method further comprises: acquiring node distribution information of the graphics processor; generating a node optimization strategy according to the node distribution information of the graphics processor; generating link connection information of the graphics processor after optimization according to the node optimization strategy; determining the optimization strategy according to the node optimization strategy and the link connection information of the graphics processor after optimization.
7. The method of claim 1, wherein the performance test of the graphic processor is performed by a graphic driver. after the performance test result of the graphics processor is generated according to the data transmission time information and the performance information of the graphics processor, the method further comprises: if the performance test result is that the performance test fails, acquiring a target graphics processor according to the performance test result; determining processor node type information according to the target graphics processor; generating a node optimization strategy according to the node type information of the processor; optimizing the target graphics processor according to the node optimization strategy, and performing data retransmission on the target graphics processor after optimization to generate performance test data after optimization.
8. The method of claim 1, wherein, before the first test data and the second test data are generated according to the performance test instruction, the method further comprises: acquiring link connection information between the plurality of graphics processors; determining a link type according to the link connection information; calculating a corresponding data hierarchical bandwidth according to the link type; generating a node allocation strategy of the graphics processor according to the data hierarchical bandwidth.
9. The method of claim 8, wherein, the formula for calculating the data hierarchical bandwidth is: wherein represents the bandwidth of the data hierarchy i; represents the total bandwidth; represents the exponential priority of the data hierarchy i; represents the exponential priority of the total data hierarchy j.
10. The method of claim 1, wherein, before the root node, the branch node and the leaf node are divided according to the forward transmission buffer division strategy generated according to the first test data, the first test data is transmitted from the sending buffer of the root node to the receiving buffer of the branch node according to the first transmission instruction, and the first test data is transmitted from the receiving buffer of the branch node to the receiving buffer of the leaf node according to the second transmission instruction, the method further comprises: inputting the first test data and the second test data into a pre-trained task classifier to generate task classification information of the first test data and the second test data; before the first test information is generated, the method further comprises: inputting the first test data and the second test data into a pre-trained task classifier to generate task classification information of the first test data and the second test data; According to the task classification information, a mapping table of transmission tasks and transmission modes is queried to obtain a transmission mode corresponding to the task classification information; According to the transmission mode, transmission resources of the first test data and the second test data are divided to generate resource division information.
11. The method of claim 10, wherein, After the step of dividing the transmission resources of the first test data and the second test data according to the transmission mode and generating the resource division information, the method further comprises: obtaining resource occupation information of executing the transmission tasks of the first test data and the second test data by a resource monitoring probe; determining whether the resource occupation information exceeds a preset resource threshold; if the resource occupation information exceeds the preset resource threshold, generating a resource division optimization strategy, and optimizing the mapping table of transmission tasks and transmission modes according to the resource division optimization strategy.
12. An electronic device, comprising: comprise: a memory for storing a computer program; a processor for executing the computer program to implement the steps of the performance test method of the graphics processor according to any one of claims 1 to 11.
13. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, wherein the computer program is executed by the processor to implement the steps of the performance test method of the graphics processor according to any one of claims 1 to 11.
14. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the performance test method of the graphics processor according to any one of claims 1 to 11.
Citation Information
Patent Citations
Data transmission test method and device, equipment and storage medium
CN113609056A
Large model test method and device, computer equipment, readable storage medium and program product
CN119829453A