Heterogeneous GPU card management method, device, equipment, medium and product

By centrally managing and coordinating the communication capabilities of each GPU node, the problem of communication difficulties between heterogeneous GPU cards is solved, and efficient inter-node communication connections are achieved.

CN121125609APending Publication Date: 2025-12-12CHINA MOBILE COMM LTD RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510187572.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

The lack of mature communication solutions between existing heterogeneous GPU cards makes communication between heterogeneous GPU cards from different manufacturers difficult, and manually configuring connections between nodes is complex and prone to errors.

Method used

The central management system obtains communication capability information from multiple GPU nodes, plans communication paths, negotiates connection strategies, and coordinates communication capabilities between GPU cards to form communication connections between nodes.

Benefits of technology

It enables effective communication between heterogeneous GPU cards, simplifies the connection configuration process between nodes, and improves communication efficiency and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125609A_ABST
    Figure CN121125609A_ABST
Patent Text Reader

Abstract

The invention provides a heterogeneous GPU card management method and device, equipment, a medium and a product, and the method comprises the steps that a center management end obtains communication capability information of a plurality of GPU node ends participating in communication, and the GPU node ends comprise GPU cards; the central management end plans communication paths among GPU cards in the plurality of GPU node ends according to the communication capability information of the plurality of GPU node ends, obtains communication path negotiation information meeting a preset communication path connection strategy, and sends the communication path negotiation information to the GPU node ends; and the central management end receives feedback information of the communication path negotiation information sent by the GPU node end, and the feedback information comprises communication path negotiation success information, or communication path negotiation failure information and latest communication capability information of the GPU node end. According to the method, the communication capability between the GPU cards is negotiated in a centralized, unified and cooperative manner, and the communication connection between the GPU cards between the nodes is established as required by issuing negotiation information to the communication nodes, so that the communication connection method of the nodes is formed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of computing power network, in particular to a management method and device of heterogeneous GPU cards, equipment, medium and product. BACKGROUND

[0002] Currently, in the application scene of a graphics processing unit (GPU), a same-vendor homogeneous GPU card is usually used to build a GPU computing power center, but the communication is limited between the same-vendor GPUs. There is no mature and available communication scheme between heterogeneous GPU cards of different vendors. The communication between heterogeneous GPU cards of different vendors is currently being explored and verified, and has great limitations.

[0003] In addition, because the products of different vendors must have multiple specifications and support capabilities, it is difficult to use a unified link and protocol to solve the communication problem between heterogeneous cards. The links and protocols supported between different nodes in the network may be different or have multiple combinations and selections. In a large-scale computing power node, manually selecting and configuring the connection between nodes will be a complex, time-consuming and error-prone process. SUMMARY

[0004] Embodiments of the present application provide a management method, device, equipment, medium and product of heterogeneous GPU cards to solve the problem that there is no mature and available communication scheme between existing heterogeneous GPU cards.

[0005] To solve the above technical problems, the present application is implemented as follows:

[0006] In a first aspect, the embodiments of the present application provide a management method of heterogeneous GPU cards, comprising:

[0007] The central management end obtains the communication capability information of a plurality of GPU node ends participating in communication, wherein the GPU node ends include GPU cards;

[0008] The central management end plans a communication path between the GPU cards in the plurality of GPU node ends according to the communication capability information of the plurality of GPU node ends, obtains communication path negotiation information satisfying a preset communication path connection strategy, and sends the communication path negotiation information to the GPU node ends;

[0009] The central management end receives feedback information of the communication path negotiation information sent by the GPU node ends, wherein the feedback information includes communication path negotiation success information, or communication path negotiation failure information and the latest communication capability information of the GPU node ends.

[0010] Optionally, the communication capability information comprises at least one of the following: node identification, GPU card type, link type, communication protocol information, addressing type, node address, GPU communication protocol or library, collective communication capability and bandwidth information.

[0011] Optionally, the central management end obtains communication capability information of a plurality of GPU node ends participating in communication, comprising:

[0012] The central management end queries whether the communication capability information of each GPU node end exists in a pre-stored communication capability database;

[0013] If the communication capability information of the GPU node end exists, the communication capability information is obtained from the communication capability database;

[0014] If the communication capability information of the GPU node end does not exist, the central management end sends a request for obtaining communication capability information to the GPU node end and obtains the communication capability information of the GPU node end.

[0015] Optionally, after the central management end obtains the communication capability information of the plurality of GPU node ends participating in communication, the method further comprises:

[0016] If the central management end cannot obtain communication path negotiation information satisfying a preset communication path connection strategy according to the communication capability information, it is prompted that the communication capability is insufficient.

[0017] Optionally, if the feedback information is communication path negotiation success information, the central management end sends communication path negotiation success confirmation information to the GPU node end sending the communication path negotiation success information;

[0018] If the feedback information is communication path negotiation failure information, the central management end updates the communication capability information of the GPU node end sending the communication path negotiation failure information, re-plans the communication path between the plurality of GPU node ends, again obtains communication path negotiation information satisfying the preset communication path connection strategy and sends the communication path negotiation failure information to the GPU node end.

[0019] In a second aspect, an embodiment of the present application provides a management method of a heterogeneous GPU card, comprising:

[0020] A GPU node end sends communication capability information of the GPU node end to the central management end, wherein the GPU node end comprises a GPU card;

[0021] The GPU node end receives communication path negotiation information of a communication path between GPU cards in the GPU node end satisfying a preset communication path connection strategy sent by the central management end.

[0022] The GPU node end performs verification processing on the communication path in the communication path negotiation information according to the current network state or communication capability, obtains feedback information, and sends the feedback information to the central management end, wherein the feedback information includes communication path negotiation success information or communication path negotiation failure information and the latest communication capability information of the GPU node end.

[0023] Optionally, the communication capability information includes at least one of the following: node identifier, GPU card type, link type, communication protocol information, addressing type, node address, GPU communication protocol or library, collective communication capability and bandwidth information, and other capability information.

[0024] Optionally, the GPU node end receives the communication path negotiation success confirmation information sent by the central management end,

[0025] or receives the communication path negotiation information that satisfies the preset communication path connection strategy again obtained by the central management end, and performs verification processing on the communication path in the communication path negotiation information according to the current network state or communication capability, obtains feedback information, and sends the feedback information to the central management end, wherein the feedback information includes communication path negotiation success information or communication path negotiation failure information and the latest communication capability information of the GPU node end.

[0026] Optionally, the verification processing includes at least one of the following: availability verification and delay test on the communication path in the communication path negotiation information, and judgment on the network state of the GPU node end.

[0027] If at least one of the verification processing fails, communication path negotiation failure information is obtained as the feedback information.

[0028] If all of the verification processing is successful, communication path negotiation success information is obtained as the feedback information.

[0029] In a third aspect, an embodiment of the present application provides a central management end, comprising:

[0030] A first obtaining module is configured to obtain communication capability information of a plurality of GPU node ends participating in communication, wherein the GPU node ends include GPU cards.

[0031] A first processing module is configured to plan a communication path between the GPU cards in the plurality of GPU node ends according to the communication capability information of the plurality of GPU node ends, obtain communication path negotiation information satisfying a preset communication path connection strategy, and send the communication path negotiation information to the GPU node ends.

[0032] The first receiving module is configured to receive feedback information of the communication path negotiation information sent by the GPU node end, wherein the feedback information comprises communication path negotiation success information, or communication path negotiation failure information and the latest communication capability information of the GPU node end.

[0033] Optionally, the communication capability information comprises at least one of the following: node identifier, GPU card type, link type, communication protocol information, addressing type, node address, GPU communication protocol or library, collective communication capability and bandwidth information.

[0034] Optionally, the first obtaining module comprises:

[0035] The querying module is configured to query whether the communication capability information of each GPU node end exists in a pre-stored communication capability database; if the communication capability information of the GPU node end exists, the communication capability information is obtained from the communication capability database; if the communication capability information of the GPU node end does not exist, the central management end sends a communication capability information obtaining request to the GPU node end and obtains the communication capability information of the GPU node end.

[0036] Optionally, the method further comprises:

[0037] The prompting module is configured to prompt insufficient communication capability if the central management end cannot obtain the communication path negotiation information satisfying the preset communication path connection strategy according to the communication capability information.

[0038] Optionally, if the feedback information is the communication path negotiation success information, the central management end sends communication path negotiation success confirmation information to the GPU node end sending the communication path negotiation success information.

[0039] If the feedback information is the communication path negotiation failure information, the central management end updates the communication capability information of the GPU node end sending the communication path negotiation failure information, re-plans the communication path between the plurality of GPU node ends, and again obtains the communication path negotiation information satisfying the preset communication path connection strategy and sends the communication path negotiation information to the GPU node end sending the communication path negotiation failure information.

[0040] In a fourth aspect, an embodiment of the present application provides a GPU node end, comprising:

[0041] The first sending module is configured to send communication capability information of the GPU node end to the central management end, wherein the GPU node end comprises a GPU card.

[0042] The second receiving module is configured to receive communication path negotiation information of a communication path between GPU cards in the GPU node end satisfying the preset communication path connection strategy sent by the center management end.

[0043] The second processing module is configured to perform verification processing on the communication path in the communication path negotiation information according to a current network state or a communication capability, obtain feedback information, and send the feedback information to the center management end, wherein the feedback information comprises communication path negotiation success information or communication path negotiation failure information and latest communication capability information of the GPU node end.

[0044] Optionally, the communication capability information comprises at least one of the following: node identification, GPU card type, link type, communication protocol information, addressing type, node address, GPU communication protocol or library, collective communication capability and bandwidth information, and other capability information.

[0045] Optionally, the method further comprises:

[0046] The third processing module is configured to receive communication path negotiation success confirmation information sent by the center management end,

[0047] Optionally, the fourth processing module is configured to receive again communication path negotiation information satisfying the preset communication path connection strategy sent by the center management end, perform verification processing on the communication path in the communication path negotiation information according to a current network state or a communication capability, obtain feedback information, and send the feedback information to the center management end, wherein the feedback information comprises communication path negotiation success information or communication path negotiation failure information and latest communication capability information of the GPU node end.

[0048] Optionally, the verification processing comprises at least one of the following: availability verification and delay testing on the communication path in the communication path negotiation information, and judgment on a self network state.

[0049] If at least one verification in the verification processing fails, communication path negotiation failure information is obtained as the feedback information.

[0050] If all verifications in the verification processing are successful, communication path negotiation success information is obtained as the feedback information.

[0051] In a fifth aspect, an electronic device is provided, comprising a processor, a memory, and a program or instructions stored in the memory and executable on the processor, and the program or instructions are executed by the processor to implement the management method of the heterogeneous GPU card according to any one of the first aspect, or to implement the steps in the management method of the heterogeneous GPU card according to any one of the second aspect.

[0052] In a sixth aspect, an embodiment of the present application provides a readable storage medium, wherein a program or instruction is stored on the readable storage medium, and the program or instruction is executed by a processor to implement the management method of the heterogeneous GPU card according to any one of the first aspect, or to implement the steps in the management method of the heterogeneous GPU card according to any one of the second aspect.

[0053] In a seventh aspect, an embodiment of the present application provides a computer program product, comprising computer instructions, which are executed by a processor to implement the management method of the heterogeneous GPU card according to any one of the first aspect, or to implement the steps in the management method of the heterogeneous GPU card according to any one of the second aspect.

[0054] In the present application, the central management end obtains communication capability information of a plurality of GPU node ends participating in communication, wherein the GPU node ends comprise GPU cards; the central management end plans a communication path between the GPU cards in the plurality of GPU node ends according to the communication capability information of the plurality of GPU node ends, obtains communication path negotiation information satisfying a preset communication path connection strategy, and sends the communication path negotiation information to the GPU node ends; the central management end receives feedback information of the communication path negotiation information sent by the GPU node ends, wherein the feedback information comprises communication path negotiation success information, or communication path negotiation failure information and latest communication capability information of the GPU node ends. The communication capability between the GPU cards is collectively and uniformly negotiated, the negotiation information is sent to each communication node, the GPU card interconnection between nodes is established on demand, the communication connection method of each node is formed, and the problem that there is no mature and available communication scheme between the existing heterogeneous GPU cards is solved. BRIEF DESCRIPTION OF DRAWINGS

[0055] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments, and are not intended to limit the scope of the application. Moreover, like reference numerals designate like parts throughout the several views in the drawings. In the drawings:

[0056] Figure 1 FIG. 1 is a flowchart of a management method of a heterogeneous GPU card according to an embodiment of the present application applied to a central management end;

[0057] Figure 2 FIG. 2 is a general flowchart of a management method of a heterogeneous GPU card according to an embodiment of the present application;

[0058] Figure 3 FIG. 3 is a communication capability matrix diagram of a management method of a heterogeneous GPU card according to an embodiment of the present application;

[0059] Figure 4is another communication capability matrix schematic diagram of the heterogeneous GPU card management method provided by the embodiment of the present application;

[0060] Figure 5 is another communication capability matrix schematic diagram of the heterogeneous GPU card management method provided by the embodiment of the present application;

[0061] Figure 6 is another communication capability matrix schematic diagram of the heterogeneous GPU card management method provided by the embodiment of the present application;

[0062] Figure 7 is another communication capability matrix schematic diagram of the heterogeneous GPU card management method provided by the embodiment of the present application;

[0063] Figure 8 is a flowchart of the heterogeneous GPU card management method applied to the GPU node end provided by the embodiment of the present application;

[0064] Figure 9 is a structure schematic diagram of the center management end provided by the embodiment of the present application;

[0065] Figure 10 is a structure schematic diagram of the GPU node end provided by the embodiment of the present application;

[0066] Figure 11 is a structure schematic diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION

[0067] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0068] Please refer to Figure 1 The embodiment of the present application provides a heterogeneous GPU card management method, which comprises the following steps:

[0069] Step 11: The center management end obtains the communication capability information of the plurality of GPU node ends participating in communication, wherein the GPU node ends comprise GPU cards;

[0070] In the embodiment of the present application, the communication between the plurality of GPU node ends is cross-node GPU card intercommunication, as shown in the figure, the communication connection between the heterogeneous GPU cards is uniformly and cooperatively managed by the center management end, when the application or service needs the intercard cooperation of the heterogeneous GPU cards, the center management end is requested to communicate and cooperate, and the appropriate connection is established. Figure 2 ​

[0071] Optionally, the communication capability information comprises at least one of the following: node identifier, GPU card type, link type, communication protocol information, addressing type, node address, GPU communication protocol or library, collective communication capability and bandwidth information.

[0072] In the embodiment of the application, the link type can be, but is not limited to, at least one of the following: IB_link and ETH_link, etc.; the communication protocol information can be, but is not limited to, at least one of the following: IB, IPoIB, RoCEv1, RoCEv2, iWARP and TCP / IP; the GPU communication protocol or library can be, but is not limited to, at least one of the following: NCCL, XHCCL and HCCL; the collective communication capability can be, but is not limited to, at least one of the following: reduce, broadcast, scatter, gather and all-to-all.

[0073] Step 12: The central management end plans the communication path between the GPU cards in the plurality of GPU node ends according to the communication capability information of the plurality of GPU node ends, obtains communication path negotiation information satisfying a preset communication path connection strategy, and sends the communication path negotiation information to the GPU node ends;

[0074] In the embodiment of the application, the central management end collects the communication capability of each participating GPU node end, and constructs a communication capability matrix between nodes as shown in Figure 3 、 Figure 4 and Figure 5 , that is, a pre-stored communication capability database, wherein the GPU node end is each type of GPU node end, including at least one GPU card; the node identifier is used to identify the identifier of each node by the management end, such as using an in-band management port IP as the node identifier in the embodiment; the node address is used for the service address of communication, and the type is different according to the addressing type; the communication capability between nodes can contain various link capabilities, such as supporting IB link and also supporting Ethernet link; for the same link mode, there can be multiple links, such as multiple ETH links.

[0075] In the embodiment of the application, the connection between each participating GPU node end is planned according to the request information, the best connection mode between nodes is selected or a suitable connection is selected according to a specified strategy as the communication path negotiation information, and each GPU node end is issued, wherein the strategy can be, but is not limited to, set based on bandwidth, link, communication protocol type, GPU communication protocol and the like requirements.

[0076] Optionally, in the embodiment of the present application, after the central management terminal obtains the communication capability information of the plurality of GPU node terminals participating in the communication, the central management terminal further comprises:

[0077] If the central management terminal cannot obtain the communication path negotiation information satisfying the preset communication path connection strategy according to the communication capability information, the central management terminal prompts that the communication capability is insufficient.

[0078] In the embodiment of the present application, if the current capability matrix information cannot satisfy the requested capability of the application or the service, the central management terminal prompts that the communication capability is insufficient, and feeds back the information that the capability does not satisfy to the service or application request terminal.

[0079] In the embodiment of the present application, the central management terminal cooperatively plans and issues each GPU node terminal to negotiate, thereby forming the communication connection between each GPU card and solving the communication capability connection and configuration problem between GPU cards of different manufacturers.

[0080] Step 13: The central management terminal receives the feedback information of the communication path negotiation information sent by the GPU node terminal, and the feedback information comprises communication path negotiation success information or communication path negotiation failure information and the latest communication capability information of the GPU node terminal.

[0081] In the embodiment of the present application, the central management terminal receives the feedback information of each GPU node terminal, and the feedback information comprises communication path negotiation success information or communication path negotiation failure information and the latest communication capability information of the GPU node terminal; for each connection, the related nodes of the negotiation success connection send negotiation success confirmation information, and each node receives the confirmation information, and constructs the communication connection matrix of the node according to the negotiation success connection information, as shown in Figure 6 and Figure 7 The communication connection matrix information is constructed by taking the GPU card as the object; the application service number is used as the identification request multi-card cooperation, and the application service intercommunication between cards is dynamically obtained by the central management terminal. The communication capability information of each GPU card node is constructed to form the communication capability matrix between nodes; when there is a communication demand between heterogeneous GPU cards, the central management terminal cooperatively plans and issues each GPU node terminal to negotiate, thereby forming the communication connection between each GPU card and solving the communication capability connection and configuration problem between GPU cards of different manufacturers.

[0082] In the embodiment of the present application, the central management end obtains communication capability information of a plurality of GPU node ends participating in communication, the GPU node ends including GPU cards; the central management end plans a communication path between the GPU cards in the plurality of GPU node ends according to the communication capability information of the plurality of GPU node ends, obtains communication path negotiation information satisfying a preset communication path connection strategy, and sends the communication path negotiation information to the GPU node ends; the central management end receives feedback information of the communication path negotiation information sent by the GPU node ends, the feedback information including communication path negotiation success information or communication path negotiation failure information and latest communication capability information of the GPU node ends. The communication capability between the GPU cards is collectively and uniformly negotiated, the negotiation information is sent to each communication node, the GPU card interconnection between nodes is established on demand, the communication connection method of each node is formed, and the problem that there is no mature and available communication scheme between existing heterogeneous GPU cards is solved.

[0083] In the embodiment of the present application, the central management end obtains communication capability information of a plurality of GPU node ends participating in communication, including:

[0084] The central management end queries whether the communication capability information of each GPU node end exists in the pre-stored communication capability database;

[0085] If the communication capability information of the GPU node end exists, the communication capability information is obtained from the communication capability database;

[0086] If the communication capability information of the GPU node end does not exist, the central management end sends a communication capability information acquisition request to the GPU node end and obtains the communication capability information of the GPU node end.

[0087] In the embodiment of the present application, the central management end first queries whether the communication capability information of the current request to each communication node exists in the communication capability matrix, that is, the pre-stored communication capability database. If the communication capability information exists, the communication capability information of the plurality of GPU node ends participating in communication is directly obtained, and subsequent steps are executed. If the communication capability information of the GPU node end does not exist, a communication capability information acquisition request is sent to the GPU node end, and corresponding information is queried from the network management platform in the case of the network management platform, so as to collectively and uniformly negotiate the communication capability between the GPU cards.

[0088] In the embodiment of the present application, if the feedback information is the communication path negotiation success information, the central management end sends communication path negotiation success confirmation information to the GPU node end sending the communication path negotiation success information;

[0089] If the feedback information is the communication path negotiation failure information, the central management end updates the communication capability information of the GPU node end sending the communication path negotiation failure information, and re-plans the communication path between the plurality of GPU node ends, and obtains again the communication path negotiation information satisfying the preset communication path connection strategy and sends to the GPU node end of the communication path negotiation failure information.

[0090] In the embodiment of the application, the central management end receives the feedback information responded by each GPU node end, the feedback information including the communication path negotiation success information or the communication path negotiation failure information and the latest communication capability information of the GPU node end; for each connection, the related nodes of the connection with the feedback information being the communication path negotiation success information are sent the negotiation success confirmation information, and after each node receives the confirmation information, the communication connection matrix of the node is constructed according to the negotiation success connection information, as shown in Figure 6 and Figure 7 If the feedback information is the communication path negotiation failure information, the communication capability matrix information is refreshed according to the failure reason and the capability information, and the above steps are repeated according to the requirement to re-plan the related nodes and re-issue the negotiation information.

[0091] Please refer to Figure 8 The embodiment of the application provides a management method of heterogeneous GPU cards, including:

[0092] Step 81: The GPU node end sends the communication capability information of the GPU node end to the central management end, the GPU node end including a GPU card;

[0093] In the embodiment of the application, the communication between the plurality of GPU node ends is the cross-node GPU card communication, as shown in Figure 2 The communication connection between the heterogeneous GPU cards is uniformly and cooperatively negotiated and managed by the central management end, and when the application or the business needs the inter-card cooperation of the heterogeneous GPU cards, the central management end is requested to communicate and negotiate to establish a suitable connection.

[0094] In the embodiment of the application, the communication capability information includes at least one of the following: node identification, GPU card type, link type, communication protocol information, addressing type, node address, GPU communication protocol or library, collective communication capability and bandwidth information.

[0095] In the embodiment of the present application, the link type can be but is not limited to at least one of IB_link and ETH_link, etc.; the communication protocol information can be but is not limited to at least one of IB, IPoIB, RoCEv1, RoCEv2, iWARP and TCP / IP; the GPU communication protocol or library can be but is not limited to at least one of NCCL, XHCCL and HCCL; and the collective communication capability can be but is not limited to at least one of reduce, broadcast, scatter, gather and all-to-all.

[0096] Step 82: The GPU node end receives the communication path negotiation information of the communication path between GPU cards in the GPU node end satisfying the preset communication path connection strategy sent by the center management end.

[0097] In the embodiment of the present application, the center management end selects the best connection mode between nodes or selects a suitable connection as the communication path negotiation information according to the request information, and then issues each GPU node end, wherein the strategy can be but is not limited to set based on bandwidth, link, communication protocol type, GPU communication protocol, etc. The way of collaborative planning and issuing each GPU node end by the center management end to negotiate forms the communication connection between each GPU card, and solves the communication capability connection and configuration problem between GPU cards of different manufacturers.

[0098] Step 83: The GPU node end verifies the communication path in the communication path negotiation information according to the current network state or communication capability, obtains feedback information, and sends the feedback information to the center management end, wherein the feedback information includes communication path negotiation success information, or communication path negotiation failure information and the latest communication capability information of the GPU node end.

[0099] In the embodiment of the present application, the verification processing includes at least one of the following: availability verification and delay test of the communication path in the communication path negotiation information, and judgment of the network state of itself.

[0100] If at least one of the verification processing fails, the communication path negotiation failure information is obtained as the feedback information.

[0101] If all items in the verification processing are verified successfully, the communication path negotiation success information is obtained as the feedback information.

[0102] The verification processing of the negotiation information sent by the center end by each GPU node end includes at least one of the following: availability verification and delay test on the communication path in the communication path negotiation information, and judgment on the network state of the GPU node end itself.

[0103] Specifically, if the network state of each GPU node end does not change or the change does not affect the current negotiation, the GPU node end can perform availability verification and delay test on the negotiated communication path according to needs or strategies, such as sending test information.

[0104] When the test fails, the GPU node end returns negotiation failure information to the center management end, requests the center management end to update the capability matrix information, and performs step 12 to re-negotiate; if there is no problem, the communication path negotiation information is confirmed, and communication path negotiation success information and necessary related information are returned; the availability verification of the communication path can be one-way or two-way verification.

[0105] If the network state of the GPU node end changes and the change causes the current negotiation to be unavailable, the capability information is updated, the center management end is requested to update the node capability matrix, and step 12 is performed to re-negotiate; by dynamically obtaining the communication capability information of each GPU card node from the center management end, the communication capability matrix between nodes is constructed to increase the verification step and improve the negotiation efficiency, and the communication capability connection and configuration problems between GPU cards of different manufacturers are solved.

[0106] In the embodiment of the application, the GPU node end receives the communication path negotiation success confirmation information sent by the center management end,

[0107] Alternatively, the GPU node end receives the communication path negotiation information again obtained by satisfying the preset communication path connection strategy sent by the center management end, and performs verification processing on the communication path in the communication path negotiation information according to the current network state or communication capability, obtains feedback information, and sends the feedback information to the center management end; the feedback information includes communication path negotiation success information, or communication path negotiation failure information and the latest communication capability information of the GPU node end.

[0108] In the embodiment of the application, the center management end receives the feedback information responded by each GPU node end; the feedback information includes communication path negotiation success information, or communication path negotiation failure information and the latest communication capability information of the GPU node end; the related nodes of the communication path with successful negotiation in each connection send negotiation success confirmation information; after each node receives the confirmation information, the communication connection matrix of the node is constructed according to the information of the communication path with successful negotiation. Figure 6 and Figure 7As shown, the communication connection matrix information is constructed for the GPU card as an object; the service number is used to identify the application service of requesting multi-card cooperation and intercommunication between cards; the communication capability information of each GPU card node is dynamically acquired through the central management end to construct the communication capability matrix between nodes; when there is a communication demand between heterogeneous GPU cards, the communication connection between the GPU cards is formed in a collaborative planning and delivery manner by the central management end to negotiate each GPU node end, and the communication capability connection and configuration problems between GPU cards of different manufacturers are solved.

[0109] Please refer to Figure 9 The embodiment of the application provides a central management end, which comprises:

[0110] The first acquisition module 91 is configured to acquire communication capability information of a plurality of GPU node ends participating in communication, wherein the GPU node ends comprise GPU cards.

[0111] The first processing module 92 is configured to plan a communication path between the GPU cards in the plurality of GPU node ends according to the communication capability information of the plurality of GPU node ends, obtain communication path negotiation information meeting a preset communication path connection strategy, and send the communication path negotiation information to the GPU node ends.

[0112] The first receiving module 93 is configured to receive feedback information of the communication path negotiation information sent by the GPU node ends, wherein the feedback information comprises communication path negotiation success information or communication path negotiation failure information and latest communication capability information of the GPU node ends.

[0113] Optionally, the communication capability information comprises at least one of the following: node identifier, GPU card type, link type, communication protocol information, addressing type, node address, GPU communication protocol or library, collective communication capability and bandwidth information.

[0114] In the embodiment of the application, the first acquisition module comprises:

[0115] The query module is configured to query whether the communication capability information of each GPU node end exists in a pre-stored communication capability database; if the communication capability information of the GPU node end exists, the communication capability information is acquired from the communication capability database; if the communication capability information of the GPU node end does not exist, the central management end sends a communication capability information acquisition request to the GPU node end and acquires the communication capability information of the GPU node end.

[0116] In the embodiment of the application, the central management end further comprises:

[0117] The prompting module is configured to prompt a communication capability shortage if the central management end cannot obtain the communication path negotiation information satisfying the preset communication path connection strategy according to the communication capability information.

[0118] In the embodiment of the application, optionally, if the feedback information is the communication path negotiation success information, the central management end sends communication path negotiation success confirmation information to the GPU node end sending the communication path negotiation success information.

[0119] If the feedback information is the communication path negotiation failure information, the central management end updates the communication capability information of the GPU node end sending the communication path negotiation failure information, and re-plans the communication path between the plurality of GPU node ends, and again obtains the communication path negotiation information satisfying the preset communication path connection strategy and sends it to the GPU node end sending the communication path negotiation failure information.

[0120] The central management end provided by the embodiment of the application can realize Figure 1 the various processes realized by the method embodiment and achieve the same technical effects. To avoid repetition, details are not repeated here.

[0121] Please refer to Figure 10 The embodiment of the application provides a GPU node end, comprising:

[0122] The first sending module 101 is configured to send the communication capability information of the GPU node end to the central management end, wherein the GPU node end comprises a GPU card.

[0123] The second receiving module 102 is configured to receive the communication path negotiation information of the communication path between the GPU cards in the GPU node end satisfying the preset communication path connection strategy sent by the central management end.

[0124] The second processing module 103 is configured to verify the communication path in the communication path negotiation information according to the current network state or the communication capability, obtain feedback information, and send the feedback information to the central management end, wherein the feedback information comprises communication path negotiation success information, or communication path negotiation failure information and the latest communication capability information of the GPU node end.

[0125] In the embodiment of the application, optionally, the communication capability information comprises at least one of the following: node identifier, GPU card type, link type, communication protocol information, addressing type, node address, GPU communication protocol or library, collective communication capability and bandwidth information, and other capability information.

[0126] In the embodiment of the application, optionally, the GPU node end further comprises:

[0127] The third processing module is used to receive the communication path negotiation success confirmation information sent by the central management terminal.

[0128] Alternatively, the fourth processing module is used to receive communication path negotiation information that satisfies the preset communication path connection strategy again sent by the central management terminal, and to verify the communication path in the communication path negotiation information according to the current network status or communication capability, to obtain feedback information, and to send the feedback information to the central management terminal. The feedback information includes communication path negotiation success information, or communication path negotiation failure information and the latest communication capability information of the GPU node.

[0129] In this embodiment of the invention, optionally, the verification process includes at least one of the following: performing availability verification and latency testing on the communication path in the communication path negotiation information, and judging its own network status;

[0130] If at least one of the verification processes fails, a communication path negotiation failure message is obtained as feedback information.

[0131] If all items in the verification process are successfully verified, a successful communication path negotiation message is obtained as feedback information.

[0132] The GPU node provided in this embodiment of the invention can achieve Figure X The various processes implemented in the method embodiments achieve the same technical effect, and will not be described again here to avoid repetition.

[0133] This invention provides an electronic device 110, see [link to relevant documentation]. Figure 11 As shown, Figure 11 This is a schematic block diagram of an electronic device 110 according to an embodiment of the present invention, including a processor 111, a memory 112, and a program or instructions stored in the memory 112 and executable on the processor 111. When the program or instructions are executed by the processor, they implement the steps in any of the heterogeneous GPU card management methods of the present invention.

[0134] This invention provides a readable storage medium on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements the various processes of the embodiments of the heterogeneous GPU card management method as described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0135] This application also provides a computer program product, including computer instructions. When executed by a processor, these computer instructions implement the various processes of the method embodiments shown in the figures above and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0136] Computer-readable media includes permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0137] It should be noted that in the technical solutions of the present disclosure, the collection, collection, update, analysis, processing, use, transmission, storage and other aspects of user personal information are in line with relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken to prevent illegal access to user personal information data, and user personal information security and network security are maintained.

[0138] It should be noted that in this paper, the term "includes", "contains" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such process, method, article or device. Without more limitations, the element defined by the statement "includes a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0139] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0140] Those skilled in the art can clearly understand the above-mentioned embodiment method can be realized by means of software and necessary general hardware platform, of course, also can be through hardware, but in many cases the former is the better embodiment. Based on such understanding, the technical scheme of the present application essentially or the part which contributes to the prior art can be embodied in the form of software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions to make a service classification device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute the method described in each embodiment of the present application.

[0141] The above only describes the preferred embodiments of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which should be considered as the protection scope of the present application.

Claims

1. A method for managing heterogeneous GPU cards, characterized in that, include: The central management terminal obtains communication capability information of multiple GPU nodes participating in the communication, wherein the GPU nodes include GPU cards; The central management terminal plans the communication path between the GPU cards in the multiple GPU nodes based on the communication capability information of the multiple GPU nodes, obtains the communication path negotiation information that satisfies the preset communication path connection strategy, and sends the communication path negotiation information to the GPU node. The central management terminal receives feedback information from the GPU node regarding the communication path negotiation information. The feedback information includes either a successful communication path negotiation or a failed communication path negotiation, along with the latest communication capability information of the GPU node.

2. The management method for heterogeneous GPU cards according to claim 1, characterized in that, The communication capability information includes at least one of the following: node identifier, GPU card type, link type, communication protocol information, addressing type, node address, GPU communication protocol or library, collective communication capability, and bandwidth information.

3. The management method for heterogeneous GPU cards according to claim 1, characterized in that, The central management terminal obtains communication capability information from multiple GPU nodes participating in the communication, including: The central management terminal queries the pre-stored communication capability database to check whether the communication capability information of each GPU node exists; If the communication capability information of the GPU node exists, then the communication capability information is obtained from the communication capability database; If the communication capability information of the GPU node does not exist, the central management terminal sends a request to the GPU node to obtain the communication capability information and obtains the communication capability information of the GPU node.

4. The management method for heterogeneous GPU cards according to claim 1, characterized in that, After the central management terminal obtains the communication capability information of the multiple GPU nodes participating in the communication, it also includes: If the central management terminal cannot obtain communication path negotiation information that satisfies the preset communication path connection strategy based on the communication capability information, it will prompt that the communication capability is insufficient.

5. The management method for heterogeneous GPU cards according to claim 1, characterized in that, If the feedback information is a successful communication path negotiation message, then the central management terminal sends a successful communication path negotiation confirmation message to the GPU node that sent the successful communication path negotiation message. If the feedback information is a communication path negotiation failure, the central management terminal updates the communication capability information of the GPU node that sent the communication path negotiation failure information, replans the communication path between the multiple GPU nodes, obtains communication path negotiation information that satisfies the preset communication path connection strategy again, and sends it to the GPU node that sent the communication path negotiation failure information.

6. A method for managing heterogeneous GPU cards, characterized in that, include: The GPU node sends its communication capability information to the central management terminal, wherein the GPU node includes a GPU card. The GPU node receives communication path negotiation information from the central management terminal, which specifies the communication path between GPU cards in the GPU node that satisfies a preset communication path connection strategy. The GPU node verifies the communication path in the communication path negotiation information based on the current network status or communication capabilities, obtains feedback information, and sends the feedback information to the central management terminal. The feedback information includes communication path negotiation success information, or communication path negotiation failure information and the latest communication capability information of the GPU node.

7. The method for managing heterogeneous GPU cards according to claim 6, characterized in that, The communication capability information includes at least one of the following: node identifier, GPU card type, link type, communication protocol information, addressing type, node address, GPU communication protocol or library, collective communication capability and bandwidth information, and other capability information.

8. The method for managing heterogeneous GPU cards according to claim 6, characterized in that, The GPU node receives a confirmation message from the central management terminal indicating that the communication path negotiation was successful. Alternatively, the system receives communication path negotiation information that satisfies the preset communication path connection strategy again, sent by the central management terminal, and verifies the communication path in the communication path negotiation information according to the current network status or communication capability to obtain feedback information. The system then sends the feedback information to the central management terminal. The feedback information includes communication path negotiation success information, or communication path negotiation failure information and the latest communication capability information of the GPU node.

9. The method for managing heterogeneous GPU cards according to claim 6, characterized in that, The verification process includes at least one of the following: performing availability verification and latency testing on the communication path in the communication path negotiation information, and judging its own network status; If at least one of the verification processes fails, a communication path negotiation failure message is obtained as feedback information. If all items in the verification process are successfully verified, a successful communication path negotiation message is obtained as feedback information.

10. A central management terminal, characterized in that, include: The first acquisition module is used to acquire communication capability information of multiple GPU nodes participating in the communication, wherein the GPU node includes a GPU card; The first processing module is used to plan the communication path between the GPU cards in the multiple GPU nodes according to the communication capability information of the multiple GPU nodes, obtain the communication path negotiation information that satisfies the preset communication path connection strategy, and send the communication path negotiation information to the GPU node. The first receiving module is used to receive feedback information of the communication path negotiation information sent by the GPU node. The feedback information includes communication path negotiation success information, or communication path negotiation failure information and the latest communication capability information of the GPU node.

11. A GPU node, characterized in that, include: The first sending module is used to send communication capability information of the GPU node to the central management terminal, wherein the GPU node includes a GPU card; The second receiving module is used to receive communication path negotiation information sent by the central management terminal, which specifies the communication path between GPU cards in the GPU node terminal that satisfies a preset communication path connection strategy. The second processing module is used to verify the communication path in the communication path negotiation information according to the current network status or communication capability, obtain feedback information, and send the feedback information to the central management terminal. The feedback information includes communication path negotiation success information, or communication path negotiation failure information and the latest communication capability information of the GPU node.

12. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein when the program or instructions are executed by the processor, they implement the management method of a heterogeneous GPU card as described in any one of claims 1 to 5, or implement the steps in the management method of a heterogeneous GPU card as described in any one of claims 6 to 9.

13. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions, which, when executed by a processor, implement the heterogeneous GPU card management method as described in any one of claims 1 to 5, or implement the steps in the heterogeneous GPU card management method as described in any one of claims 6 to 9.

14. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the management method for a heterogeneous GPU card as described in any one of claims 1 to 5, or implement the steps in the management method for a heterogeneous GPU card as described in any one of claims 6 to 9.