Network card scheduling method and device, equipment, storage medium and program product

By screening and binding the most suitable network card resources in a multi-network card architecture, the problem of low resource utilization in traditional methods is solved, the rational allocation and efficient utilization of network card resources are achieved, and the overall performance of the computing cluster is improved.

CN120675993APending Publication Date: 2025-09-19SUGON INFORMATION IND +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511012887.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional network card scheduling methods have the problem of low resource utilization in multi-network card architectures, resulting in communication bottlenecks between computing nodes in multi-network card systems, limiting the improvement of large model training efficiency.

Method used

By screening candidate network cards from the computing cluster's resource pool based on the target processor's scheduling requirements, and performing multiple screenings based on the transmission link's attribute information, the transmission performance is evaluated, and the target network card is finally scheduled to the target processor for binding and operation, ensuring the rational allocation and efficient utilization of network card resources.

Benefits of technology

It improves the utilization rate of network card resources, avoids the problem of some high-performance network cards being idle, fully taps the potential of each network card, realizes the call of all network cards, and improves the overall performance of the computing cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120675993A_ABST
    Figure CN120675993A_ABST
Patent Text Reader

Abstract

The invention relates to a network card scheduling method and device, equipment, a storage medium and a program product, and the method comprises the steps: screening out a needed network card from a resource pool of a calculation cluster according to the scheduling demand of a target processor in a target node, and obtaining a plurality of candidate network cards; according to the attribute information of the transmission links between the target processor and the plurality of candidate network cards, the required network cards are screened out from the plurality of candidate network cards to obtain a target network card, and finally the target network card is scheduled to the target processor for binding operation. According to the method, the functions of the network cards are considered when the plurality of candidate network cards are obtained through primary screening, the transmission performance is considered when the target network cards are obtained through secondary screening, the target network cards obtained through multiple times of screening integrate various factors, the actual requirements of the target processor can be better matched, reasonable distribution of network card resources can be achieved, and the user experience is improved. And thus, the network card resource utilization rate in the whole computing cluster is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication technology, and in particular to a network card scheduling method, apparatus, device, storage medium and program product. Background Art

[0002] In the current fields of machine learning and deep learning, as the scale of models and datasets continues to increase, traditional single-NIC architectures can no longer meet the needs of large-scale model training. Most current computing systems adopt multi-NIC architectures. Multi-NIC architectures distribute computing tasks across multiple NICs or compute nodes, leveraging the advantages of parallel computing to accelerate model training. Currently, while multi-NIC architectures hold great potential for training large models, they face more challenges than single-NIC architectures in practical applications. One of these challenges is the communication bottleneck between compute nodes in multi-NIC systems. Parameters and gradient information must be frequently exchanged between nodes to maintain model synchronization and consistency. Therefore, optimizing the communication efficiency between multiple NICs is a critical issue and plays a significant role in improving the efficiency of large model training.

[0003] However, when using traditional network card scheduling methods to schedule network cards in a multi-network card architecture for communication, there is a problem of low resource utilization. Summary of the Invention

[0004] Based on this, it is necessary to provide a network card scheduling method, device, equipment, storage medium and program product that can improve resource utilization in response to the above technical problems.

[0005] In a first aspect, the present application provides a network card scheduling method, applied to a target node in a computing cluster, the method comprising:

[0006] According to the scheduling requirements of the target processor in the target node, the required network cards are screened from the resource pool of the computing cluster to obtain multiple candidate network cards;

[0007] Based on the attribute information of the transmission link between the target processor and multiple candidate network cards, the desired network card is screened out from the multiple candidate network cards to obtain the target network card; the attribute information is used to evaluate the transmission performance of the transmission link from multiple dimensions;

[0008] Schedule the target network card to the target processor for binding and running.

[0009] In some embodiments, the method of selecting a desired network card from multiple candidate network cards based on attribute information of a transmission link between a target processor and multiple candidate network cards to obtain a target network card includes:

[0010] Scoring each candidate network card based on attribute information of a transmission link between a target processor and multiple candidate network cards to obtain a score value for each candidate network card;

[0011] The candidate network card corresponding to the largest score value is determined as the target network card.

[0012] In some embodiments, scoring each candidate network card based on attribute information of a transmission link between a target processor and multiple candidate network cards to obtain a score value for each candidate network card includes:

[0013] Evaluate the transmission performance of the transmission link in each dimension and obtain the score of each candidate network card in each dimension;

[0014] According to the score values ​​of each candidate network card in each dimension, the score value of each candidate network card is obtained.

[0015] In some embodiments, the attribute information includes rate information of the candidate network card, distance information between the candidate network card and the target processor, and channel information between the candidate network card and the target processor. The transmission performance of the transmission link is evaluated in various dimensions to obtain a score value for each candidate network card in each dimension, including:

[0016] Based on the rate information of each candidate network card, the transmission performance of the transmission link is evaluated to obtain the rate dimension score item;

[0017] Based on the distance information between each candidate network card and the target processor, the transmission performance of the transmission link is evaluated to obtain a distance dimension score item;

[0018] Based on the channel information between each candidate network card and the target processor, the transmission performance of the transmission link is evaluated to obtain the channel dimension score item;

[0019] Based on the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item, the score value of each candidate network interface card in each dimension is determined.

[0020] In some embodiments, the transmission performance of the transmission link is evaluated based on the rate information of each candidate network interface card to obtain a rate dimension score item, including:

[0021] According to the rate information of each candidate network card, the maximum rate corresponding to the network card with the maximum rate capability among all candidate network cards is obtained;

[0022] Based on the rate information and maximum rate of each candidate network card, the transmission performance of the transmission link is evaluated to obtain the rate dimension score item.

[0023] In some embodiments, the transmission performance of the transmission link is evaluated based on the distance information between each candidate network card and the target processor to obtain a distance dimension score item, including:

[0024] According to the distance information between each candidate network card and the target processor, the shortest distance corresponding to the network card with the shortest distance to the target processor among all candidate network cards is obtained;

[0025] Based on the rate information and shortest distance of each candidate network card, the transmission performance of the transmission link is evaluated to obtain the distance dimension score item.

[0026] In some embodiments, based on the channel information between each candidate network card and the target processor, the transmission performance of the transmission link is evaluated to obtain a channel dimension score item, including:

[0027] If the channel information indicates that a target channel exists between the corresponding candidate network card and the target processor, determining the channel dimension score item to be a first value;

[0028] If the channel information indicates that there is no target channel between the corresponding candidate network card and the target processor, the channel dimension score item is determined to be a second value.

[0029] In some embodiments, the score value of each candidate network interface card in each dimension is determined based on the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item, including:

[0030] Assign corresponding weights to the rate dimension scoring items, distance dimension scoring items, and channel dimension scoring items based on scheduling requirements;

[0031] Based on the weighted rate, distance, and channel dimension scoring items, the score values ​​of each candidate network interface card in each dimension are determined.

[0032] In some embodiments, based on the scheduling requirements of the target processor in the target node, the required network cards are screened from the resource pool of the computing cluster to obtain multiple candidate network cards, and the multiple candidate network cards include:

[0033] Determine the required network card type based on the scheduling requirements of the target processor in the target node; the network card type includes the sending type and the receiving type;

[0034] A network card of a type corresponding to the network card type is screened out from a resource pool of the computing cluster to obtain multiple candidate network cards.

[0035] The method described in the embodiments of the present application determines the required network card type based on the scheduling requirements of the target processor, thereby ensuring that the selected network card is highly matched with the functional requirements of the processor, ensuring that the network card efficiently completes the data transmission task, and avoiding problems such as poor performance or inability to implement functions due to the use of an incompatible network card.

[0036] In some embodiments, the method further comprises:

[0037] According to the configuration information of the computing cluster, the binding relationship between the target processor and each network card in the resource pool, as well as the communication link information between the target processor and each network card are obtained;

[0038] When scheduling the target processor, the scheduling requirements of the target processor are determined based on the binding relationship and communication link information.

[0039] In some embodiments, the target network card includes a target sending network card and / or a target receiving network card, and scheduling the target network card to the target processor for binding operation includes:

[0040] Schedule the target sending network card to the sending port of the target processor for binding and running;

[0041] And / or, dispatching the target receiving network card to the receiving port of the target processor for binding and operation.

[0042] In a second aspect, the present application further provides a network card scheduling device, the device comprising:

[0043] The first screening module is used to screen out the required network cards from the resource pool of the computing cluster according to the scheduling requirements of the target processor in the target node, and obtain multiple candidate network cards;

[0044] A second screening module is configured to screen a desired network card from the plurality of candidate network cards based on attribute information of a transmission link between the target processor and the plurality of candidate network cards, thereby obtaining a target network card; the attribute information is used to evaluate the transmission performance of the transmission link from multiple dimensions;

[0045] The scheduling module is used to schedule the target network card to the target processor for binding and running.

[0046] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0047] According to the scheduling requirements of the target processor in the target node, the required network cards are screened from the resource pool of the computing cluster to obtain multiple candidate network cards;

[0048] Based on the attribute information of the transmission link between the target processor and multiple candidate network cards, the desired network card is screened out from the multiple candidate network cards to obtain the target network card; the attribute information is used to evaluate the transmission performance of the transmission link from multiple dimensions;

[0049] Schedule the target network card to the target processor for binding and running.

[0050] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps:

[0051] According to the scheduling requirements of the target processor in the target node, the required network cards are screened from the resource pool of the computing cluster to obtain multiple candidate network cards;

[0052] Based on the attribute information of the transmission link between the target processor and multiple candidate network cards, the desired network card is screened out from the multiple candidate network cards to obtain the target network card; the attribute information is used to evaluate the transmission performance of the transmission link from multiple dimensions;

[0053] Schedule the target network card to the target processor for binding and running.

[0054] In a fifth aspect, the present application further provides a computer program product, the computer program product comprising a computer program, which, when executed by a processor, implements the following steps:

[0055] According to the scheduling requirements of the target processor in the target node, the required network cards are screened from the resource pool of the computing cluster to obtain multiple candidate network cards;

[0056] Based on the attribute information of the transmission link between the target processor and multiple candidate network cards, the desired network card is screened out from the multiple candidate network cards to obtain the target network card; the attribute information is used to evaluate the transmission performance of the transmission link from multiple dimensions;

[0057] Schedule the target network card to the target processor for binding and running.

[0058] The above-mentioned network card scheduling method, apparatus, device, storage medium and program product screen the required network cards from the resource pool of the computing cluster according to the scheduling requirements of the target processor in the target node to obtain multiple candidate network cards. Then, based on the attribute information of the transmission link between the target processor and the multiple candidate network cards, the required network card is screened from the multiple candidate network cards to obtain the target network card. Finally, the target network card is scheduled to the target processor for binding and operation. The above-mentioned method considers the network card function when initially screening multiple candidate network cards, and considers the transmission performance when secondary screening to obtain the target network card. The target network card obtained through multiple screenings integrates multiple factors and can better match the actual needs of the target processor, thereby achieving reasonable allocation of network card resources. This avoids the problem of traditional methods that only consider a single factor and continuously call some high-performance network cards, causing other network cards to be idle and unused. The above-mentioned method can fully tap the potential of each network card, realize the call of all network cards, and thus improve the utilization rate of network card resources in the entire computing cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 A diagram showing an application environment of a network card scheduling method in some embodiments;

[0060] Figure 2 This is one of the flowcharts of the network card scheduling method in some embodiments;

[0061] Figure 3This is a second flow chart of a network card scheduling method in some embodiments;

[0062] Figure 4 This is a third flow chart of a network card scheduling method in some embodiments;

[0063] Figure 5 This is a fourth flowchart of a network card scheduling method in some embodiments;

[0064] Figure 6 This is a fifth flowchart of a network card scheduling method in some embodiments;

[0065] Figure 7 This is a sixth flowchart of a network card scheduling method in some embodiments;

[0066] Figure 8 This is a seventh flowchart of a network card scheduling method in some embodiments;

[0067] Figure 9 This is an eighth flowchart of a network card scheduling method in some embodiments;

[0068] Figure 10 This is a ninth flowchart of a network card scheduling method in some embodiments;

[0069] Figure 11 This is a tenth flowchart of a network card scheduling method in some embodiments;

[0070] Figure 12 11 is a flowchart of a network card scheduling method in some embodiments;

[0071] Figure 13 Schematic diagram showing the effects before and after using the network card scheduling method in some embodiments;

[0072] Figure 14 is a structural block diagram of a network card scheduling device in some embodiments;

[0073] Figure 15 1 is a diagram of the internal structure of a computer device in some embodiments. DETAILED DESCRIPTION

[0074] In the embodiments of this application, the term "and / or" describes the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0075] In the embodiments of the present application, the term "plurality" refers to two or more than two, and other quantifiers are similar.

[0076] In the embodiments of the present application, the term "at least one" means one or more. For example, at least one of A, B and C can mean the following six situations: A exists alone, B exists alone, C exists alone, A and B exist at the same time, A and C exist at the same time, B and C exist at the same time, and A, B and C exist at the same time.

[0077] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0078] In the current fields of machine learning and deep learning, large model training is a major focus. Large models typically have a large number of parameters and require processing large datasets to improve model performance and accuracy. However, as the size of models and datasets continues to increase, traditional single-NIC architectures are no longer sufficient for large-scale model training. To address the shortcomings of single-NIC architectures, most current computing systems adopt multi-NIC architectures. Multi-NIC architectures distribute computing tasks across multiple NICs or compute nodes, leveraging the advantages of parallel computing to accelerate the training process. By fully utilizing multiple computing resources, multi-NIC architectures can significantly reduce training time and improve model training efficiency.

[0079] Currently, although multi-NIC architectures have great potential in large-scale model training, they face more challenges than single-NIC architectures in practical applications. One of them is the communication bottleneck between computing nodes in multi-NIC systems. Parameters and gradient information need to be frequently exchanged between nodes to maintain model synchronization and consistency. This communication process can become a bottleneck in the training process, limiting the improvement of overall training performance. Therefore, optimizing the communication efficiency between multiple NICs is an important issue and is of great significance to improving the efficiency of large-scale model training. However, when the NICs in a multi-NIC architecture are scheduled for communication using traditional NIC scheduling methods, there is a problem of low NIC resource utilization.

[0080] In view of this, the embodiments of the present application propose a network card scheduling method, device, equipment, storage medium and program product, which can comprehensively consider various information, perform multiple screening operations from the network card resource pool, obtain the target network card, and schedule the target network card to improve the utilization rate of network card resources.

[0081] It should be noted that the beneficial effects or technical problems solved by the embodiments of the present application are not limited to this one, but may also include other implicit or related problems. For details, please refer to the description of the following embodiments.

[0082] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0083] In some embodiments, the network card scheduling method provided in the embodiments of the present application can be applied to Figure 1 In the application environment shown, the computing cluster in the application environment includes multiple computing nodes 10 (two nodes are used as an example in the figure), and each computing node 10 has multiple communication links. Each communication link 101 includes multiple processors 1011 and multiple network cards 1012. Figure 1 In the example, multiple processors are specifically a~h) including input processor (see Figure 1 a, e in ), multiple intermediate processors (see Figure 1 b, c, g, f) and output processors (see Figure 1 d, h in); multiple network cards 1012 (in Figure 1 In the example, multiple network cards are A to H, including input network cards associated with input processors and output network cards associated with output processors. Figure 1Taking each full-duplex network card as an example (i.e., the network card can be used for sending and receiving at the same time), in actual applications, the network card may not be a full-duplex network card, i.e., it can only be used for receiving or sending. Each computing node can communicate with each other through the network card to form multiple ring transmission links. Taking the computing node on the left as an example to illustrate the specific composition structure of the communication link, the communication links include AabcdD, DabcdA, CabcdB, BabcdC; the specific composition structure of the ring transmission link between two computing nodes is illustrated, and the ring transmission link includes: AabcdDEerghHA, DabcdAHerghED and other 4 transmission links. Among them, the computing node 10 is used to screen out the best input network card and output network card for the input position and output position of each communication link from the network card resource pool, and then schedule the best input network card and output network card to the corresponding input position and output position to obtain the input network card and output network card. Optionally, a default network card can be bound to an input location or output location. In this case, the computing node 10 only needs to select the optimal network card from the network card resource pool and bind it to the input location or output location that is not bound to a default network card. The computing node 10 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. The computing node 10 can also be implemented as a standalone server or a server cluster consisting of multiple servers. The processor 1011 can be a central processing unit (CPU) or a graphics processing unit (GPU).

[0084] Those skilled in the art will understand that Figure 1 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the application environment to which the solution of the present application is applied. The specific application environment may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0085] In some embodiments, as Figure 2 As shown, a network card scheduling method is provided, which is applied to Figure 1 Taking any computing node in the computing cluster (hereinafter referred to as the target node) as an example, the following steps are included:

[0086] S201 , according to the scheduling requirements of the target processor in the target node, screen out the required network cards from the resource pool of the computing cluster to obtain multiple candidate network cards.

[0087] The target processor is the input processor or output processor of the communication link in the target node. The scheduling requirements include functional requirements, which include receiving functional requirements and sending functional requirements. If the target processor is an input processor, then the scheduling requirement is that a network card with receiving function is required (i.e., receiving functional requirement). If the target processor is an output processor, then the scheduling requirement is that a network card with sending function is required (i.e., sending functional requirement). The resource pool is a network card resource pool, which is shared by multiple computing nodes in the computing cluster. The resource pool includes network cards with receiving function, network cards with sending function, and network cards with both sending and receiving functions. The candidate network card is the network card obtained by the target node through initial screening from the resource pool.

[0088] In an embodiment of the present application, the target node may first determine a target processor from a plurality of processors. Specifically, the target node may determine a plurality of target processors at a time, and then schedule a network card for each target processor in parallel. Optionally, the target node may determine the target processor step by step in sequence, and then schedule a network card for the target processor in sequence. The method of determining the target processor step by step in sequence may be: determining the target processor according to the number of the processor, for example, first scheduling the network card for the processor with a higher number, and then scheduling the network card for the processor with a lower number; optionally, the target processor may also be determined according to the performance of the processor, for example, first scheduling the network card for the processor with better performance, and then scheduling the network card for the processor with slightly lower performance; optionally, the target processor may also be determined according to the type of processor, for example, first calling the network card for the input processor, and then calling the network card for the output processor.

[0089] After determining the target processor, the target node can further analyze the network card function type required by the target processor based on the system architecture information of the computing cluster, and then search for all normal available network cards from the resource pool of the computing cluster, and then traverse and search for network cards that meet the network card function type required by the target processor based on the network card information of all available network cards as preliminary screening network cards, and then determine the network cards that are currently in an idle state from the preliminary screening network cards as candidate network cards. It should be noted that the target node can obtain the system architecture information of the computing cluster in advance from a preset path or a preset storage location. The system architecture information includes information about each communication link between the processor and the network card in the computing node, and information about each transmission link between each computing node. The system architecture information of the computing cluster is stored in a fixed file uploaded by the user in advance, and can also be stored in a dynamic file updated by the user during the scheduling process.

[0090] S202 , based on the attribute information of the transmission link between the target processor and the multiple candidate network cards, select a required network card from the multiple candidate network cards to obtain a target network card.

[0091] Among them, the attribute information is used to evaluate the transmission performance of the transmission link from multiple dimensions and may include information from multiple dimensions. The attribute information includes at least one of the rate information of the candidate network card, the distance information between the candidate network card and the target processor, and the channel information between the candidate network card and the target processor. The transmission link is part of the communication link, specifically referring to the link between the target processor and the network card. For example, it may include the link between the input processor and the input network card, and the link between the output processor and the output network card. The target network card is the network card obtained by the target node through secondary screening from multiple candidate network cards.

[0092] In an embodiment of the present application, a target node may obtain attribute information of a transmission link between a target processor and multiple candidate network cards based on the system architecture information of a computing cluster, and then filter out a desired network card from the multiple candidate network cards based on the attribute information and preset screening conditions to obtain a target network card. The preset screening conditions may include at least one of the following: a rate of the candidate network card is less than a preset rate threshold, a distance between the candidate network card and the target processor is less than a preset distance threshold, and a channel between the candidate network card and the target processor satisfies a preset channel. Specifically, when the attribute information includes rate information, if the rate of the candidate network card is less than a preset rate threshold, it is determined that the candidate network card meets the preset screening condition; if the rate of the candidate network card is not less than the preset rate threshold, it is determined that the candidate network card does not meet the preset screening condition; when the attribute information includes distance information, if the distance between the candidate network card and the target processor is less than a preset distance threshold, it is determined that the candidate network card meets the preset screening condition; if the distance between the candidate network card and the target processor is not less than the preset distance threshold, it is determined that the candidate network card does not meet the preset screening condition; when the attribute information includes channel information, if the candidate network card and the target processor support a preset channel, it is determined that the candidate network card meets the preset screening condition; if the candidate network card and the target processor do not support the preset channel, it is determined that the candidate network card does not meet the preset screening condition. When the attribute information includes two or all of rate information, distance information, and channel information, if all of them meet the corresponding conditions, it is determined that the candidate network card meets the preset screening condition; if one of them does not meet the corresponding conditions, it is determined that the candidate network card does not meet the preset screening condition.

[0093] S203, dispatching the target network card to the target processor for binding and running.

[0094] In an embodiment of the present application, after the target node obtains the target network card based on the above steps, it can schedule the target network card to the target processor, associate the target network card with the target processor, and update the status information of the target network card at the same time. Specifically, the status of the target network card under the currently occupied function is updated from the idle state to the occupied state. For example, the target network card itself can receive and send at the same time, and the function type required by the target processor is reception. Then the update content is to update the status of the receiving function of the target network card from the idle state to the occupied state.

[0095] The network card scheduling method provided in the embodiment of the present application is to screen out the required network card from the resource pool of the computing cluster according to the scheduling requirements of the target processor in the target node, obtain multiple candidate network cards, and then screen out the required network card from the multiple candidate network cards according to the attribute information of the transmission link between the target processor and the multiple candidate network cards, obtain the target network card, and finally schedule the target network card to the target processor for binding operation. The above method takes into account the function of the network card when obtaining multiple candidate network cards in the initial screening, and considers the transmission performance when obtaining the target network card in the secondary screening. The target network card obtained by multiple screenings integrates various factors, can better match the actual needs of the target processor, and can achieve reasonable allocation of network card resources. It avoids the problem that the traditional method only considers a single factor and continuously calls some network cards with good performance, causing other network cards to be idle and unused. The above method can fully tap the potential of each network card, realize the call of all network cards, and thus improve the utilization rate of network card resources in the entire computing cluster.

[0096] In some embodiments, a specific implementation method for selecting a desired network card from multiple candidate network cards is also provided, such as Figure 3 As shown, the above-mentioned S202 of "screening out a desired network card from the multiple candidate network cards according to the attribute information of the transmission link between the target processor and the multiple candidate network cards to obtain the target network card" includes:

[0097] S301 , scoring each candidate network card according to attribute information of a transmission link between a target processor and multiple candidate network cards, to obtain a score value of each candidate network card.

[0098] In an embodiment of the present application, the target node can obtain the attribute information of the transmission link between the target processor and multiple candidate network cards based on the system architecture information of the computing cluster, and then input the attribute information into a preset scoring model, and score the comprehensive performance of each candidate network card through the scoring model to obtain a score value for each candidate network card. Among them, the scoring model is used to score the comprehensive performance of each candidate network card based on the attribute information of the transmission link between the target processor and multiple candidate network cards. The scoring model can be obtained by training an initial neural network model in advance based on the attribute sample information of the transmission link between the target processor and multiple candidate network cards, a model constructed based on a mathematical relationship, or a machine learning algorithm model.

[0099] Optional, such as Figure 4 As shown, the target node can score each candidate network card based on attribute information of multiple dimensions. Specifically, the above S301 may include:

[0100] S3011: Evaluate the transmission performance of the transmission link in each dimension to obtain a score value of each candidate network card in each dimension.

[0101] The attribute information of each dimension includes the rate information of the candidate network card, the distance information between the candidate network card and the target processor, and the channel information between the candidate network card and the target processor.

[0102] In an embodiment of the present application, the target node can obtain attribute information of each dimension on the transmission link between the target processor and multiple candidate network cards based on the system architecture information of the computing cluster, and then input the attribute information of each dimension into a preset scoring model. The scoring model is used to score the comprehensive performance of each candidate network card in each dimension to obtain a score value for each candidate network card.

[0103] S3012: Obtain a score value of each candidate network card according to the score value of each candidate network card in each dimension.

[0104] In an embodiment of the present application, the target node obtains the score value of each candidate network card in each dimension based on the above steps, and can sum the score value of each candidate network card in each dimension to obtain the total score value of each candidate network card, and use the total score value as the score value of each candidate network card.

[0105] S302: Determine the candidate network card corresponding to the largest score value as the target network card.

[0106] In an embodiment of the present application, after the target node obtains the score values ​​of each candidate network card based on the above steps, it can compare the score values ​​of each candidate network card to obtain the maximum score value, and then determine the candidate network card corresponding to the maximum score value as the target network card. Optionally, the target node can sort the score values ​​of each candidate network card in order from large to small, and then use the first score as the maximum score value. Optionally, the target node can sort the score values ​​of each candidate network card in order from small to large, and then use the last score as the maximum score value.

[0107] The method described in the embodiment of the present application can comprehensively and objectively evaluate the performance of each candidate network card by scoring based on the attribute information of the transmission link between the target processor and multiple candidate network cards. Different transmission link attributes will affect the performance of the network card in actual use. Scoring based on these attributes avoids the one-sidedness that may be caused by considering only a single factor, thereby more accurately screening out the most suitable network card. The transmission performance of the transmission link is evaluated from multiple dimensions such as the rate information of the candidate network card, the distance information between the candidate network card and the target processor, and the channel information, further refining the scoring criteria. Each dimension reflects the characteristics of the network card in a certain aspect. The multi-dimensional scoring can more accurately capture the advantages and disadvantages of the network card in different aspects, so that the target network card finally screened out can achieve a better balance in all aspects, thereby improving the accuracy of the screening results.

[0108] In some embodiments, the attribute information of each dimension includes the rate information of the candidate network card, the distance information between the candidate network card and the target processor, and the channel information between the candidate network card and the target processor. On this basis, a specific implementation method for evaluating the transmission performance of the transmission link in each dimension is also provided, such as Figure 5 As shown, the "evaluating the transmission performance of the transmission link in each dimension to obtain the score value of each candidate network interface card in each dimension" in S3011 includes:

[0109] S401 : Evaluate the transmission performance of the transmission link based on the rate information of each candidate network interface card to obtain a rate dimension score item.

[0110] The rate dimension score item represents the performance score of the candidate network card in the rate dimension.

[0111] In an embodiment of the present application, the target node can extract the rate information of each candidate network card from the resource pool information, and then input the rate information of each candidate network card into a preset scoring model. The scoring model is used to score the transmission performance of the transmission link of each candidate network card under the rate dimension to obtain a rate dimension scoring item.

[0112] Optional, such as Figure 6 As shown, the above S401 may further include:

[0113] S4011: Obtain the maximum rate corresponding to the network card with the maximum rate capability among all candidate network cards according to the rate information of each candidate network card.

[0114] In an embodiment of the present application, after obtaining the rate information of each candidate network card, the target node can filter out the maximum rate value from the rate information of each candidate network card as the maximum rate, and use the candidate network card corresponding to the maximum rate as the network card with the maximum rate capability.

[0115] S4012: Evaluate the transmission performance of the transmission link based on the rate information and maximum rate of each candidate network card to obtain a rate dimension score item.

[0116] In the embodiment of the present application, after obtaining the rate information and maximum rate of each candidate network card, the target node can input the rate information and maximum rate of each candidate network card into the rate scoring sub-model in the scoring model, and use the rate scoring sub-model to evaluate the transmission performance of the transmission link to obtain the rate dimension score item. The rate scoring sub-model is used to divide the rate information and maximum rate of each candidate network card, which can be expressed by the following relationship:

[0117] ;

[0118] in, represents the rate dimension scoring item, Indicates the rate information of each candidate network card. Indicates the maximum rate.

[0119] S402 : Evaluate the transmission performance of the transmission link based on the distance information between each candidate network card and the target processor to obtain a distance dimension score item.

[0120] The distance dimension score item represents the performance score of the candidate network card in the distance dimension.

[0121] In an embodiment of the present application, the target node can extract the distance information between each candidate network card and the target processor from the system architecture information, and then input the distance information between each candidate network card and the target processor into a preset scoring model. The scoring model is used to score the transmission performance of the transmission link of each candidate network card in the distance dimension to obtain a distance dimension scoring item.

[0122] Optional, such as Figure 7 As shown, the above S402 may further include:

[0123] S4021 , according to the distance information between each candidate network card and the target processor, obtain the shortest distance corresponding to the network card with the shortest distance to the target processor among all the candidate network cards.

[0124] In an embodiment of the present application, after obtaining the distance information between each candidate network card and the target processor, the target node can filter out the minimum distance value from the distance information of each candidate network card as the shortest distance, and use the candidate network card corresponding to the shortest distance as the network card with the shortest distance to the target processor.

[0125] S4022: Evaluate the transmission performance of the transmission link based on the rate information and the shortest distance of each candidate network card to obtain a distance dimension score item.

[0126] In the embodiment of the present application, after obtaining the rate information and shortest distance of each candidate network card, the target node can input the rate information and shortest distance of each candidate network card into the distance scoring sub-model in the scoring model, and use the distance scoring sub-model to evaluate the transmission performance of the transmission link to obtain a distance dimension score item. The distance scoring sub-model is used to divide the shortest distance and the rate information of each candidate network card, which can be expressed by the following relationship:

[0127] ;

[0128] in, represents the rate dimension scoring item, Indicates the rate information of each candidate network card. Indicates the shortest distance.

[0129] S403 : Evaluate the transmission performance of the transmission link based on the channel information between each candidate network card and the target processor to obtain a channel dimension score item.

[0130] The channel information indicates whether the candidate network card and the target processor support GPU Remote Direct Memory Access (GPUDirect RDMA). The channel dimension score item indicates the performance score of the candidate network card in the channel dimension.

[0131] In an embodiment of the present application, the target node can extract the channel information between each candidate network card and the target processor from the system architecture information, and then input the channel information between each candidate network card and the target processor into a preset scoring model. The scoring model is used to score the transmission performance of the transmission link of each candidate network card in the channel dimension to obtain a channel dimension scoring item.

[0132] Optional, such as Figure 8 As shown, the above S403 may further include:

[0133] S4031: If the channel information indicates that a target channel exists between the corresponding candidate network card and the target processor, determine that the channel dimension score item is a first value.

[0134] S4032: If the channel information indicates that there is no target channel between the corresponding candidate network card and the target processor, determine that the channel dimension score item is a second value.

[0135] The first value is greater than the second value, for example, the first value is 1 and the second value is 0.

[0136] In an embodiment of the present application, after obtaining the channel information between each candidate network card and the target processor, the target node can input the channel information between each candidate network card and the target processor into the channel scoring sub-model in the scoring model, evaluate the transmission performance of the transmission link through the channel scoring sub-model, and obtain a channel dimension scoring item. Specifically, if the channel information indicates that there is a target channel between the corresponding candidate network card and the target processor, the channel dimension scoring item is determined to be a first value; if the channel information indicates that there is no target channel between the corresponding candidate network card and the target processor, the channel dimension scoring item is determined to be a second value. Among them, the channel scoring sub-model is used to score whether there is a target channel between the candidate network card and the target processor, which can be expressed by the following relationship:

[0137] ;

[0138] in, represents the channel dimension scoring item, is the first value or the second value.

[0139] S404 : Determine the score value of each candidate network interface card in each dimension according to the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item.

[0140] In the embodiment of the present application, after the target node obtains the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item based on the above steps, it can sum the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item to obtain the score value of each candidate network card in each dimension. Specifically, it can be expressed by the following relationship:

[0141] ;

[0142] in, Indicates the score of each candidate network card in each dimension.

[0143] Optional, such as Figure 9 As shown, the above S404 may further include:

[0144] S4041 , assign corresponding weights to the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item according to the scheduling requirements.

[0145] Among them, the scheduling requirements also include transmission scenario requirements, and the transmission scenario requirements include at least one of high-speed data transmission scenarios, short-range communication scenarios, complex network topology scenarios, and low-latency scenarios.

[0146] In an embodiment of the present application, the target node can obtain the transmission scenario requirements during the computing cluster transmission process, and then assign corresponding weights to the rate dimension scoring items, distance dimension scoring items, and channel dimension scoring items according to different scenario requirements.

[0147] Specifically, in high-speed data transmission scenarios, when scheduling requirements prioritize high-speed data transmission (such as real-time data processing in large data centers or high-definition video streaming), the rate dimension is the most critical and should be assigned a higher weight. The distance dimension has a certain impact on data transmission latency, but is less important than the rate. The channel dimension, as long as it meets basic data transmission channel requirements, can have a relatively lower weight. Therefore, the rate dimension weight can be set to 0.7, the distance dimension weight to 0.2, and the channel dimension weight to 0.1.

[0148] In short-range communication scenarios, if scheduling requirements primarily focus on short-range communications (such as communications between devices within an enterprise LAN or within a computer room), the impact of distance is relatively small, and rate remains the key factor in ensuring fast data transmission. The channel dimension plays a role in ensuring data transmission stability and parallelism, so its weight can be appropriately increased. For this purpose, the weight of rate can be set to 0.6, the weight of distance to 0.1, and the weight of channel to 0.3.

[0149] In complex network topology scenarios (such as multi-device interconnection in large data centers and distributed system communication in cloud computing environments), where multiple channels are required for data distribution and transmission, the channel dimension becomes crucial and should be assigned a higher weight. The rate and distance dimensions also need to be considered, but their weighting is slightly lower than that of the channel dimension. Therefore, the rate dimension can be weighted 0.3, the distance dimension 0.2, and the channel dimension 0.5.

[0150] For low-latency scenarios with extremely high latency requirements, such as real-time gaming and financial transactions, both the rate and distance dimensions have a direct impact on latency and should be assigned higher weights. The channel dimension has a relatively small impact on latency and can be weighted lower. For this reason, the rate dimension weight can be set to 0.5, the distance dimension weight to 0.4, and the channel dimension weight to 0.1.

[0151] S4042 : Determine the score value of each candidate network interface card in each dimension according to the weighted rate dimension scoring item, distance dimension scoring item, and channel dimension scoring item.

[0152] In an embodiment of the present application, the target node obtains the weights corresponding to the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item based on the above steps, and then multiplies the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item with the corresponding weights to obtain the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item after the weights are assigned, and then sums the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item after the weights are assigned to obtain the score value of each candidate network card in each dimension.

[0153] The method described in the embodiments of the present application, on the one hand, comprehensively considers the rate information of the candidate network card, the distance information from the target processor, and the channel information between the two; the rate information directly reflects the data transmission capability of the network card, the distance information affects the delay and loss of signal transmission, and the channel information reflects the specific conditions of the network connection; this multi-dimensional evaluation method can comprehensively and accurately reflect the actual performance of the transmission link, avoiding the one-sidedness that may be caused by evaluating only from a single indicator. On the other hand, according to the scheduling requirements, the scoring items of each dimension are assigned corresponding weights. Different transmission scenarios have different performance requirements for network cards. By flexibly adjusting the weights, the evaluation results can be made more consistent with the actual scheduling requirements, improving the adaptability and performance of the system in different scenarios. On the other hand, through accurate evaluation and comprehensive scoring, the network card that best suits the current scheduling requirements can be selected from multiple candidate network cards, which helps to improve the utilization efficiency of network card resources and avoid resource waste caused by improper selection; and the above method can dynamically adjust the weights and scores according to real-time scheduling requirements, thereby optimizing the selection of network cards in real time. This dynamic adjustment mechanism enables the system to better adapt to changing environments and maintain efficient resource utilization.

[0154] In some embodiments, a specific implementation method for filtering out the required network cards from the resource pool of the computing cluster is also provided, such as Figure 10 As shown, the above-mentioned step S201 of "screening the required network cards from the resource pool of the computing cluster according to the scheduling requirements of the target processor in the target node to obtain multiple candidate network cards" includes:

[0155] S501 : Determine the required network card type according to the scheduling requirements of the target processor in the target node.

[0156] The network card type includes the sending type and the receiving type.

[0157] In an embodiment of the present application, after determining the target processor, the target node can further analyze the scheduling requirements of the target processor based on the system architecture information of the computing cluster, that is, analyze the network card function type required by the target processor. If the scheduling requirement of the target processor is a sending function requirement, the required network card type is determined to be a sending type. If the scheduling requirement of the target processor is a receiving function requirement, the required network card type is determined to be a receiving type.

[0158] S502 , filtering out network cards of a type corresponding to the network card type from a resource pool of the computing cluster to obtain multiple candidate network cards.

[0159] In the embodiment of the present application, after the target points obtain the required network card type based on the above steps, network cards of the type corresponding to the network card type can be screened out from the resource pool of the computing cluster to obtain multiple candidate network cards.

[0160] The method described in the embodiments of the present application determines the required network card type based on the scheduling requirements of the target processor, thereby ensuring that the selected network card is highly matched with the functional requirements of the processor, ensuring that the network card efficiently completes the data transmission task, and avoiding problems such as poor performance or inability to implement functions due to the use of an incompatible network card.

[0161] In some embodiments, as Figure 11 As shown, the above network card scheduling method further includes:

[0162] S601 , according to the configuration information of the computing cluster, obtain the binding relationship between the target processor and each network card in the resource pool, and the communication link information between the target processor and each network card.

[0163] The computing cluster configuration information includes system architecture information, the binding relationship between the target processor and each network card in the resource pool, and the communication link information between the target processor and each network card. The binding relationship represents the corresponding relationship between the functions of the target processor and the required network card. The communication link information includes the input processor, the input location connected to the input processor, the output processor, and the output location connected to the output processor. Optionally, the communication link information also includes whether the input location and output location are bound to the default network card.

[0164] In an embodiment of the present application, the target node can obtain the configuration information of the computing cluster in advance from a preset path or a preset storage location. Then, after determining the target processor, the target node can extract the binding relationship between the target processor and each network card in the resource pool, as well as the communication link information between the target processor and each network card from the configuration information.

[0165] S602 , when scheduling the target processor, determining the scheduling requirements of the target processor according to the binding relationship and the communication link information.

[0166] Among them, scheduling requirements include receiving function requirements and sending function requirements.

[0167] In an embodiment of the present application, during the process of the target processor scheduling the network card, the target node can determine the network card function type required by the target processor based on the binding relationship between the target processor and each network card in the resource pool obtained in the above steps, as well as the communication link information between the target processor and each network card, to determine whether the scheduling requirement of the target processor is a receiving function requirement or a sending function requirement.

[0168] The method described in the embodiments of this application can clearly understand the usage status and connection status of each processor and network card by obtaining the binding relationship and communication link information between the target processor and each network card. This allows tasks to be accurately assigned to the network card with the highest matching degree when scheduling the target processor, avoiding resource mismatch and waste, and thus improving the utilization of network card resources.

[0169] In some embodiments, a specific implementation method for scheduling the target network card to the target processor for binding and running is further provided. The "scheduling the target network card to the target processor for binding and running" in the above S203 includes:

[0170] The target sending network card is dispatched to the sending port of the target processor for binding and running, and / or the target receiving network card is dispatched to the receiving port of the target processor for binding and running.

[0171] The target network card includes a target sending network card and / or a target receiving network card.

[0172] In an embodiment of the present application, the target node can determine the target sending network card for the input processor and the target receiving network card for the input processor based on the steps described in the above embodiment, and then schedule the target sending network card to the sending port of the target processor for binding operation, and / or schedule the target receiving network card to the receiving port of the target processor for binding operation.

[0173] The method described in the embodiments of this application achieves the goal of binding the sending port and receiving port of each computing node's communication link to different network interface card devices, thereby establishing multiple efficient ring communication paths between the computing nodes and utilizing multiple network interface cards. Furthermore, the input and output network interface cards in each communication link are relatively independent, making data exchange between nodes more efficient, increasing the parallelism of data transmission between nodes and bandwidth utilization, and thus improving overall communication efficiency.

[0174] Based on all the above embodiments, a network card scheduling method is also provided, such as Figure 12 As shown, the method includes:

[0175] S701 , according to the configuration information of the computing cluster, obtain the binding relationship between the target processor and each network card in the resource pool, and the communication link information between the target processor and each network card.

[0176] S702 , when scheduling the target processor, determining the scheduling requirements of the target processor according to the binding relationship and the communication link information.

[0177] S703: Determine the required network card type based on the scheduling requirements of the target processor in the target node. The network card type includes a sending type and a receiving type.

[0178] S704: Filter out network cards of a type corresponding to the network card type from a resource pool of the computing cluster to obtain multiple candidate network cards.

[0179] S705 , evaluating the transmission performance of the transmission link in each dimension, and obtaining a score value of each candidate network card in each dimension.

[0180] S706 , obtaining the maximum rate corresponding to the network card with the maximum rate capability among all the candidate network cards according to the rate information of each candidate network card.

[0181] S707 : Evaluate the transmission performance of the transmission link based on the rate information and maximum rate of each candidate network card to obtain a rate dimension score item.

[0182] S708 , obtaining the shortest distance corresponding to the network card with the shortest distance to the target processor among all the candidate network cards according to the distance information between each candidate network card and the target processor.

[0183] S709 : Evaluate the transmission performance of the transmission link based on the rate information and the shortest distance of each candidate network card to obtain a distance dimension score item.

[0184] S710: If the channel information indicates that a target channel exists between the corresponding candidate network card and the target processor, determine that the channel dimension score item is a first value.

[0185] S711: If the channel information indicates that there is no target channel between the corresponding candidate network card and the target processor, determine that the channel dimension score item is a second value.

[0186] S712: Assign corresponding weights to the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item according to the scheduling requirements.

[0187] S713 , determining the score value of each candidate network interface card in each dimension according to the weighted rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item.

[0188] S714: The candidate network card corresponding to the largest score value is determined as the target network card. The attribute information is used to evaluate the transmission performance of the transmission link from multiple dimensions.

[0189] S715 , dispatching the target sending network card to the sending port of the target processor for binding and running, and / or dispatching the target receiving network card to the receiving port of the target processor for binding and running.

[0190] In the embodiment of the present application, the method is applied to the ROCm Collective Communication Library (RCCL) as an example for explanation. RCCL is a communication library for GPU (Graphic Processing Unit) communication, which can realize data exchange between processes. The specific process is as follows:

[0191] First, for each available network card found, we need to add the data structure attributes use_send and use_recv to indicate whether the card has been assigned for sending and receiving. By reading the system configuration file, we can determine whether the network card supports full-duplex operation. If the network card supports full-duplex, it can be used for simultaneous sending and receiving; otherwise, the network card can only be used for one-way sending / receiving. When a network card is determined to be used for sending / receiving, use_send and use_recv need to set the determined operation to 1, and the other to -1, indicating that the network card is only used for one-way sending / receiving.

[0192] Then, you need to add a network card search algorithm to try to call more network card devices when the number of communication channels is greater than 1. The algorithm logic is as follows: (1) Determine whether the network card searched in this round is used for sending or receiving, obtain the GPU index used for sending / receiving in this node, and set the index number of the first searched network card to 0; (2) According to the index number of the currently searched network card, first determine whether the network card has been used for sending / receiving. If it has been used for sending / receiving, search for the next network card; (3) Check the network card speed and the distance between the network card and the current GPU through the system file; (4) Check whether the GDR operation is supported between the network card and the current GPU, and record it with useGDR, 1 for support and 0 for non-support; (5) Save the search information and determine whether the network card is the last network card. If not, add 1 to the network card index number, return to step (2), and continue searching for the next network card; (6) After all network cards have been searched, use the network card scoring mechanism to select the network card used for sending / receiving in this round of search based on the search information, and set the corresponding use_send and use_recv to the corresponding values.

[0193] After finding the communication paths between the GPU and each network card based on the above search algorithm, we need to design a network card scoring mechanism. The rating score is Indicates that factors involved in the scoring include network card speed , the distance between the network card and the current GPU , network card and current GPU support GDR operation. The maximum network card rate found in this round is ,use Divide by , get the score of the network card speed; the shortest distance found in this round of search is called ,use Divide by , get the distance score; whether the network card and the current GPU support GDR operation, if they support GDR, then The score is 1, otherwise it is 0. Therefore, the scoring formula for the network card is as follows. Based on this scoring mechanism, the algorithm can select the network card with the highest possible speed, the closest possible distance to the current GPU, and the greatest support for GDR operations.

[0194]

[0195] After implementing the above optimization strategy, you can specify four communication channels when calling the RCCL communication library. The added network card search algorithm will then perform eight rounds of searches, selecting the optimal network card in each round for sending and receiving data with the GPU. This optimized topology improves the system's connectivity structure, fully utilizing all network card resources in a multi-NIC system, significantly improving the parallelism and efficiency of data communication.

[0196] Through the above optimization strategy, the sending port and receiving port of each computing node's communication channel are bound to different network card devices, and multiple efficient ring communication paths are constructed to achieve the purpose of utilizing multiple network cards. At this time, the input network card and output network card in each communication channel are relatively independent, and the data exchange between nodes becomes more efficient, which improves the parallelism of data transmission between nodes and the utilization of bandwidth, thereby improving the overall communication efficiency. The following Table 1 and Figure 13 The results of using RCCL before and after optimization to communicate with other collective communication libraries (AllReduce, AllGather, Broadcast, Reduce, and ReduceScatter) are shown. The optimization effect is about 30%.

[0197]

[0198] The methods described in the above steps are all described in the above embodiments. Please refer to the above description for details and will not be repeated here.

[0199] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0200] Based on the same inventive concept, embodiments of the present application also provide a network card scheduling device for implementing the aforementioned network card scheduling method. The implementation solution provided by this device is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations of one or more network card scheduling device embodiments provided below can be found in the above-mentioned limitations of the network card scheduling method and will not be repeated here.

[0201] In some embodiments, as Figure 14 As shown, a network card scheduling device is provided, comprising:

[0202] The first screening module 11 is configured to screen out required network cards from a resource pool of a computing cluster according to a scheduling requirement of a target processor in a target node, thereby obtaining a plurality of candidate network cards.

[0203] The second screening module 12 is used to screen out a desired network card from multiple candidate network cards based on attribute information of the transmission link between the target processor and the multiple candidate network cards to obtain a target network card; the attribute information is used to evaluate the transmission performance of the transmission link from multiple dimensions.

[0204] The scheduling module 13 is used to schedule the target network card to the target processor for binding and running.

[0205] In some embodiments, the second screening module includes:

[0206] The scoring unit is used to score each candidate network card according to the attribute information of the transmission link between the target processor and the multiple candidate network cards, and obtain a score value of each candidate network card.

[0207] The first determining unit is configured to determine the candidate network card corresponding to the largest score value as the target network card.

[0208] In some embodiments, the scoring unit includes:

[0209] The evaluation subunit is used to evaluate the transmission performance of the transmission link in various dimensions and obtain the score value of each candidate network card in each dimension.

[0210] The determination subunit is configured to obtain a score value of each candidate network card according to the score value of each candidate network card in each dimension.

[0211] In some embodiments, the above-mentioned evaluation sub-unit is specifically used to evaluate the transmission performance of the transmission link according to the rate information of each candidate network card, and obtain a rate dimension scoring item; evaluate the transmission performance of the transmission link according to the distance information between each candidate network card and the target processor, and obtain a distance dimension scoring item; evaluate the transmission performance of the transmission link according to the channel information between each candidate network card and the target processor, and obtain a channel dimension scoring item; determine the score value of each candidate network card in each dimension according to the rate dimension scoring item, the distance dimension scoring item and the channel dimension scoring item.

[0212] In some embodiments, the above-mentioned evaluation sub-unit is further specifically used to obtain the maximum rate corresponding to the network card with the maximum rate capability among all candidate network cards based on the rate information of each candidate network card; based on the rate information and maximum rate of each candidate network card, evaluate the transmission performance of the transmission link to obtain a rate dimension score item.

[0213] In some embodiments, the above-mentioned evaluation sub-unit is further specifically used to obtain the shortest distance corresponding to the network card with the shortest distance to the target processor among all candidate network cards based on the distance information between each candidate network card and the target processor; and evaluate the transmission performance of the transmission link based on the rate information and shortest distance of each candidate network card to obtain a distance dimension score item.

[0214] In some embodiments, the evaluation subunit is further configured to determine that the channel dimension score item is a first value if the channel information indicates that a target channel exists between the corresponding candidate network card and the target processor;

[0215] If the channel information indicates that there is no target channel between the corresponding candidate network card and the target processor, the channel dimension score item is determined to be a second value.

[0216] In some embodiments, the above-mentioned determination subunit is specifically used to assign corresponding weights to the rate dimension scoring item, the distance dimension scoring item and the channel dimension scoring item according to the scheduling requirements; and determine the score value of each candidate network card in each dimension based on the rate dimension scoring item, the distance dimension scoring item and the channel dimension scoring item after the weights are assigned.

[0217] In some embodiments, the first screening module includes:

[0218] The second determining unit is configured to determine a required network card type according to a scheduling requirement of a target processor in a target node; the network card type includes a sending type and a receiving type.

[0219] The screening unit is used to screen out network cards of a type corresponding to the network card type from a resource pool of the computing cluster to obtain multiple candidate network cards.

[0220] In some embodiments, the network card scheduling device further includes:

[0221] The acquisition module is used to obtain the binding relationship between the target processor and each network card in the resource pool, as well as the communication link information between the target processor and each network card according to the configuration information of the computing cluster.

[0222] The determination module is used to determine the scheduling requirements of the target processor according to the binding relationship and communication link information when the target processor is scheduled.

[0223] In some embodiments, the scheduling module is specifically used to schedule the target sending network card to the sending port of the target processor for binding operation, and / or schedule the target receiving network card to the receiving port of the target processor for binding operation.

[0224] Each module in the network card scheduling device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0225] In some embodiments, a computer device is provided. The computer device can be a terminal or a server. The internal structure diagram thereof can be as follows: Figure 15As shown, the computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, while the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operating system and computer programs stored in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means may be implemented via Wi-Fi, a mobile cellular network, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a network card scheduling method. The display unit of the computer device is used to produce a visually visible image and may be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.

[0226] Those skilled in the art will understand that Figure 15 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0227] In some embodiments, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the network card scheduling method described in any of the above embodiments when executing the computer program.

[0228] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the network card scheduling method described in any of the above embodiments are implemented.

[0229] In some embodiments, a computer program product is provided, including a computer program, which, when executed by a processor, implements the steps of the network card scheduling method described in any of the above embodiments.

[0230] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processors (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0231] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A network card scheduling method, characterized in that: Applied to a target node in a computing cluster, the method includes: According to the scheduling requirements of the target processor in the target node, the required network cards are screened from the resource pool of the computing cluster to obtain multiple candidate network cards; According to attribute information of the transmission link between the target processor and the plurality of candidate network cards, a desired network card is screened out from the plurality of candidate network cards to obtain a target network card; the attribute information is used to evaluate the transmission performance of the transmission link from multiple dimensions; The target network card is dispatched to the target processor for binding and running.

2. The method according to claim 1, characterized in that The step of filtering out a desired network card from the plurality of candidate network cards based on the attribute information of the transmission link between the target processor and the plurality of candidate network cards to obtain the target network card includes: Scoring each of the candidate network cards according to attribute information of the transmission link between the target processor and the plurality of candidate network cards to obtain a score value for each of the candidate network cards; The candidate network card corresponding to the largest score value is determined as the target network card.

3. The method according to claim 2, characterized in that Scoring each candidate network card according to the attribute information of the transmission link between the target processor and the plurality of candidate network cards to obtain a score value for each candidate network card includes: Evaluate the transmission performance of the transmission link in each dimension to obtain a score value of each candidate network interface card in each dimension; The score value of each candidate network card is obtained according to the score value of each candidate network card in each dimension.

4. The method according to claim 3, characterized in that The attribute information includes rate information of the candidate network card, distance information between the candidate network card and the target processor, and channel information between the candidate network card and the target processor. The evaluating the transmission performance of the transmission link in each dimension to obtain a score value of each candidate network card in each dimension includes: Evaluate the transmission performance of the transmission link according to the rate information of each candidate network card to obtain a rate dimension score item; Evaluate the transmission performance of the transmission link according to the distance information between each candidate network card and the target processor to obtain a distance dimension score item; Evaluate the transmission performance of the transmission link according to the channel information between each candidate network card and the target processor to obtain a channel dimension score item; According to the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item, a score value of each candidate network interface card in each dimension is determined.

5. The method according to claim 4, characterized in that The transmission performance of the transmission link is evaluated based on the rate information of each candidate network card to obtain a rate dimension score item, including: According to the rate information of each candidate network card, obtaining the maximum rate corresponding to the network card with the maximum rate capability among all the candidate network cards; The transmission performance of the transmission link is evaluated based on the rate information of each candidate network card and the maximum rate to obtain a rate dimension score item.

6. The method according to claim 4, characterized in that The step of evaluating the transmission performance of the transmission link based on the distance information between each candidate network card and the target processor to obtain a distance dimension score item includes: According to the distance information between each candidate network card and the target processor, the shortest distance corresponding to the network card with the shortest distance to the target processor among all the candidate network cards is obtained; The transmission performance of the transmission link is evaluated according to the rate information of each candidate network card and the shortest distance to obtain a distance dimension score item.

7. The method according to claim 4, characterized in that The step of evaluating the transmission performance of the transmission link based on the channel information between each candidate network card and the target processor to obtain a channel dimension score item includes: If the channel information indicates that a target channel exists between the corresponding candidate network card and the target processor, determining that the channel dimension score item is a first value; If the channel information indicates that there is no target channel between the corresponding candidate network card and the target processor, the channel dimension score item is determined to be a second value.

8. The method according to claim 4, characterized in that The determining, based on the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item, of a score value of each candidate network interface card in each dimension includes: Assign corresponding weights to the rate dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item according to the scheduling requirements; According to the weighted speed dimension scoring item, the distance dimension scoring item, and the channel dimension scoring item, the score value of each candidate network interface card in each dimension is determined.

9. The method according to any one of claims 1 to 8, characterized in that The method further comprises screening the required network cards from the resource pool of the computing cluster according to the scheduling requirements of the target processor in the target node to obtain multiple candidate network cards, including: Determining the required network card type according to the scheduling requirements of the target processor in the target node; the network card type includes a sending type and a receiving type; Network cards of a type corresponding to the network card type are screened out from a resource pool of the computing cluster to obtain the multiple candidate network cards.

10. The method according to any one of claims 1 to 8, characterized in that The method further comprises: According to the configuration information of the computing cluster, obtaining the binding relationship between the target processor and each network card in the resource pool, and the communication link information between the target processor and each network card; When the target processor is scheduled, a scheduling requirement of the target processor is determined according to the binding relationship and the communication link information.

11. The method according to any one of claims 1 to 8, characterized in that The target network card includes a target sending network card and / or a target receiving network card, and scheduling the target network card to the target processor for binding operation includes: Dispatching the target sending network card to the sending port of the target processor for binding operation; And / or, the target receiving network card is dispatched to the receiving port of the target processor for binding operation.

12. A network card scheduling device, characterized in that: The device comprises: The first screening module is used to screen out the required network cards from the resource pool of the computing cluster according to the scheduling requirements of the target processor in the target node, and obtain multiple candidate network cards; a second screening module, configured to screen a desired network card from the plurality of candidate network cards based on attribute information of a transmission link between the target processor and the plurality of candidate network cards, thereby obtaining a target network card; wherein the attribute information is used to evaluate the transmission performance of the transmission link from multiple dimensions; The scheduling module is used to schedule the target network card to the target processor for binding and running.

13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.