Server scheduling method based on Fat-Tree topology in large model, equipment medium and product
By obtaining server deployment schemes and constraints in the Fat-Tree topology in the large model and determining the scheduling scheme, the problem that the Fat-Tree topology in the existing technology is not effectively utilized, and the efficiency of server scheduling and system reliability are improved.
Patent Information
- Application Number
- CN202510126867.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, when performing server scheduling, no targeted scheduling scheme is proposed for the Fat-Tree topology in the data center network (DCN), resulting in low scheduling efficiency of the server.
A server scheduling method based on the Fat-Tree topology in a large model is provided. By obtaining the deployment scheme of the server in the topology of high bandwidth domain (HBD), combining the first and second constraints of the Fat-Tree topology, the scheduling scheme is determined to optimize the scheduling of the server.
Through the scheduling scheme for Fat-Tree topology, the efficiency of server scheduling is improved, communication delay is reduced, resource utilization is optimized, and system reliability and fault tolerance are improved.
Smart Images

Figure CN119966828A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a server scheduling method, device medium and product based on Fat-Tree topology in a large model. Background Art
[0002] The artificial intelligence data center (AIDC) is a high-performance computing and data processing center designed for artificial intelligence applications. Its main equipment includes intelligent computing chips. It contains traditional data center networks (DCNs) and high-bandwidth domains (HBDs) to meet different communication needs. Optimizing communication latency is crucial to improving resource utilization, and server scheduling based on the characteristics of large language model (LLM) training tasks is an effective way to reduce latency.
[0003] In related technologies, some manufacturers have developed server scheduling technology based on DCN topology.
[0004] However, the inventors at least found that: in the related art, when executing server scheduling, no targeted scheduling solution is proposed for the characteristics of the Fat-Tree topology structure of the DCN, which results in low scheduling efficiency of the server. Summary of the invention
[0005] One purpose of the present application is to provide a server scheduling method, device medium and product based on the Fat-Tree topology in a large model, at least to solve the technical problem in the related art that no targeted scheduling solution is proposed for the characteristics of the Fat-Tree topology structure of the DCN, resulting in low scheduling efficiency of the server.
[0006] To achieve the above objectives, some embodiments of the present application provide the following aspects:
[0007] In a first aspect, some embodiments of the present application provide a server scheduling method based on a Fat-Tree topology in a large model, the method being applied to a Fat-Tree topology structure, the method comprising: obtaining a deployment plan of a server in a topology structure of an HBD; in the deployment plan, the servers are arranged in a number of sub-lines, each sub-line corresponding to a group of servers; the last server in each sub-line is connected to the last server in the next sub-line, the first server in the next sub-line is connected to the first server in the next sub-line, and so on, forming an interlaced "zigzag" connection pattern; obtaining a first constraint and a second constraint of the Fat-Tree topology structure; the first constraint is used to ensure that a TP group does not span multiple backbone switches, and the second constraint is used to ensure that the TP group is aligned; determining a scheduling plan based on the deployment plan, the first constraint and the second constraint, so as to schedule the server according to the scheduling plan.
[0008] In a second aspect, some embodiments of the present application further provide an electronic device, comprising: one or more processors; and a memory storing computer program instructions, wherein the computer program instructions, when executed, cause the processor to perform the steps of the method described above.
[0009] In a third aspect, some embodiments of the present application further provide a computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method as described above.
[0010] In a fourth aspect, some embodiments of the present application further provide a computer program product, comprising a computer program / instruction, which implements the steps of the method described above when executed by a processor.
[0011] Compared with the related art, the solution provided in the embodiment of the present application provides a server scheduling method based on the Fat-Tree topology in a large model, and the method is applied to the Fat-Tree topology structure, by obtaining the deployment plan of the server in the topology structure of the HBD; in the deployment plan, the servers are arranged in several sub-lines, and each sub-line corresponds to a group of servers; the last server in each sub-line is connected to the last server in the next sub-line, and the first server in the next sub-line is connected to the first server in the next sub-line, and so on, forming an interlaced "zigzag" connection pattern; and obtaining the first constraint condition and the second constraint condition of the Fat-Tree topology structure; the first constraint condition is used to ensure that the TP group does not span multiple backbone switches, and the second constraint condition is used to ensure that the TP group is aligned, so as to determine the scheduling plan according to the deployment plan, the first constraint condition and the second constraint condition, so as to schedule the server according to the scheduling plan. In this application, the deployment scheme based on the special "zigzag" connection mode, combined with the characteristics of the Fat-Tree topology structure, provides two constraints. The first constraint and the second constraint can essentially divide the large-scale HBD-based topology structure into several sub-HBDs, and then determine the scheduling scheme based on the sub-HBDs, and its time complexity is O(nlog n). BRIEF DESCRIPTION OF THE DRAWINGS
[0012] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0013] Figure 1An exemplary flow chart of a server scheduling method based on a Fat-Tree topology in a large model provided according to some embodiments of the present application;
[0014] Figure 2 An exemplary schematic diagram of the deployment scheme in a server scheduling method based on a Fat-Tree topology in a large model provided according to some embodiments of the present application;
[0015] Figure 3 An exemplary schematic diagram of TP group alignment in a server scheduling method based on a Fat-Tree topology in a large model provided according to some embodiments of the present application;
[0016] Figure 4 An exemplary flow chart of step S101 in a server scheduling method based on a Fat-Tree topology in a large model provided according to some embodiments of the present application;
[0017] Figure 5 An exemplary flow chart of step S103 in a server scheduling method based on a Fat-Tree topology in a large model provided according to some embodiments of the present application;
[0018] Figure 6 An exemplary schematic diagram of the length of the sub-line in a server scheduling method based on a Fat-Tree topology in a large model provided according to some embodiments of the present application;
[0019] Figure 7 This is an exemplary structural diagram of an electronic device provided according to some embodiments of the present application. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0021] The following terms are used in this article.
[0022] Intelligent computing chips are chips used for artificial intelligence computing tasks, including intelligent computing chips, NPU, TPU, FPGA and other chips.
[0023] Artificial Intelligence Data Center, AIDataCenter, referred to as AIDC, refers to a data center designed to provide high-performance computing and large-scale data processing for artificial intelligence applications. It is used for large model training or reasoning tasks. The main computing equipment includes intelligent computing chips.
[0024] High-bandwidth domain, or HBD for short, is a network architecture used to meet the high bandwidth requirements of large model training. It is mainly used to support parallel dimensions with large communication volumes, such as tensor parallelism.
[0025] Data Center Network, or DCN for short, is a network architecture used to meet general communication needs. It mainly supports parallel dimensions with relatively small communication volumes, such as data parallelism and pipeline parallelism.
[0026] Large Language Model, the full English name is Large Language Model, abbreviated as LLM.
[0027] Tensor parallelism, also known as TP, is a model parallel technology.
[0028] First embodiment
[0029] The first embodiment of the present application relates to a server scheduling method based on a Fat-Tree topology in a large model. The method is applied to a Fat-Tree topology structure, such as Figure 1 As shown, the method may include the following steps:
[0030] Step S101, obtaining a deployment scheme of servers in a topological structure of an HBD; in the deployment scheme, the servers are arranged in a plurality of sub-lines, each sub-line corresponding to a group of servers; the last server in each sub-line is connected to the last server in the next sub-line, the first server in the next sub-line is connected to the first server in the next sub-line, and so on, forming an interlaced "zigzag" connection mode;
[0031] Step S102, obtaining a first constraint and a second constraint of the Fat-Tree topology structure; the first constraint is used to ensure that the TP group does not span multiple backbone switches, and the second constraint is used to ensure that the TP group is aligned;
[0032] Step S103: Determine a scheduling plan according to the deployment plan, the first constraint condition and the second constraint condition, so as to schedule the server according to the scheduling plan.
[0033] The following is a detailed description of each of the above steps.
[0034] For step S101, specifically, in some examples, see Figure 2 As shown, a deployment scheme is shown, which is a network topology consisting of 32 servers. The servers are numbered from 1 to 32, and they are connected together in a certain pattern. Specifically, the servers are arranged in rows and columns, with 8 servers in each row and a total of 4 rows. The last server in each row is connected in sequence, that is, each server is connected to the server on its right, and, except for the last row, the last server in each row is also connected to the first server in the next sub-line, forming an interlaced "zigzag" connection pattern. The details are as follows:
[0035] Row 1: 1-5-9-13-17-21-25-29;
[0036] Row 2: 2-6-10-14-18-22-26-30;
[0037] Row 3: 3-7-11-15-19-23-27-31;
[0038] Row 4: 4-8-12-16-20-24-28-32;
[0039] Among them, servers numbered 1, 5, 9, 13, 17, 21, 25, 29 form a sub-line, servers numbered 2, 6, 10, 14, 18, 22, 26, 30 form a sub-line, servers numbered 3, 7, 11, 15, 9, 23, 27, 31 form a sub-line, and servers numbered 4, 8, 12, 16, 20, 24, 28, 32 form a sub-line. Server No. 29 in the first row is connected to Server No. 2 in the second row, Server No. 30 in the second row is connected to Server No. 3 in the third row, and Server No. 31 in the third row is connected to Server No. 4 in the fourth row, forming a closed loop.
[0040] It should be noted that in Figure 2 In the example shown, the sub-lines are specifically parallel to each other. However, in some other examples, the sub-lines may be non-parallel to each other; or the sub-lines themselves may be broken lines, which is not specifically limited in the present embodiment.
[0041] Regarding step S102, specifically, in some examples, a group of backbone switches has limited coverage. Therefore, if a TP group spans multiple backbone switches, cross-track traffic may be generated. The cross-track traffic will cause network congestion and degraded communication performance. In order to avoid this situation, it is necessary to avoid letting the TP group span multiple backbone switches. Furthermore, communication between servers in different states may cause cross-track traffic. However, after the implementation of the deployment scheme, this problem is actually simplified: the TP group should avoid crossing sub-lines. As long as the servers in the TP group are located on the same sub-line and are in the same state, the DP communication between these TP groups will not generate cross-track traffic. Therefore, the first constraint condition is used to ensure that the TP group does not span multiple backbone switches.
[0042] Specifically, in some examples, under the Fat-Tree topology of DCN, a CP group is composed of TP groups between multiple sub-lines. In order to ensure that the traffic of the CP group is completely processed inside the ToR switch, these TP groups need to be aligned, such as Figure 3 As shown. That is to say, if any server under the ToR switch fails, all servers under the ToR switch must be abandoned. This means that a failure on one sub-line will affect all sub-lines through the broadcast mechanism, causing the impact range of the failure to expand by p times. The alignment constraint of the TP group is based on the granularity of the ToR switch. If this constraint is enabled, the failure of any server under the corresponding ToR switch will be considered to have spread to all servers under the ToR switch.
[0043] With respect to step S103, specifically, in some examples, based on the deployment scheme, in combination with the first constraint and the second constraint, a scheduling scheme that satisfies the first constraint and the second constraint can be determined. Since the scheduling scheme is based on the first constraint and the second constraint, it aims to optimize the use of network resources while ensuring the reliability and fault tolerance of the network. By aligning constraints and coverage constraints, the spread of failures of servers in the entire switch domain caused by a single server failure can be reduced, thereby improving the stability and efficiency of the entire network.
[0044] It is not difficult to find that in an embodiment of the present application, a server scheduling method based on the Fat-Tree topology in a large model is provided, and the method is applied to the Fat-Tree topology structure, by obtaining the deployment plan of the server in the topology structure of the HBD; in the deployment plan, the servers are arranged in several sub-lines, and each sub-line corresponds to a group of servers; the last server in each sub-line is connected to the last server in the next sub-line, and the first server in the next sub-line is connected to the first server in the next sub-line, and so on, forming an interlaced "zigzag" connection pattern; and obtaining the first constraint condition and the second constraint condition of the Fat-Tree topology structure; the first constraint condition is used to ensure that the TP group does not span multiple backbone switches, and the second constraint condition is used to ensure that the TP group is aligned, so as to determine the scheduling plan according to the deployment plan, the first constraint condition and the second constraint condition, so as to schedule the server according to the scheduling plan. In this application, the deployment scheme based on the special "zigzag" connection mode, combined with the characteristics of the Fat-Tree topology structure, provides two constraints. The first constraint and the second constraint can essentially divide the large-scale HBD-based topology structure into several sub-HBDs, and then determine the scheduling scheme based on the sub-HBDs, and its time complexity is O(nlog n).
[0045] Second embodiment
[0046] The second embodiment of the present application relates to a server scheduling method based on the Fat-Tree topology in a large model. The second embodiment is an improvement on the first embodiment, and the specific improvement is that: in the first embodiment, the deployment plan can be set manually; while in the second embodiment of the present application, the deployment plan can be automatically generated, providing a specific implementation method for obtaining the deployment plan of the server in the topological structure of the HBD.
[0047] Specifically, in some embodiments, the step of obtaining the deployment scheme of the server in the topology structure of the HBD, that is, step S101, may further include the following steps: Figure 4 As shown:
[0048] Step S1011, determining the first serial number information of the server in the Fat-Tree topology structure;
[0049] Step S1012, determining the number of servers directly connected to each ToR switch according to the Fat-Tree topology structure;
[0050] Step S1013, determining second numbering information according to the first numbering information and the number of servers directly connected to each ToR switch; the second numbering information is used to indicate the serial number of the server in the topology structure of the HBD; it can be seen that there is a mapping relationship between the first numbering information and the second numbering information;
[0051] Step S1014: Obtain a deployment plan of the servers in the topology of the HBD according to the second numbering information. In this way, the servers are arranged in each sub-line according to the second numbering information.
[0052] Specifically, in some examples, the type of the DCN topology structure may be a Fat-Tree topology structure. The first numbering information is used to represent an ordered set of numbers corresponding to each server. Different servers in the Fat-Tree topology structure have their own numbers, which may be represented by letters such as A, B, and C, or by numbers, or a combination of numbers and letters, etc. This embodiment does not limit this. For example: server 1, server 2, ..., server n, where n is a positive integer representing the total number of servers in the Fat-Tree topology structure.
[0053] Specifically, in some examples, the number of servers directly connected to each ToR switch is determined based on the number of servers under each layer of switches. This can be achieved by the following formula:
[0054] p=a
[0055] in, a Indicates the number of servers directly connected to each top-of-rack switch, p Indicates the number of servers directly connected to each ToR switch.
[0056] Specifically, in some examples, the second numbering information in the topology structure of the server HBD is determined based on the first numbering information and the number of servers directly connected to each ToR switch. Here, the type of the topology structure of the HBD can be any ring topology structure or a linear topology structure, and this embodiment does not specifically limit this. It can be understood that in a ring topology structure, the server connections form a closed loop, and each server is directly connected to two other servers to form a continuous path. In a linear topology structure, servers or nodes are arranged linearly, and each node is usually only connected to adjacent nodes. It can be seen that through this step, the mapping relationship between the server numbering in the Fat-Tree topology structure and the HBD topology structure can be determined. In this way, the position and connection relationship of the server in the topology structure of the HBD can be determined based on the server numbering in the Fat-Tree topology structure.
[0057] Specifically, in some examples, the second numbering information represents a deployment plan of the server in the HBD topology structure, and relevant personnel can deploy the corresponding server in the HBD topology structure according to the second numbering information. The deployment plan includes the location relationship and connection relationship of the server carried in the second numbering information.
[0058] Optionally, in some embodiments, determining the deployment scheme of servers in the topological structure of the HBD based on the second numbering information includes: determining the traffic pattern of the HBD topological structure; the traffic pattern is used to indicate that between the sub-lines, the last server in each sub-line is connected to the last server in the next sub-line, and the first server in the next sub-line is connected to the first server in the next-next sub-line, and so on, forming an interlaced "zigzag" connection pattern; determining the deployment scheme of servers in the topological structure of the HBD based on the "zigzag" connection pattern and the second numbering information.
[0059] Optionally, in some embodiments, the determining according to the first numbering information and the number of servers directly connected to each ToR switch may further include the following steps:
[0060] Determine a first ordered set of servers deployed in the Fat-Tree topology structure according to the total number of servers in the Fat-Tree topology structure and the number of servers directly connected to each ToR switch;
[0061] Determine the length of each sub-line according to the first ordered set and the number of servers directly connected to each ToR switch; wherein the length of the sub-line is used to characterize the number of servers that each sub-line should include;
[0062] The second numbering information is determined according to the length of the sub-line.
[0063] Specifically, in some examples, the first ordered set can be used S Indicates that if the total number of servers in the Fat-Tree topology is n, then the first ordered set S In the example, the server numbers range from 1 to n. The server numbers are determined by the positions of the servers in the Fat-Tree topology.
[0064] Specifically, in some examples, determining the length of each sub-line according to the first ordered set and the number of servers directly connected to each ToR switch can be implemented by the following formula:
[0065]
[0066] Wherein, l represents the length of each sub-line, S represents the first ordered set, the p Indicates the number of servers directly connected to each ToR switch.
[0067] Optionally, in some embodiments, determining the second numbering information according to the length of the sub-line may further include the following steps:
[0068] Determine the data range according to the number of servers directly connected to each ToR switch;
[0069] Determine the index of each of the sub-lines according to the data range;
[0070] Determine the position index within the target sub-line according to the index of the sub-line and the length of the sub-line; the target sub-line is determined according to the index of the sub-line;
[0071] Determine, according to the index of the subline, the location index, and the number of servers directly connected to each ToR switch, a second ordered set for characterizing servers deployed in the topology structure of the HBD;
[0072] The second numbering information is determined according to the second ordered set.
[0073] Specifically, in some examples, if the number of servers directly connected to each ToR switch is p , the data range can be from 0 to p-1 .
[0074] Specifically, in some examples, according to the data range from 0 to p-1 By traversing, the index of each subline can be determined. It can be understood that each subline index i corresponds to a subline. In other words, the subline index i can be used to determine which subline is currently being processed, so that it is convenient to allocate servers to different sublines. For example, p=4 , then the value of the sub-line index i can be 0, 1, 2, 3, and different index values represent different sub-lines.
[0075] Specifically, in some examples, j can be used to represent the position index located in the target sub-line, and the target sub-line is the sub-line corresponding to the index i of the sub-line. j , used to determine the position of the server in the subline corresponding to the index i of the subline. In some examples, the traversal range of the position index j located in the target subline may be different depending on the parity of the index i of the subline.
[0076] Optionally, in some embodiments, determining the position index within the target sub-line based on the index of the sub-line and the length of the sub-line may further include the following steps: determining the parity of the index of the sub-line; and determining the position index within the target sub-line based on the parity and the length of the sub-line.
[0077] Optionally, in some embodiments, determining the position index within the target sub-line based on the parity and the length of the sub-line may further include the following steps: if the index of the sub-line is an even number, the position index increases in sequence based on the length of the sub-line; if the index of the sub-line is an odd number, the position index increases in reverse order based on the length of the sub-line.
[0078] Specifically, if the index i of the sub-line is an even number, the position index j can be increased in sequence based on the length l of the target sub-line, which means that the server numbers are increased in sequence, such as starting from 0; if the index i of the sub-line is an odd number, the position index j can be increased in reverse order based on the length l of the target sub-line, which means that the server numbers are increased in reverse order, such as starting from l-1. For example, if the index i of the sub-line is an even number, the position index j ranges from 0 to l-1; if the index i of the sub-line is an odd number, the position index j ranges from l-1 to 0. It can be seen that the position index j is used to traverse the servers in each sub-line to determine the new position of each server in the deployment scheme of the server in the topology of the HBD, and the new position corresponds to a new serial number.
[0079] Optionally, in some embodiments, determining the second ordered set for characterizing the servers deployed in the topology structure of the HBD according to the index of the subline, the location index, and the number of servers directly connected to each ToR switch can be specifically implemented by the following formula:
[0080] i+j·p
[0081] Among them, the i represents the index of the sub-line, the j represents the position index, and the p represents the number of servers directly connected to each ToR switch.
[0082] Specifically, in some examples, we can first initialize an empty ordered set S deploy= [], the empty ordered set is used to store the second numbering information. In combination with the context, it can be known that the second numbering information is obtained by converting the first numbering information of the server in the topological structure based on the Fat-Tree through a preset rule. Specifically, the first numbering information of the server in the topological structure based on the DCN can be converted by combining the index i of each sub-line, the position index j and the number of servers directly connected to each ToR switch p to obtain the new serial number of each server deployed in the topological structure of the HBD, that is, the second ordered set S deploy .
[0083] Exemplarily, each i+j·p obtained can be added to the originally empty ordered set initialized above to obtain the second ordered set S deploy The second ordered set S deploy It is an ordered set of servers in the topology of the HBD, which is composed of the serial numbers of the servers in the topology of the HBD. As mentioned above, the index i of the subline is used to determine which subline is currently being processed; the position index j is used to determine the position of the server in the subline corresponding to the index i of the subline, and p represents the number of servers directly connected to each ToR switch.
[0084] like Figure 2 As shown, assuming that the number of servers directly connected to each ToR switch is p, there are p sub-lines L1, L2, ..., L p , then the i-th subline consists of the servers whose serial numbers corresponding to the position index are divided by the number of servers directly connected to each ToR switch p plus i-1, that is:
[0085] Li={j|(j-1)%p==i-1}
[0086] The L i Indicates the sub-line corresponding to the index i of the sub-line.
[0087] Furthermore, a series of the p sub-lines are connected end to end to form a new arrangement method for deploying servers in the topological structure of the HBD. Figure 2 In the second ordered set S deploy The composition is as follows:
[0088] S deploy ={1,5,9,13,17,21,25,29,
[0089] 30,26,22,18,14,10,6,2,
[0090] 3,7,11,15,19,23,27,31,
[0091] 32,28,24,20,16,12,8,4}
[0092] Optionally, in some embodiments, determining the second numbering information based on the second ordered set may further include the following steps: determining an edge set representing the connection relationship between servers deployed in the topological structure of the HBD based on the second ordered set; determining the second numbering information based on the second ordered set and the edge set.
[0093] Specifically, the second number information can be represented by G deploy Indicates that the second ordered set can be represented by S deploy Indicates that the edge set can be represented by E deploy Indicates that the second number information G deploy = deploy ,E depoly >.
[0094] Optionally, in some embodiments, determining the edge set for characterizing the connection relationships of servers deployed in the topological structure of the HBD based on the second ordered set may further include the following steps: obtaining the number of servers directly connected to the server in a single direction in the topological structure of the HBD; determining the target server based on the second ordered set and the number of servers directly connected to the server in a single direction in the topological structure of the HBD; determining the edge set for characterizing the connection relationships of servers deployed in the topological structure of the HBD based on the second ordered set and the target server.
[0095] Specifically, in some examples, the number of servers directly connected to the server in a single direction in the HBD topology can be represented by φ. Exemplarily, the number of servers directly connected to the server in a single direction in the HBD topology is φ. For example, when φ=2, the server S deploy [i] Can be directly connected to S deploy [i-2], S deploy [i-1], S deploy [i+1] and S deploy [i+2]. Combine Figure 2 As shown, server 9 can be directly connected to servers 1, 5, 13, and 17.
[0096] Specifically, in some examples, the second ordered set S can be traversed deploy , according to the second ordered set S deploy Each server in And the number of servers φ directly connected to the server in a single direction in the HBD topology, find all target servers that satisfy the following formula
[0097] ji≤φ
[0098] Specifically, in some examples, the second ordered set and the target server may form For every pair that satisfies the condition Can be found in E deploy Create an edge in the Can connect directly to the server The E deploy Represents the edge set.
[0099] It should be noted that the edge set E deploy The following conditions must also be met:
[0100] 1≤i≤j≤|S deploy |
[0101] By using the above formula, it can be ensured that the solution only considers the case where the sub-line index i is less than or equal to the position index j, so as to avoid repeatedly creating the same connection relationship. It can be understood that if Connect to Then there is no need to make Then connect to
[0102] Optionally, in some embodiments, determining the edge set used to characterize the connection relationship of the servers deployed in the topology structure of the HBD according to the second ordered set and the target server is specifically implemented by the following formula:
[0103]
[0104] Wherein, i represents the index of the sub-line, j represents the position index, S deploy represents the second ordered set, the The second ordered set S deploy The server in represents the target server, φ represents the number of servers directly connected to the server in a single direction in the topology of the HBD, and E deploy Represents the edge set.
[0105] It can be seen that the edge set E deploy Including all satisfying |1≤i≤j≤|S deploy | and ji≤φ Yes, so far, the second ordered set S is completed. deploy Establishment of a connection relationship between the target server and the target server.
[0106] It is not difficult to find that, compared with the related art, in the solution provided by the embodiment of the present application, based on the collaborative design of the HBD topology and the Fat-Tree topology, a server deployment method based on a large model is proposed; by determining the first numbering information of the server in the Fat-Tree topology, and according to the Fat-Tree topology, the number of servers directly connected to each ToR switch is determined; and then according to the first numbering information and the number of servers directly connected to each ToR switch, the second numbering information is determined, so as to determine the deployment scheme of the server in the HBD topology according to the second numbering information. Wherein, the second numbering information is used to indicate the serial number of the server in the HBD topology; in the deployment scheme, the servers are arranged in each sub-line based on the second numbering information, the last server in each sub-line is connected to the last server in the next sub-line, the first server in the next sub-line is connected to the first server in the next sub-line, and so on, forming an interlaced "zigzag" connection mode. It can be seen that the present application can simultaneously perceive the topology of HBD and the Fat-Tree topology; since the number of servers directly connected to each ToR switch is determined according to the Fat-Tree topology, it is beneficial to obtain a deployment scheme for servers in the topology structure adapted to the HBD according to the characteristics of the Fat-Tree topology. Since the deployment scheme has strong adaptability, it is beneficial to optimize the use of network resources, improve the performance of training tasks based on large models, and facilitate subsequent server scheduling.
[0107] Third embodiment
[0108] The third embodiment of the present application relates to a server scheduling method based on a Fat-Tree topology in a large model. The third embodiment is an improvement on the first embodiment, and the specific improvement is: in the third embodiment of the present application, a specific implementation method for obtaining the first constraint condition and the second constraint condition of the Fat-Tree topology structure is provided.
[0109] Specifically, in some embodiments, the method for determining the first constraint condition may include:
[0110] Step S201, determining the number of backbone switch coverage constraints; the number of backbone switch coverage constraints is used to indicate that each sub-line needs to be covered by at least one backbone switch;
[0111] Step S202: Obtain the first constraint condition according to the number of coverage constraints of the backbone switch.
[0112] Specifically, in some examples, the number of backbone switch coverage constraints can be represented by n_subline; the first constraint condition can specifically be such that the number of intelligent computing chips in the scheduling scheme is greater than or equal to the number of intelligent computing chips required for the task scale of the large model.
[0113] Optionally, in some embodiments, determining the number of backbone switch coverage constraints may include: obtaining a total constraint number and a maximum number of sub-lines; determining a relatively smaller value between the total constraint number and the maximum number of sub-lines; and determining the number of backbone switch coverage constraints based on the relatively smaller value between the total constraint number and the maximum number of sub-lines.
[0114] Specifically, in some examples, the following formula can be seen:
[0115] n_subline=min(n_maxsubline,n_constraints)
[0116] Among them, the n_subline represents the number of coverage constraints of the backbone switch, which can also be called the number of sub-lines; the n_constraints represents the total number of constraints applied, and the total number of constraints may include alignment constraints, backbone switch coverage constraints, etc.; the n_maxsubline represents the maximum number of sub-lines allowed in the Fat-Tree topology structure.
[0117] According to the above formula, the values of the maximum number of sublines n_maxsubline and the total number of constraints n_constraints can be compared, and the smaller value between the two can be selected as the value of the number of backbone switch coverage constraints n_subline.
[0118] By determining the number of backbone switch coverage constraints by the above method, it can be ensured that the number of backbone switch coverage constraints will not exceed the maximum number of sub-lines allowed in the Fat-Tree topology, and will not exceed the number limited by the constraints. The value of the number n_subline of the backbone switch coverage constraints determines how many sub-lines will be considered in the algorithm for server scheduling. For example, if the maximum number of sub-lines allowed in the Fat-Tree topology is 10 sub-lines, but the total number of constraints is only 5, then the number n_subline of the backbone switch coverage constraints will be set to 5, because this is a smaller value limited by the number of constraints. This ensures that the scheduling method will not exceed the number of sub-lines allowed by the constraints when scheduling servers.
[0119] Specifically, in some embodiments, the method for determining the second constraint condition may include:
[0120] Step S301, determining the number of alignment constraints; the number of alignment constraints is used to characterize the number of backbone switch domains that need to be aligned;
[0121] Step S302: Obtain the second constraint condition according to the number of the alignment constraints.
[0122] Specifically, in some examples, the number of alignment constraints can be represented by n_align, which refers to the number of backbone switch domains that need to be aligned. It can be understood that within these switch domains, if a server fails, then all servers on the sub-lines within the switch domain will be considered affected. For example, if server 5 fails, then servers 6, 7, and 8 in the same switch domain as server 5 will also be marked as failed, because servers 6, 7, and 8 are in the same switch domain as server 5 and therefore need to remain aligned.
[0123] Optionally, in some embodiments, determining the number of alignment constraints may include: acquiring a total number of constraints and a maximum number of sub-lines; and determining the number of alignment constraints according to a difference between the total number of constraints and the maximum number of sub-lines.
[0124] Specifically, in some examples, the following formula can be seen:
[0125] n_align=max(0,n_constraints-n_maxsubline)
[0126] Among them, the n_align represents the number of the alignment constraints; the n_constraints represents the total number of constraints applied, and the total number of constraints may include alignment constraints, backbone switch coverage constraints, etc.; the n_maxsubline represents the maximum number of sub-lines allowed in the Fat-Tree topology structure.
[0127] According to the above formula, subtracting the maximum number of sublines n_maxsubline from the total number of constraints n_constraints will give a difference. The max function is used to ensure that the result is not less than 0. This means that if the total number of constraints n_constraints is less than or equal to the maximum number of sublines n_maxsubline, the number of alignment constraints n_align will be set to 0, because there can be no alignment constraints exceeding the maximum number of sublines n_maxsubline.
[0128] By determining the number of alignment constraints by the above method, it can be ensured that the number of alignment constraints obtained is reasonable and does not exceed the actual constraint conditions. In this embodiment, the alignment constraints are used to ensure that the status of all servers in a specific backbone switch domain remains consistent to reduce the impact range of the failure.
[0129] Exemplarily, in the first constraint and the second constraint, the first n_maxsubline constraints are backbone switch coverage constraints, which means that each subline needs to be covered by at least one backbone switch to ensure the connectivity of the network. These constraints ensure that each subline has at least one backbone switch as its uplink. The last n_domain constraints can be alignment constraints, which ensure that the TP groups (Top-of-Rack groups) of server scheduling are aligned at the ToR switch level. The purpose of the alignment constraint is to reduce the spread of faults of all servers under the entire ToR switch caused by server failures, thereby improving the reliability and efficiency of the network. Simply put, the backbone switch coverage constraint ensures that each subline is covered by at least one backbone switch, while the alignment constraint ensures that when scheduling servers, either all servers under the same ToR switch are selected or not selected at all, so as to reduce the scope of the impact of the failure. These two constraints work together to optimize the server scheduling strategy of the network.
[0130] It should be noted that this embodiment may also be an improvement based on the second embodiment.
[0131] It is not difficult to find that compared with the related technology, this embodiment provides a specific implementation method for determining the scheduling plan based on the deployment plan, the first constraint and the second constraint, which can find the scheduling plan that minimizes the cross-track flow and satisfies the second constraint while meeting the requirements of the task scale represented by the first constraint, and is more efficient.
[0132] Fourth embodiment
[0133] The fourth embodiment of the present application relates to a server scheduling method based on a Fat-Tree topology in a large model. The fourth embodiment is an improvement on the first embodiment, and the specific improvement is: in the fourth embodiment of the present application, a specific implementation method for determining a scheduling scheme based on a deployment scheme, the first constraint condition and the second constraint condition is provided.
[0134] Specifically, in some embodiments, the scheduling scheme is determined according to the deployment scheme, the first constraint condition and the second constraint condition, that is, step S103 may further include the following steps: Figure 5 As shown:
[0135] Step S1031, determining a subline ID according to the number of alignment constraints;
[0136] Step S1032: if the subline ID belongs to the faulty server set, the faulty server set is updated, and servers on all sublines related to the subline ID are marked as faulty, to obtain an updated target faulty server set;
[0137] Step S1033, according to the number of coverage constraints of the backbone switch, perform the following operations on each sub-line: according to the length of the sub-line and the deployment plan, obtain the server set, edge set and target faulty server set corresponding to each sub-line; according to the server set, edge set and target faulty server set corresponding to each sub-line, obtain the scheduling plan of each sub-line; each sub-line scheduling plan is a scheduling plan for each sub-line;
[0138] Step S1034: determine a scheduling plan according to the scheduling plans of each sub-line.
[0139] Optionally, in some embodiments, the determining of the subline IDsid according to the number of the alignment constraints, that is, step S1031, may further include the following steps:
[0140] Step S10311, determining a first traversal range according to the number of the alignment constraints;
[0141] Step S10312, traversing in a second traversal range according to each first index in the first traversal range; wherein the second traversal range is determined according to the number of servers under a group of backbone switches;
[0142] Step S10313: determine a sub-line ID according to the first index, the second index in the second traversal range, and the number of servers under the group of backbone switches.
[0143] Specifically, in some examples, an empty placement scheme set placement_scheme={} may be initialized. Then, the number of alignment constraints n_align and the number of backbone switch coverage constraints n_subline may be calculated.
[0144] Further, a first traversal range may be determined according to the number n_align of the alignment constraints, and an outer loop operation may be performed according to the first traversal range. The first traversal range may be from 0 to (n_align-1). The outer loop operation is used to ensure that each alignment constraint is traversed.
[0145] Further, for each first index i in the first traversal range, traverse in the second traversal range, and perform an inner loop operation according to the second traversal range. If d represents the number of servers under the group of backbone switches, the second traversal range can be from 1 to d. That is, for each first index i, traverse j from 1 to d. The j represents the second index. The inner loop operation is used to generate all possible sub-line IDs.
[0146] Optionally, in some embodiments, the determining of the subline ID according to the first index, the second index in the second traversal range, and the number of servers under the group of backbone switches, that is, step S10313 may further include the following steps:
[0147] Step S103131, determining the product of the first index and the number of servers under the group of backbone switches;
[0148] Step S103132, determining a sub-line ID according to the multiplication result and the sum of the second index.
[0149] Specifically, in some examples, determining the current subline ID according to the first index, the second index in the second traversal range, and the number of servers under the group of backbone switches can be implemented by the following formula:
[0150] sid=i×d+j
[0151] Among them, the sid represents the current sub-line ID, the i represents the first index, the j represents the second index, and the d represents the number of servers under the group of backbone switches.
[0152] It can be seen from the above formula that the sub-line ID is achieved by multiplying the first index i of the outer loop by the number d of servers under the group of backbone switches and adding the second index j of the inner loop.
[0153] Further, in some examples, if the calculated subline ID belongs to the faulty server set F, the following operations may be performed: updating the faulty server set F, marking servers on all sublines associated with the subline ID as faulty, and obtaining an updated target faulty server set.
[0154] Optionally, in some embodiments, the updating of the faulty server set, marking the servers on all sub-lines associated with the sub-line ID as faulty, and obtaining an updated target faulty server set can be specifically implemented by the following formula:
[0155]
[0156] Wherein, F represents the known set of faulty servers, sid represents the sub-line ID currently considered, and p represents the number of servers directly connected to each ToR switch.
[0157] Specifically, the quotient obtained by calculating sid-1 divided by p indicates how many times p is the number of the first server in the group where the subline ID (sid) is located. Rounding down ensures that the maximum integer not greater than the quotient can be obtained. For example, if the subline ID (sid) is 5, and the number of servers directly connected to each ToR switch is 3, then (sid-1) / p is (5-1) / 3=4 / 3. After rounding down, the result is 1, which means that the subline ID (sid) is in the second group (because the group index starts from 0, so 1 is added after rounding down).
[0158] Furthermore, the above rounded-down result can be multiplied by p to calculate the starting server number of the group, that is, the starting index of the first sub-line of the parallel group where the sub-line ID (sid) is located, and then generate a continuous sub-line ID set starting from the starting index of the first sub-line of the parallel group where sid is located plus 1 and ending at the last sub-line of the group. The above-generated sub-line ID set is combined with the original faulty server set F, that is, all servers on these sub-lines are added to the faulty server set to obtain an updated target faulty server set. In this way, if a server on a sub-line sid fails, all servers under the same ToR switch are marked as faulty and added to the faulty server set F. The purpose of this is to meet the alignment constraints and ensure that these servers will not be selected during scheduling, thereby avoiding the spread of faults.
[0159] Further, in some examples, for each constraint i in the number n_subline of the backbone switch coverage constraints, from 1 to n_subline, the following operations are performed: according to the deployment scheme G deploy = deploy ,E depoly > and the length of the sub-line, from the second ordered set S deploy A server set S_subline corresponding to the subline pops up, that is, the server set S_subline corresponding to the subline = S deploy .pop(l'). Wherein, l' represents the length of the sub-line. It should be noted that the sub-line length l' has a different meaning from the length l of each sub-line mentioned above. The length l of each sub-line will be divided into sub-lines l' by the backbone switch boundary. Figure 6 As shown, {1,5} is a sub-line.
[0160] Exemplarily, the length of the sub-line can be obtained by the following formula:
[0161]
[0162] Wherein, l' represents the length of the sub-line, d represents the number of servers under the group of backbone switches, and p represents the number of servers directly connected to each ToR switch.
[0163] Further, the edge set E_subline and the target faulty server set F_subline can be determined according to the server set S_subline corresponding to the subline. The faulty server set F_subline is the intersection of the target faulty server set F and the server set S_subline corresponding to the subline. Then, the subline scheduling scheme can be obtained according to the server set S_subline corresponding to the subline, the edge set E_subline and the faulty server set F_subline.
[0164] Further, according to the second ordered set S deploy The remaining sublines are popped out, and the remaining edge set E_res and the remaining faulty server set F_res are constructed according to the above idea. Then, the remaining subline scheduling scheme can be obtained according to the server set corresponding to the remaining sublines, the remaining edge set E_res and the remaining faulty server set F_res.
[0165] Furthermore, the scheduling scheme may be determined by merging the scheduling schemes of the sub-lines.
[0166] It can be seen that in this embodiment, by processing the sub-lines and the remaining parts of each sub-line in the deployment plan in stages, it can be ensured that while satisfying the first constraint and the second constraint, the optimal scheduling plan can be found for the server, thereby helping to improve the scheduling efficiency and reliability of the server.
[0167] It should be noted that this embodiment may also be an improvement based on the second embodiment and / or the third embodiment.
[0168] It is not difficult to find that in this embodiment, by gradually relaxing the constraints, the alignment constraints are processed first and then the coverage constraints are processed to achieve effective scheduling of servers. The time complexity of the method is O(nlog n), which shows that it is relatively efficient in processing large-scale problems.
[0169] Fifth embodiment
[0170] The fifth embodiment of the present application relates to a server scheduling method based on the Fat-Tree topology in a large model. The fifth embodiment is an improvement based on the fourth embodiment, and the specific improvement is: in the fifth embodiment of the present application, a specific implementation method for obtaining the scheduling scheme of each sub-line is provided according to the server set, edge set and target fault server set corresponding to each sub-line.
[0171] Specifically, in some embodiments, obtaining the scheduling scheme for each sub-line according to the server set, edge set and target faulty server set respectively corresponding to each sub-line may further include the following steps:
[0172] Get the number of servers included in each TP group;
[0173] In combination with the number of servers included in each TP group, a scheduling plan for each sub-line is obtained according to the server set, edge set and target fault server set respectively corresponding to each sub-line.
[0174] Specifically, in some examples, m can be used to represent the number of servers included in the TP group. Furthermore, the sub-line scheduling scheme can be obtained based on the server set S_subline, edge set E_subline and target fault server set F_subline corresponding to the sub-line and the number of servers m included in each TP group, and the sub-line scheduling scheme can be merged into the scheduling scheme placement_scheme. Similarly, the remaining sub-line scheduling schemes can be merged into the scheduling scheme placement_scheme. At this point, the merged scheduling scheme placement_scheme can be returned, and the scheduling scheme placement_scheme can contain the optimal scheduling schemes for all servers.
[0175] Optionally, in some embodiments, obtaining the number of servers included in each TP group may include the following steps:
[0176] Get the number of intelligent computing chips included in a TP group and the number of intelligent computing chips included in each server;
[0177] The number of servers included in each TP group is determined according to the number of intelligent computing chips included in the TP group and the number of intelligent computing chips included in each server.
[0178] Specifically, in some examples, the number of servers included in each TP group is determined according to the number of intelligent computing chips included in the TP group and the number of intelligent computing chips included in each server, which can be specifically implemented by the following formula:
[0179]
[0180] Among them, m represents the number of servers included in the TP group, t represents the number of intelligent computing chips included in the TP group, and r represents the number of intelligent computing chips included in each server.
[0181] It should be noted that this embodiment may also be an improvement based on any one or more of the first to third embodiments.
[0182] It is not difficult to find that compared with the related art, in this embodiment, by obtaining the number of servers included in each TP group; combined with the number of servers included in each TP group, the scheduling plan of each sub-line is obtained according to the server set, edge set and target fault server set corresponding to each sub-line, and a specific implementation method for obtaining the scheduling plan of each sub-line according to the server set, edge set and target fault server set corresponding to each sub-line is provided.
[0183] Sixth embodiment
[0184] The sixth embodiment of the present application relates to a server scheduling method based on a Fat-Tree topology in a large model. The sixth embodiment is an improvement based on the fifth embodiment, and the specific improvement is: in the sixth embodiment of the present application, a specific implementation method of obtaining the scheduling scheme of each sub-line is provided according to the server set, edge set and target fault server set corresponding to each sub-line, in combination with the number of servers included in each TP group.
[0185] Specifically, in some embodiments, the obtaining of the scheduling scheme for each sub-line according to the server set, edge set and target fault server set respectively corresponding to each sub-line in combination with the number of servers included in each TP group may further include the following steps:
[0186] Determine a healthy server set according to the server set corresponding to each of the sub-lines and the target faulty server set;
[0187] Determine a healthy edge set for characterizing a connection relationship between healthy servers in the healthy server set according to the healthy server set and the edge sets corresponding to each of the sub-lines;
[0188] Determine a healthy server subgraph according to the healthy server set and the healthy edge set;
[0189] Determine a connected component according to the healthy server subgraph; the connected component is used to represent an independent part obtained by segmenting a network connection based on the HBD topology structure;
[0190] The scheduling scheme for each sub-line is determined according to the connected components and the number of servers included in each TP group.
[0191] Specifically, all target faulty server sets can be removed from the server sets corresponding to each of the sub-lines in each sub-line, and the healthy server set H can be obtained. Further, a healthy edge set HE between the healthy servers can be constructed. Exemplarily, the healthy edge set HE can be initialized to all node pairs (u, v) in the healthy server set H. At this point, the healthy server subgraph HealthyHBD can be composed of the healthy server set H and the edge set HE.
[0192] Specifically, in some examples, the connected components are determined based on the healthy server subgraph by traversing each healthy server in the healthy server set and performing the following operation on each unvisited healthy server in the healthy server set: starting from the current healthy server, searching the healthy server subgraph for a healthy server connected to the current healthy server to obtain the connected components.
[0193] Optionally, in some embodiments, the connected components can be determined based on a depth-first search algorithm and the healthy server subgraph. Specifically, for each unvisited node s, a depth-first search algorithm can be executed to identify the connected components containing the node s, and all healthy servers connected to the node s can be traversed through the depth-first search algorithm.
[0194] Optionally, in some embodiments, determining the connected components according to the depth-first search algorithm and the healthy server subgraph may include the following steps: determining all neighbor servers of the target server in the healthy server subgraph according to the depth-first search algorithm; and determining the connected components based on the neighbor servers.
[0195] Specifically, in some examples, a stack can be initialized, and the stack is used to store the nodes to be visited. Among them, the initial state of the stack only contains the starting node node in the healthy server set H. Exemplarily, an empty list component can also be initialized, and the empty list component is used to store all the nodes in the found connected components. Further, the depth-first search algorithm can be executed, and the depth-first search algorithm enters a loop and continues to execute as long as the stack is not empty. In each loop, a node current can be popped out of the stack, and the node current is the next node to be processed. Further, it can be checked whether the popped node current has been visited. This can be achieved by checking whether the node current is in the visited set. If the node current has not been visited, the following operations are performed: the node current is added to the visited set and marked as visited; the node current is added to the component list to indicate that it is part of the connected component. Furthermore, all neighbor nodes of the node current in the healthy server subgraph HealthyHBD can be traversed, and for each neighbor node, if it has not been visited, it is pushed into the stack so that it can be visited in subsequent loops. When the stack is empty, it means that all reachable nodes have been visited and added to the connected component, and the depth-first search algorithm ends the loop. Finally, the depth-first search algorithm can return a component list, which can contain all nodes reachable from the starting node node, that is, a complete connected component.
[0196] Optionally, in some embodiments, determining a sub-line scheduling scheme for characterizing a scheduling scheme for each sub-line based on the connected components and the number of servers included in each TP group may include: detecting a size relationship between each of the connected components and the number of servers included in each TP group; and determining the sub-line scheduling scheme based on the size relationship.
[0197] Optionally, in some embodiments, determining a scheduling scheme according to the size relationship may further include the following steps:
[0198] If the size of the connected component is greater than or equal to the number of servers included in the TP group, a number of servers equal to the number of servers included in the TP group are taken out from the connected component, and the number of servers equal to the number of servers included in the TP group taken out from the connected component is taken as a TP group until all servers in the connected component are assigned to the TP group; the sub-line scheduling scheme is determined according to the TP group obtained after the connected component is assigned.
[0199] Specifically, in some examples, each connected component in the connected component list component_list can be traversed. For each connected component, check whether the size of the connected component is greater than or equal to the number of servers included in the TP group. If the size of the connected component is greater than or equal to the number of servers included in the TP group, m servers are taken out of the connected component, and the m servers taken out of the connected component are added to the scheduling scheme placement_scheme as a TP group. The above process will be repeated until all servers in the connected component are assigned to the TP group. After that, the generated scheduling scheme placement_scheme can be returned.
[0200] This embodiment may also be an improvement based on any one or more of the first to fourth embodiments.
[0201] It is not difficult to find that compared with the related art, in this embodiment, by traversing each healthy server in the healthy server set, the following operations are performed on each unvisited healthy server in the healthy server set: starting from the current healthy server, searching the healthy server connected to the current healthy server in the healthy server subgraph to obtain the connected component, providing a specific implementation method for determining the connected component based on the healthy server subgraph.
[0202] Seventh embodiment
[0203] The seventh embodiment of the present application relates to a server scheduling method based on a Fat-Tree topology in a large model. The seventh embodiment is an improvement based on the fourth embodiment, and the specific improvement is: in the seventh embodiment of the present application, a specific implementation method for determining the scheduling scheme according to the scheduling scheme of each sub-line is provided.
[0204] Specifically, in some embodiments, determining the scheduling scheme according to the scheduling schemes of the sub-lines, that is, step S1034, may further include the following steps:
[0205] Step S10341, determining an initial scheduling scheme according to the scheduling schemes of the sub-lines;
[0206] Step S10342, determining a lower limit value, and determining an upper limit value according to the number of backbone switch groups in the network in the Fat-Tree topology structure and the maximum number of sub-lines; the upper limit value and the lower limit value are used to characterize the search range;
[0207] Step S10343, determining an intermediate value according to the upper limit value and the lower limit value;
[0208] Step S10344: determine the scheduling scheme according to the intermediate value and the initial scheduling scheme.
[0209] Specifically, in some examples, the upper limit value may be represented by high, and the lower limit value may be represented by low. In some examples, the upper limit value and the lower limit value may be initialized. Exemplarily, the lower limit value may be initialized to 0; the number of backbone switch groups may be represented by n_domain, the maximum number of sub-lines may be represented by n_maxsubline, and the upper limit value may be the sum of the number of backbone switch groups n_domain and the maximum number of sub-lines n_maxsubline.
[0210] In addition, it should be noted that the upper limit value high may represent the total number of constraints, and these constraints may include backbone switch coverage constraints and alignment constraints.
[0211] Optionally, in some embodiments, the number of backbone switch groups n_domain may be obtained by the following formula:
[0212] n_domain=n / d;
[0213] Wherein, n represents the total number of servers in the Fat-Tree topology structure, and d represents the number of servers under the group of backbone switches.
[0214] Optionally, in some embodiments, the maximum number of sub-lines n_maxsubline may be obtained by the following formula:
[0215] n_maxsubline=n_domain×l;
[0216] Among them, the n_domain represents the number of backbone switch groups, and the l represents the length of the sub-line.
[0217] Specifically, in some examples, the middle value may be determined according to the upper limit and the lower limit. When the lower limit is less than or equal to the upper limit, the calculation formula of the middle value mid may refer to the following:
[0218]
[0219] Specifically, in some examples, the intermediate value can be used as a constraint variable, the intermediate value is the current constraint quantity, and a search is performed between the upper limit value and the lower limit value, and the scheduling scheme with the minimum constraint quantity that satisfies the job scale is found by continuously adjusting the intermediate value mid. If the scheduling scheme meets the requirements of the job scale, the scheduling scheme can be output; otherwise, the constraint quantity can be further adjusted through the intermediate value mid, and a new scheduling scheme can be re-determined.
[0220] Specifically, the size of the initial scheduling scheme |placement_scheme| (i.e., the number of servers in the current scheduling scheme), the number of servers m included in each TP group, and the number of intelligent computing chips r in each server can be determined. The product of the size of the initial scheduling scheme, the number of servers included in each TP group, and the number of intelligent computing chips in each server is: |placement_scheme|·m·r; the product result represents the total number of intelligent computing chips provided by the current scheduling scheme. The intermediate value can be updated according to the product result and the first constraint condition to obtain a new upper limit value and a new lower limit value.
[0221] Optionally, in some embodiments, the first constraint condition is specifically to make the number of intelligent computing chips in the scheduling scheme greater than or equal to the number of intelligent computing chips required for the task scale of the large model. Assuming that the number of intelligent computing chips required for the task scale is s, the product result and the number of intelligent computing chips s can be compared, and the intermediate value can be updated according to the comparison result to obtain a new upper limit value and a new lower limit value.
[0222] Specifically, in some examples, if the multiplication result |placement_scheme|·m·r≥s, it is considered that the current number of constraints may be too many, so try to reduce the number of constraints, and the lower limit value can be updated to the middle value plus 1, that is: low=mid+1. Otherwise, it is considered that the current number of constraints may be too few, and the number of constraints needs to be increased to better meet the job scale. The upper limit value can be updated to the middle value minus 1, that is: high=mid-1. And according to the new upper limit value and the new lower limit value, a new scheduling scheme is determined, and the cycle is repeated until a scheduling scheme that still satisfies the minimum number of constraints and still satisfies: |placement_scheme|·m·r≥s is obtained, and the final scheduling scheme is obtained. Exemplarily, if after the loop ends, if a scheduling scheme that meets the above conditions is still not found, the returned result can be empty (None). It can be seen that in this step, by adjusting the middle value mid, the search range can be gradually narrowed until a scheduling scheme that meets the conditions is found, or a result that no such scheduling scheme exists is obtained.
[0223] It should be noted that this embodiment may also be an improvement based on any one or more of the first to fourth embodiments.
[0224] It is not difficult to find that, compared with the related art, this embodiment provides a specific implementation method for determining the scheduling scheme based on the intermediate value, the initial scheduling scheme and the first constraint condition, so as to balance the job scale and network traffic to find the optimal server scheduling scheme. By gradually adjusting the number of constraints, the method can gradually narrow the search range, while meeting the job requirements, minimize the cross-track traffic in the network, thereby improving network efficiency.
[0225] The step division of the above methods is only for clear description. When implemented, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the scope of protection of this application.
[0226] It is understandable that as the computing power of intelligent computing chips continues to upgrade and the scale of large models continues to expand, communication has gradually become one of the bottlenecks in large model training tasks. Therefore, fully improving communication performance under existing network resources through software and hardware collaborative design is of great significance to the development of the artificial intelligence industry.
[0227] This application is the first in the industry to propose a server scheduling method for large model tasks that can simultaneously perceive HBD topology and Fat-Tree topology, and provides relevant algorithm implementation. Compared with related technologies, the technical solution provided by this application has at least the following beneficial effects:
[0228] Optimize resource utilization: By considering the Fat-Tree topology of HBD and DCN at the same time, hardware resources can be used more effectively and resource waste can be avoided. The HBD topology helps to identify the uneven distribution of hardware resources, while the Fat-Tree topology of DCN provides an optimization solution for network connections, jointly achieving efficient resource utilization.
[0229] Improve system reliability: By comprehensively considering the Fat-Tree topology of HBD and DCN, server scheduling can be optimized, thereby improving system reliability.
[0230] Reduced operating costs: This application can help enterprises reduce operating costs by optimizing resource utilization and improving system reliability to achieve cost-effectiveness. More efficient resource utilization can reduce hardware procurement and maintenance costs, while higher system reliability can reduce troubleshooting and system downtime time, further reducing related costs.
[0231] Enhanced scalability: This application can help enterprises more easily expand their server infrastructure to cope with growing demand. By comprehensively considering the Fat-Tree topology of HBD and DCN, this technology can provide more flexible scheduling options to adapt to different workloads and traffic patterns.
[0232] Improve communication performance: This application can improve the communication performance of the server by optimizing resource utilization and network connections. More efficient resource utilization can reduce resource contention and bottleneck problems, while the optimization of network connections can reduce latency and increase communication throughput.
[0233] In summary, it can be seen that the technology of this application is highly advanced and has great protection value.
[0234] In addition, some embodiments of the present application also provide an electronic device. The electronic device may be a digital computer in various forms, such as a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, etc. The electronic device may also be a mobile device in various forms, such as a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices.
[0235] The electronic device includes: one or more processors; and a memory storing computer program instructions, wherein when the computer program instructions are executed, the processor executes the steps of the method provided in any one or more of the above embodiments. Figure 7 An exemplary structural diagram of the electronic device is disclosed. Figure 7 As shown, the electronic device includes: one or more processors 1101, memory 1102, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Among them, the components shown in this article, their connections and relationships, and their functions are only used as examples, and are not intended to limit the implementation of the present application described and / or required herein.
[0236] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103 and the output device 1104 may be connected via a bus or other means. Figure 7 The example of connecting through bus is taken in the following.
[0237] The input device 1103 can receive input digital or character information, and generate key signal input related to the user settings and function control of the electronic device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator rod, one or more mouse buttons, a trackball, a joystick and other input devices. The output device 1104 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The display device may include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display and a plasma display. In some embodiments, the display device may be a touch screen.
[0238] In order to provide interaction with the user, the electronic device may be a computer. The computer has: a display device (e.g., a cathode ray tube (CRT) or an LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0239] In the embodiments of the present application, a computer program / instruction is stored on a computer-readable medium, and when the computer program / instruction is executed by a processor, the steps of the method provided by any one or more of the above embodiments are implemented. The computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently without being assembled into the device. The above computer-readable medium carries one or more computer-readable instructions.
[0240] The memory 1102 can be used as a non-transient computer-readable storage medium, which can be used to store non-transient software programs, non-transient computer executable programs and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transient software programs, instructions and modules stored in the memory 1102, so as to implement the program instructions / modules corresponding to the method provided by any one or more embodiments in the embodiments of the present application.
[0241] The memory 1102 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 1102 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory 1102 may optionally include a memory remotely arranged relative to the processor 1101, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0242] It should be noted that the computer-readable medium described in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM, Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.
[0243] Computer readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change random-access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0244] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0245] In the above-described embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. For example, it can be implemented by using a dedicated integrated circuit (ASIC, Application-Specific Integrated Circuit), a general-purpose computer or any other similar hardware device. In certain embodiments, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including relevant data structures) can be stored in a computer-readable recording medium, for example, a RAM memory, a magnetic or optical drive or a floppy disk and similar devices. In addition, some steps or functions of the present application can be implemented by hardware, for example, as a circuit that cooperates with a processor to perform each step or function.
[0246] The computer program product provided in the embodiment of the present application includes one or more computer programs / instructions, and when the computer program / instructions are executed by the processor, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from a computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL, Digital Subscriber Line)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive, SSD, solid statedisk), etc.
[0247] The flow chart or block diagram in the accompanying drawings shows the possible architecture, function and operation of the equipment, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0248] The scope of the present application is defined by the appended claims rather than the above description, and it is intended that all changes that fall within the meaning and scope of the equivalent elements of the claims are included in the present application. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the device claim may also be implemented by one unit or device through software or hardware. The words "first", "second", etc. are only used to distinguish the description, and do not indicate any particular order, nor can they be understood as indicating or implying relative importance.
[0249] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily mention changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims, and the above embodiments should be regarded as exemplary and non-restrictive.
Claims
1. A server scheduling method based on Fat-Tree topology in a large model, characterized in that: The method is applied to a Fat-Tree topology structure, and the method comprises: Obtain a deployment scheme of servers in a topological structure of an HBD; in the deployment scheme, the servers are arranged in a plurality of sub-lines, each sub-line corresponds to a group of servers; the last server in each sub-line is connected to the last server in the next sub-line, the first server in the next sub-line is connected to the first server in the next sub-line, and so on, forming an interlaced "zigzag" connection mode; Acquire a first constraint and a second constraint of the Fat-Tree topology structure; the first constraint is used to ensure that the TP group does not span multiple backbone switches, and the second constraint is used to ensure that the TP group is aligned; A scheduling scheme is determined according to the deployment scheme, the first constraint condition and the second constraint condition, so as to schedule the server according to the scheduling scheme.
2. The method according to claim 1, characterized in that The step of obtaining the deployment scheme of the server in the topology structure of the HBD includes: Determine first serial number information of the server in the Fat-Tree topology structure; According to the Fat-Tree topology, determine the number of servers directly connected to each ToR switch; Determine second numbering information according to the first numbering information and the number of servers directly connected to each ToR switch; the second numbering information is used to indicate the sequence number of the server in the topology structure of the HBD; A deployment plan of the servers in the topology structure of the HBD is obtained according to the second numbering information.
3. The method according to claim 1, characterized in that The first constraint condition and the second constraint condition for obtaining the Fat-Tree topological structure include: Determine the number of backbone switch coverage constraints; the number of backbone switch coverage constraints is used to characterize the number of backbone switch coverage constraints, so that each sub-line needs to be covered by at least one backbone switch; obtain the first constraint condition according to the number of backbone switch coverage constraints; The number of alignment constraints is determined; the number of alignment constraints is used to characterize the number of backbone switch domains that need to be aligned; and the second constraint condition is obtained according to the number of alignment constraints.
4. The method according to claim 3, characterized in that Determining the number of backbone switch coverage constraints includes: obtaining a total constraint number and a maximum number of sub-lines; determining a relatively smaller value between the total constraint number and the maximum number of sub-lines; determining the number of backbone switch coverage constraints based on the relatively smaller value between the total constraint number and the maximum number of sub-lines; The determining the number of alignment constraints includes: obtaining a total number of constraints and a maximum number of sub-lines; and determining the number of alignment constraints according to a difference between the total number of constraints and the maximum number of sub-lines.
5. The method according to claim 3, characterized in that: The determining of the scheduling scheme according to the deployment scheme, the first constraint condition and the second constraint condition comprises: Determine a subline ID according to the number of alignment constraints; If the subline ID belongs to the faulty server set, the faulty server set is updated, and servers on all sublines related to the subline ID are marked as faulty, to obtain an updated target faulty server set; According to the number of coverage constraints of the backbone switch, the following operations are performed on each sub-line: according to the length of the sub-line and the deployment plan, the server set, the edge set and the target faulty server set corresponding to each sub-line are obtained; according to the server set, the edge set and the target faulty server set corresponding to each sub-line, the scheduling plan of each sub-line is obtained; the scheduling plan of each sub-line is the scheduling plan of each sub-line; A scheduling scheme is determined according to the scheduling schemes of the sub-lines.
6. The method according to claim 5, characterized in that Determining the subline ID according to the number of alignment constraints includes: Determining a first traversal range according to the number of the alignment constraints; According to each first index in the first traversal range, traverse in the second traversal range; wherein the second traversal range is determined according to the number of servers under a group of backbone switches; A subline ID is determined according to the first index, a second index in the second traversal range, and the number of servers under the group of backbone switches.
7. The method according to claim 6, characterized in that The determining of the subline ID according to the first index, the second index in the second traversal range, and the number of servers under the group of backbone switches includes: Determine a product of the first index and the number of servers under the group of backbone switches; A subline ID is determined according to the multiplication result and the sum of the second index.
8. The method according to claim 5, characterized in that The obtaining of the scheduling scheme for each sub-line according to the server set, the edge set and the target faulty server set respectively corresponding to each sub-line comprises: Get the number of servers included in each TP group; In combination with the number of servers included in each TP group, a scheduling plan for each sub-line is obtained according to the server set, edge set and target fault server set respectively corresponding to each sub-line.
9. The method according to claim 8, characterized in that The obtaining of the number of servers included in each TP group includes: Get the number of intelligent computing chips included in a TP group and the number of intelligent computing chips included in each server; The number of servers included in each TP group is determined according to the number of intelligent computing chips included in the TP group and the number of intelligent computing chips included in each server.
10. The method according to claim 8, characterized in that In combination with the number of servers included in each TP group, and according to the server set, edge set and target faulty server set respectively corresponding to each sub-line, obtaining the scheduling scheme for each sub-line includes: Determine a healthy server set according to the server set corresponding to each of the sub-lines and the target faulty server set; Determine a healthy edge set for characterizing a connection relationship between healthy servers in the healthy server set according to the healthy server set and the edge sets corresponding to each of the sub-lines; Determine a healthy server subgraph according to the healthy server set and the healthy edge set; Determine a connected component according to the healthy server subgraph; the connected component is used to represent an independent part obtained by segmenting a network connection based on the HBD topology structure; The scheduling scheme for each sub-line is determined according to the connected components and the number of servers included in each TP group.
11. The method according to claim 5, characterized in that Determining the scheduling scheme according to the scheduling scheme of each sub-line includes: Determine an initial scheduling plan according to the scheduling plans of each of the sub-lines; Determine a lower limit value, and determine an upper limit value according to the number of backbone switch groups in the Fat-Tree topology and the maximum number of sub-lines allowed in the Fat-Tree topology; the upper limit value and the lower limit value are used to characterize the search range; Determine an intermediate value according to the upper limit value and the lower limit value; The scheduling scheme is determined according to the intermediate value and the initial scheduling scheme.
12. The method according to claim 11, characterized in that Determining the scheduling scheme according to the intermediate value and the initial scheduling scheme includes: Determine the size of the initial scheduling solution, the number of servers included in each TP group, and the number of intelligent computing chips in each server; According to the size of the initial scheduling scheme, the product of the number of servers included in each TP group and the number of intelligent computing chips in each server, and the first constraint condition, the intermediate value is updated to obtain a new upper limit value and a new lower limit value; The scheduling scheme is determined according to the new upper limit value and the new lower limit value.
13. An electronic device, characterized in that: The electronic device comprises: one or more processors; and A memory storing computer program instructions, which when executed cause the processor to perform the steps of the method as claimed in any one of claims 1 to 12.
14. A computer readable medium having a computer program / instructions stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.