Large model-based server scheduling method, device, medium and product
By forming an interlaced "zigzag" connection pattern in the HBD topology, determining the healthy server subgraph and connected components, and generating a scheduling plan, the problem of insufficient consideration of the synergy between DCN and HBD is solved, and the efficient utilization of intelligent computing chips and the improvement of network communication performance are achieved.
Patent Information
- Application Number
- CN202510126869.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-01-27
AI Technical Summary
In the existing technology, the server scheduling method of artificial intelligence data center does not fully consider the synergy between the data center network (DCN) and the high-bandwidth domain (HBD), resulting in the high complexity of the scheduling method.
By numbering the physical connections in the HBD topology according to the DCN topology, an interlaced "zigzag" connection pattern is formed, and the healthy server subgraph and connected components are determined. According to the connected components and the number of servers included in each TP group, a scheduling plan is generated.
It reduces the complexity of server scheduling, balances network load, improves the utilization and communication performance of intelligent computing chips, maximizes network communication performance, and improves overall resource utilization and task execution efficiency.
Smart Images

Figure CN119967063B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and particularly relates to a server scheduling method based on a large model, a device, a medium and a product. BACKGROUND
[0002] An artificial intelligence data center (AIDataCenter, AIDC for short) refers to a data center designed for high-performance computing and large-scale data processing for artificial intelligence applications. It is used for large model training or inference tasks, and the main computing devices include intelligent computing chips. In order to be compatible with the design requirements of the traditional data center network (DCN) and the communication needs of the large model, the computing network of the AIDC usually includes the DCN and the HBD. The DCN can realize remote direct memory access communication, has excellent scalability and low cost. And the HBD can provide higher bandwidth and lower latency to meet the needs of large-scale data transmission.
[0003] In the related art, optimizing communication delay is crucial to improving the overall resource utilization of the AIDC. In the AIDC, scheduling servers according to the characteristics of large language model (LLM) training tasks is an important means to reduce the overall network communication delay. Some manufacturers have developed server scheduling techniques based on the topology of the DCN.
[0004] However, the inventors have found that at least the following technical problems exist in the related art: In the related art, the computing power network is often divided into two independent networks: DCN and HBD. And when designing the network topology and scheduling tasks, more attention is paid to optimizing these two networks separately, without fully considering their synergies, resulting in high complexity of the scheduling method. SUMMARY
[0005] One object of the present application is to provide a server scheduling method based on a large model, a device, a medium and a product, at least to solve the technical problem that the complexity of the scheduling method is high due to the lack of full consideration of the synergies between the DCN and the HBD in the related art.
[0006] To achieve the above object, some embodiments of the present application provide the following aspects:
[0007] In a first aspect, some embodiments of the present application also provide a server scheduling method based on a large model. The method is applied to the deployment scheme obtained by the server deployment method described above. The scheduling method includes: determining a healthy server subgraph according to the deployment scheme; the healthy server subgraph includes a set of healthy servers and the connection relationship between each healthy server in the set of healthy servers; determining a connected component according to the healthy server subgraph; and determining a scheduling scheme according to the connected component and the number of servers included in each TP group.
[0008] In a second aspect, some embodiments of the present application further provide an electronic device, comprising: one or more processors; and a memory storing computer program instructions which, when executed, cause the processor to perform the steps of the method as described above.
[0009] In a third aspect, some embodiments of the present application further provide a computer readable medium having stored thereon computer program instructions executable by a processor to implement the method as described above.
[0010] In a fourth aspect, some embodiments of the present application further provide a computer program product comprising computer program / instructions which, when executed by a processor, implement the steps of the method as described above.
[0011] Compared with the related art, in the scheme provided by the embodiments of the present application, the method is applied to a preset server deployment scheme, in the deployment scheme, the servers are numbered in the topology structure of HBD according to the physical connection order in the topology structure of DCN; the servers are arranged in the manner of rows and columns, the last server of the target row is connected to the last server of the next row, the first server of the next row is connected to the first server of the row after the next row, and so on, thus forming an interlaced zigzag connection mode; the scheduling method comprises: determining a healthy server subgraph according to the deployment scheme; the healthy server subgraph comprises a healthy server set and a connection relationship between each healthy server in the healthy server set; determining a connected component according to the healthy server subgraph; and determining a scheduling scheme according to the connected component and the number of servers contained in each TP group. Since the deployment scheme in the present application is an interlaced zigzag connection mode, a certain redundancy can be provided, because even if a server in a certain row fails, data can still be transmitted through servers in other rows. At the same time, the zigzag connection mode also helps to balance the load in the network, because data can be transmitted in multiple paths. It can be understood that the generation of the traditional scheduling scheme based on large model servers is a problem of a complex process, in the present application, by first finding the connected component, and then reasonably determining the scheduling scheme for allocating servers according to the connected component and the number of servers contained in each TP group, a groundbreaking 、The large model server scheduling problem under the actual DCN is ingeniously divided into several large model scheduling sub-problems under ideal conditions. The problem decomposition of the complex process into smaller sub-problems can reduce the complexity of the problem and make it easier to solve. Therefore, the scheduling scheme can balance the utilization rate of the intelligent algorithm chip and the communication performance under the premise of meeting the task demand, maximize the network communication performance, and improve the overall resource utilization rate and task execution efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0012] One or more embodiments are illustrated by way of example in the figures that are part of this document and which illustrate key principles of the embodiments. The drawings are not to scale, and are provided merely for explanatory purposes. The drawings should not be considered to limit the scope of the embodiments in any way.
[0013] Figure 1 An exemplary flowchart of a large model-based server scheduling method according to some embodiments of the present application;
[0014] Figure 2 An exemplary flowchart of a large model-based server scheduling method according to some embodiments of the present application;
[0015] Figure 3 An exemplary flowchart of step S101 in a large model-based server scheduling method according to some embodiments of the present application;
[0016] Figure 4 An exemplary flowchart of step S102 in a large model-based server scheduling method according to some embodiments of the present application;
[0017] Figure 5 An exemplary flowchart of step S103 in a large model-based server scheduling method according to some embodiments of the present application;
[0018] Figure 6 An exemplary flowchart of a large model-based server scheduling method according to some embodiments of the present application;
[0019] Figure 7 An exemplary structural diagram of an electronic device according to some embodiments of the present application. DETAILED DESCRIPTION
[0020] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0021] The following terms are used herein.
[0022] AI chip, a chip used for artificial intelligence computing tasks, including AI chips, NPUs, TPUs, FPGAs, etc.
[0023] High-bandwidth domain, English full name High-bandwidth domain, abbreviated as HBD. A network architecture used to meet the high-bandwidth requirements in large model training, mainly used to support tensor parallelism and other communication-intensive parallel dimensions.
[0024] Data center network, English full name Data Center Network, abbreviated as DCN. A network architecture used to meet general communication needs, mainly supporting data parallelism, pipeline parallelism and other communication-intensive parallel dimensions.
[0025] Large language model, English full name Large Language Model, abbreviated as LLM.
[0026] Tensor parallelism, English full name Tensor Parallelism, abbreviated as TP. A model parallel technology.
[0027] Data parallelism, English full name Data Parallelism, abbreviated as DP, is a parallel computing strategy.
[0028] Pipeline parallelism, English full name Pipeline Parallelism, abbreviated as PP, is a parallel computing strategy.
[0029] Context parallel, English full name Context Parallel, abbreviated as CP, is a parallel computing strategy.
[0030] First embodiment
[0031] The first embodiment of the present application relates to a server scheduling method based on a large model. The method is applied to a preset server deployment scheme, as shown in Figure 1As shown, in the deployment scheme, the servers are numbered in the topology of HBD according to the physical connection order in the topology of DCN; the servers are arranged in the manner of rows and columns, the last server of a target row is connected to the last server of the next row, the first server of the next row is connected to the first server of the row after the next row, and so on, forming an interlaced "zigzag" connection mode. Figure 1 In the example shown, p=4, indicating that there are 4 rows of servers.
[0032] Specifically, in some examples, the deployment scheme is a server deployment scheme generated based on the topology of DCN and the topology of HBD.
[0033] Specifically, in some embodiments, the scheduling method can include the following steps, such as Figure 2 As shown:
[0034] Step S101, determining a healthy server subgraph according to the deployment scheme; the healthy server subgraph includes a healthy server set and the connection relationship between each healthy server in the healthy server set;
[0035] Step S101, determining a connected component according to the healthy server subgraph; the connected component is used to represent an independent part obtained by dividing the network connection based on the topology of HBD;
[0036] Step S103, determining a scheduling scheme according to the connected component and the number of servers contained in each TP group.
[0037] For step S101, specifically, in some examples, a healthy server subgraph including healthy servers and their connection relationships can be constructed according to the deployment scheme. Exemplarily, the healthy server subgraph can be represented as HealthyHBD. The healthy server subgraph HealthyHBD can be composed of a healthy server set H and an edge set HE.
[0038] For step S102, specifically, in some examples, connected components are identified in the health server subgraph. A component_list can be used to represent the set of connected components; it can be appreciated that when the topology of a large-scale HBD is split due to some reasons (e.g., failed servers), the originally complete network connection is broken, resulting in the originally continuous network structure being divided into several independent parts. Each independent part, i.e., a sub-HBD, can be regarded as a smaller, independent network unit, which loses direct network connection with each other. The sub-HBDs correspond to the connected components formed after the split. Also, a group of servers for performing tensor parallelism (TP), i.e., servers within a P group, are to be kept within the same sub-HBD and cannot span into other sub-HBDs to ensure that efficient communication and data exchange between them can be performed. It can be appreciated that if the TP group spans into other sub-HBDs, the communication between the servers within the group would need to be relayed through higher-level network devices (e.g., core switches), which would increase the communication delay, reduce the training efficiency, and possibly affect the performance and stability of the model training. In some examples, a depth-first search algorithm can be employed to identify the connected components.
[0039] For step S103, specifically, in some examples, the number of servers included in each TP group, which is used to represent a group of servers for tensor parallelism (TP), can be understood as follows: in tensor parallelism (TP), the number of servers included in each TP group refers to the allocation of different parts of a model to different servers or GPUs for simultaneous processing when processing tasks in parallel. This can speed up the calculation process, as multiple servers can work in parallel to process large data sets or complex models. In this application, each TP group can be regarded as a subset of servers deployed in the deployment scheme, which collectively process a portion of the task. For example, if a TP group consists of two servers, then the two servers will process the assigned task portions in parallel. Exemplarily, the number of servers included in each TP group can be represented by m; the connected components can be divided into several TP groups according to the connected components and the number of servers included in each TP group. Based on the excellent physical properties of the large-scale HBD topology, the TP groups can be sequentially arranged within each connected component to generate a scheduling scheme that can maximize the utilization of the intelligent computing chip. Exemplarily, the scheduling scheme can be represented by placement_scheme, which is a set of TP groups; for example, if the set of connected components component_list = [[1, 2, 3, 4, 5, 6, 7, 8, 9], [10, 11, 12, 13, 14, 15]], and the number of servers included in the TP group m = 4, then the scheduling scheme placement_scheme = {{1, 2, 3, 4}, {5, 6, 7, 8}, {10, 11, 12, 13}}.
[0040] Exemplarily, the scheduling scheme can be generated by a large model task deployer. The scheduling scheme is used to guide how servers allocate tasks and how to communicate in the topology of DCN and HBD to minimize network communication delay and maximize GPU utilization. It can be understood that after obtaining the scheduling scheme, the scheduling scheme can be applied to the actual server scheduling process.
[0041] It should be noted that the scheduling method can be applied to large-scale HBD, such as how to plan the Rank of servers in tasks to minimize communication delay under the known topology of HDB and the topology of DCN for specific LLM training tasks. Since each of the above steps only needs to traverse the edges and nodes in the healthy server subgraph once, the overall time complexity of the scheduling method is O(n).
[0042] It can be seen that in the embodiment of the present application, the scheduling scheme refers to a set of available TP groups, and the scheduling scheme can meet the needs of higher-level parallel strategies (such as data parallelism DP, context parallelism CP) and specific tasks. Since the utilization rate of intelligent computing chips is generally considered to be a more critical indicator than communication performance, the large model task deployer will give priority to maximizing the utilization rate of intelligent computing chips while ensuring that the tasks can run normally, and will also consider maximizing communication performance. It should be noted that in the process of determining the entire scheduling scheme, the scale of the task is regarded as a fixed parameter. That is to say, it is assumed that when executing the scheduling scheme, the scale of the task will not change, but a better server allocation scheme will be found under a given scale.
[0043] In some embodiments, the task scheduler can ignore the communication performance limitations of the DCN and focus on maximizing the utilization of the intelligent computing chip. In this case, the time complexity of the scheduling solution generation is O(n), which is linear in the number of servers.
[0044] It should be noted that the scheduling solution provided in the embodiment of the present application can be applied to different topologies.
[0045] It is not difficult to find that compared with the related art, in the embodiment of the present application, a healthy server subgraph is determined according to the deployment plan; the healthy server subgraph includes a healthy server set and the connection relationship between each healthy server in the healthy server set, and then the connected component is determined according to the healthy server subgraph, and then the scheduling plan is determined according to the connected component and the number of servers included in each TP group. It can be understood that the generation of the traditional scheduling plan based on the large model server is a complex process. In the present application, the scheduling plan for allocating servers is determined reasonably according to the connected component and the number of servers included in each TP group by first finding the connected component, which is groundbreaking. 、 It cleverly proposes to divide the large-model server scheduling problem under actual DCN into several large-model scheduling sub-problems under ideal conditions, so that the problem of the complex process can be decomposed into smaller sub-problems, which can reduce the complexity of the problem and make it easier to solve. In this way, the scheduling scheme can balance the utilization and communication performance of the intelligent computing chip while meeting the task requirements, maximize the network communication performance, and improve the overall resource utilization and task execution efficiency.
[0046] Second embodiment
[0047] The second embodiment of the present application relates to a large model-based server scheduling method. The second embodiment is an improvement based on the first embodiment, and the specific improvement is that in the second embodiment of the present application, a specific implementation manner of determining a healthy server subgraph according to the deployment scheme is provided.
[0048] Specifically, in some embodiments, the step S101 of determining a healthy server subgraph according to the deployment scheme can further include the following steps, as shown in the following. Figure 3
[0049] Step S1011, determining the healthy server set according to the deployment scheme;
[0050] Step S1011, determining a healthy edge set for representing the connection relationship between each healthy server in the healthy server set according to the healthy server set;
[0051] Step S1013, determining a healthy server subgraph according to the healthy server set and the healthy edge set.
[0052] Optionally, in some embodiments, the step S1011 of determining the healthy server set according to the deployment scheme can further include the following steps:
[0053] Step S10111, obtaining a topology structure graph based on HBD according to the deployment scheme;
[0054] Step S10112, determining the healthy server set according to the topology structure graph and the failed server set.
[0055] Specifically, in some examples, the topology structure graph of HBD corresponding to the deployment scheme can be, but is not limited to, an undirected graph. Illustratively, the topology structure graph can be represented by infHBD. In the topology structure graph, a node S represents a server, and E represents the connection relationship between servers. In this way, the topology structure graph infHBD can be composed of a server set S and an edge set E. Among them, the server set can be an ordered set, and the nodes are numbered according to the physical connection order of the servers in the topology structure of the DCN, and the connection relationship can include a primary link and a backup link.
[0056] Specifically, in some examples, the failed server set can be represented by F, and the healthy server set can be represented by H. The edge set for representing the connection relationship between each healthy server in the healthy server set can be represented by HE.
[0057] Exemplarily, all the faulty server nodes, i.e., the faulty server set F, can be removed from the topology graph infHBD, and the healthy server set H can be obtained. Further, a healthy edge set HE between the healthy servers can be constructed. Exemplarily, the healthy edge set HE can be initialized as all the node pairs (u, v) in the healthy server set H. It can be understood that the node pairs (u, v) are edges connecting healthy servers in the edge set E in the topology of the original topology graph infHBD. Thus far, the healthy server subgraph HealthyHBD can be composed of the healthy server set H and the edge set HE.
[0058] It can be found that, compared with the related art, in the embodiments of the present application, the healthy server set is determined according to the deployment scheme; the healthy edge set used to represent the connection relationship between each healthy server in the healthy server set is determined according to the healthy server set; and the healthy server subgraph is determined according to the healthy server set and the healthy edge set. A specific implementation manner of determining the healthy server subgraph according to the deployment scheme is provided.
[0059] Third Embodiment
[0060] The third embodiment of the present application relates to a server scheduling method based on a large model. The third embodiment is an improvement on the basis of the first embodiment, and the specific improvement is that in the third embodiment of the present application, a specific implementation manner of determining a connected component according to the healthy server subgraph is provided.
[0061] Specifically, in some embodiments, the step S102 of determining the connected component according to the healthy server subgraph can further include the following steps, as shown in Figure 4
[0062] Step S1021, traversing each healthy server in the healthy server set, and performing the following operation on each unvisited healthy server in the healthy server set:
[0063] Step S1022, starting from the current healthy server, searching for healthy servers connected to the current healthy server in the healthy server subgraph to obtain the connected component.
[0064] In particular, in some examples, before determining the connected components according to the health server subgraph, three empty data structures can be initialized, respectively: a connected component list component_list, a visited node set visited, and a final placement scheme placement_scheme. Specifically, the connected component list component_list is initialized as an empty list, the visited node set visited is initialized as an empty set, and the placement scheme placement_scheme is initialized as an empty set.
[0065] Further, each server in the health server set H, i.e., a node s, can be traversed, for each unvisited node s, a connected component containing the node s can be identified, and all health servers connected to the node s can be traversed; further, all health servers connected to the node s can be marked as visited, and the found connected component can be added to the connected component list component_list. Embodiments of the present application do not limit the specific algorithm for determining the connected component, and any algorithm capable of obtaining the connected component is within the protection scope of the present application.
[0066] Optionally, in some embodiments, the connected component can be determined according to a depth-first search algorithm and the health server subgraph. Specifically, for each unvisited node s, a depth-first search algorithm can be executed to identify a connected component containing the node s, and all health servers connected to the node s can be traversed through the depth-first search algorithm.
[0067] Optionally, in some embodiments, the determining the connected component according to the depth-first search algorithm and the health server subgraph can include the following steps: determining all neighbor servers of a target server in the health server subgraph according to the depth-first search algorithm; and determining the connected component according to the neighbor servers.
[0068] In some examples, a stack stack can be initialized to store nodes to be visited. The stack stack can initially contain only the starting node node in the healthy server set H. In some examples, an empty list component can also be initialized to store all nodes in the found connected component. In some examples, the depth-first search algorithm can be executed in a loop until the stack stack is empty. In each loop, a node current can be popped from the stack stack, which is the node to be processed next. In some examples, it can be checked whether the node current has been visited. This can be done by checking whether the node current is in the visited set. If the node current has not been visited, the node current can be added to the visited set and marked as visited. In some examples, the node current can be added to the component list to indicate that it is part of a connected component. In some examples, all neighbor nodes of the node current in the healthy server subgraph HealthyHBD can be traversed. For each neighbor node, if it has not been visited, it can be pushed into the stack stack to be visited in a subsequent loop. When the stack stack is empty, it means that all reachable nodes have been visited and added to the connected component, and the depth-first search algorithm can end the loop. Finally, the depth-first search algorithm can return the component list, which can contain all nodes reachable from the starting node node, i.e., a complete connected component.
[0069] Optionally, in some embodiments, after determining the connected component from the healthy server subgraph, i.e., after step S202, the method can further include ranking the connected component to obtain a ranked target connected component. Correspondingly, in step S203, the scheduling scheme is determined according to the connected component and the number of servers included in each TP group, specifically, the scheduling scheme is determined according to the target connected component and the number of servers included in each TP group.
[0070] In some examples, to ensure the rationality of the scheduling scheme, each connected component can be ranked to obtain a ranked target connected component, and the target connected component can be added to the connected component list component_list.
[0071] It should be noted that the embodiments of the present application can also be improved on the basis of the second embodiment.
[0072] It can be found that, compared with the related art, in the embodiment, by traversing each health server in the health server set, for each health server in the health server set that has not been accessed, the following operation is performed: starting from the current health server, searching for health servers connected to the current health server in the health server subgraph to obtain the connected component, thereby providing a specific implementation manner of determining a connected component according to the health server subgraph.
[0073] Fourth embodiment
[0074] The fourth embodiment of the present application relates to a server scheduling method based on a large model. The fourth embodiment is an improvement on the basis of the first embodiment, and the specific improvement is that in the fourth embodiment of the present application, a specific implementation manner of determining a scheduling scheme according to the connected component and the number of servers contained in each TP group is provided.
[0075] Specifically, in some embodiments, determining the scheduling scheme according to the connected component and the number of servers contained in each TP group, i.e., step S103, can further include the following steps, as shown in Figure 5
[0076] Step S1031, detecting the size relationship between each connected component and the number of servers contained in each TP group;
[0077] Step S1032, determining a scheduling scheme according to the size relationship.
[0078] Optionally, in some embodiments, determining the scheduling scheme according to the size relationship, i.e., step S2032, can further include the following steps:
[0079] If the size of the connected component is greater than or equal to the number of servers in the TP group, a number of servers equal to the number of servers in the TP group are taken out from the connected component, and the servers taken out from the connected component are taken as a TP group, until all servers in the connected component are allocated to TP groups.
[0080] Determining a scheduling scheme according to the TP groups obtained after the connected component is allocated.
[0081] In particular, in some examples, each connected component in the connected component list component_list can be traversed. For each connected component, it is checked whether the size of the connected component is greater than or equal to the number m of servers in the TP group. If the size of the connected component is greater than or equal to the number m of servers in the TP group, m nodes are taken from the connected component, and the m nodes taken from the connected component are added to the scheduling scheme placement_scheme as a TP group. The above process is repeated until all nodes in the connected component are assigned to a TP group. After that, the generated scheduling scheme placement_scheme can be returned.
[0082] It should be noted that the embodiments of the present application can also be improved on the basis of the second embodiment and / or the third embodiment.
[0083] Fifth Embodiment
[0084] The fifth embodiment of the present application relates to a server scheduling method based on a large model. The fifth embodiment is an improvement on the basis of the first embodiment, and the specific improvement is that in the fifth embodiment of the present application, a determination method of the deployment scheme is provided.
[0085] In particular, in some embodiments, the determination method of the deployment scheme can include the following steps, as shown in Figure 6
[0086] Step S201, determining first numbering information of servers in the topology structure of the DCN;
[0087] Step S202, determining parallel parameters according to the topology structure of the DCN; the parallel parameters are used to represent the number of sub-lines, and each sub-line corresponds to a group of servers;
[0088] Step S203, determining second numbering information according to the first numbering information and the parallel parameters; the second numbering information is used to represent the serial number of the servers in the topology structure of the HBD; it can be seen that there is a mapping relationship between the first numbering information and the second numbering information;
[0089] Step S204, determining a server deployment scheme based on the topology structure of the HBD in the topology structure of the HBD according to the second numbering information.
[0090] Specifically, in some examples, the DCN topology may be a Rail-Optimized topology or a Fat-Tree topology, which is not specifically limited in this embodiment. The first numbering information is used to represent an ordered set of numbers corresponding to each server. Different servers in the DCN topology have their own numbers, which can be represented by letters such as A, B, and C, or by numbers, or a combination of numbers and letters, which is not limited in this embodiment. For example: Server 1, Server 2, ..., Server n, where n is a positive integer representing the total number of servers in the DCN topology.
[0091] Specifically, in some examples, different types of topologies correspond to different parallel parameters. The parallel parameters are the number of sub-lines. Figure 2 The figure shows a network topology consisting of 32 servers. The servers are numbered from 1 to 32 and are connected together in a certain pattern. Specifically, the servers are arranged in rows and columns, with 8 servers in each row and a total of 4 rows. The last server in each row is connected in sequence, that is, each server is connected to the server to its right, and, except for the last row, the last server in each row is also connected to the first server in the next sub-line, forming an interlaced "zigzag" connection pattern. The details are as follows:
[0092] Row 1: 1-5-9-13-17-21-25-29;
[0093] Row 2: 2-6-10-14-18-22-26-30;
[0094] Row 3: 3-7-11-15-19-23-27-31;
[0095] Row 4: 4-8-12-16-20-24-28-32;
[0096] Among them, servers numbered 1, 5, 9, 13, 17, 21, 25, and 29 form a sub-line; servers numbered 2, 6, 10, 14, 18, 22, 26, and 30 form a sub-line; servers numbered 3, 7, 11, 15, 9, 23, 27, and 31 form a sub-line; and servers numbered 4, 8, 12, 16, 20, 24, 28, and 32 form a sub-line. Server 29 in the first row is connected to server 2 in the second row; server 30 in the second row is connected to server 3 in the third row; and server 31 in the third row is connected to server 4 in the fourth row, forming a closed loop.
[0097] It should be noted that, in Figure 2In the shown example, the sub-lines are specifically parallel to each other. However, in some other examples, the sub-lines can also be non-parallel to each other; or the sub-lines can be broken lines themselves, which are not specifically limited in the embodiments of the present application.
[0098] Specifically, in some examples, the second numbering information in the topology of the HBD is determined according to the first numbering information and the parallel parameter. Here, the type of the topology of the HBD can be any ring topology or linear topology, which is not specifically limited in the embodiments. It can be understood that in a ring topology, the servers are connected to form a closed ring, and each server is directly connected to another two servers to form a continuous path. In a linear topology, the servers or nodes are linearly arranged, and each node is usually connected to only an adjacent node. It can be seen that through this step, the mapping relationship between the numbering of the servers in the topology of the DCN and the numbering of the servers in the topology of the HBD can be determined, so that the position and connection relationship of the servers in the topology of the HBD can be determined according to the numbering of the servers in the topology of the DCN.
[0099] Specifically, in some examples, the second numbering information represents a deployment scheme of the servers in the topology of the HBD, and relevant personnel can deploy the corresponding servers in the topology of the HBD according to the second numbering information. The deployment scheme includes the position relationship and connection relationship of the servers carried in the second numbering information.
[0100] It can be found that, compared with the related art, in the scheme provided by the embodiments of the present application, based on the collaborative design of the topology of the HBD and the topology of the DCN, a server deployment method based on a large model is proposed. The first numbering information of the servers in the topology of the DCN is determined, and the parallel parameter is determined according to the topology of the DCN. The parallel parameter is used to represent the number of sub-lines. Then, the second numbering information is determined according to the first numbering information and the parallel parameter, so that the deployment scheme of the servers in the topology of the HBD is determined according to the second numbering information. The second numbering information is used to represent the serial number of the servers in the topology of the HBD. In the deployment scheme, the second numbering information is used as the basis for the head-to-tail connection between the parallel sub-lines. It can be seen that the present application can simultaneously perceive the topology of the HBD and the topology of the DCN. Since the parallel parameter is determined according to the topology of the DCN, it is beneficial to obtain a deployment scheme of the servers in the topology of the HBD according to the characteristics of the topology of the DCN. Since the deployment scheme has strong adaptability, it is beneficial to optimize the use of network resources, improve the performance of the training task based on a large model, and facilitate subsequent server scheduling.
[0101] Specifically, in some embodiments, the DCN topology includes at least a Rail-Optimized based topology or a Fat-Tree based topology; and in the case that the DCN topology is a Rail-Optimized based topology, the determining the parallelism parameter according to the DCN topology can further include the following steps:
[0102] obtaining the number of rails of the Rail-Optimized based topology and the number of chips in a single server;
[0103] determining the parallelism parameter according to the number of rails and the number of chips in a single server;
[0104] Specifically, in some examples, the number of rails is used to identify the “rails” or “layers” in the Rail-Optimized based topology. In this way, the path of data flow and potential bottlenecks can be understood through the number of rails.
[0105] Specifically, in some examples, the number of chips in a single server is the number of chips included in each server in the Rail-Optimized based topology.
[0106] Optionally, in some embodiments, the determining the parallelism parameter according to the number of rails and the number of chips in a single server can be achieved through the following formula:
[0107] p = k / r
[0108] wherein, k represents the number of rails, r represents the number of chips in a single server, p represents the parallelism parameter.
[0109] Optionally, in some embodiments, in the case that the DCN topology is a Fat-Tree based topology, the determining the parallelism parameter according to the DCN topology can further include the following steps: obtaining the number of servers under each layer of switches in the Rail-Optimized based topology; and determining the parallelism parameter according to the number of servers under each layer of switches.
[0110] Specifically, in some examples, the number of servers under each layer of switches in the Rail-Optimized based topology is the number of servers directly connected to each layer of switches, which includes the number of servers directly connected to the top-of-rack switches, the size of the backbone switch domain, and the number of servers under higher layer switches.
[0111] Optionally, in some embodiments, the determining the parallel parameter according to the number of servers under each layer switch can be achieved by the following formula:
[0112] p = a
[0113] wherein, a represents the number of servers directly connected to each rack-top switch, p represents the parallel parameter.
[0114] Specifically, in some embodiments, the determining the second numbering information according to the first numbering information and the parallel parameter can further comprise the following steps: determining a first ordered set of servers deployed in the topology of the DCN according to the total number of servers in the topology of the DCN and the first numbering information; determining the length of each parallel sub-line according to the first ordered set and the parallel parameter; wherein the length of each parallel sub-line represents the number of servers each parallel sub-line should contain; and determining the second numbering information according to the length of each parallel sub-line.
[0115] Specifically, in some examples, the first ordered set can be represented by S wherein, if the total number of servers in the topology of the DCN is n, the first ordered set S wherein, the numbering of the servers is from 1 to n. Wherein, the numbering of the servers is determined by the position of the servers in the topology of the DCN.
[0116] Specifically, in some examples, the determining the length of each parallel sub-line according to the first ordered set and the parallel parameter can be achieved by the following formula:
[0117]
[0118] wherein, l represents the length of each parallel sub-line, the S represents the first ordered set, the p represents the parallel parameter.
[0119] Optionally, in some embodiments, the determining the second numbering information according to the length of each parallel sub-line can further comprise the following steps:
[0120] determining a data range according to the parallel parameter;
[0121] determining the index of each parallel sub-line according to the data range;
[0122] determining a position index within a target parallel sub-line according to the index of the parallel sub-line and a length of the parallel sub-line; the target parallel sub-line is determined according to the index of the parallel sub-line;
[0123] determining a second ordered set of servers deployed in a topology of HBD according to the index of the parallel sub-line, the position index and the parallel parameter;
[0124] determining the second number information according to the second ordered set.
[0125] Specifically, in some examples, if the parallel parameter is p , the data range can be from 0 to p-1 .
[0126] Specifically, in some examples, the index of each parallel sub-line can be determined according to the data range, i.e. from 0 to p-1 . It can be understood that the index i of each parallel sub-line corresponds to a parallel sub-line. That is, the index i of the parallel sub-line can be used to determine which parallel sub-line is currently being processed, so as to facilitate the allocation of servers to different parallel sub-lines. For example, p=4 , the index i of the parallel sub-line can be 0, 1, 2, 3, and different index values each represent a different parallel sub-line.
[0127] Specifically, in some examples, j can be used to represent the position index within the target parallel sub-line, i.e. the parallel sub-line corresponding to the index i of the parallel sub-line. That is, the position index j is used to determine the position of the server in the parallel sub-line corresponding to the index i of the parallel sub-line. In some examples, according to the parity of the index i of the parallel sub-line, the traversal range of the position index j within the target parallel sub-line will be different.
[0128] Optionally, in some embodiments, the determining the position index within the target parallel sub-line according to the index of the parallel sub-line and the length of the parallel sub-line can further include the following steps:
[0129] determining the parity of the index of the parallel sub-line;
[0130] determining the position index within the target parallel sub-line according to the parity and the length of the parallel sub-line.
[0131] Optionally, in some embodiments, the determining the position index within the target parallel sub-line according to the parity and the length of the parallel sub-line can further include:
[0132] If the index of the parallel sub-line is even, the position index is sequentially increased based on the length of the parallel sub-line;
[0133] If the index of the parallel sub-line is odd, the position index is sequentially decreased based on the length of the parallel sub-line.
[0134] Specifically, if the index i of the parallel sub-line is even, the position index j can be sequentially increased based on the length l of the target parallel sub-line, which means that the number of the server is sequentially increased, such as starting from 0; if the index i of the parallel sub-line is odd, the position index j can be sequentially decreased based on the length l of the target parallel sub-line, which means that the number of the server is sequentially decreased, such as starting from l-1. For example, if the index i of the parallel sub-line is even, the range of the position index j is from 0 to l-1; if the index i of the parallel sub-line is odd, the range of the position index j is from l-1 to 0. It can be seen that the position index j is used to traverse the server in each parallel sub-line to determine the new position of each server based on the server deployment scheme of the topology structure of HBD in the topology structure of HBD, and the new position corresponds to a new serial number.
[0135] Optionally, in some embodiments, the determining, according to the index of the parallel sub-line, the position index and the parallel parameter, the second ordered set for characterizing the servers deployed in the topology structure of HBD can be realized by the following formula:
[0136] i+j·p
[0137] Wherein, the i represents the index of the parallel sub-line, the j represents the position index, and the p represents the parallel parameter.
[0138] Specifically, in some examples, an empty ordered set S deploy can be initialized first, which is used to store the second serial number information. It can be known from the context that the second serial number information is obtained by converting the first serial number information of the servers in the topology structure of DCN through a preset rule, and specifically, the first serial number information of the servers in the topology structure of DCN can be converted to obtain the new serial number of each server deployed in the topology structure of HBD, i.e. the second ordered set S deploy , by combining the index i of each parallel sub-line, the position index j and the parallel parameter p.
[0139] For example, each i+j·p obtained can be added to the above-mentioned initialized empty ordered set to obtain the second ordered set S deploy . The second ordered set Sdeploy is an ordered set of servers in the topology of HBD, which is composed of the serial numbers of servers in the topology of HBD. It is mentioned above that the index i of the parallel sub-line is used to determine which parallel sub-line is currently being processed; the position index j is used to determine the position of the server in the parallel sub-line corresponding to the index i of the parallel sub-line, and p represents the parallel parameter.
[0140] Referring to Figure 1 Fig. 1, assuming that the parallel parameter is p, there are p parallel sub-lines L1, L2,..., L p , then the i-th parallel sub-line is composed of servers whose serial numbers corresponding to the position index are divided by the parallel parameter p with a remainder of i-1, that is:
[0141] Li={j|(j-1)%p==i-1}
[0142] The L i represents the parallel sub-line corresponding to the index i of the parallel sub-line.
[0143] Further, a series of the p parallel sub-lines are connected head to tail to form a new arrangement of deploying servers in the topology of HBD. Figure 2 In the topology of HBD, the second ordered set S deploy is composed as follows:
[0144] S deploy ={1,5,9,13,17,21,25,29,
[0145] 30,26,22,18,14,10,6,2,
[0146] 3,7,11,15,19,23,27,31,
[0147] 32,28,24,20,16,12,8,4}
[0148] Optionally, in some embodiments, the determining the second numbering information according to the second ordered set can further include the following steps:
[0149] determining an edge set representing the connection relationship of the servers deployed in the topology of HBD according to the second ordered set;
[0150] determining the second numbering information according to the second ordered set and the edge set.
[0151] Specifically, the second numbering information can be represented as G deploy , and the second ordered set can be represented as S deployIndicates that the edge set can be represented by E deploy Indicates that the second number information G deploy = deploy ,E depoly >.
[0152] Specifically, in some embodiments, determining, based on the second ordered set, an edge set used to characterize connection relationships between servers deployed in the topology structure of the HBD may further include the following steps:
[0153] Obtain the number of servers directly connected to the server in a single direction in the HBD topology;
[0154] determining a target server according to the second ordered set and the number of servers directly connected to the server in a single direction in a topology structure of the HBD;
[0155] An edge set for characterizing connection relationships between servers deployed in a topology structure of an HBD is determined according to the second ordered set and the target server.
[0156] Specifically, in some examples, the number of servers that the server is directly connected to in a single direction in the HBD topology can be represented by φ. For example, the number of servers that the server is directly connected to in a single direction in the HBD topology is φ. For example, when φ=2, the server S deploy [i]Can be directly connected to S deploy [i-2], S deploy [i-1], S deploy [i+1] and S deploy [i+2]. Combine Figure 2 As shown, server 9 can be directly connected to servers 1, 5, 13, and 17.
[0157] Specifically, in some examples, the second ordered set S can be traversed deploy , according to the second ordered set S deploy Each server in And the number of servers directly connected to the server in a single direction in the HBD topology structure φ, find all target servers that meet the following formula
[0158] ji≤φ
[0159] Specifically, in some examples, the second ordered set and the target server may form For every pair that satisfies the condition Can be in E deploy Create an edge in the server Can be directly connected to the server The E deploy Indicates the edge set.
[0160] It should be noted that the edge set E deploy Also meet the following conditions:
[0161] 1≤i≤j≤|S deploy |
[0162] Through the above formula, it can be ensured that the scheme only considers the case that the index i of the parallel sub-line is less than or equal to the position index j, so that the same connection relationship is avoided to be repeatedly created. It can be understood that if Connected to Then there is no need to make Connected to
[0163] Optionally, in some embodiments, the edge set used to represent the connection relationship between the servers deployed in the topology of the HBD is determined according to the second ordered set and the target server, specifically through the following formula:
[0164]
[0165] Wherein, the i represents the index of the parallel sub-line, the j represents the position index, the S deploy Indicates the second ordered set, the Indicates the server in the second ordered set S deploy , the Indicates the target server, the φ indicates the number of servers directly connected in a single direction in the topology of the HBD, and the E deploy Indicates the edge set.
[0166] It can be seen that the edge set E deploy , including all pairs satisfying |1≤i≤j≤|S deploy | and j-i≤φ. Thus, the establishment of the connection relationship between the second ordered set S deploy And the target server is completed.
[0167] The step division of the above various methods is only for clear description, and can be combined into one step or some steps can be split and decomposed into multiple steps when implemented, as long as the same logical relationship is included, all are within the protection scope of the present application; adding irrelevant modifications or introducing irrelevant designs in the algorithm or process, but not changing the core design of the algorithm and process are within the protection scope of the present application.
[0168] It can be understood that, under the background of increasing computing power of intelligent chips and continuous expansion of large model scale, communication gradually becomes one of the bottlenecks of large model training tasks. Therefore, fully improving the communication performance under the existing network resources through the collaborative design of software and hardware is of great significance to the development of the artificial intelligence industry.
[0169] The present application is the first proposed server deployment and large model task server scheduling method that can simultaneously perceive HBD topology and DCN topology, and provides the implementation of the related system. Compared with the related art, the technical solution provided by the present application has at least the following beneficial effects:
[0170] Optimize resource utilization: by considering the topology of HBD and DCN at the same time, hardware resources can be more effectively utilized, and resource waste can be avoided. HBD topology helps to identify the uneven distribution of hardware resources, while DCN topology provides an optimization scheme for network connection, together realizing efficient utilization of resources.
[0171] Improve system reliability: by considering the topology of HBD and DCN, the deployment and scheduling of servers can be optimized, thereby improving the reliability of the system. For example, by deploying servers in a DCN topology with redundant network connections, the impact of network failures on system operation can be reduced.
[0172] Reduce operating costs: the present application optimizes resource utilization and improves system reliability to achieve cost-effectiveness, which can help enterprises reduce operating costs. More efficient resource utilization can reduce hardware procurement and maintenance costs, while higher system reliability can reduce troubleshooting and system downtime, thereby further reducing related costs.
[0173] Enhance scalability: the present application can help enterprises more easily expand their server infrastructure to meet growing demand. By considering the topology of HBD and DCN, the technology can provide more flexible deployment and scheduling options to adapt to different workloads and traffic patterns.
[0174] Improve communication performance: the present application can improve the communication performance of servers by optimizing resource utilization and network connections. More efficient resource utilization can reduce resource contention and bottleneck problems, while network connection optimization can reduce latency and improve communication throughput.
[0175] In summary, the present application has high technical advancement and great protection value.
[0176] The arrows in the above flowchart represent the execution order and dependency relationship between these steps, and the later completed steps depend on the earlier completed steps.
[0177] In addition, some embodiments of the present application also provide an electronic device. The electronic device can be various forms of digital computers, such as a laptop computer, a desktop computer, a workstation, a personal digital assistant, a server, a blade server, a mainframe computer, and the like. The electronic device can also be various forms of mobile devices, such as a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices.
[0178] The electronic device includes one or more processors, and a memory storing computer program instructions which, when executed, cause the processor to perform the steps of the method provided by any one or more embodiments described above. Figure 7 An exemplary structural diagram of the electronic device is disclosed. As shown in Figure 7 The electronic device includes one or more processors 1101, a memory 1102, and an interface for connecting components, including a high-speed interface and a low-speed interface. Various components are connected to each other by different buses, and can be installed on a common motherboard or otherwise installed as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information on a GUI on an external input / output device, such as a display device coupled to the interface. In some other embodiments, multiple processors and / or buses can be used with multiple memories and multiple storage, if needed. Also, multiple electronic devices can be connected, each device providing part of the necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Among them, the components shown herein, their connections and relationships, and their functions are only examples, and are not intended to limit the implementation of the present application described and / or claimed herein.
[0179] The electronic device can also include an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103, and the output device 1104 can be connected by a bus or other means, Figure 7 for example, by a bus connection.
[0180] The input device 1103 can receive input of a number or character information, and generate a key signal corresponding to a user command input with respect to a function or setting of the electronic device, for example, a touch screen, a key pad, a mouse, a track ball, a touch pad, a jog wheel, one or more mouse buttons, a track ball, a joystick, etc. The output device 1104 can include a display device, a sub-lit device (e.g., an LED), a haptic feedback device (e.g., a vibration motor), etc. The display device can include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display, and a plasma display. In some embodiments, the display device can be a touch screen.
[0181] To provide for interaction with a user, the electronic device can be a computer. The computer has a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0182] In the embodiments of the present application, the computer program / instruction is stored on the computer readable medium, and the computer program / instruction is executed by the processor to implement the steps of the method provided by any one or more of the embodiments. The computer readable medium can be included in the electronic device described in the above embodiments, or can exist separately and not be assembled into the device. The computer readable medium carries one or more computer readable instructions.
[0183] The memory 1102 can be used as a non-transitory computer readable storage medium for storing non-transitory software programs, non-transitory computer executable programs and modules. The processor 1101 executes various functions and data processing of the server by running the non-transitory software programs, instructions and modules stored in the memory 1102, so as to implement the program instructions / modules corresponding to the method provided by any one or more of the embodiments in the present application.
[0184] The memory 1102 can include a program region that can store an operating system, an application program required for at least one function, and a data region that can store data created according to usage of the electronic device, etc. In addition, the memory 1102 can include a high-speed random access memory, and also can include a non-transitory memory such as at least one of a magnetic disk storage device, a flash memory device, or other non-transitory solid state storage device. In some embodiments, the memory 1102 can optionally include a memory disposed remotely from the processor 1101, which can be connected to the electronic device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0185] It should be noted that the computer-readable medium described in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, device or apparatus, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus.
[0186] The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage device, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0187] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0188] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. For example, the embodiments can be implemented by an application specific integrated circuit (ASIC), a general purpose computer or any other similar hardware device. In some embodiments, the software programs of the embodiments can be executed by a processor to implement the above steps or functions. Similarly, the software programs (including related data structures) of the embodiments can be stored in a computer readable recording medium, such as a RAM memory, a magnetic or optical drive or a floppy disk and the like. In addition, some steps or functions of the embodiments can be implemented by hardware, such as a circuit cooperating with a processor to perform the steps or functions.
[0189] The computer program product provided by the embodiments of the present application includes one or more computer programs / instructions, which, when executed by a processor, generate all or part of the processes or functions described in the embodiments of the present application. The computer can be a general purpose computer, a special purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center and the like integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD) or a semiconductor medium (for example, a solid state disk (SSD) and the like.
[0190] The computer program product of the present application can be a computer program embedded in a computer readable medium such as hard disk, CD, DVD, flash memory, etc. The computer readable medium can be directly readable by a computer or the computer can read the computer readable medium through a drive unit. The computer program product can also be a downloadable computer program transmitted via a network, e.g., the Internet, a local area network, etc. The downloadable computer program can be transmitted via a wired network or a wireless network, e.g., a satellite network, a cellular network, etc. The computer program product can also be a computer program product that is directly loadable into the internal memory of the computer.
[0191] The scope of the present application is defined by the appended claims rather than the description preceding them, therefore, all changes and modifications that come within the meaning and range of equivalents of the claims are to be embraced by the application. No feature of the application is considered critical unless the claims expressly state the indispensable feature. Furthermore, no embodiment of the application is considered essential unless the claims expressly state the essential nature of the embodiment. In addition, no single feature or combination of features should be considered critical unless expressly stated in the claims.
[0192] The above description is only specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be defined by the protection scope of the claims, and the above embodiments should be considered as exemplary and non-limiting.
Claims
1. A server scheduling method based on a large model, characterized in that: The method is applied to a preset server deployment scheme, in which the servers are numbered in a topology of a high-bandwidth domain (HBD) according to a physical connection order in a topology of a data center network (DCN); the servers are arranged in rows and columns, with the last server in a target row connected to the last server in the next row, and the first server in the next row connected to the first server in the next row, and so on, forming an interlaced "zigzag" connection pattern; the scheduling method includes: Determine a healthy server subgraph according to the deployment plan; the healthy server subgraph includes a healthy server set and a connection relationship between healthy servers in the healthy server set; Determining connected components based on the healthy server subgraph; wherein the connected components are used to represent independent parts obtained by segmenting the network connection based on the HBD topology structure; A scheduling scheme is determined according to the connectivity components and the number of servers included in each TP group.
2. The method according to claim 1, characterized in that Determining the healthy server subgraph according to the deployment scheme includes: Determining the healthy server set according to the deployment plan; determining, according to the healthy server set, a healthy edge set for characterizing a connection relationship between healthy servers in the healthy server set; A healthy server subgraph is determined according to the healthy server set and the healthy edge set.
3. The method according to claim 2, characterized in that According to the deployment scheme, determining the healthy server set includes: According to the deployment plan, obtain a topology diagram based on HBD; The healthy server set is determined according to the topology diagram and the faulty server set.
4. The method according to claim 1, wherein Determining the connected components according to the health server subgraph includes: Traverse each healthy server in the healthy server set, and perform the following operations on each healthy server in the healthy server set that has not been visited: Starting from the current healthy server, the healthy server connected to the current healthy server is searched in the healthy server subgraph to obtain the connected component.
5. The method according to claim 1, wherein Determining the connected components according to the health server subgraph includes: Connected components are determined based on a depth-first search algorithm and the healthy server subgraph.
6. The method according to claim 5, characterized in that Determining the connected components according to the depth-first search algorithm and the healthy server subgraph includes: Determine all neighbor servers of the target server in the healthy server subgraph according to the depth-first search algorithm; The connected components are determined according to the neighbor servers.
7. The method according to claim 1, characterized in that After determining the connected components according to the healthy server subgraph, the method further includes: Sorting the connected components to obtain sorted target connected components; Correspondingly, determining the scheduling scheme according to the connectivity component and the number of servers included in each TP group is specifically: determining the scheduling scheme according to the target connectivity component and the number of servers included in each TP group.
8. The method according to claim 1, characterized in that Determining a scheduling solution according to the connected components and the number of servers included in each TP group includes: Detecting a size relationship between each of the connected components and the number of servers included in each TP group; A scheduling plan is determined based on the size relationship.
9. The method according to claim 8, characterized in that Determining a scheduling scheme according to the size relationship includes: If the size of the connectivity component is greater than or equal to the number of servers in the TP group, extracting a number of servers equal to the number of servers in the TP group from the connectivity component, and treating the number of servers equal to the number of servers in the TP group extracted from the connectivity component as a TP group until all servers in the connectivity component are assigned to TP groups; A scheduling scheme is determined according to the TP groups obtained after the connected components are allocated.
10. The method according to any one of claims 1 to 9, characterized in that The method for determining the deployment solution includes: Determine first number information of a server in a topology structure of a DCN; Determine a parallel parameter according to the DCN topology; the parallel parameter is used to represent the number of sub-lines, and each sub-line corresponds to a group of servers; Determining second numbering information according to the first numbering information and the parallel parameter; the second numbering information is used to indicate a sequence number of the server in a topology structure of the HBD; The deployment scheme of the servers in the topology structure of the HBD is determined according to the second numbering information.
11. The method according to claim 10, characterized in that The determining, according to the first numbering information and the parallel parameter, the second numbering information includes: Determine a first ordered set of servers deployed in the topology structure of the DCN according to the total number of servers in the topology structure of the DCN and the first numbering information; Determine the length of each parallel sub-line according to the first ordered set and the parallel parameter; wherein the length of the parallel sub-line is used to represent the number of servers that each parallel sub-line should include; The second numbering information is determined according to the length of the parallel sub-lines.
12. The method according to claim 11, characterized in that The determining the second numbering information according to the length of the parallel sub-lines includes: determining a data range according to the parallel parameter; Determine an index of each of the parallel sub-lines according to the data range; Determining a position index within a target parallel sub-line according to the index of the parallel sub-line and the length of the parallel sub-line; the target parallel sub-line is determined according to the index of the parallel sub-line; Determining a second ordered set for characterizing servers deployed in a topology structure of an HBD according to the index of the parallel subline, the position index, and the parallel parameter; The second numbering information is determined according to the second ordered set.
13. An electronic device, characterized in that: The electronic device comprises: one or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method according to any one of claims 1 to 12.
14. A computer readable medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Separation decoupling resource allocation method and device, equipment and medium
CN119271406A
Apparatus and method to select transmission sources and destinations allowing all-to-all communication without link congestion in a network including plural topological structures
US20180113739A1