Resource allocation method and device and cloud server cluster architecture
By acquiring and reconfiguring the topological connection relationship in the cloud server cluster, the problem of idle GPU waste and cross-node communication performance loss when users request resources is solved, and efficient resource allocation and stable system operation are achieved.
Patent Information
- Application Number
- CN202510020923.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-09
AI Technical Summary
In an existing cloud server cluster, when the number of resources requested by users is smaller than the number of GPU resources in a single node, GPUs in some nodes are idle and performance losses are caused by cross-node communication.
By obtaining node information of each server node in the cloud server cluster, including the number of optical transceivers and topological connection relationships, the user's target topological connections are determined, and the topological connection relationships between server nodes are reconfigured to achieve resource allocation.
It reduces idle and waste of GPUs, simplifies cluster resource scheduling and management, improves resource allocation efficiency and system stability, reduces network overhead, and improves network stability and reliability.
Smart Images

Figure CN119966918A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of cloud computing technology, and are related to but not limited to a resource allocation method, device, and cloud server cluster architecture. Background Art
[0002] Currently, a cloud server cluster may be configured with multiple nodes, and the multiple nodes may be combined to provide resources to users. For example, a user may send a resource request to a cloud server cluster, and the cloud server cluster allocates resources based on the resource request.
[0003] However, in the related art, multiple graphics processors (GPUs) are set in each node, and the GPU resources in a single node are limited. When the number of resources requested by the user is less than the number of GPU resources in a single node, the GPUs in some nodes will be idle and wasted, and cross-node communication will also cause performance loss. Therefore, how to reduce resource waste in cloud service clusters is an urgent problem to be solved. Summary of the invention
[0004] In order to solve the problems existing in the related technologies, the embodiments of the present application provide a resource allocation method, device and cloud server cluster architecture.
[0005] In a first aspect, the present application provides a resource allocation method, the resource allocation method comprising:
[0006] In response to a resource allocation request initiated by a user, node information of each server node in the cloud server cluster is obtained; wherein a graphics processor is set in each server node, the resource allocation request includes at least the number of node requests, and the node information includes at least the number of optical transceivers in each server node and the topological connection relationship between the server nodes;
[0007] Determining a target topology connection corresponding to the user based on the node request quantity and the quantity of the optical transceivers;
[0008] Based on the target topological connection, the topological connection relationship between the server nodes is reconfigured to allocate resources to the user.
[0009] In a second aspect, an embodiment of the present application provides a resource allocation device, characterized in that the resource allocation device includes:
[0010] An acquisition module, configured to obtain node information of each server node in the cloud server cluster in response to a resource allocation request initiated by a user; wherein a graphics processor is provided in each server node, the resource allocation request includes at least the number of node requests, and the node information includes at least the number of optical transceivers in each server node and the topological connection relationship between the server nodes;
[0011] A determination module, configured to determine a target topology connection corresponding to the user based on the number of node requests and the number of optical transceivers;
[0012] A reconfiguration module is used to reconfigure the topological connection relationship between the server nodes based on the target topological connection to allocate resources to the user.
[0013] In a third aspect, an embodiment of the present application provides a cloud server cluster architecture, characterized in that the cloud server cluster architecture includes:
[0014] A cloud server cluster, wherein the cloud server cluster includes a plurality of server nodes, each of which is provided with a graphics processor;
[0015] An optical switch connected to a plurality of server nodes;
[0016] Multiple operating systems, each operating system is connected to a server node for receiving resource allocation requests initiated by users;
[0017] A management server, wherein the management network is connected to the multiple server nodes, and is used to respond to a resource allocation request initiated by a user and obtain node information of each server node in the cloud server cluster; wherein a graphics processor is set in each server node, the resource allocation request at least includes the number of node requests, and the node information at least includes the number of optical transceivers in each server node and the topological connection relationship between the server nodes; based on the number of node requests and the number of optical transceivers, determining the target topological connection corresponding to the user; based on the target topological connection, reconfiguring the topological connection relationship between the server nodes to allocate resources to the user.
[0018] In the embodiment of the present application, a GPU is set in each server node, and there is no need to determine whether each server node has unused idle resources, nor is there a need to schedule idle resources in multiple server nodes, which reduces the problem of idle and wasted GPUs in some nodes when the number of resources requested by the user is less than the GPU resources in a single node, can simplify the computational complexity of cluster resource scheduling and management, and improve resource allocation efficiency; at the same time, the GPU resources in each server node are independent, which can reduce interference between different tasks, thereby improving the overall stability of the system; and after determining the target topological connection corresponding to the user, the embodiment of the present application reconfigures the topological connection relationship between the server nodes, and can use existing connections to reduce the number of connections that need to be reconfigured, make full use of existing resources, simplify network management tasks, avoid frequent addition or deletion of connections, thereby reducing network overhead, and improving network stability and reliability.
[0019] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is an optional architecture diagram of a cloud server cluster architecture provided in an embodiment of the present application;
[0021] Figure 2 This is an optional process diagram of the resource allocation method provided in the embodiment of the present application. Figure 1 ;
[0022] Figure 3 This is an optional process diagram of the resource allocation method provided in the embodiment of the present application. Figure 2 ;
[0023] Figure 4 is a schematic diagram of solving a target topological relationship provided in an embodiment of the present application;
[0024] Figure 5 It is a structural diagram of the cloud server cluster architecture provided in an embodiment of the present application;
[0025] Figure 6 It is a structural diagram of a one-to-one connection of optical transceivers provided in an embodiment of the present application;
[0026] Figure 7 is a flowchart of a resource allocation process provided in an embodiment of the present application;
[0027] Figure 8 It is a structural diagram of a resource allocation device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0029] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as those commonly understood by those skilled in the art of the technical field of the embodiments of the present application. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0030] GPU cloud computing is a rapidly growing market, and workloads such as artificial intelligence, machine learning, and graphics rendering require GPU acceleration.
[0031] At present, in order to improve the communication bandwidth between GPUs, the main architecture of GPU cloud servers in the existing technology is to use NVLink inside the server to interconnect GPUs at high speed, and use Infiniband and RoceV2 and other technologies to interconnect different servers. When the GPU cloud server allocates resources to users, it can use preemptive resource allocation to suspend or check lower priority GPU jobs to free up resources for higher priority tasks. GPU job priorities may be differentiated by user type. You can also use elastic GPU pools, and use virtual GPU (such as vGPU) technology to create elastic GPU resource pools, which can dynamically allocate resources to users according to workload requirements. In addition, multiple tasks can also share the use of GPU technology to make full use of GPU resources.
[0032] In the existing GPU cloud server cluster architecture, NVLink is used inside the server to provide high-speed bandwidth between GPUs. However, even if Infiniband, RoceV2 and other technologies are used, the bandwidth obtained for communication between different server nodes is still much smaller than that inside the server. To ensure communication efficiency, cloud server providers generally rent GPUs in the same server to users when the number of GPUs rented by users is less than or equal to the number of GPUs on a single server.
[0033] However, the number of GPUs on a single server is limited, generally up to 8, and scattered allocation of GPU resources may cause waste. For example, for a 4-GPU server, when multiple users rent a 3-GPU node, one GPU will be idle and unused, resulting in resource waste. Even when multiple users rent a 2-GPU node, it may still cause waste. For example, on two 4-GPU servers, when users 1, 2, and 3 rent the computing power of two GPUs respectively, among which user 1 and user 2 rent 4 GPUs of server 1, and user 3 rents 2 GPUs of server 2, user 2 releases resources after a period of time. At this time, if user 4 wants to rent 4-GPU computing power, although there are 4 remaining GPUs, it is impossible to allocate resources for it on a single server. Or when the number of GPUs is greater than the number that a single server can carry (generally 8), cross-node communication will also cause certain performance losses and will also be affected by the fragmentation of user resource allocation. The method of sharing GPUs for multiple tasks is mainly aimed at the resource utilization on a single GPU, which will be affected by resource competition and has high requirements for performance interference analysis technology.
[0034] In order to solve the problems existing in the related technology, an embodiment of the present application provides a resource allocation method, which responds to a resource allocation request initiated by a user and obtains the node information of each server node in a cloud server cluster; wherein a graphics processor is set in each server node, the resource allocation request includes at least the number of node requests, the node information includes at least the number of optical transceivers in each server node and the topological connection relationship between the server nodes, based on the number of node requests and the number of optical transceivers, the target topological connection corresponding to the user is determined, and based on the target topological connection, the topological connection relationship between the server nodes is reconfigured to allocate resources to the user.
[0035] In this way, in the embodiment of the present application, a GPU is set in each server node, and there is no need to determine whether each server node has unused idle resources, nor is there a need to schedule idle resources in multiple server nodes, which reduces the problem of idle and wasted GPUs in some nodes when the number of resources requested by the user is less than the number of GPU resources in a single node, can simplify the computational complexity of cluster resource scheduling and management, and improve resource allocation efficiency; at the same time, the GPU resources in each server node are independent, which can reduce interference between different tasks, thereby improving the overall stability of the system; and after determining the target topological connection corresponding to the user, the embodiment of the present application reconfigures the topological connection relationship between the server nodes, and can use existing connections to reduce the number of connections that need to be reconfigured, make full use of existing resources, simplify network management tasks, avoid frequent addition or deletion of connections, thereby reducing network overhead, and improving network stability and reliability.
[0036] The embodiment of the present application flexibly utilizes the GPU to solve the problem of GPU idleness and waste caused by fragmented resource allocation. The resource allocation method in the present application can be used for small and medium-sized machine learning tasks and some high-performance computing tasks with relatively fixed communication modes.
[0037] Figure 1 is an optional architecture diagram of a cloud server cluster architecture provided in an embodiment of the present application, such as Figure 1 As shown, the cloud server cluster architecture includes a cloud server cluster 101, an optical switch 102, an operating system 103 and a management server (not shown in the figure).
[0038] The cloud server cluster 101 includes a plurality of server nodes 1011, each of which is provided with a graphics processor 1011-1 and an optical transceiver 1011-2, the number of optical transceivers in each server node is the same, and the optical transceivers in different server nodes are connected one-to-one; an optical switch 102 is connected to the plurality of server nodes 1011, and the optical switch may be an optical circuit optical switch or an optoelectronic cross-connection panel; a plurality of operating systems 103, each of which is connected to a server node 1011, and is used to receive a resource allocation request initiated by a user; a management server, a management network connected to the plurality of server nodes, is used to respond to a resource allocation request initiated by a user and obtain node information of each server node in the cloud server cluster; a graphics processor is provided in each server node, the resource allocation request includes at least the number of node requests, and the node information includes at least the number of optical transceivers in each server node and the topological connection relationship between the server nodes; based on the number of node requests and the number of optical transceivers, the target topological connection corresponding to the user is determined; based on the target topological connection, the topological connection relationship between the server nodes is reconfigured to allocate resources to the user.
[0039] Here, the operating system 103 may be software that provides a basic platform and services for the server node 1011. The user interacts with the operating system 103 through a cluster controller or a node controller to apply for and manage server nodes in the cluster.
[0040] In some embodiments, the optical switch 102 cannot forward data like a traditional electrical switch, and each interface of the optical transceiver can only be connected to the optical transceiver interface of another server node through optical switching, so that the optical transceiver interface of each server node in the cluster can only send data to the optical transceiver interface of another server node. That is, the optical switch 102 provides a one-to-one connection.
[0041] Here, each server node may also be provided with hardware devices such as a central processing unit (CPU), a storage unit, and a network card. The network card is used for the server node to connect to the management server.
[0042] The storage unit includes a solid-state memory, a hard disk drive, an optical disk drive, etc. The storage unit optionally includes one or more storage devices that are physically far away from the processor. The storage unit includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The storage unit described in the embodiments of the present application is intended to include any suitable type of memory.
[0043] The central processing unit can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., among which the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0044] Here, the management server can be implemented as any terminal with data processing function, such as a laptop computer, a tablet computer, a desktop computer, a mobile device (e.g., a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device), an intelligent robot, an intelligent home appliance, and an intelligent vehicle-mounted device; in another implementation, the management server provided in the embodiment of the present application can also be implemented as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected by wired or wireless communication, which is not limited in the embodiment of the present application.
[0045] See also Figure 2 , Figure 2 This is an optional process diagram of the resource allocation method provided in the embodiment of the present application. Figure 1 , the following will combine Figure 2 The steps shown are explained, and it should be noted that Figure 2The resource allocation method in the embodiment is described by taking the management server as an execution subject as an example, and the resource allocation method can be implemented through steps S201 to S203.
[0046] Step S201, in response to a resource allocation request initiated by a user, obtaining node information of each server node in the cloud server cluster; wherein a graphics processor is set in each server node, the resource allocation request includes at least the number of node requests, and the node information includes at least the number of optical transceivers in each server node and the topological connection relationship between the server nodes.
[0047] In some embodiments, when performing machine learning tasks or computing tasks, users need to apply for GPU resources to accelerate model training, or improve computing efficiency through data parallelism or model parallelism. At this time, the user can submit an application to the management server based on a job management system (such as Linux resource management tool (Slurm, Simple Linux Utility for Resource Management), job scheduling system (PBS, Portable Batch System), etc.), that is, issue a resource allocation request, which can include the number of GPUs required by the user, that is, the number of node requests.
[0048] In the embodiment of the present application, a GPU is set in each server node, and there is no need to deal with resource allocation and conflict issues between multiple GPUs in a server node, which reduces the complexity of resource allocation. Users can be more efficiently allocated to any available GPU node, thereby improving overall resource utilization. A single GPU configuration can reduce the number of idle GPUs, thereby reducing overall power consumption.
[0049] In an embodiment of the present application, when a resource allocation request from a user is received, the current node information of the cluster, that is, the usage status, will be obtained, such as the number of optical transceivers in each server node and the topological connection relationship between the server nodes. The topological connection relationship between the server nodes may refer to the connection relationship between the nodes in the current cluster.
[0050] In the embodiment of the present application, a server node ID is assigned to each server node, and the interface between node k and the i-th optical transceiver of this node is called interface (k, i), where 1≤i≤d, and d is the number of interfaces of the optical transceiver. Here, for each interface (k, i), the interface (m, j) of the corresponding optical transceiver connected to it through the optical switch is recorded, as well as the resource allocation ID: id to which this topological connection belongs. When the topological connection does not belong to a resource in use or in allocation, the resource allocation ID is set to -1.
[0051] Here, the information recorded by the management server on the real-time topological connection relationship can be given in the form of a triple [(k,i), (m,j), id], and all interfaces corresponding to the topological connection form a triple queue. When the cluster is initialized, it is assumed that all interfaces have no connected interfaces. For the interface (k, j) that is not connected to other interfaces, the management network record is [(k,j), (k,j), id].
[0052] In some embodiments, the node information may also be idle server nodes among current server nodes, and resources are allocated to users in the idle server nodes.
[0053] Step S202: Determine a target topology connection corresponding to the user based on the node request quantity and the quantity of the optical transceivers.
[0054] In some embodiments, when receiving a resource allocation application from a user, it is first determined whether the user inputs a custom topology connection (i.e., a topology connection designed by the user himself). If not, the target topology connection corresponding to the user is determined according to the number of node requests and the number of optical transceivers. The number of server nodes in the target topology connection is the same as the number of node requests.
[0055] Here, the target topology connection refers to the connection relationship between the number of server nodes requested by the node, for example, full connection, grid connection, tree connection or ring connection.
[0056] In some embodiments, when the number of node requests is less than or equal to the number of optical transceivers in each server node + 1, the target topology connection can be a full connection of the assigned server nodes, and the assigned server nodes are the number of idle server nodes assigned by the node requests; when the number of node requests is greater than the number of optical transceivers in each server node + 1, the target topology connection can be a topology connection mode calculated by any feasible topology connection algorithm for the number of idle server nodes assigned by the node requests.
[0057] Step S203: Based on the target topological connection, reconfigure the topological connection relationship between the server nodes to allocate resources to the user.
[0058] In the embodiment of the present application, after obtaining the target topological connection, when the server nodes in the cluster release resources (i.e., no longer use certain server nodes), the topological connection of the cluster is not reset, so the existing connections can be used to reduce the number of interfaces that need to be reconfigured. In this way, it is only necessary to find the subgraph that is most similar to the target topological connection in the topological connection relationship of the entire cluster, modify the topological connection based on the subgraph, modify it to be the same as the target topological connection, and allocate several server nodes after the modified topological connection to the user, so as to realize resource allocation to the user and minimize the number of connections that need to be modified for reconfiguration.
[0059] In the embodiment of the present application, a GPU is set in each server node, and there is no need to determine whether each server node has unused idle resources, nor is there a need to schedule idle resources in multiple server nodes, which reduces the problem of idle and wasted GPUs in some nodes when the number of resources requested by the user is less than the GPU resources in a single node, can simplify the computational complexity of cluster resource scheduling and management, and improve resource allocation efficiency; at the same time, the GPU resources in each server node are independent, which can reduce interference between different tasks, thereby improving the overall stability of the system; and after determining the target topological connection corresponding to the user, the embodiment of the present application reconfigures the topological connection relationship between the server nodes, and can use existing connections to reduce the number of connections that need to be reconfigured, make full use of existing resources, simplify network management tasks, avoid frequent addition or deletion of connections, thereby reducing network overhead, and improving network stability and reliability.
[0060] Figure 3 This is an optional process diagram of the resource allocation method provided in the embodiment of the present application. Figure 2 ,like Figure 3 As shown, step S203 can be implemented through steps S301 to S304:
[0061] Step S301: Determine currently unoccupied idle server nodes in the cloud server cluster.
[0062] In an embodiment of the present application, the management server can directly obtain idle server nodes that are currently not occupied in the cloud server cluster, and the idle server nodes refer to server nodes that are not allocated to users.
[0063] Step S302: In response to the number of the idle server nodes being greater than the node request number, based on the topological connection relationship between the server nodes, determining the idle topological connection relationship between the idle server nodes.
[0064] In some embodiments, when the number of idle server nodes in the cluster is less than the number of node requests, it means that the current idle server nodes are insufficient to realize the user's calculation, so resources may not be allocated to the user, and the result of insufficient resources is fed back to the user. When the number of idle server nodes is equal to the number of node requests, it means that the current idle server nodes can realize the user's calculation, the topological connection is calculated based on the algorithm, and the topological connection relationship of the current idle nodes is modified, and the idle nodes after the modified topological connection are allocated to the user; when the number of idle server nodes is greater than the number of node requests, the idle topological connection relationship between the current idle server nodes is determined at this time, that is, the connection relationship between each idle server node and the connection relationship between each idle server node and other server nodes.
[0065] Step S303: In the idle topology connection relationship, based on the target topology connection, determine a sub-topology connection relationship that meets the reconfiguration condition.
[0066] In an embodiment of the present application, the idle topology connection relationship refers to the connection relationship between each optical receiver in the idle server node in the cluster and other nodes, which can constitute an idle topology connection relationship graph, and the target topology connection can also be a target topology connection graph. The sub-topology connection relationship that meets the reconfiguration conditions can be a sub-graph in the idle topology connection relationship graph that is closest to the target topology connection graph. For example, the target topology connection and the sub-topology connection relationship are both four server nodes, and 80% of the connection relationships between the four server nodes (which can be the edges connecting two server nodes) are the same, and 20% of the connection relationships in the sub-topology connection relationships are different.
[0067] Step S304: Based on the target topology connection, reconfigure the sub-topology connection relationship to allocate resources to the user.
[0068] Here, after obtaining the sub-topology connection relationship, the connection relationship in the sub-topology connection relationship that is different from the target topology connection can be reconfigured, that is, modified, and the connection between the server nodes in the sub-topology connection relationship can be modified to be the same as the target topology connection, and the modified server nodes with the connection relationship can be allocated to the user to allocate resources to the user.
[0069] In the embodiment of the present application, the connection relationship between the current server nodes is utilized to find a subgraph that is closest to the target topology connection graph. Reconfiguration based on the subgraph can minimize the number of connections that need to be modified for reconfiguration, reduce the waiting time for users to wait for reconfiguration, and improve allocation efficiency.
[0070] In some embodiments, step S303 may be implemented by steps S3031 to S3033:
[0071] Step S3031: Based on the target topology connection, determine at least one connection relationship in the idle topology connection relationship, the number of nodes being the same as the requested number of nodes.
[0072] In some embodiments, after the target topology connection is determined, a connection relationship in which the number of multiple nodes is the same as the number of node requests can be determined in the idle topology connection relationship. For example, if the number of node requests is 4, a connection relationship in which the number of multiple server nodes is 4 can be determined in the idle topology connection relationship, that is, a subgraph in which the number of multiple server nodes is 4.
[0073] Step S3032: Determine the degree of difference between each connection relationship and the target topological connection.
[0074] In some embodiments, the degree of difference is used to characterize the similarity between each connection relationship (i.e., subgraph) and the target topological connection, that is, the degree of overlap between the connection edges between server nodes in each connection relationship and the connection edges between server nodes in the target topological connection. The higher the degree of overlap, the higher the similarity between the connection relationship and the target topological connection, and the fewer connections need to be modified when assigning the connection relationship to the user.
[0075] Step S3033: Among the at least one connection relationship, determine the connection relationship with the smallest degree of difference from the target topological connection as the sub-topological connection relationship.
[0076] In some embodiments, the connection relationship with the smallest degree of difference from the target topological connection among at least one connection relationship may be determined as a sub-topological connection relationship.
[0077] The embodiment of the present application determines the connection relationship with the least difference from the target topology connection based on the idle topology connection relationship as the topology connection allocated to the user. In this way, the number of connections that need to be modified before allocation to the user can be minimized, thereby improving allocation efficiency.
[0078] In some embodiments, step S3032 may be implemented by steps S1 to S2:
[0079] Step S1, determine the first number of connection edges between the server node and the external node in each connection relationship, and the second number of different connection edges between each connection relationship and the target topology connection; wherein the external node is a server node in the cloud server cluster other than the server node in each connection relationship.
[0080] In some embodiments, Figure 4 is a schematic diagram of solving the target topological relationship provided in the embodiment of the present application, such as Figure 4As shown in Figure (a), the target topological relationship G(a * ,E * ), (b) is an idle topological connection relationship, and the dotted box is a connection relationship G(a, E(a)) in the idle topological connection relationship, where line 1 is the connection line between the server node and the external node in the connection relationship, based on which the first quantity d(a) can be determined, line 2 is the edge in the connection relationship that is more than the target topological relationship, and line 3 is the edge in the connection relationship that is less than the target topological relationship. At this time, the second quantity d(G(a * ,E * ),G(a,E(a))) is the sum of the number of lines 2 and 3.
[0081] Step S2: Determine the degree of difference between each connection relationship and the target topological connection based on the first number, the second number and a preset weight.
[0082] The preset weight w is a weight value greater than 0, which can be selected based on actual experience. The degree of difference can be expressed by formula (1):
[0083] J(a)=d(G(a * ,E * ),G(a,E(a)))+d(a)*w (1);
[0084] Among them, G(a * ,E * ) is the target topological relationship, and G(a,E(a)) is the connection relationship.
[0085] In some embodiments, step S3031 may be implemented by steps S11 to S14:
[0086] Step S11: setting an initial node set, wherein the initial node set is an empty set.
[0087] In an embodiment of the present application, an initial node set S can be set first, and S is initially an empty set. In an embodiment of the present application, server nodes can be added to S until the number of server nodes in the initial node set is equal to the number of node requests, and a connection relationship is obtained.
[0088] Step S12: adding the server nodes that meet the connection conditions in the idle topological connection relationship to the initial node set to obtain an updated node set.
[0089] In some embodiments, among the server nodes in the target topology connection, the minimum number of connections of the server nodes is d *, the set of idle server nodes in the cluster is P, and there are at most |P|*d nodes in P\S that may be connected to the nodes in S (where |P| is the number of idle server nodes in P, and d is the number of optical transceivers). Randomly select nodes whose number of connections to nodes in S is greater than or equal to d * If the node p cannot be found, then select the node with the largest number of connections with the nodes in S in P\S as point p, add p to the initial node set S, and obtain the updated node set.
[0090] Step S13, executing the update process of the update node set again until the number of target nodes in the update node set is equal to the number of node requests, so as to determine one of the connection relationships based on the target nodes in the update node set.
[0091] In an embodiment of the present application, the update process of updating the node set in step S12 is repeated until the number of target nodes in the updated node set is equal to the number of node requests. At this time, the connection between the server nodes and the server nodes in the updated node set can constitute a connection relationship.
[0092] Step S14: execute the connection relationship determination process again to obtain at least one connection relationship.
[0093] In the embodiment of the present application, by executing step S12 and step S13 multiple times, at least one connection relationship can be obtained. Here, the server nodes in at least one connection relationship can be partially different, that is, there can be server nodes 5 in different connection relationships.
[0094] In some embodiments, step S12 may be implemented by steps S121 to S123:
[0095] Step S121: Obtain the minimum number of connections of server nodes in the target topology connection.
[0096] In some embodiments, the minimum number of connections of server nodes in the target topology connection refers to the minimum number of connections of server nodes in the target topology connection. For example, there are four server nodes in the target topology connection, 2 server nodes are connected to 3 server nodes respectively, one server node is connected to 2 server nodes, and one server node is connected to 1 server node. At this time, the minimum number of connections is 1.
[0097] Step S122: Based on the minimum number of connections, a server node among the idle server nodes whose number of connections with the server nodes in the initial node set is greater than the minimum number of connections is determined as a target node; or, in the case that there is no server node among the idle server nodes whose number of connections is greater than the minimum number of connections, a server node among the idle server nodes whose number of connections with the server nodes in the initial node set is the largest, is determined as the target node.
[0098] In an embodiment of the present application, after determining the minimum number of connections, server nodes among the idle server nodes whose number of connections with server nodes in the initial node set is greater than the minimum number of connections can be determined as target nodes that can be added to the initial node set.
[0099] In some embodiments, if there is no server node with a connection number greater than the minimum connection number among the idle server nodes, the server node with the largest number of connections with the server nodes in the initial node set among the idle server nodes is determined as the target node that can be added to the initial node set.
[0100] Step S123: Add the target node to the initial node set to obtain the updated node set.
[0101] In the embodiment of the present application, the obtained target node may be added to the initial node set to obtain an updated node set.
[0102] The embodiment of the present application obtains at least one connection relationship through a random algorithm, and can quickly determine at least one connection relationship from idle topological connection relationships, thereby improving reconfiguration efficiency.
[0103] In some embodiments, step S202 may be implemented by steps S2021 to S2023:
[0104] Step S2021: In response to not receiving the custom topology connection, determining a comparison result between the node request quantity and a preset quantity; the preset quantity is obtained based on the quantity of the optical transceivers.
[0105] In some embodiments, users can define topological connections by themselves. The topological connections are given by files uploaded by users (such as in the form of triples as described above), and the data format is the same as the topological connection information format of the server nodes in the cluster recorded by the management server. In the case where the user provides a custom topological connection, the custom topological connection is determined as the target topological connection.
[0106] In some embodiments, the preset number is equal to the number of optical transceivers in each server node + 1. When no custom topology connection is received, a comparison result between the node request number and the preset number is determined, and the comparison result includes two results: less than or equal to and greater than.
[0107] Step S2022: In response to the comparison result indicating that the number of node requests is less than or equal to the preset number, a fully connected topology is used as the target topology connection.
[0108] In some embodiments, if the comparison result indicates that the number of node requests is less than or equal to a preset number, then the fully connected topology may be used as the target topology connection.
[0109] Step S2023: In response to the comparison result indicating that the number of node requests is greater than the preset number, generating the target topology connection based on a preset algorithm.
[0110] In some embodiments, if the comparison result indicates that the number of node requests is greater than a preset number, then any topological connection method calculated by a topological connection algorithm applicable to distributed machine learning on an optical network can be used as the target topological connection.
[0111] The following is an explanation of an exemplary application of the embodiments of the present application in a practical application scenario.
[0112] Machine learning training tasks, such as deep neural network training, require high-bandwidth, low-latency interconnection between servers and GPU accelerators. When GPU servers are provided on cloud servers, if traditional electrical networks are used, the network transmission bandwidth between server nodes is much larger than the GPU interconnection transmission bandwidth within the server nodes. Compared with electrical networks, optical networks using wavelength division multiplexing (WDM) can provide orders of magnitude improvements in bandwidth and latency.
[0113] This application applies the all-optical network to the cloud computing cluster, and provides small and medium-sized GPU server customization. The server can be pre-installed with a collective communication program suitable for the network architecture in this application. The collective communication program provides an application programming interface (API) that complies with the message passing interface (MPI) standard, which can be directly called by the existing mainstream machine learning framework. When users use the mainstream machine learning framework, they will automatically use the collective communication program customized by this application without additional steps.
[0114] This application uses a one-time reconfiguration of the optical network for each resource allocation for cloud servers that provide GPU hardware. It can be used for small and medium-scale machine learning or GPU high-performance computing tasks with relatively fixed communication modes. It is suitable for situations with a large number of fragmented users, as well as medium-scale tasks where the number of GPUs is greater than the number that a single server can carry (generally 8).
[0115] The use of an all-optical network will make the communication mode relatively fixed, and the communication of distributed machine learning is suitable for such a communication network. This is because the communication modes of distributed machine learning are mostly relatively fixed. For example, when performing data parallelism, you only need to use All Reduce to calculate and synchronize the average value of model parameters (that is, to calculate the average value of model parameters on all nodes with the help of communication).
[0116] Figure 5 It is a structural diagram of a cloud server cluster architecture provided by an embodiment of the present application, which includes a management network (not shown in the figure), multiple server nodes 501 connected to the management network, an optical switch 502 connected to the multiple server nodes 501, and multiple operating systems 503, each operating system 503 corresponds to a server node 501, and the operating system 503 is software that provides a basic platform and services for the server node 501. Users interact with these operating systems through a cluster controller or a node controller to apply for and manage server nodes.
[0117] In some embodiments, the optical switch 502 may be an optical circuit switch (OCS) or an optical patch panel. The optical switch 502 cannot forward data like a traditional electrical switch, and each interface can only be connected to another fixed interface, so that the interface of the optical transceiver of each server node in the cluster can only send data to the interface of the optical transceiver of another server node. That is, the optical switch 502 provides a one-to-one connection. Figure 6 is a schematic diagram of a one-to-one connection structure of optical transceivers provided in an embodiment of the present application, such as Figure 6 As shown, the interface of the optical transceiver 601 can only send data to the interface of the optical transceiver 601 of another server node.
[0118] The embodiment of the present application directly transmits optical signals through the optical switch 502, which has no "optical-electrical-optical" conversion compared to the electrical switch, and has lower latency; it can provide higher bandwidth support compared to the electrical network; because the optical switch can achieve one-to-one connection, the communication path is fixed, which avoids packet queuing and congestion, and can reduce invalid bandwidth.
[0119] In order to cope with the fragmented resource redistribution of users, the optical switch in the embodiment of the present application has a reconfiguration function, that is, it can suspend communication and reconfigure the connection object of the corresponding interface of the above one-to-one connection. Both the OCS and the optoelectronic cross-connection panel support the reconfiguration function. Here, the optoelectronic cross-connection panel takes 10 seconds to several minutes according to the number of interfaces that need to be reconfigured. The present application uses a non-blocking optoelectronic cross-connection panel, such as the Telescent reconfigurable optoelectronic cross-connection panel. The optoelectronic cross-connection panel has the advantages of lower optical loss, cheaper price, and more interfaces. For example, the Telescent reconfigurable optoelectronic cross-connection panel has 1008 duplex interfaces.
[0120] In the embodiment of the present application, reconfiguration only needs to be performed when the user applies for a cloud server, allowing the user to wait for a few minutes when applying for a machine, and this is performed synchronously with other deployment work of the machine.
[0121] For an all-optical network, the embodiments of the present application also require optical fibers, optical transceivers, and other equipment with low optical loss and sufficient bandwidth. Optical transceivers can efficiently convert between light and electricity, connect GPUs and optical networks, and provide terabit per second (Tbps) level transmission bandwidth.
[0122] In an embodiment of the present application, unlike a common GPU server cluster in which a single node contains multiple GPUs, the present application uses a single GPU in a single server node. Multiple optical transceivers (generally 4-8) are used in each server to connect to the GPU, and the other end is connected to the optoelectronic cross-connect panel through optical fiber. In the present application, the number of optical transceivers in each server node is the same, and a full-duplex network connection mode is adopted. For each server node in the cluster, the configuration of the central processing unit (CPU, Central Processing Unit), storage, etc. is the same as that of a normal server cluster node, and a network card is used for the server node to connect to the Internet and the cluster management network. In addition, the hardware of each server, including the GPU, has the same specifications.
[0123] Figure 7 is a flow chart of a resource allocation process provided in an embodiment of the present application, such as Figure 7 As shown, the resource allocation process when a user applies for a server node includes steps S701 to S707 to implement:
[0124] In an embodiment of the present application, a server node ID is assigned to each server node, and the i-th transceiver interface between node k and this node is called interface (k, i), where 1≤i≤d, and d is the number of interfaces of the optical transceiver.
[0125] The management network can record the following information: the total number of server nodes N; the real-time resource allocation ID (i.e., if the server node is allocated to the 10th user, the resource allocation ID of the server node is 10) and the resource allocation status, including in use and released (the resource allocation ID is -1 at this time); real-time topology connection information, that is, for each interface (k, i), record the corresponding interface (m, j) connected to it through the optoelectronic cross-connection panel, and the resource allocation ID to which this topology connection belongs: id. When the topology connection does not belong to a resource in use or being allocated, the resource allocation ID is set to -1.
[0126] The information recorded by the management network on real-time topology connection information is given in the form of a triplet [(k,i),(m,j),id], and all interfaces corresponding to the topology connections form a triplet queue.
[0127] When the cloud server is initialized, it is assumed that all interfaces are not connected. For an interface (k, j) that is not connected to other interfaces, the management network record is [(k, j), (k, j), id].
[0128] S701: Whether the user inputs a custom topology.
[0129] In an embodiment of the present application, when a new user server allocation request is received, the user's resource allocation ID is recorded, and the number of server GPUs applied for by the user is set to n (when applying for server resources, n is required to be less than or equal to the number of remaining GPUs in the current cluster).
[0130] When receiving a resource allocation application from a user, it is first determined whether the user inputs a custom topology, and if so, step S702 is executed, and if not, step S703 is executed.
[0131] Here, users can define the optical network topology by themselves. The network topology is given by the file uploaded by the user. The data format is the same as the topology connection information format recorded by the cluster management system. The node ID is 1-n, representing the number of the node applied for, and n is the number of server GPUs applied by the user. At this time, the user needs to provide the corresponding collective communication algorithm.
[0132] S702: The user uploads a topology file.
[0133] S703: Whether the number of GPUs applied for is greater than the number of optical transceiver interfaces of a single node + 1.
[0134] When the user does not customize the topology, the topology connection mode to be used is determined based on whether the number of GPUs requested by the user is greater than the number of optical transceiver interfaces of a single node + 1.
[0135] If yes, execute step S704; if no, execute step S705.
[0136] S704. Use full connection as topology.
[0137] Here, when n≤d+1, the allocated server nodes are fully connected, and at this time, there is no need to modify the collective communication algorithm of the operating system.
[0138] S705: Generate topology using an existing algorithm.
[0139] When n>d+1, any topological connection algorithm applicable to distributed machine learning on optical networks is used to design topological connections for the nodes of the entire resource allocation. At this time, the collective communication algorithm corresponding to the topological connection algorithm is used.
[0140] S706. Output topology connection.
[0141] S707 , reconfigure the connection network based on the topological connection.
[0142] In some embodiments, when multiple users apply for resources at the same time, the algorithm should be run in a blocking order to calculate the topological information of resource allocation. Since this application is aimed at small and medium-sized machine learning tasks and some high-performance computing tasks with relatively fixed communication modes, the scale is small and the time consumed by using the algorithm is negligible.
[0143] For the reconnection process during resource allocation, since a non-blocking optoelectronic cross-connect panel is used and different resource allocations have different nodes, the reconnection process can be performed synchronously.
[0144] In some embodiments, when a user applies to release a resource, the corresponding ID status of the released resource is set to released, the status of all nodes in the released resource is set to idle, and the resource IDs of all connections in the released resource are set to -1.
[0145] After the network topology connection design is obtained, the topology connection is not reset when the resources are released, so the existing connection can be used to reduce the number of interfaces that need to be reconfigured. In this way, it is only necessary to find similar subgraphs in the topology connection of the entire cluster to minimize the number of connections that need to be modified for reconfiguration.
[0146] The embodiment of the present application can also isolate the subgraph from the remaining graph as much as possible. For example, when a fully connected subgraph with 3 nodes needs to be found, there are both a fully connected subgraph with 3 nodes and a fully connected subgraph with 4 nodes. If a subgraph consisting of 3 nodes in the fully connected subgraph with 4 nodes is used, this subgraph has more connections with the outside, which will affect the efficiency of subsequent topology reconfiguration. Therefore, when searching for a subgraph, the number of connections between the subgraph and the remaining nodes should be minimized as much as possible.
[0147] The embodiment of the present application defines the problem of how to determine which server nodes to allocate to a user as the following optimization problem:
[0148] For a set A, |A| is the number of elements in the set. For a user's request for n GPU server resources, this application defines a set consisting of n nodes as An. If all nodes in a represent an idle server node in the cluster, then a∈A′ is recorded. n .
[0149] For a∈A′ n , define G(a,E) as a graph, where E is the set of edges connecting certain points in a.
[0150] For two graphs G(a1,E1) and G(a2,E2), the distance between them is defined as d(G(a1,E1), G(a2,E2)) = |E′2 / E1| + |E1 / E′2|, where after mapping the i-th node in a2 to the i-th node in a1 (1≤i≤n), the edges in E2 are also mapped accordingly to obtain E′2, |E′2 / E1| refers to the number of edges in E′2 except E1, and |E1 / E′2| refers to the number of edges in E1 except E′2.
[0151] For the idle server node set a∈A′ n , define d(a) as the number of edges formed by points in the cluster other than the midpoint of a and the points in a, and let the edge set corresponding to a be E(a) (which can be obtained from the cluster management network).
[0152] Suppose the topology graph (i.e. target topology connection) that needs to be searched is G(a * ,E * ), that is, the server nodes finally allocated and the corresponding topological connections, the problem solved by this application is abstracted into formula (2):
[0153]
[0154] Among them, J(a) is the set of points in the subgraph with the smallest difference between the number of user applications and the topological connection in the current topological connection. w is a weight value greater than 0, which can be selected according to actual experience.
[0155] Please continue to refer to Figure 4 , (a) is the target topology graph to be obtained, (b) is the current topology connection graph in the cluster, and the dotted box is the candidate subgraph, where line 1 is the connection line between the candidate subgraph and the nodes outside the candidate subgraph, line 2 is the line where the candidate subgraph has more nodes than the target topology graph, and line 3 is the line where the candidate subgraph has less nodes than the target topology graph. At this time, d(G(a * ,E *),G(a,E(a))) is the sum of the number of lines 2 and 3, which is 3, and d(a) is the number of line 1, which is 4.
[0156] Solving formula (2) is an integer programming problem, which is an NP-hard problem. If the total number of nodes N or the number of search nodes n is large, it is difficult to solve the problem efficiently even with advanced integer programming solvers. This application does not require the solution to be optimal, but only needs to find a relatively good solution. Therefore, this application uses the following random algorithm:
[0157] First, suppose there is a set S, which is initially an empty set. This application will add nodes to S to construct a feasible solution.
[0158] Secondly, let G(a * ,E * ) in the nodes, the minimum number of connections between nodes is d * , let the idle node set in the cluster be P, the number of nodes in P\S that may be connected to the nodes in S is at most |P|d, and randomly select nodes with a number of connections greater than or equal to d nodes in S * If the node p cannot be found, then select the node with the largest number of connections with the nodes in S in P\S as point p. Add p to the set S and repeat the above steps until the number of nodes in S is n, and a feasible solution S is obtained.
[0159] Repeat the above method of obtaining feasible solutions k times to obtain k solutions, and take the solution that minimizes the objective function J(a) as the final solution. Here, the integer k is selected based on experience, because each solution search has a time complexity of only O(n*d), so consider taking a larger k to obtain better results.
[0160] The embodiment of the present application uses an all-optical network to build a cloud cluster system that is easy to allocate resources, making full use of the GPU on the cluster; the process and algorithm of reconfiguring the network topology reduce the average time required for reconfiguration.
[0161] The embodiment of the present application uses an all-optical network connection method to provide high-speed and flexible communication on the GPU cloud server, overcoming the resource waste caused by fragmented user resource allocation in the existing GPU cloud server.
[0162] The embodiment of the present application uses a reconfigurable all-optical network, which efficiently provides a mechanism for allocating resources to new users without affecting existing users. Since all server nodes use a single GPU, there is no waste of resources in a cluster composed of multiple GPU nodes. Among them, the present application uses a subgraph search algorithm, which reduces the waiting time for new users to wait for optical network reconfiguration to a certain extent.
[0163] Under the premise of ensuring efficient use of GPU computing resources, the embodiments of the present application utilize the high bandwidth, low latency, and appropriate network topology of the optical network to ensure the efficiency of communication between GPUs. For the all-optical network topology, when the number of GPUs applied by the user is less than or equal to the number of optical transceiver interfaces of each node (such as 4) + 1, the present application uses a fully connected graph to ensure that all communications are carried out efficiently. When the number of GPUs applied by the user is greater than the number of optical transceiver interfaces of each node + 1, a suitable topology connection is used to ensure the efficiency of the collective communication primitives commonly used in distributed machine learning such as All Reduce, and therefore the efficiency of communication of distributed machine learning tasks can also be guaranteed.
[0164] In some embodiments, Figure 8 is a schematic diagram of the structure of the resource allocation device provided in the embodiment of the present application, such as Figure 8 As shown, the resource allocation device 80 includes an acquisition module 801, a determination module 802 and a reconfiguration module 803, wherein the acquisition module 801 is used to obtain the node information of each server node in the cloud server cluster in response to a resource allocation request initiated by a user; wherein a graphics processor is set in each server node, the resource allocation request includes at least the number of node requests, and the node information includes at least the number of optical transceivers in each server node and the topological connection relationship between the server nodes; the determination module 802 is used to determine the target topological connection corresponding to the user based on the number of node requests and the number of optical transceivers; the reconfiguration module 803 is used to reconfigure the topological connection relationship between the server nodes based on the target topological connection to allocate resources to the user.
[0165] In some embodiments, the reconfiguration module 803 is also used to determine currently unoccupied idle server nodes in the cloud server cluster; in response to the number of idle server nodes being greater than the number of node requests, based on the topological connection relationship between the server nodes, determine the idle topological connection relationship between the idle server nodes; among the idle topological connection relationship, based on the target topological connection, determine the sub-topological connection relationship that meets the reconfiguration condition; based on the target topological connection, reconfigure the sub-topological connection relationship to allocate resources to the user.
[0166] In some embodiments, the reconfiguration module 803 is also used to determine, based on the target topology connection, at least one connection relationship in the idle topology connection relationship in which the number of nodes is the same as the number of nodes requested; respectively determine the degree of difference between each connection relationship and the target topology connection; and among the at least one connection relationship, determine the connection relationship with the smallest degree of difference with the target topology connection as the sub-topology connection relationship.
[0167] In some embodiments, the reconfiguration module 803 is also used to determine a first number of connection edges between the server node and the external node in each connection relationship, and a second number of different connection edges between each connection relationship and the target topological connection; wherein the external node is a server node in the cloud server cluster other than the server node in each connection relationship; based on the first number, the second number and the preset weight, the degree of difference between each connection relationship and the target topological connection is determined.
[0168] In some embodiments, the reconfiguration module 803 is also used to set an initial node set, which is an empty set; add the server nodes that meet the connection conditions in the idle topological connection relationship to the initial node set to obtain an updated node set; execute the update process of the updated node set again until the number of target nodes in the updated node set is equal to the number of node requests, so as to determine one of the connection relationships based on the target nodes in the updated node set; and execute the connection relationship determination process again to obtain at least one connection relationship.
[0169] In some embodiments, the reconfiguration module 803 is also used to obtain the minimum number of connections of the server nodes in the target topology connection; based on the minimum number of connections, determine the server node among the idle server nodes whose number of connections with the server nodes in the initial node set is greater than the minimum number of connections as the target node; or, in the case that there is no server node among the idle server nodes whose number of connections is greater than the minimum number of connections, determine the server node among the idle server nodes whose number of connections with the server nodes in the initial node set is the largest as the target node; add the target node to the initial node set to obtain the updated node set.
[0170] In some embodiments, the determination module 802 is also used to determine a comparison result between the number of node requests and a preset number in response to not receiving a custom topology connection; the preset number is obtained based on the number of optical transceivers; in response to the comparison result indicating that the number of node requests is less than or equal to the preset number, the fully connected topology is used as the target topology connection; in response to the comparison result indicating that the number of node requests is greater than the preset number, the target topology connection is generated based on a preset algorithm.
[0171] The description of the resource allocation device in the embodiment of the present application is similar to the description of the resource allocation method described above, and has similar beneficial effects as the resource allocation method, so it is not repeated here. For technical details not disclosed in the embodiment, please refer to the description of the resource allocation method of the present application for understanding.
[0172] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the serial number of each step / process mentioned above does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The serial numbers of the embodiments of the present application mentioned above are for description only and do not represent the advantages and disadvantages of the embodiments.
[0173] The present application uses descriptions such as "upper", "lower", "top", "bottom", "front", "back", "inside" and "outside" to indicate directions or positional relationships. This is only for the convenience of describing the present application, and does not indicate or imply that the device referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, it should not be understood as limiting the scope of protection of the present application.
[0174] In the description of this application, it should also be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be directly connected or indirectly connected through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to the specific circumstances.
[0175] It should be noted that, in this application, the terms "comprises", "includes" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0176] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0177] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, the functional units in the embodiments of the present application may be all integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0178] The above are only embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A resource allocation method, characterized in that: The resource allocation method comprises: In response to a resource allocation request initiated by a user, node information of each server node in the cloud server cluster is obtained; wherein a graphics processor is set in each server node, the resource allocation request includes at least the number of node requests, and the node information includes at least the number of optical transceivers in each server node and the topological connection relationship between the server nodes; Determining a target topology connection corresponding to the user based on the node request quantity and the quantity of the optical transceivers; Based on the target topological connection, the topological connection relationship between the server nodes is reconfigured to allocate resources to the user.
2. The resource allocation method according to claim 1, characterized in that: The reconfiguring the topological connection relationship between the server nodes based on the target topological connection to allocate resources to the user includes: In the cloud server cluster, determining idle server nodes that are not currently occupied; In response to the number of the idle server nodes being greater than the node request number, based on the topological connection relationship between the server nodes, determining the idle topological connection relationship between the idle server nodes; In the idle topology connection relationship, based on the target topology connection, determine a sub-topology connection relationship that meets the reconfiguration condition; Based on the target topology connection, the sub-topology connection relationship is reconfigured to allocate resources to the user.
3. The resource allocation method according to claim 2, characterized in that: The step of determining, in the idle topology connection relationship, a sub-topology connection relationship that satisfies a reconfiguration condition based on the target topology connection includes: Based on the target topological connection, determining at least one connection relationship in the idle topological connection relationship having the same number of nodes as the requested number of nodes; Determining the degree of difference between each connection relationship and the target topological connection respectively; Among the at least one connection relationship, the connection relationship with the smallest degree of difference from the target topological connection is determined as the sub-topological connection relationship.
4. The resource allocation method according to claim 3, characterized in that: The respectively determining the difference between each connection relationship and the target topological connection includes: Determine a first number of connection edges between the server node and the external node in each connection relationship, and a second number of different connection edges between each connection relationship and the target topological connection; wherein the external node is a server node in the cloud server cluster other than the server node in each connection relationship; Based on the first number, the second number and a preset weight, a degree of difference between each connection relationship and the target topological connection is determined.
5. The resource allocation method according to claim 3, characterized in that: The determining, based on the target topological connection, in the idle topological connection relationship, at least one connection relationship having the same number of nodes as the requested number of nodes comprises: Setting an initial node set, wherein the initial node set is an empty set; Adding the server nodes that meet the connection conditions in the idle topological connection relationship to the initial node set to obtain an updated node set; The updating process of the updating node set is executed again until the number of target nodes in the updating node set is equal to the number of node requests, so as to determine one of the connection relationships based on the target nodes in the updating node set; The connection relationship determination process is executed again to obtain at least one connection relationship.
6. The resource allocation method according to claim 5, characterized in that: The step of adding the server nodes that meet the connection condition in the idle topological connection relationship to the initial node set to obtain the updated node set includes: Obtaining the minimum number of connections of server nodes in the target topology connection; Based on the minimum number of connections, determining, as a target node, a server node among the idle server nodes whose number of connections with the server nodes in the initial node set is greater than the minimum number of connections; or, In the case that there is no server node with a connection number greater than the minimum connection number among the idle server nodes, determining the server node with the largest number of connections with the server nodes in the initial node set among the idle server nodes as the target node; The target node is added to the initial node set to obtain the updated node set.
7. The resource allocation method according to any one of claims 1 to 6, characterized in that: The determining, based on the node request quantity and the quantity of the optical transceivers, a target topology connection corresponding to the user comprises: In response to not receiving the custom topology connection, determining a comparison result between the node request quantity and a preset quantity; the preset quantity is obtained based on the quantity of the optical transceivers; In response to the comparison result indicating that the number of node requests is less than or equal to the preset number, connecting the fully connected topology as the target topology; In response to the comparison result indicating that the number of node requests is greater than the preset number, the target topology connection is generated based on a preset algorithm.
8. A resource allocation device, characterized in that: The resource allocation device comprises: An acquisition module, configured to obtain node information of each server node in the cloud server cluster in response to a resource allocation request initiated by a user; wherein a graphics processor is provided in each server node, the resource allocation request includes at least the number of node requests, and the node information includes at least the number of optical transceivers in each server node and the topological connection relationship between the server nodes; A determination module, configured to determine a target topology connection corresponding to the user based on the number of node requests and the number of optical transceivers; A reconfiguration module is used to reconfigure the topological connection relationship between the server nodes based on the target topological connection to allocate resources to the user.
9. A cloud server cluster architecture, characterized in that: The cloud server cluster architecture includes: A cloud server cluster, wherein the cloud server cluster includes a plurality of server nodes, each of which is provided with a graphics processor; An optical switch connected to a plurality of server nodes; Multiple operating systems, each operating system is connected to a server node for receiving resource allocation requests initiated by users; A management server, wherein the management network is connected to the multiple server nodes, and is used to respond to a resource allocation request initiated by a user and obtain node information of each server node in the cloud server cluster; wherein a graphics processor is set in each server node, the resource allocation request at least includes the number of node requests, and the node information at least includes the number of optical transceivers in each server node and the topological connection relationship between the server nodes; based on the number of node requests and the number of optical transceivers, determining the target topological connection corresponding to the user; based on the target topological connection, reconfiguring the topological connection relationship between the server nodes to allocate resources to the user.
10. The cloud server cluster architecture according to claim 9, characterized in that: The server node also includes an optical receiver, the number of optical transceivers in each server node is the same, and the optical transceivers in different server nodes are connected one-to-one; The optical switch is an optical circuit optical switch or an optoelectronic cross-connection panel.