Supercomputing task execution method, control device, storage medium and cluster system
By establishing an SSH protocol communication mechanism between the workstation and the server, generating a resource polling table and an arbitrator, the problem of low resource utilization in the distributed system is solved, and the reasonable allocation of supercomputing cluster resources and the balance of task execution are achieved.
Patent Information
- Application Number
- CN202510934900.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-07
AI Technical Summary
In distributed systems, existing resource allocation methods lead to low utilization of server resources, uneven resource use, and unreasonable execution of high- and low-priority tasks. Polling algorithms perform poorly when faced with tasks of different priorities or with large differences in resource requirements.
By establishing a communication mechanism based on the SSH protocol between the workstation and the server, the supercomputing cluster resource information is obtained, and a resource polling table and resource arbitrator are generated to ensure the rationality of resource selection, including the time arbitrator and data throughput arbitrator, and reasonably allocate supercomputing tasks.
It achieves the rational allocation of all server resources in the supercomputing cluster, improves resource utilization and the balance of task execution, and ensures the effective utilization of high-priority tasks and resource performance.
Smart Images

Figure CN120803723A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of supercomputing, and particularly relates to a supercomputing task execution method, a control device, a storage medium and a cluster system. BACKGROUND
[0002] In a distributed system, a supercomputing application may simultaneously request large-scale resources, such as computing resources, storage resources, network bandwidth of servers, etc., for supercomputing. Related resource allocation methods, such as static allocation or simple random allocation, cannot adapt to complex and variable load conditions, and are prone to cause low utilization of resources such as servers and unbalanced use of resources in a large-scale cluster.
[0003] In related technologies, although the round-robin algorithm can allocate resources in sequence and the resource allocation is relatively balanced, it performs poorly when facing tasks with different priorities or large differences in resource requirements. For example, high-performance resources may need to perform high-priority tasks due to their excellent performance needs, but due to the limitation of the round-robin order, they need to wait until the low-performance resource round-robin ends before allocating high-performance tasks for execution. SUMMARY
[0004] The application aims to provide a supercomputing task execution method, a control device, a storage medium and a cluster system, and aims to solve the problems of low utilization of resources such as servers, unbalanced use of resources, and unreasonable execution of high and low priority tasks in a large-scale cluster.
[0005] According to a first aspect of the application, a supercomputing task execution method is provided, applied to a workstation, and the supercomputing task execution method comprises: sending request information to a server end based on a communication mechanism established with the server end; receiving resource information of a supercomputing cluster sent by the server end in response to the request information; generating a resource polling table and a resource arbitrator in accordance with the resource information of the supercomputing cluster, and synchronizing the resource polling table and the resource arbitrator to the server end; and after the supercomputing cluster enters a ready state, batch-downloading supercomputing tasks to the server end according to application requirements, and executing corresponding supercomputing tasks by the supercomputing cluster according to the resource polling table and the resource arbitrator.
[0006] A communication mechanism based on an SSH protocol is established between the workstation and the computing nodes of the supercomputing cluster on the server side; based on the established communication mechanism, the workstation actively requests the supercomputing cluster to obtain all resource information of the supercomputing cluster; the server side responds to the request information of the workstation and returns all resource information of the supercomputing cluster; the workstation calculates a resource polling table and a resource arbiter that meet the supercomputing cluster and synchronizes to the supercomputing cluster; the supercomputing cluster receives the information and confirms; after the supercomputing cluster enters a ready state, the workstation batch-delivers supercomputing task execution requests according to application requirements, and the server side executes corresponding supercomputing tasks according to the resource polling table and the resource arbiter. The resource polling table and the resource arbiter can ensure the rationality of resource selection, so that all server resources of the supercomputing cluster can be reasonably allocated.
[0007] In an optional embodiment, before the request information is sent to the server side, the supercomputing task execution method further includes: applying for a storage space to the server side to establish a temporary data folder on the server side and generate a uuid folder with a unique identification bit.
[0008] The resource information is saved in the form of a file in the uuid folder, which is used for data verification and result checking operations.
[0009] In an optional embodiment, the resource polling table and the resource arbiter that meet the supercomputing cluster are generated according to the resource information of the supercomputing cluster, including: extracting at least one characteristic value of the resource information of the supercomputing cluster, including: extracting at least one characteristic value of the resource information of the supercomputing cluster, including: the peak performance of the computing nodes of the supercomputing cluster, the number and size of the memory of the computing nodes, the type and size of the storage of the computing nodes, and the type and support rate of the network of the computing nodes; and generating the resource polling table and the resource arbiter according to at least one characteristic value of the resource information of the supercomputing cluster.
[0010] In an optional embodiment, the resource polling table includes two resource polling tables, each resource polling table in the two resource polling tables includes a table entry index and corresponding table entry entity data, and the two resource polling tables are configured to be independent, not to share data, and not to repeat table entries.
[0011] In an optional embodiment, the resource arbitrator comprises a time arbitrator and a data throughput arbitrator, and the executing corresponding supercomputing tasks by the supercomputing cluster according to the resource polling table and the resource arbitrator comprises: determining a maximum timeout time for a corresponding computing node to execute a supercomputing task according to the time arbitrator, and the computing node will not be selected as a server for executing a supercomputing task again after waiting for the supercomputing task to be executed and completing execution, until a new supercomputing execution task is assigned when all entries in the resource polling table are polled and a next polling stage is entered; determining a maximum data throughput for a corresponding computing node to execute a supercomputing task according to the data throughput arbitrator, and the computing node will not be selected as a server for executing a supercomputing task again after waiting for the supercomputing task to be executed and completing execution, until a new supercomputing execution task is assigned when all entries in the resource polling table are polled and a next polling stage is entered, wherein the time arbitrator and the data throughput arbitrator do not work at the same time.
[0012] The specific resource polling table and the resource arbitrator can ensure the rationality of resource selection, so that all server resources of the supercomputing cluster can be reasonably allocated.
[0013] In an optional embodiment, the supercomputing task execution method further comprises: receiving a supercomputing task execution result sent by the server end; and displaying the supercomputing task execution result on the user end.
[0014] According to a second aspect of the present application, a supercomputing task execution method is provided, applied to a server end, and the supercomputing task execution method comprises: receiving request information sent by a workstation based on a communication mechanism established with the workstation; in response to the request information, acquiring resource information of a supercomputing cluster, and sending the resource information of the supercomputing cluster to the workstation; receiving a resource polling table and a resource arbitrator generated according to the resource information of the supercomputing cluster and sent by the workstation; and after the supercomputing cluster enters a ready state, receiving supercomputing tasks according to application requirements sent in batches by the workstation, and executing corresponding supercomputing tasks according to the resource polling table and the resource arbitrator.
[0015] In an optional embodiment, before the receiving the request information sent by the workstation, the supercomputing task execution method further comprises: receiving a storage space application of the workstation, establishing a temporary data folder, and generating a uuid folder with a unique identification bit; and after the acquiring the resource information of the supercomputing cluster, storing the acquired resource information in the uuid folder.
[0016] In an optional embodiment, after receiving the resource polling table and the resource arbitrator sent by the workstation, the supercomputing task execution method further comprises: checking whether the number of servers as supercomputing task execution is accurate and whether the configuration of each piece of data is accurate for the resource polling table; and determining the effective mode for the resource arbitrator.
[0017] In an optional embodiment, the resource polling table comprises two resource polling tables, each of the two resource polling tables comprises a table entry index and corresponding table entry entity data, and the two resource polling tables are configured to be independent in table entry, not to share data, and not to repeat table entries.
[0018] In an optional embodiment, the resource arbitrator comprises a time arbitrator and a data throughput arbitrator, and the execution of the corresponding supercomputing task according to the resource polling table and the resource arbitrator comprises: determining the maximum timeout time for the corresponding computing node to execute the supercomputing task according to the time arbitrator, the computing node executes the supercomputing task with the determined maximum timeout time, and after waiting for the supercomputing task to be executed, the computing node will not be selected as a server for executing the supercomputing task again until the next polling phase is entered after all table entries in the resource polling table are polled and executed, and a new supercomputing execution task is allocated; determining the maximum data throughput for the corresponding computing node to execute the supercomputing task according to the data throughput arbitrator, the computing node executes the supercomputing task with the determined maximum data throughput, and after waiting for the supercomputing task to be executed, the computing node will not be selected as a server for executing the supercomputing task again until the next polling phase is entered after all table entries in the resource polling table are polled and executed, and a new supercomputing execution task is allocated, wherein the time arbitrator and the data throughput arbitrator do not work at the same time.
[0019] In an optional embodiment, after receiving the supercomputing task execution request according to the application requirement in batches issued by the workstation, the supercomputing task execution method further comprises: placing the supercomputing task execution request in the middleware of the execution queue, so that the middleware screens a suitable computing node according to the resource polling table and the resource arbitrator to execute the corresponding supercomputing task.
[0020] According to a second aspect of the present application, a workstation is provided, which comprises a control device, the control device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the computer program to implement the supercomputing task execution method applied to the workstation as described above.
[0021] According to a third aspect of the present application, a server is provided, the server comprising a control device, the control device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the computer program to implement the above-mentioned supercomputing task execution method applied to a server side.
[0022] According to a fourth aspect of the present application, a machine-readable storage medium is provided, the machine-readable storage medium storing instructions that cause a machine to execute the above-mentioned supercomputing task execution method applied to a workstation or the above-mentioned supercomputing task execution method applied to a server side.
[0023] According to a fifth aspect of the present application, a multi-data center cluster system is provided, the multi-data center cluster system comprising a plurality of clusters, the above-mentioned server electrically connected to each cluster of the plurality of clusters, and the above-mentioned workstation electrically connected to the server, each cluster comprising a plurality of computing nodes.
[0024] Through the above technical solution, the supercomputing task execution method provided by the embodiments of the present application establishes a communication mechanism based on the SSH protocol between the workstation and the computing nodes of the supercomputing cluster at the server side; based on the established communication mechanism, the workstation actively requests the supercomputing cluster to obtain all resource information of the supercomputing cluster; the server side responds to the request information of the workstation and replies all resource information of the supercomputing cluster; the workstation calculates a resource polling table and a resource arbiter that meet the supercomputing cluster and synchronizes to the supercomputing cluster; the supercomputing cluster receives the information and confirms; after the supercomputing cluster enters a ready state, the workstation batch issues a supercomputing task execution request according to application requirements, and the server side executes the corresponding supercomputing task according to the resource polling table and the resource arbiter. The embodiments of the present application can ensure the rationality of resource selection through the resource polling table and the resource arbiter, so as to realize that all server resources of the supercomputing cluster will be reasonably allocated.
[0025] Other features and advantages of the present application will be described in the following description, and become apparent from the description, or be learned by practice of the present application. The purposes and other advantages of the present application can be achieved and obtained by the structures and processes indicated in the description and drawings. BRIEF DESCRIPTION OF DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the embodiments or related art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0027] Figure 1 is a topological schematic diagram of a multi-data center cluster system architecture provided by an example embodiment of the present application.
[0028] Figure 2 is a flow schematic diagram of a supercomputing task execution method provided by an example embodiment of the present application.
[0029] Figure 3 is a schematic diagram of a resource polling table and a resource arbiter of an example embodiment of the present application.
[0030] Figure 4 is a flow schematic diagram of a supercomputing task execution method provided by another example embodiment of the present application.
[0031] Figure 5 is a workflow schematic diagram of a multi-data center cluster system of an example embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0033] As described above, in the related art, although the polling algorithm can allocate resources in sequence and the resource allocation is relatively balanced, it performs poorly when facing tasks with different priorities or large differences in resource demand. For example, high-performance resources may need to perform high-priority tasks due to their excellent performance needs, but due to the limitation of the polling order, they need to wait until the low-performance resources finish polling before they can be allocated to perform high-performance tasks. And the arbitration mechanism is usually based on a single judgment standard, such as the resource request of first come first served or the remaining amount of available resources, and cannot comprehensively consider multiple factors such as the high and low priority of tasks, the performance efficiency of resources, and the overall load balancing of distributed systems, making it difficult to achieve reasonable allocation of resources. In view of this, the embodiments of the present application provide a supercomputing task execution method.
[0034] Before explaining the embodiments of the present application in detail, a multi-data center cluster system architecture to which the embodiments of the present application are applied will be introduced. Figure 1The supercomputer large-scale multi-data center cluster system architecture (topology) is shown. In the supercomputer cluster, there can be multiple data centers, and switches or routers are used to build communication rules between multiple data centers; within the same data center, several switches can be used to build communication rules between servers. Please refer to Figure 1 , 101-116 indicate the computing nodes (clusters, also referred to as clusters) in the supercomputer cluster, 120-123 indicate the switch devices in the supercomputer cluster, 124 indicates the router communication device between data centers, 125 indicates the jump machine for accessing the supercomputer cluster through the server, and 126 indicates the workstation device. Among them, 101-104 indicate four clusters in data centerl, and each cluster is deployed with multiple server resources, the number of which depends on the actual use demand. For example, the communication between clusters 101-104 depends on the switch device 120, that is, the server communication between clusters 101-104 is performed through the switch device 120 for data transceiving exchange. Similarly, Figure 1 The clusters in data centers data center2, data center3 and data center4 in perform data exchange as above. The data transceiving exchange between data centers depends on the routing device 124, that is, the communication between each data center is not directly performed, but needs to pass through the 124 routing device. The server device 125 is independent of the supercomputer cluster and does not belong to the supercomputer node, and is connected with the 124 routing device. The workstation 126 establishes communication with the server device 125, and the end user can use the workstation 126 to access the supercomputer cluster, query the supercomputer cluster computing, storage and network resources, and issue a supercomputer application task through the workstation 126. At the same time, after the supercomputer application is executed, the supercomputer cluster returns the supercomputer application execution result to the workstation 126, and the end user can view the execution result through the workstation 126.
[0035] Please refer to Figure 2 The supercomputer task execution method provided by the embodiments of the present application can be applied to the workstation 126 of Figure 1 , and the supercomputer task execution method can include the following steps:
[0036] Step S210: Based on the communication mechanism established with the server side, request information is sent to the server side.
[0037] In the embodiments of the present application, before the communication mechanism between the workstation and the server is established, the workstation has stored the relevant information of the server, for example, the necessary information can include IP address, login account, login password, etc. In addition, the server can be configured to have the ability to communicate with the data center of the supercomputing cluster; it can also be configured to query and modify the information including but not limited to the following: 1) all server information of the data center of the supercomputing cluster (for example, including IP address, CPU type, storage, network, etc.); 2) all server node administrators and remote access permissions; 3) supercomputing task issuing, transferring, collecting, etc. queue middleware plug-ins.
[0038] In the preferred embodiments of the present application, before sending the request information to the server, the supercomputing task execution method can further include: applying for storage space to the server to establish a temporary data folder on the server and generate a uuid folder with a unique identification bit.
[0039] For example, the workstation and the server can establish communication based on the Secure Shell (SSH) protocol. The SSH is a protocol for secure remote login and other secure network services over an insecure network, and the communication mechanism based on the SSH protocol has the ability of bidirectional data transmission. After the communication is established, the workstation can apply for storage space to the server to establish a temporary data folder and generate a uuid folder with a unique identification bit. The uuid folder can be used for data verification and result checking in the following operations. Based on the established communication mechanism, the workstation sends request information to the server to request the resource information of all servers in the data center of the supercomputing cluster.
[0040] Step S220: receiving the resource information of the supercomputing cluster sent by the server in response to the request information.
[0041] Based on the above example, after receiving the request information, the server can obtain the resource information of all servers in the data center of the supercomputing cluster (for example, including IP address, CPU type, storage, network, etc.) from the information base, save the resource information in the form of a file in the uuid folder, and send the resource information to the workstation in the form of a response message.
[0042] Step S230: generating a resource polling table and a resource arbiter in accordance with the resource information of the supercomputing cluster, and synchronizing the resource polling table and the resource arbiter to the server.
[0043] In the preferred embodiment of the present application, the resource polling table and the resource arbiter are generated according to the resource information of the supercomputing cluster, which can include: extracting at least one of the following characteristic values of the resource information of the supercomputing cluster: the peak performance of the computing nodes of the supercomputing cluster, the number and size of the memory of the computing nodes, the type and size of the storage of the computing nodes, and the type and support rate of the network of the computing nodes; and generating the resource polling table and the resource arbiter according to at least one of the characteristic values of the resource information of the supercomputing cluster.
[0044] For example, after the workstation receives the resource information of the supercomputing cluster returned by the server, the following characteristic values in the resource information are extracted: 1) the peak performance of the computing nodes of the supercomputing cluster (theoretical); 2) the number and size of the memory of the computing nodes; 3) the type and size of the storage of the computing nodes; and 4) the type and support rate of the network of the computing nodes. The peak performance of the computing nodes = the number of CPU (core) of the computing nodes x the number of floating point operations per cycle x the clock frequency. For example, the peak performance of a computing node with 8 cores, a clock frequency of 3 GHz, and a floating point operation number of 4 is 8 x 4 x 3.0 = 96 GFLOPS.
[0045] According to the above characteristic values, the resource polling table and the resource arbiter are generated. For example, the resource information of the supercomputing cluster can be input into a pre-configured algorithm or function, and a specific resource polling table and a resource arbiter are output.
[0046] The preferred resource polling table of the embodiment of the present application can include two resource polling tables, and each of the two resource polling tables includes a table entry index and corresponding table entry entity data; the two resource polling tables are configured to be independent in table entry, not to share data, and not to repeat table entries.
[0047] Please refer to the contents indicated in Figure 3 The resource polling table can include A table and B table, each of which includes a plurality of table entry indexes (index items) and table entry entity data (entry items), and each of which can include key information. Meanwhile, the data contents of the A table and the B table are not repeated, the data of the A table and the B table are not shared, and the independence of each is maintained. The index item can be a serial number generated in sequence, and the entry item can store the resource information of the corresponding server.
[0048] In the embodiment of the present application, the resource information can be distributed in the two resource polling tables (A table and B table) in an approximately equal manner, and the allocation manner tends to be randomly selected without referring to any characteristic attribute. The two resource polling tables can be configured as A table and B table, and the table entries are consistent and include index and entry items. The table entry items of the A table and the B table are not fixed and can be allocated according to the actual computing nodes of the supercomputer cluster. In the A table and the B table, the table entry items are independent, the data is not shared, and all table entries are not repeated.
[0049] For details, please refer to the content indicated in 304 of Figure 3 In the embodiment of the present application, the entry item can include:
[0050] IP field: This field refers to the specific IPV4 address of the computing node network;
[0051] weight field, configured to indicate the weight value of the corresponding computing node, which can be used to describe the computing capability of the corresponding computing node;
[0052] throughput field: configured to indicate the data throughput of the corresponding computing node, which can be used to describe the storage capability value of the corresponding computing node. The value can be configured by the user;
[0053] timer field: configured to indicate the time arbitrator default unit value of the corresponding computing node, which can be used to describe the network capability value of the corresponding computing node. The value can be configured by the user;
[0054] tag field: configured to indicate the state of the corresponding computing node, which can be divided into online, running and offline. The online state indicates that the computing node can be allocated for task execution; the running state indicates that the computing node is performing a task; and the offline state indicates that the computing node cannot be allocated.
[0055] The preferred resource arbitrator of the embodiment of the present application can include a time arbitrator and a data throughput arbitrator. That is, the embodiment of the present application provides two arbitration methods, time arbitration or data throughput arbitration. The time arbitrator and the data throughput arbitrator do not work at the same time.
[0056] For example, after the workstation generates the resource polling table and the resource arbitrator according to the resource information of the supercomputer cluster by the above-mentioned method, the resource polling table and the resource arbitrator can be sent to the server side through the communication mechanism of the SSH protocol. After the server side receives the corresponding message, it confirms the successful reception.
[0057] In a preferred embodiment of the present application, after the server receives the resource polling table and resource arbitrator sent by the workstation, the supercomputing task execution method may also include: for the resource polling table, verifying whether the number of servers executing the supercomputing task is accurate, and verifying whether the configuration of each data is accurate; and for the resource arbitrator, determining the effectiveness method.
[0058] In an embodiment of the present application, there is a time difference between the server sending the resource information of the supercomputing cluster and receiving the resource polling table and the resource arbitrator. Therefore, after receiving the resource polling table and the resource arbitrator sent by the workstation, the server can verify whether the server entry in the resource polling table is correct (for example, the server is shut down). If there is a difference, the latest server configuration is updated, for example, the shut down server and the corresponding configuration are deleted. For example, as described above, after receiving the request information, the server can obtain the resource information of all servers in the data center of the supercomputing cluster from the information library (for example, including IP address, CPU type, storage, network and other information), and save the resource information in the form of a file in the uuid folder. After receiving the resource polling table and resource arbitrator from the workstation, the server performs verification operations on the corresponding data. For example, this includes: verifying that all server entries in resource polling table A and resource polling table B are consistent with all server entries in the supercomputing cluster's data center; confirming that each entry in the table is consistent with the actual server configuration; if not, updating the latest server configuration; and confirming the resource arbitrator's effectiveness mode (i.e., time arbitrator or data throughput arbitrator). After completing the above verification operations, the server persists the resource polling table and resource arbitrator to a local uuid folder, in a file format such as .csV.
[0059] Step S240: After the supercomputing cluster enters the ready state, supercomputing tasks are sent to the server in batches according to application requirements, so that the supercomputing cluster executes the corresponding supercomputing tasks according to the resource polling table and resource arbitrator.
[0060] For example, after a supercomputing cluster (the server resources of all computing nodes) enters the ready state, a workstation can send a supercomputing application request via SSH to batch-deliver supercomputing tasks to the server. Upon receiving the supercomputing application request from the workstation, the server places the supercomputing task into the execution queue of the middleware. The middleware then intelligently selects the appropriate computing node for supercomputing task execution based on the source polling table and resource arbitrator.
[0061] In a preferred embodiment of the present application, the supercomputing cluster may perform the corresponding supercomputing task according to the resource polling table and the resource arbitrator, including: determining the maximum timeout for the corresponding computing node to perform the supercomputing task according to the time arbitrator, and after the computing node performs the supercomputing task with the determined maximum timeout and waits for the supercomputing task to be completed, it will no longer be selected as the server to perform the supercomputing task, until all items in the resource polling table are polled and executed, and a new supercomputing task is assigned when the next polling phase is entered; determining the maximum data throughput for the corresponding computing node to perform the supercomputing task according to the data throughput arbitrator, and after the computing node performs the supercomputing task with the determined maximum data throughput and waits for the supercomputing task to be completed, it will no longer be selected as the server to perform the supercomputing task, until all items in the resource polling table are polled and executed, and a new supercomputing task is assigned when the next polling phase is entered. Wherein, the time arbitrator and the data throughput arbitrator do not work at the same time.
[0062] Please refer to Figure 3 As indicated by 303, the resource arbitrator in the embodiment of the present application includes two arbitration modes, namely, a time arbitrator and a data throughput arbitrator. In the embodiment of the present application, the maximum timeout time for a computing node to execute a supercomputing task can be expressed as (weight*the default value of the time arbitrator), and the weight value can be obtained through the resource polling table. For example, the weight corresponding to the computing node selected as the server is 3, and the default value is 60 seconds. Then, the maximum timeout time for the corresponding computing node to execute a supercomputing task is 3*60=180 seconds. It can be interpreted as: after the computing node executes a supercomputing task for a maximum of 180 seconds and waits for the task to be executed, it will no longer be selected as a task execution server until all items in the resource polling table are polled and executed, and then enters the next polling stage, and a new supercomputing execution task will be assigned. In the embodiment of the present application, the maximum data throughput of a computing node to execute a supercomputing task can be expressed as (weight*default value), and the weight value can be obtained through the resource polling table. For example, if the weight of a compute node selected as a server is 3 and the default value is 1MB, the maximum data throughput of the corresponding compute node for supercomputing tasks is 3*1MB=3MB. This means that after the compute node executes a supercomputing task with a maximum data throughput of 3MB and waits for the task to complete, it will not be selected as a supercomputing task execution server again until all entries in the resource polling table have been polled and executed, and then enters the next polling phase, at which point a new supercomputing execution task will be assigned. The two aforementioned arbitrators can arbitrate from two dimensions: time and data. End users can choose based on their actual usage.
[0063] In the preferred embodiments of the present application, the supercomputing task execution method can further include: receiving the supercomputing task execution result sent by the server side; and displaying the supercomputing task execution result on the user side.
[0064] According to the above example, after all the supercomputing tasks in the to-be-queued list of the server side are executed, the server side can feed back the execution results of all the supercomputing tasks to the workstation. After receiving the execution results of the tasks, the workstation can display the execution results to the terminal user, so that the terminal user can analyze the supercomputing tasks according to the execution results.
[0065] In the preferred embodiments of the present application, the server side can return the supercomputing task execution results in batches, release the corresponding resources, and update the resource polling table and the resource arbiter.
[0066] Accordingly, the supercomputing task execution method provided in the embodiments of the present application establishes a communication mechanism based on the SSH protocol between the workstation and the (computing nodes of the supercomputing cluster of) server side; based on the established communication mechanism, the workstation actively requests the supercomputing cluster to obtain all the resource information of the supercomputing cluster; the server side responds to the request information of the workstation and replies to all the resource information of the supercomputing cluster; the workstation calculates the resource polling table and the resource arbiter that meet the supercomputing cluster and synchronizes them to the supercomputing cluster; the supercomputing cluster receives the information and confirms it; after the supercomputing cluster enters the ready state, the workstation batches the supercomputing task execution requests according to the application requirements, and the server side executes the corresponding supercomputing tasks according to the resource polling table and the resource arbiter. The embodiments of the present application can ensure the rationality of resource selection through the resource polling table and the resource arbiter, so as to realize the rational allocation of all the server resources of the supercomputing cluster.
[0067] For reference Figure 4 The supercomputing task execution method provided in the embodiments of the present application can be applied to the server side 125 of the supercomputing cluster 100, and the supercomputing task execution method can include the following steps: Figure 1
[0068] Step S410: receiving the request information sent by the workstation based on the communication mechanism established with the workstation.
[0069] In the preferred embodiments of the present application, before receiving the request information sent by the workstation, the supercomputing task execution method can further include: receiving the storage space application of the workstation, establishing a temporary data folder, and generating a uuid folder with a unique identification bit; and after obtaining the resource information of the supercomputing cluster, storing the obtained resource information in the uuid folder.
[0070] As described above, the workstation and the server side can establish communication based on the communication mechanism of the SSH protocol, for example. After the communication is established, the workstation can apply for storage space from the server side, establish a temporary data folder, and generate a uuid folder with a unique identification bit. Based on the communication mechanism established with the server side, the workstation sends request information to the server side, requesting resource information of all servers in the data center of the supercomputing cluster.
[0071] Step S420: In response to the request information, the resource information of the supercomputing cluster is obtained, and the resource information of the supercomputing cluster is sent to the workstation.
[0072] In the above example, after receiving the request information, the server side can obtain the resource information of all servers in the data center of the supercomputing cluster from the information library, save the resource information in the form of a file in the uuid folder, and send the resource information to the workstation in the form of a response message.
[0073] Step S430: Receive the resource polling table and resource arbiter generated according to the resource information of the supercomputing cluster sent by the workstation.
[0074] The preferred resource polling table of the embodiment of the application can include two resource polling tables, each of which includes a table entry index and corresponding table entry entity data; the two resource polling tables are configured to be independent, data is not shared, and table entries are not repeated.
[0075] The preferred resource arbiter of the embodiment of the application can include a time arbiter and a data throughput arbiter. That is, the embodiment of the application provides two arbitration methods, time arbitration or data throughput arbitration. The time arbiter and the data throughput arbiter do not work at the same time.
[0076] In the above example, after receiving the resource information of the supercomputing cluster returned by the server side, the workstation generates a specific resource polling table and resource arbiter according to the following characteristic values in the resource information: 1) the peak performance of the computing node (theoretical) of the supercomputing cluster; 2) the number and size of the memory in the computing node; 3) the type and size of the storage of the computing node; 4) the type and support rate of the network of the computing node. The resource polling table can refer to the content indicated in Figure 3 The resource arbiter can refer to the content indicated in Figure 3 The detailed configuration of the resource polling table and the resource arbiter can refer to the above, which will not be described here.
[0077] In the preferred embodiment of the present application, after step S430, the supercomputing task execution method can further include: for the resource polling table, checking whether the number of servers as the supercomputing task execution is accurate and checking whether the configuration of each piece of data is accurate; and for the resource arbiter, determining the effective mode.
[0078] Taking the above example, after the server end receives the resource polling table and the resource arbiter sent by the workstation, the corresponding data is checked, for example, including: verifying whether all server entries in the resource polling table A and the resource polling table B are consistent with all server entries of the data center of the supercomputing cluster; and confirming whether each entry in the table entry is consistent with the actual configuration of the server, if not, updating the latest server configuration in real time; confirming the effective mode of the resource arbiter (i.e., time arbiter or data throughput arbiter). After completing the above verification operation, the server end persists the resource polling table and the resource arbiter to the local uuid folder, and the file format is, for example,.csv.
[0079] Step S440: After the supercomputing cluster enters the ready state, the supercomputing task according to the application requirement sent by the workstation in batches is received, and the corresponding supercomputing task is executed according to the resource polling table and the resource arbiter.
[0080] In the preferred embodiment of the present application, after receiving the supercomputing task execution request sent by the workstation in batches according to the application requirement, the supercomputing task execution method can further include: putting the supercomputing task execution request into the middleware of the execution queue, so that the middleware screens the appropriate computing node according to the resource polling table and the resource arbiter to execute the corresponding supercomputing task.
[0081] Among them, as described above, the middleware is an important component for realizing asynchronous communication and data transmission. Taking the above example, after the supercomputing cluster (server resources of all computing nodes) enters the ready state, the workstation can communicate through the SSH protocol to send a supercomputing application request to batch the supercomputing task to the server end. After the server end receives the supercomputing application request sent by the workstation, the supercomputing task is put into the middleware of the execution queue, and the middleware intelligently screens the appropriate computing node to execute the supercomputing task according to the source polling table and the resource arbiter.
[0082] In the preferred embodiment of the present application, the super-computing cluster can execute corresponding super-computing tasks according to the resource polling table and the resource arbitrator, which can include: determining the maximum timeout time for the corresponding computing node to execute the super-computing task according to the time arbitrator, executing the super-computing task by the computing node with the determined maximum timeout time, and after waiting for the super-computing task to be executed, the server will not be selected again as the server for executing the super-computing task until all entries in the resource polling table are executed in the next polling phase, and a new super-computing execution task is assigned; determining the maximum data throughput for the corresponding computing node to execute the super-computing task according to the data throughput arbitrator, executing the super-computing task by the computing node with the determined maximum data throughput, and after waiting for the super-computing task to be executed, the server will not be selected again as the server for executing the super-computing task until all entries in the resource polling table are executed in the next polling phase, and a new super-computing execution task is assigned. The time arbitrator and the data throughput arbitrator do not work at the same time.
[0083] Please refer to Figure 3 The resource arbitrator in the embodiment of the present application includes two arbitration modes, i.e., the time arbitrator and the data throughput arbitrator. In the embodiment of the present application, the maximum timeout time for the computing node to execute the super-computing task can be represented as (weight*default value of the time arbitrator), and the weight value can be obtained through the resource polling table. For example, the weight of the selected computing node as the server is 3, and the default value is 60 seconds, so the maximum timeout time for the corresponding computing node to execute the super-computing task is 3*60=180 seconds. It can be explained that the computing node performs the super-computing task for a maximum of 180 seconds and waits for the task to be executed, and after that, the server will not be selected again as the server for executing the task until all entries in the resource polling table are executed in the next polling phase, and a new super-computing execution task is assigned. In the embodiment of the present application, the maximum data throughput for the computing node to execute the super-computing task can be represented as (weight*default value), and the weight value can be obtained through the resource polling table. For example, the weight of the selected computing node as the server is 3, and the default value is 1MB, so the maximum data throughput for the corresponding computing node to execute the super-computing task is 3*1MB=3MB. It can be explained that the computing node performs the super-computing task for a maximum data throughput of 3MB and waits for the task to be executed, and after that, the server will not be selected again as the server for executing the super-computing task until all entries in the resource polling table are executed in the next polling phase, and a new super-computing execution task is assigned. The above two arbitrators can be arbitrated from two dimensions, i.e., the time dimension and the data dimension, and the end user can select according to the actual use.
[0084] The embodiment of the present application further provides a workstation, the workstation comprising a control device, the control device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the supercomputing task execution method applied to the workstation.
[0085] The embodiment of the present application further provides a server, the server comprising a control device, the control device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the supercomputing task execution method applied to the server.
[0086] The embodiment of the present application further provides a machine readable storage medium, the machine readable storage medium storing instructions, and the instructions cause a machine to execute the supercomputing task execution method applied to the workstation or the supercomputing task execution method applied to the server.
[0087] It should be noted that the control device and the machine readable storage medium can implement the supercomputing task execution method provided by the above embodiment, and the specific implementation manner can refer to the description of the supercomputing task execution method in the above embodiment, and will not be described here.
[0088] The embodiment of the present application further provides a multi-data center cluster system, the multi-data center cluster system can comprise a plurality of clusters, the server electrically connected with each cluster in the plurality of clusters and the workstation electrically connected with the server, and each cluster comprises a plurality of computing nodes.
[0089] In the embodiment of the present application, the topology of the data center cluster system can be as shown in Figure 1 The flow of the multi-data center cluster system executing the supercomputing task can be as shown in Figure 5
[0090] In step 501, the workstation and the server can establish communication using the SSH protocol, and based on the SSH protocol, the workstation and the server have the ability of bidirectional data transmission. After the communication is established, the workstation applies for a storage space to the server, establishes a temporary data folder, and generates a uuid folder with a unique identification bit. The content in the uuid folder is convenient for subsequent data verification, result checking and other operations.
[0091] In step 502, based on the SSH communication mechanism created in step 501, the workstation sends a request to the server, requesting the server node information in all data centers in the supercomputing cluster.
[0092] In step 503, the server receives the request information, obtains all server information of the supercomputing cluster data center (such as IP address, CPU type, storage, network, etc.) from the information base, saves the resource information in the form of a file in the uuid folder, and sends the resource information to the workstation in the form of a response message.
[0093] In step 504, the workstation receives the resource information of the supercomputing cluster returned by the server, and generates a specific resource polling table and a resource arbiter according to the following characteristic values of the resource information: 1) the number of CPU of the computing node * the frequency of each CPU; 2) the number and size of the memory of the computing node; 3) the type and size of the storage of the computing node; and 4) the type and support rate of the network of the computing node.
[0094] The workstation sends the resource polling table and the resource arbiter to the server, and the server confirms the successful reception of the corresponding information after receiving the information.
[0095] In step 505, the server receives the resource polling table and the resource arbiter sent by the workstation, and performs verification. After the server information is confirmed, the information is stored persistently in the local path. At the same time, all server resources are in a ready state, waiting for task execution and returning to the workstation to issue supercomputing tasks.
[0096] In step 506, the workstation can issue supercomputing application tasks in batches after receiving the state request returned by the server.
[0097] In step 507, the server receives the supercomputing application request sent by the workstation, puts the supercomputing application request into the execution queue middleware, and the middleware intelligently selects suitable computing nodes for supercomputing task execution according to the source polling table and the resource arbiter.
[0098] In step 508, after all supercomputing tasks in the standby queue of the server are executed, the server returns all supercomputing application task execution results to the workstation. After receiving the task execution results, the workstation displays the results to the terminal user, and the terminal user analyzes the results according to the execution results.
[0099] It can be understood that the circuit structure, name and parameter described in the above embodiment are only examples. Those skilled in the art can also easily combine and adjust the structural characteristics of the above multiple embodiments according to the use needs, and the concept of the present application should not be limited to the specific details of the above examples.
[0100] Although the present application has been described in detail with reference to the foregoing embodiments, it should be understood that modifications can be made to the foregoing embodiments, or additional implementations can be implemented, without departing from the spirit and scope of the embodiments.
Claims
1. A method for executing a supercomputing task, characterized in that: Applied to a workstation, the supercomputing task execution method includes: Based on the communication mechanism established with the server, sending request information to the server; Receiving resource information of the supercomputing cluster in response to the request information sent by the server; Generate a resource polling table and a resource arbitrator that conform to the supercomputing cluster according to the resource information of the supercomputing cluster, and synchronize the resource polling table and the resource arbitrator to the server side; and After the supercomputing cluster enters the ready state, supercomputing tasks are sent to the server end in batches according to application requirements, so that the supercomputing cluster executes the corresponding supercomputing tasks according to the resource polling table and the resource arbitrator.
2. The supercomputing task execution method according to claim 1, characterized in that: Before sending the request information to the server, the supercomputing task execution method further includes: Apply for storage space from the server to create a temporary data folder on the server and generate a uuid folder with a unique identification bit.
3. The supercomputing task execution method according to claim 1, characterized in that: The step of generating a resource polling table and a resource arbitrator that conform to the supercomputing cluster according to the resource information of the supercomputing cluster includes: Extracting at least one of the following characteristic values of the resource information of the supercomputing cluster: peak performance of computing nodes of the supercomputing cluster, number and size of computing node memories, type and size of computing node storage, and type and supported rate of computing node networks; and The resource polling table and the resource arbitrator are generated according to at least one characteristic value of the resource information of the supercomputing cluster.
4. The supercomputing task execution method according to claim 1 or 3, characterized in that: The resource polling table includes two resource polling tables, each of the two resource polling tables includes a table entry index and corresponding table entry entity data, The two resource polling tables are configured such that table entries are independent, data is not shared, and table entries are not repeated.
5. The supercomputing task execution method according to claim 1 or 3, characterized in that: The resource arbitrator includes a time arbitrator and a data throughput arbitrator, and the supercomputing cluster executes the corresponding supercomputing task according to the resource polling table and the resource arbitrator, including: Determine, according to the time arbitrator, the maximum timeout for the corresponding computing node to execute the supercomputing task. After the computing node executes the supercomputing task within the determined maximum timeout and waits for the supercomputing task to be completed, it will no longer be selected as a server to execute the supercomputing task until all items in the resource polling table are polled and executed, and a new supercomputing execution task is assigned when entering the next polling phase; The data throughput arbitrator determines the maximum data throughput of the corresponding computing node for executing the supercomputing task. After the computing node executes the supercomputing task with the determined maximum data throughput and waits for the supercomputing task to be completed, it will no longer be selected as a server for executing the supercomputing task until all items in the resource polling table are polled and executed, and a new supercomputing execution task is assigned when entering the next polling phase. The time arbitrator and the data throughput arbitrator do not work at the same time.
6. The supercomputing task execution method according to claim 1, characterized in that: The supercomputing task execution method further includes: Receiving the supercomputing task execution result sent by the server; and The execution results of the supercomputing task are displayed on the user side.
7. A supercomputing task execution method, characterized in that: Applied to the server side, the supercomputing task execution method includes: Based on the communication mechanism established with the workstation, receiving the request information sent by the workstation; In response to the request information, obtaining resource information of the supercomputing cluster and sending the resource information of the supercomputing cluster to the workstation; receiving a resource polling table and a resource arbitrator that are generated according to the resource information of the supercomputing cluster and are in compliance with the supercomputing cluster and sent by the workstation; and After the supercomputing cluster enters the ready state, it receives supercomputing tasks issued in batches by the workstations according to application requirements, and executes corresponding supercomputing tasks according to the resource polling table and the resource arbitrator.
8. The supercomputing task execution method according to claim 7, characterized in that: Before receiving the request information sent by the workstation, the supercomputing task execution method further includes: Receive a storage space application from the workstation, create a temporary data folder, and generate a uuid folder with a unique identifier; and After obtaining the resource information of the supercomputing cluster, the obtained resource information is stored in the uuid folder.
9. The supercomputing task execution method according to claim 8, characterized in that: After receiving the resource polling table and the resource arbitrator sent by the workstation, the supercomputing task execution method further includes: For the resource polling table, verify whether the number of servers executing the supercomputing task is accurate, and verify whether the configuration of each data is accurate; and For the resource arbitrator, a validation mode is determined.
10. The supercomputing task execution method according to claim 7, characterized in that: The resource polling table includes two resource polling tables, each of the two resource polling tables includes a table entry index and corresponding table entry entity data, The two resource polling tables are configured such that table entries are independent, data is not shared, and table entries are not repeated.
11. The supercomputing task execution method according to claim 7, characterized in that: The supercomputing task execution method according to claim 1 or 3, wherein the resource arbitrator includes a time arbitrator and a data throughput arbitrator, and executing the corresponding supercomputing task according to the resource polling table and the resource arbitrator includes: Determine, according to the time arbitrator, the maximum timeout for the corresponding computing node to execute the supercomputing task. After the computing node executes the supercomputing task within the determined maximum timeout and waits for the supercomputing task to be completed, it will no longer be selected as a server to execute the supercomputing task until all items in the resource polling table are polled and executed, and a new supercomputing execution task is assigned when entering the next polling phase; The data throughput arbitrator determines the maximum data throughput of the corresponding computing node for executing the supercomputing task. After the computing node executes the supercomputing task with the determined maximum data throughput and waits for the supercomputing task to be completed, it will no longer be selected as a server for executing the supercomputing task until all items in the resource polling table are polled and executed, and a new supercomputing execution task is assigned when entering the next polling phase. The time arbitrator and the data throughput arbitrator do not work at the same time.
12. The supercomputing task execution method according to claim 7, characterized in that: After receiving the supercomputing task execution requests issued in batches by the workstations according to application requirements, the supercomputing task execution method further includes: The middleware places the supercomputing task execution request into the execution queue, and uses the middleware to screen appropriate computing nodes according to the resource polling table and the resource arbitrator to execute the corresponding supercomputing task.
13. A workstation, characterized in that: The workstation includes a control device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the supercomputing task execution method according to any one of claims 1 to 6.
14. A server, characterized in that: The server includes a control device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the supercomputing task execution method according to any one of claims 7 to 12.
15. A machine-readable storage medium, characterized in that The machine-readable storage medium stores instructions, which enable the machine to execute the supercomputing task execution method according to any one of claims 1 to 6 or the supercomputing task execution method according to any one of claims 7 to 12.
16. A multi-data center cluster system, characterized in that: The multi-data center cluster system includes a plurality of clusters, the server according to claim 14 electrically connected to each of the plurality of clusters, and the workstation according to claim 13 electrically connected to the server. Each of the clusters includes multiple computing nodes.
Citation Information
Patent Citations
Speed-measuring resource dynamic distributing method and system for network speed-measuring system
CN101068171A
Cloud scheduling method of supercomputing resources, cloud scheduling center and system
CN109951558A
High-availability distributed concurrent task scheduling system and method
CN116010079A
Dynamic priority weighted polling arbitration method and arbiter
CN116627870A
Method and device for accessing supercomputing cluster to computing power network
CN117827433A