A storage service resource management method for high performance computing

By adopting a storage service resource management method with a hierarchical management structure on the high-performance computing platform, and using multi-process concurrency to perform query and scheduling tasks, the problem of low mapping relationship query and scheduling efficiency between the computing node and the storage service resource is solved, and efficient and scalable storage service resource management is achieved.

CN114217914BActive Publication Date: 2025-06-06JIANGNAN INST OF COMPUTING TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110387037.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-12
Publication Date
2025-06-06
Estimated Expiration
2041-04-12

AI Technical Summary

Technical Problem

In high-performance computing platforms, the mapping relationship query and scheduling between the computing node and the storage service resource is inefficient, especially when the forwarding node fails or needs to temporarily improve I/O performance, management is difficult and inefficient.

Method used

The storage service resource management method based on a hierarchical management structure is adopted, including the management node layer, the CE node layer and the computing node layer. Through mapping queries from the computing node to the storage service resource, mapping queries from the storage service resource to the computing node and storage service resource scheduling, query and scheduling tasks are performed using multi-process concurrency.

Benefits of technology

It realizes efficient storage service resource management, rapid query and scheduling, improves the scalability and universality of the computing platform, and solves the problem of rapid query and scheduling of storage service resources and computing nodes in high-performance computing platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114217914B_ABST
    Figure CN114217914B_ABST
Patent Text Reader

Abstract

The present invention discloses a storage service resource management method for high-performance computing, including query of mapping from computing node to storage service resource, query of mapping from storage service resource to computing node and scheduling of storage service resource; the management node is used to assign query tasks to designated CE nodes, and is also used to select scheduling strategies and calculate mapping relationships, and dispatch scheduling tasks to designated CE nodes; the CE node layer is used to log in to multiple computing nodes in a multi-process manner on the CE node, execute specific query tasks, and is also used to log in to forwarding nodes in a multi-process manner on the CE node, and then obtain specific mapping information on the forwarding node and execute specific scheduling tasks; the computing node layer is a usage layer of storage service resources. The present invention solves the problem of fast query and scheduling of storage service resources and computing nodes, is fast and efficient, and has strong scalability and versatility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a storage service resource management method oriented to high performance computing, and belongs to the field of high performance computing. Background Art

[0002] In the field of high-performance computing, as the computing performance of high-performance computers continues to improve, the storage scale is also expanding. At present, computing performance is mainly improved by the continuous expansion of computing resources, while the storage system adopts a three-layer forwarding architecture of computing nodes-forwarding nodes-global storage. By increasing the number of forwarding nodes, the pressure on the underlying distributed storage is effectively alleviated, and the storage scale is expanded.

[0003] In P-class high-performance computers, the number of computing nodes has reached tens of thousands, and the number of forwarding nodes has reached hundreds. In some high-performance computing platforms, a fixed service relationship is formed between computing nodes and forwarding nodes in a static mapping manner. In the following two scenarios, the disadvantages of the above fixed mapping relationship are particularly prominent: First, when a forwarding node fails, its corresponding computing node file system will be unavailable. In order to ensure that computing resources are not wasted, other forwarding nodes must be used instead. However, remapping has poor operability, and the existing mapping relationship changes, which will lead to confusion in the mapping relationship and increase the difficulty of management exponentially. Second, in order to temporarily improve the I / O performance of user applications, one of the most direct methods is to allocate more storage service resources to the computing nodes in the user queue. However, it is difficult to re-establish the mapping relationship between computing nodes and storage service resources.

[0004] In ultra-large-scale environments, there is no efficient mapping query method, and storage service resource scheduling modifies mapping relationships one by one in a single process, which has very low execution efficiency. Summary of the invention

[0005] The purpose of the present invention is to provide a storage service resource management method for high performance computing to solve the problems of mapping relationship query and storage service resource scheduling between computing nodes and storage service resources in a high performance computing platform.

[0006] To achieve the above-mentioned purpose, the technical solution adopted by the present invention is: to provide a storage service resource management method for high-performance computing, based on a hierarchical management structure consisting of a management node layer, a CE node layer and a computing node layer, including computing node to storage service resource mapping query, storage service resource to computing node mapping query and storage service resource scheduling;

[0007] The management node layer is used to group the computing nodes to be queried and format the query results, and is also used to assign the query task to the specified CE node, and is also used to select the scheduling strategy and calculate the mapping relationship, and send the scheduling task to the specified CE node;

[0008] The CE node layer is used to log in to multiple computing nodes in a multi-process manner on the CE node to perform specific query tasks, and is also used to log in to the forwarding node in a multi-process manner on the CE node, and then obtain specific mapping information on the forwarding node, and is also used to log in to the computing node in a multi-process manner on the CE node to perform specific scheduling tasks;

[0009] The computing node layer is a usage layer for storage service resources;

[0010] The query of the computing node to storage service resource mapping includes the following steps:

[0011] S11, the management node groups the computing nodes to be queried according to the principle of uniform distribution, and then sends the allocated computing nodes to the designated CE nodes respectively;

[0012] S12, the CE node obtains the computing node to be queried from the management node, and the CE node immediately sends the query task to the assigned computing node;

[0013] S13, the CE node receives the query result sent back by the computing node and feeds it back to the management node;

[0014] S14, the management node formats and outputs the query results, thereby completing the query task of the computing node storage service resources;

[0015] The query of mapping storage service resources to computing nodes includes the following steps:

[0016] S21, after the management node sends a query instruction to the CE node, it logs in to the forwarding node concurrently on multiple CE nodes;

[0017] S22. On the forwarding node, use the netstat command to obtain the established TCP connection.

[0018] S23, filtering the IP address and port number of the computing node corresponding to the storage service resource according to the port number specified in the query instruction of the management node;

[0019] S24, according to the naming rules of the IP addresses of the computing nodes, convert the IP addresses into computing node numbers, feed back to the management node, and format and output them at the management node;

[0020] The storage service resource scheduling comprises the following steps:

[0021] S31, the management node selects a scheduling strategy and calculates the mapping relationship between storage service resources and computing nodes;

[0022] S32, dispatching the scheduling task to the designated CE node;

[0023] S33. Log in to the computing node in a multi-process manner on the CE node to execute specific scheduling tasks.

[0024] The further improved scheme in the above technical scheme is as follows:

[0025] 1. In the above solution, the CE node can reuse the forwarding node or can be replaced by other independent nodes.

[0026] 2. In the above solution, the scheduling strategy described in S31 can be any custom rule, which is formulated according to management requirements.

[0027] Due to the application of the above technical solution, the present invention has the following advantages compared with the prior art:

[0028] The present invention proposes an efficient query and scheduling method, which makes full use of the ideas of layering and concurrency, distributes query and scheduling tasks to multiple sub-control nodes, and implements query and scheduling tasks in a multi-process manner on the sub-control nodes, thereby solving the problem of fast query and scheduling of storage service resources and computing nodes under the layered architecture of some high-performance computing platforms. The method is fast and efficient, and has strong scalability and versatility. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Attached Figure 1 A schematic diagram of a hierarchical storage architecture for high-performance computing;

[0030] Attached Figure 2 This is a schematic diagram of a hierarchical management structure;

[0031] Attached Figure 3 A schematic diagram of the query process for storing service resources to computing nodes;

[0032] Attached Figure 4 The following is a schematic diagram of the scheduling task execution process. DETAILED DESCRIPTION

[0033] Embodiment: The present invention provides a storage service resource management method for high-performance computing, based on a hierarchical management structure consisting of a management node layer, a CE node layer, and a computing node layer, including computing node to storage service resource mapping query, storage service resource to computing node mapping query, and storage service resource scheduling;

[0034] The management node layer is used to group the computing nodes to be queried and format the query results, and is also used to assign the query task to the specified CE node, and is also used to select the scheduling strategy and calculate the mapping relationship, and send the scheduling task to the specified CE node;

[0035] The CE node layer is used to log in to multiple computing nodes in a multi-process manner on the CE node to perform specific query tasks, and is also used to log in to the forwarding node in a multi-process manner on the CE node, and then obtain specific mapping information on the forwarding node, and is also used to log in to the computing node in a multi-process manner on the CE node to perform specific scheduling tasks;

[0036] The computing node layer is a usage layer for storage service resources;

[0037] The query of the computing node to storage service resource mapping includes the following steps:

[0038] S11, the management node groups the computing nodes to be queried according to the principle of uniform distribution, and then sends the allocated computing nodes to the designated CE nodes respectively;

[0039] S12, the CE node obtains the computing node to be queried from the management node, and the CE node immediately sends the query task to the assigned computing node;

[0040] S13, the CE node receives the query result sent back by the computing node and feeds it back to the management node;

[0041] S14, the management node formats and outputs the query results, thereby completing the query task of the computing node storage service resources;

[0042] The query of mapping storage service resources to computing nodes includes the following steps:

[0043] S21, after the management node sends a query command to the CE node, it logs in to the forwarding node on multiple CE nodes concurrently. Here, the CE nodes are selected in order, and the number of CE nodes is determined by the number of forwarding nodes. Generally, the ratio is 1:16, which can also be changed according to actual needs;

[0044] S22. On the forwarding node, use the netstat command to obtain the established TCP connection.

[0045] S23, filtering the IP address and port number of the computing node corresponding to the storage service resource according to the port number specified in the query instruction of the management node;

[0046] S24, according to the naming rules of the IP addresses of the computing nodes, convert the IP addresses into computing node numbers, feed back to the management node, and format and output them at the management node;

[0047] The storage service resource scheduling comprises the following steps:

[0048] S31. The management node selects a scheduling strategy and calculates the mapping relationship between storage service resources and computing nodes. With the mapping relationship, the CE node knows which storage service resources the computing node uses. The CE node is an intermediate execution node and knows what to do after receiving the task.

[0049] S32, dispatching the scheduling task to the designated CE node;

[0050] S33. Log in to the computing node in a multi-process manner on the CE node to execute specific scheduling tasks.

[0051] The CE node may reuse the forwarding node, or may be replaced by other independent nodes.

[0052] The scheduling strategy described in S31 may be any custom rule, which is formulated according to management requirements.

[0053] The further explanation of the above embodiment is as follows:

[0054] Computing node to storage service resource mapping query: Figure 2 As shown in the figure, a layered architecture is adopted. The management node is responsible for grouping the computing nodes to be queried and formatting the query results. The CE nodes (Command Control Nodes) are responsible for specific query tasks. On the CE nodes, multiple processes are used to log in to multiple computing nodes and concurrently obtain specified information. The CE nodes can reuse forwarding nodes or be replaced by other separate nodes.

[0055] The specific workflow is as follows:

[0056] 1) The management node groups the computing nodes to be queried and sends them to the designated CE nodes respectively;

[0057] 2) The CE node sends the query task to the assigned computing node;

[0058] 3) The CE node receives the query results sent back by the computing node and feeds them back to the management node;

[0059] 4) The management node formats and outputs the query results, thereby completing the query task of the computing node storage service resources.

[0060] Query the mapping of storage service resources to computing nodes: The query principle is similar to that of computing nodes querying storage service resources. The management node assigns the query task to the specified CE node, logs in to the forwarding node in a multi-process manner on the CE node, and then obtains specific mapping information on the forwarding node. The specific process is as follows: Figure 3 shown.

[0061] After the management node issues a query command, it logs in to the forwarding node concurrently on multiple CE nodes;

[0062] On the forwarding node, first, use the netstat command to obtain the established TCP connections;

[0063] Secondly, the IP address of the computing node corresponding to the storage service resource is filtered according to the port number specified by the query instruction of the management node;

[0064] Next, according to the naming rules of the IP addresses of the computing nodes, the IP addresses are converted into computing node numbers, fed back to the management node, and formatted and outputted at the management node.

[0065] Storage service resource scheduling: The scheduling principle is similar to the query principle of computing nodes. The management node first selects the scheduling strategy and calculates the mapping relationship (the scheduling strategy can be any custom rule, formulated according to management needs), and then sends the scheduling task to the specified CE node. On the CE node, log in to the computing node in a multi-process manner to perform specific scheduling tasks. The specific process is as follows: Figure 4 shown.

[0066] When adopting the above-mentioned storage service resource management method for high-performance computing, it proposes an efficient query and scheduling method, which makes full use of the ideas of layering and concurrency, distributes query and scheduling tasks to multiple sub-control nodes, and implements query and scheduling tasks in a multi-process manner on the sub-control nodes, solving the problem of fast query and scheduling of storage service resources and computing nodes under the layered architecture of some high-performance computing platforms. It is fast and efficient, and has strong scalability and versatility.

[0067] In order to facilitate a better understanding of the present invention, the terms used in this article are briefly explained below:

[0068] Forwarding node: Figure 1 As shown, it is the node corresponding to the forwarding layer.

[0069] Global file system: provides a unified large-capacity distributed file system.

[0070] Front-end file system: The client of the front-end file system is deployed on the computing node and is responsible for performing operations such as protocol analysis on the IO requests of the user application. The server is deployed on the forwarding node and is responsible for forwarding the IO requests of the computing node to the back-end global file system.

[0071] Storage service resources: The server side of the front-end file system on the forwarding node, which is differentiated by the forwarding node, network type, and port number.

[0072] Mapping: For a computing node to use the global file system, it must be forwarded through the server of the front-end file system, that is, it must establish a connection with a storage service resource. A computing node can only be connected to one storage service resource, and one storage service resource may establish connections with multiple computing nodes.

[0073] Storage Server Resources (SSR): Dynamically change the mapping relationship between computing nodes and storage service resources.

[0074] User queue: a collection of multiple computing nodes.

[0075] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable people familiar with the technology to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit of the present invention should be included in the protection scope of the present invention.

Claims

1. A storage service resource management method for high performance computing. It is characterized in that Based on the hierarchical management structure consisting of management node layer, CE node layer and computing node layer, including computing node to storage service resource mapping query, storage service resource to computing node mapping query and storage service resource scheduling; The management node layer is used to group the computing nodes to be queried and format the query results, and is also used to assign the query task to the specified CE node, and is also used to select the scheduling strategy and calculate the mapping relationship, and send the scheduling task to the specified CE node; The CE node layer is used to log in to multiple computing nodes in a multi-process manner on the CE node to perform specific query tasks, and is also used to log in to the forwarding node in a multi-process manner on the CE node, and then obtain specific mapping information on the forwarding node, and is also used to log in to the computing node in a multi-process manner on the CE node to perform specific scheduling tasks; The computing node layer is a usage layer for storage service resources; The query of the computing node to storage service resource mapping includes the following steps: S11, the management node groups the computing nodes to be queried according to the principle of uniform distribution, and then sends the allocated computing nodes to the designated CE nodes respectively; S12, the CE node obtains the computing node to be queried from the management node, and the CE node immediately sends the query task to the assigned computing node; S13, the CE node receives the query result sent back by the computing node and feeds it back to the management node; S14, the management node formats and outputs the query results, thereby completing the query task of the computing node storage service resources; The query of mapping storage service resources to computing nodes includes the following steps: S21, after the management node sends a query instruction to the CE node, it logs in to the forwarding node concurrently on multiple CE nodes; S22. On the forwarding node, use the netstat command to obtain the established TCP connection. S23, filtering the IP address and port number of the computing node corresponding to the storage service resource according to the port number specified in the query instruction of the management node; S24, according to the naming rules of the IP addresses of the computing nodes, convert the IP addresses into computing node numbers, feed back to the management node, and format and output them at the management node; The storage service resource scheduling comprises the following steps: S31, the management node selects a scheduling strategy and calculates the mapping relationship between storage service resources and computing nodes; S32, dispatching the scheduling task to the designated CE node; S33. Log in to the computing node in a multi-process manner on the CE node to execute specific scheduling tasks.

2. A storage service resource management method for high performance computing according to claim 1, Features: The CE node may reuse the forwarding node, or may be replaced by other independent nodes.

3. A storage service resource management method for high performance computing according to claim 1, Features: The scheduling strategy described in S31 is a custom rule formulated according to management requirements.

Citation Information

Patent Citations

  • Resource distribution method and device

    CN104754740A

  • A container resource management system based on Shenwei architecture

    CN109739640A