Node state detection method and device, equipment and medium

By grouping nodes and two-stage detection processes, the problem of inability to parallelize and fuzzy fault location in RDMA protocol detection is solved, efficient network connectivity detection is achieved, and the deployment efficiency of cloud platform storage clusters is improved.

CN120378331APending Publication Date: 2025-07-25JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510770938.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, RDMA protocol detection cannot detect node status in parallel, resulting in the network connectivity detection before the deployment of cloud platform storage clusters, and it is impossible to accurately judge server or client failure.

Method used

By grouping nodes, dividing them into customer nodes and service nodes, and adopting a two-stage detection process, first filtering out the target node combination, and then performing accurate pairing detection to reduce the number of invalid detections.

Benefits of technology

It significantly shortens the detection time of network connectivity between nodes, improves the detection efficiency and overall deployment efficiency before the deployment of cloud platform storage clusters, especially when the number of nodes is large, it doubles the detection time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378331A_ABST
    Figure CN120378331A_ABST
Patent Text Reader

Abstract

The invention provides a node state detection method and device, equipment and a medium, and the method comprises the steps: carrying out the grouping processing of a plurality of nodes, and obtaining a plurality of client nodes and a plurality of service nodes in a to-be-detected state; pairing the plurality of client nodes and the plurality of service nodes, performing first network detection on a plurality of paired first node combinations, and determining a plurality of target node combinations of which the node states are target states; and when the number of the multiple target nodes of the multiple target node combinations is greater than or equal to a preset number, performing second network detection on multiple second node combinations formed by pairing the multiple target nodes and the multiple first nodes, and determining node states of the multiple first nodes according to a detection result of the second network detection of the multiple second node combinations. The detection frequency of invalid pairing is reduced, the time consumption of network connectivity detection between the nodes is shortened, the node state is accurately determined, and the detection efficiency before deployment of the cloud platform storage cluster and the overall deployment efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computers, and in particular, to a method, apparatus, device, and medium for detecting node status. Background Art

[0002] The Remote Direct Memory Access (RDMA) protocol enables direct access to memory between hosts and is widely used in fields such as high-performance computing, data centers, and cloud computing. When a cloud platform deploys a storage cluster using a data network card that supports the RDMA protocol, it is necessary to detect the network connectivity of each node before deployment to identify network configuration anomalies or network card hardware failures in advance, ensuring that the storage cluster can be deployed normally and maintain stable subsequent functions.

[0003] Currently, in the related art for RDMA protocol detection, one node is selected as the server and another node as the client, and detection is performed by sending packets at both ends using a network connectivity detection command (rping command). However, when the related art detects a network anomaly between two nodes, it is unable to accurately determine whether the problem lies with the server node or the client node. At the same time, the related art uses a mode where a single node serially detects each other node one by one, resulting in a large amount of time consumption for the pre-deployment network connectivity detection during the deployment of the storage cluster, and with the increase in the number of nodes, the detection time consumption shows a significant growth trend. Summary of the Invention

[0004] The present disclosure provides a method, apparatus, device, and medium for detecting node status, aiming at the problems in the related art of being unable to determine node status and the excessive time consumption of network connectivity detection.

[0005] In a first aspect embodiment of the present disclosure, a method for detecting node status is proposed, including: grouping a plurality of nodes to obtain a plurality of client nodes and a plurality of server nodes in a to-be-detected state; pairing the plurality of client nodes and the plurality of server nodes, and performing a first network detection on the paired plurality of first node combinations to determine a plurality of target node combinations with a node status of a target state; when the number of target nodes in the plurality of target node combinations is greater than or equal to a preset number, performing a second network detection on the plurality of second node combinations after pairing the plurality of target nodes and the plurality of first nodes, and determining the node status of the plurality of first nodes according to the detection results of the second network detection of the plurality of second node combinations, where the plurality of first nodes are the nodes other than the plurality of target nodes among the plurality of client nodes and the plurality of server nodes.

[0006] In a second aspect embodiment of the present disclosure, a device for detecting node status is proposed, including:

[0007] A grouping unit for grouping multiple nodes to obtain multiple customer nodes and multiple service nodes in a to-be-detected state;

[0008] A first detection unit for pairing multiple customer nodes and multiple service nodes, and performing a first network detection on multiple first node combinations after pairing to determine multiple target node combinations with node states being target states;

[0009] A second detection unit for, when the number of multiple target nodes in multiple target node combinations is greater than or equal to a preset number, performing a second network detection on multiple second node combinations after pairing multiple target nodes and multiple first nodes, and determining the node states of multiple first nodes according to the detection results of the second network detection of multiple second node combinations, where multiple first nodes are nodes other than multiple target nodes among multiple customer nodes and multiple service nodes.

[0010] An embodiment of the third aspect of the present disclosure provides an electronic device, including: a processor and a memory for storing a computer program that can run on the processor, where the processor is configured to execute the method described in the embodiment of the first aspect of the present disclosure when running the computer program.

[0011] An embodiment of the fourth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the method described in the embodiment of the first aspect of the present disclosure.

[0012] An embodiment of the fifth aspect of the present disclosure provides a computer program product, including a computer program that implements the method described in the embodiment of the first aspect of the present disclosure when executed by a processor.

[0013] In summary, according to a node status detection method provided by the present disclosure, by grouping multiple nodes, a plurality of client nodes and a plurality of service nodes in a to-be-detected state are obtained; the plurality of client nodes and the plurality of service nodes are paired, and a first network detection is performed on the plurality of first node combinations after pairing to determine a plurality of target node combinations with the node status being the target status; when the number of a plurality of target nodes in the plurality of target node combinations is greater than or equal to a preset number, a second network detection is performed on the plurality of second node combinations after pairing the plurality of target nodes and the plurality of first nodes, and according to the detection results of the second network detection of the plurality of second node combinations, the node status of the plurality of first nodes is determined. It realizes the optimization of the randomly disordered pairing detection mode between two nodes in the related art into a two-stage detection process. First, target nodes are screened out in the first stage, and then accurate pairing detection is carried out in the second stage based on the target nodes, effectively reducing the number of detections of invalid pairings, significantly shortening the time-consuming of the network connectivity detection between nodes, especially doubling the detection time when the number of nodes is large, and at the same time being able to accurately determine the node status, significantly improving the detection efficiency and the overall deployment efficiency before the deployment of the cloud platform storage cluster.

[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0016] Figure 1 is a schematic flowchart of a node status detection method provided by an embodiment of the present disclosure;

[0017] Figure 2 is a schematic flowchart of another node status detection method provided by an embodiment of the present disclosure;

[0018] Figure 3 is a schematic diagram of a specific node status detection method provided by an embodiment of the present disclosure;

[0019] Figure 4 is a schematic structural diagram of a node status detection device provided by an embodiment of the present disclosure;

[0020] Figure 5 is a schematic diagram of the hardware composition structure of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The following describes exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0022] The Remote Direct Memory Access (RDMA) protocol is a technology that allows hosts to directly access each other's memory and is widely used in high-performance computing, data centers, cloud computing, and other fields. When a cloud platform deploys a storage cluster using a data network card that supports the RDMA protocol, in order to identify problems such as abnormal network configurations or network card hardware failures in advance, it is usually necessary to detect the network connectivity between each node before deploying the storage cluster to ensure that the storage cluster can be deployed normally as much as possible and to ensure the stability of subsequent cluster functions.

[0023] Currently, the detection method of the RDMA protocol in related technologies usually requires selecting one node as the detection server (Server side) and other nodes as clients (Client side), and sending packets for detection between the Server side and the Client side through a network connectivity detection command (rping command).

[0024] However, the detection of the RDMA protocol in related technologies cannot be detected in parallel. At the same time, when abnormal network connectivity is detected between two nodes, it is impossible to accurately determine whether the problem lies with the Server side or the Client side.

[0025] It can be seen that due to the limitations of the above RDMA protocol detection in related technologies, the current detection used by the cloud platform is to serially detect each node once with a single node and then detect each other node once. This method results in a large amount of time for the deployment of the storage cluster being consumed in the early network connectivity detection stage, especially when there are more nodes, the time consumption is longer.

[0026] To solve the problems of related technologies, embodiments of the present disclosure provide an accelerated node status detection method for RDMA protocol detection. By grouping and numbering nodes according to quantity and role, they are divided into service nodes on the Server side and client nodes on the Client side, ensuring that the role of a single node is unique during the detection process. A two-stage detection method is proposed, preferentially screening a certain proportion of target nodes and completing the detection with the fewest possible detection times. At the same time, the original detection steps are decomposed, and the detection process is divided into multiple steps such as starting the Server-side and Client-side services, two-stage detection, and closing the services, reducing the number of times of frequently starting the Server-side and Client-side service stages, which take the longest time, and improving the detection efficiency. In addition, each detection step is changed from serial to parallel execution, which can significantly reduce the time-consuming of the entire connectivity detection during the deployment stage.

[0027] Embodiments of the present disclosure will be described in detail below.

[0028] As Figure 1 shown, embodiments of the present disclosure provide a node status detection method, including the following steps:

[0029] Step 101, perform grouping processing on multiple nodes to obtain multiple client nodes and multiple service nodes in a to-be-detected state.

[0030] In some embodiments, a node refers to a host (such as a server, storage node) in a cloud platform that supports the RDMA protocol and has an independent network address and hardware resources.

[0031] In some embodiments, the grouping processing refers to dividing nodes into client nodes (Client) and service nodes (Server) according to role, and the role remains fixed during the detection.

[0032] In some embodiments, the to-be-detected state refers to a state in which the node has completed service initialization (such as starting the rping server / client), but has not participated in the detection yet.

[0033] In some embodiments, the present disclosure can evenly divide the nodes into the Client queue (client queue) and the Server queue (service queue), where the nodes in the client queue and the service queue can each account for 50%. The nodes in each queue only assume a single role during the detection (the client nodes in the Client queue only initiate connections, and the service nodes in the Server queue only listen for connections).

[0034] Among them, after grouping the nodes, the present disclosure can further send connectivity detection instructions (i.e., rping commands) to the client nodes and service nodes, so that the client nodes in the Client queue execute rping - c to start the client service, and the service nodes in the Server queue execute rping - s to start the server service. All nodes enter the state to be detected (marked as

Ready

[0035] Step 102: Pair multiple client nodes and multiple service nodes, and perform a first network detection on the multiple first - node combinations after pairing to determine multiple target - node combinations whose node states are target states.

[0036] In some embodiments, a first - node combination refers to a detection unit (Client - Server pair) formed by pairing 1 client node with 1 service node.

[0037] In some embodiments, the target state represents the state where the node combination passes the network connectivity detection (such as "Normal" being normal), that is, both nodes within the node combination are network - connected normally.

[0038] In some embodiments, a target node refers to the general term for the client node and the service node that pass the network connectivity detection (target state) in the first - node combination.

[0039] In some embodiments, when the present disclosure groups multiple nodes, each node needs to be numbered. Then, the client nodes and service nodes can be paired in the order of the numbers (such as C1 - S1, C2 - S2, C3 - S3) to form the first - node combination. Perform a network connectivity detection on each group. If the client node successfully sends and receives data packets to the service node (the target state is "Normal"), then mark this combination as a target - node combination, and the client node and the service node within this combination are target nodes (such as both C1 and S1 are target nodes); if it fails, do not mark it temporarily and keep it in the state to be detected. Among them, a single detection failure may be a temporary link failure (such as network congestion of C2 - S2). There is no rush to mark it as abnormal, but further verification is carried out in the subsequent stage. Implement the screening of target nodes (such as C1 and S1) with target states through the first - stage pairing detection (such as 3 - group detections). These nodes can be used as benchmarks for the second - stage pairing detection in the subsequent detection.

[0040] In some embodiments, the present disclosure may accumulate the number of target nodes. If the accumulated number reaches or exceeds a preset number (e.g., 1 / 3 of the total number of nodes 6, i.e., 2), then step 103 is entered; otherwise, polling continues to pair with other combinations.

[0041] Step 103: When the number of multiple target nodes in multiple combinations of target nodes is greater than or equal to the preset number, perform a second network detection on multiple combinations of second nodes after pairing the multiple target nodes and multiple first nodes, and determine the node status of the multiple first nodes according to the detection results of the second network detection of the multiple combinations of second nodes.

[0042] In the present disclosure, the multiple first nodes are the nodes other than the multiple target nodes among the multiple client nodes and multiple service nodes.

[0043] In some embodiments, the preset number is a preset threshold for the number of target nodes (e.g., 1 / 3 of the total number of nodes), which is used to determine whether to enter the second-stage detection.

[0044] In some embodiments, a combination of second nodes is a new detection unit formed by pairing a target node (with normal network connectivity detection) with the remaining nodes in a pending detection state (first nodes).

[0045] In some embodiments, the target nodes (e.g., C1, S1) are re-paired with the remaining nodes in a pending detection state (first nodes, e.g., C2, C3, S2, S3) to form combinations of second nodes (e.g., C1-S2, C1-S3, S1-C2, S1-C3). For example, using a detected target node (e.g., C1) as a client node to detect other service nodes (S2, S3); or using a detected target node (e.g., S1) as a service node to detect other client nodes (C2, C3).

[0046] In some embodiments, if the detection is successful (e.g., C1-S2 is normal), then mark the node status of S2 as the target status; if the detection fails (e.g., C1-S3 is abnormal), then directly determine S3 as an abnormal node (since C1 has passed the first-stage verification, the fault is located at S3). This realizes the problem of using the target nodes detected from the first nodes as a "trusted benchmark" and directly locating the abnormal nodes (first nodes) when the detection fails, avoiding the problem of "Client / Server fault ambiguity" in the related art. At the same time, directly determine the node status of the first node as the abnormal status after the detection fails, and delete the first node in the abnormal status from its corresponding queue, no longer participating in subsequent pairing, reducing ineffective detections.

[0047] In summary, according to the node status detection method proposed by the present disclosure, by grouping multiple nodes, multiple client nodes and multiple service nodes in a to-be-detected state are obtained; the multiple client nodes and multiple service nodes are paired, and the first network detection is performed on the multiple first node combinations after pairing to determine multiple target node combinations with the node status being the target status; when the number of multiple target nodes in the multiple target node combinations is greater than or equal to the preset number, the second network detection is performed on the multiple second node combinations after pairing the multiple target nodes and the multiple first nodes, and according to the detection results of the second network detection of the multiple second node combinations, the node status of the multiple first nodes is determined. It realizes the optimization of the random and disordered pairing detection mode between two nodes in the related art into a two-stage detection process. First, the target nodes are screened out through the first stage, and then the second stage of precise pairing detection is carried out based on the target nodes, effectively reducing the number of detections of invalid pairings, greatly shortening the time-consuming of the network connectivity detection between nodes, especially when the number of nodes is large, the detection time can be reduced by several times, and at the same time, the node status can be accurately determined, significantly improving the detection efficiency and the overall deployment efficiency before the deployment of the cloud platform storage cluster.

[0048] Figure 2 Further, a flowchart of another node status detection method proposed by the present disclosure is shown. Based on Figure 1 the embodiments shown are further explained, Figure 2 it may include the following steps.

[0049] Step 201, allocate multiple nodes to a client queue and a service queue respectively, to obtain multiple client nodes in the client queue and multiple service nodes in the service queue.

[0050] In some embodiments, the present disclosure may evenly divide the cluster nodes into two independent queues (client queue C and service queue S) according to the quantity, with fixed roles (C only serves as the client, and S only serves as the server).

[0051] Specifically, the present disclosure may evenly divide N nodes into group C and group S through a configuration file or an automated script (for example, when N = 6, C = {C1, C2, C3}, S = {S1, S2, S3}) to avoid the service restart overhead caused by the frequent role switching of nodes in the related art, and the service only needs to be globally started once.

[0052] Step 202, number the multiple client nodes and multiple service nodes to obtain multiple client node numbers of the multiple client nodes and multiple service node numbers of the multiple service nodes.

[0053] In some embodiments, the present disclosure may assign a unique number to each node (such as C1 - Cn, S1 - Sn) for subsequent pairing and polling.

[0054] Specifically, the present disclosure can be numbered in queue order. For example, the customer queue numbers are Client-1 to Client-m, and the service queue numbers are Server-1 to Server-k, to ensure no duplicate combinations during pairing (e.g., C1-S1 is only detected once), reducing ineffective detections.

[0055] Step 203: Based on multiple customer node numbers and multiple service node numbers, send connectivity detection instructions to the multiple customer nodes and multiple service nodes, so that the multiple customer nodes and multiple service nodes are in a state to be detected.

[0056] In some embodiments, the present disclosure can concurrently send connectivity detection instructions to all nodes, enabling the nodes in the C queue and S queue to execute rping–s and rping-c respectively to start the rping server and client services. After the startup is completed, set the active state of the nodes to Ready and enter the state to be detected. This realizes parallel startup of all node services and reduces the detection time consumption.

[0057] Step 204: Based on multiple customer node numbers and multiple service node numbers, pair the multiple customer nodes and multiple service nodes to obtain multiple first node combinations after pairing.

[0058] In some embodiments, the present disclosure can pair the nodes in the C queue and S queue in order of number (e.g., C1-S1, C2-S2), obtaining multiple first node combinations [(C1,S1),(C2,S2),...]. Each first node combination corresponds to one network connectivity detection. This realizes ordered pairing and avoids repeated detections caused by random combinations, improving the detection coverage rate.

[0059] Step 205: Perform a first network detection on the multiple first node combinations, determine the network connectivity between the customer node and the service node in each first node combination, and determine the first node combinations with normal network connectivity as the target node combinations with the node state being the target state.

[0060] In some embodiments, the present disclosure can preferentially pair nodes with the same number based on the numbers of the nodes in their respective queues (e.g., S1, S2... C1, C2...), that is, pair S1 with C1, S2 with C2, and so on, to form a first node combination (e.g., <S1,C1> combination). Then initiate an RDMA connectivity detection for each pair of combinations. If the customer node in each first node combination successfully sends and receives an RDMA data packet to its paired service node (e.g., the rping return loss rate is 0 or lower than a preset threshold), then mark the node states of both the service node and the customer node in this combination as the target state, that is, both are marked as Normal, and release the active states of the two nodes back to Ready (allowing them to participate in subsequent rounds of pairing detections).

[0061] In some embodiments, if the detection fails (such as timeout, connection refused, or too high packet loss rate), since it is impossible to immediately determine whether the failure is in the client node or the service node in the first node combination of the current exception, the node status of the nodes is not directly marked as abnormal. Instead, the active status of the two nodes is also released to Ready, so that they enter the next round of pairing pool and wait to be recombined with other nodes for detection (for example, S1 may be paired with C2 in the next round, and C1 with S2).

[0062] Step 206, when the number of multiple target nodes is greater than or equal to a preset number, determine the queue of the multiple target nodes.

[0063] In some embodiments, the present disclosure can accumulate the number of target nodes (such as both C1 and S1 are target nodes). If it exceeds a preset threshold (such as 1 / 3 of the total number of nodes), then enter the second detection stage.

[0064] Specifically, continuously perform multiple rounds of numbered pairing detection until the total number of nodes with the target status (i.e., Normal nodes) reaches or exceeds a preset number (for example, 1 / 3 of the total number of cluster nodes. When the total number of nodes is 30, the number of Normal nodes ≥ 10), or reaches a preset maximum polling number, then terminate the first-stage detection and enter the next stage. This realizes quickly screening out the node pairs with the target status, provides a credible benchmark for the second stage, and avoids excessive detection.

[0065] Step 207, based on the queue of the multiple target nodes, determine multiple first nodes for network pairing with the multiple target nodes.

[0066] In some embodiments, the present disclosure selects, from the peer queue, the first nodes that are not marked as target nodes based on the queue of the target nodes (client queue or service queue), and performs cross pairing with the target nodes to form the detection units (second node combinations) of the second stage. That is, use the target nodes (such as the target nodes in the C queue) to pair with the remaining nodes in the other queue (S queue). For example, if the queue of the target nodes is the client queue, then determine the first nodes for network pairing with the target nodes from the service queue. If the queue of the target nodes is the service queue, then determine the first nodes for network pairing with the target nodes from the client queue. This realizes using the verified target nodes as a benchmark to quickly locate abnormal nodes.

[0067] Step 208, pair the multiple first nodes and the multiple target nodes to obtain multiple second node combinations after pairing.

[0068] In some embodiments, during the second-phase detection, the present disclosure may traverse the set of target nodes. For each target node, according to the queue it belongs to (customer queue / service queue), the first nodes that did not participate in the first-phase detection or whose detection results did not meet the standards are filtered out from the peer queue (i.e., the nodes in the peer queue and with the active state being Ready state); according to the node numbers, the target nodes are paired with the first nodes one by one to form a second node combination in the form of <target node, first node>, ensuring that each first node only participates in one pairing, avoiding repeated detection and greatly reducing the number of detections.

[0069] Step 209: Perform a second network detection on multiple second node combinations to determine the network connectivity between the target node and the first node in each second node combination.

[0070] In some embodiments, after the pairing of the second node combinations is completed, an RDMA network connectivity detection is performed on the target node and the first node in each combination. Specifically, taking the target node (whose connectivity has been verified to be normal) as the reference endpoint, according to whether it is a customer node or a service node, an RDMA data packet is initiated or received from the first node, and indicators such as the success rate, latency, and packet loss rate of the data packet transmission are counted through the rping command to determine whether the network connectivity between the two nodes is normal. This detection process relies on the trusted state of the target node to accurately locate the connectivity problem of the first node, ensuring the reliability of the detection result and the accuracy of the fault location.

[0071] For example, if the target node is a customer node (C), then a connection is initiated to the first node (S) through rping -c. If successful, it proves that the network connectivity between the customer node (C) and the first node (S) is normal; if the target node is a service node (S), then wait for the first node (C) to connect through rping -c. If successful, it proves that the network connectivity between the customer node (S) and the first node (C) is normal.

[0072] Step 210: Determine the node state of the first node based on the network connectivity between the target node and the first node in each second node combination.

[0073] In some embodiments, when the network connectivity between the first node and the target node is normal, determine the node state of the first node as the target state; when the network connectivity between the first node and the target node is abnormal, determine the node state of the first node as the abnormal state.

[0074] In the second-phase detection, based on the network connectivity detection results between the target nodes and the first nodes within each second-node combination, determine the status of the first nodes: If the connectivity between the two is normal (such as successful data packet transmission and reception, and the packet loss rate is lower than the threshold), mark the first nodes as the target status, and release the active status back to Ready, waiting to configure and detect with other nodes; if the connectivity is abnormal (such as connection timeout, no response, or high packet loss rate), directly mark the first nodes as the abnormal status. This determination logic relies on the credibility of the target nodes (verified through previous detections), attributes the abnormal connectivity to the first nodes, and achieves precise positioning of their statuses.

[0075] Among them, when it is determined that the node status of the first nodes is the abnormal status, the first nodes with the abnormal status (nodes in the Abnormal state) can be directly removed from their corresponding queues and no longer used, so as to avoid repeated detections and reduce ineffective operations.

[0076] Step 211: Based on multiple client node numbers and multiple service node numbers, send detection shutdown instructions to multiple client nodes and multiple service nodes, so that the multiple client nodes and multiple service nodes are in the detection shutdown state.

[0077] In some embodiments, when the status detection of all nodes is completed, concurrently shut down the rping services of all nodes, summarize and display the analysis results. When the status detection of all nodes is completed, construct a concurrent execution sequence based on the client node numbers and service node numbers, and send detection shutdown instructions to all nodes in batches, so that each node terminates the RDMA connectivity detection service and enters the detection shutdown state. Subsequently, the system automatically summarizes the detection result data of each node (including indicators such as Normal / Abnormal status, packet loss rate, response time, etc.), generates a visual analysis report, and displays the overall result of the network connectivity detection and the distribution of abnormal nodes.

[0078] In summary, the node status detection method proposed in this disclosure targets the RDMA protocol deployment scenario in the cloud platform, breaks through the efficiency bottleneck of serial detection in related technologies, and realizes the parallel transformation of the detection process through node grouping, fixed roles, and two-phase detection strategies. Specifically, the nodes are pre-divided into client nodes and service nodes and their roles are fixed to avoid frequent start and stop of the RDMA service; through the first phase, the target nodes are screened out as the credible benchmark, and in the second phase, the remaining nodes are directionally detected with them as the core, significantly reducing the number of ineffective pairings. It not only solves the problems of inability to be parallel and fuzzy fault location in RDMA protocol detection, but also greatly shortens the time-consuming of network connectivity detection through the concurrent detection mechanism, especially doubling the detection efficiency in the scenario of a large number of nodes, and finally significantly optimizing the deployment efficiency of the cloud platform.

[0079] Based on Figures 1 to 2 the embodiments shown, such asFigure 3 As shown, the present disclosure provides a schematic diagram of a specific node status detection method.

[0080] Referring to Figure 3 , the present disclosure provides an accelerated node connectivity detection process. Taking the batch deployment of an 8-node storage cluster as an example, the 8 nodes are grouped and numbered. Nodes 1, 2, 3, and 4 are added to the Server node group (service queue) and numbered as S-Node1, S-Node2, S-Node3, and S-Node4; nodes 5, 6, 7, and 8 are added to the Client node group (customer queue) and numbered as C-Node1, C-Node2, C-Node3, and C-Node4.

[0081] Start the RDMA detection environment for all nodes and prepare the status uniformly, that is, traverse all nodes in the service queue and the customer queue, and start the RDMA network detection service through the connectivity detection instructions (rping –s (Server side starts listening), rping –c (Client side starts active connection) commands); and mark the active status of all nodes as Ready and enter the to-be-detected status.

[0082] Start the first-stage detection. Poll the nodes in the customer queue and the service queue in sequence according to the numbers, pair them up in pairs to get the first node combination and detect the first node combination. If the first node combination passes the detection (the network connectivity is normal), mark the status of both nodes in the first node combination as Normal, that is, the node status is the target status; if the first node combination detects an abnormality (the network connectivity fails), put the nodes in the first node combination back into the original queue and wait for other nodes to pair and detect later; when the number of nodes with the status of Normal (that is, the number of target nodes) reaches 1 / 3 of the total number of nodes (8) (that is, ≥3), the first-stage detection ends, and the target nodes are retained as the trusted benchmark.

[0083] Start the second phase of detection, using the target node of the first phase as a benchmark, efficiently verify the status of the remaining nodes, and accurately locate the anomaly. That is, prioritize the target nodes with a node status of Normal in the service queue and the customer queue, and pair them with the nodes in the peer queue that are in the state to be detected (such as S-Node1 and C-Node2, S-Node2 and C-Node3, etc.) to obtain the second node combination; if the second node combination passes the detection (that is, the network connectivity is normal), mark the peer node status as Normal (that is, the target status), and release it back to the original queue to participate in subsequent pairing (expand the benchmark node pool); if the second node combination detects an abnormality (that is, the network connectivity is abnormal), mark the peer node status as Abnormal, and remove it from the queue to which it belongs (terminate the node detection process); repeat the above pairing, detection, and marking process until the node status (Normal / Abnormal) of all nodes in the service queue and the customer queue are determined.

[0084] After the full node status detection is completed, all nodes in the service queue and customer queue are traversed, and the rping service shutdown operation is executed to release network resources and end the entire detection process.

[0085] In order to implement the node status detection method provided by the embodiment of the present disclosure, the embodiment of the present disclosure also provides a node status detection device, such as Figure 4 As shown, the node status detection device 400 includes:

[0086] A grouping unit 410 is used to group multiple nodes to obtain multiple client nodes and multiple service nodes in a state to be detected;

[0087] A first detection unit 420 is used to pair the plurality of client nodes with the plurality of service nodes, and perform a first network detection on the paired plurality of first node combinations to determine a plurality of target node combinations whose node states are target states;

[0088] The second detection unit 430 is used to perform a second network detection on a plurality of second node combinations that are formed by pairing the plurality of target nodes with the plurality of first nodes when the number of the plurality of target nodes in the plurality of target node combinations is greater than or equal to a preset number, and determine the node status of the plurality of first nodes based on the detection result of the second network detection of the plurality of second node combinations, wherein the plurality of first nodes are nodes other than the plurality of target nodes among the plurality of client nodes and the plurality of service nodes.

[0089] In some embodiments, the grouping unit 410 is configured to: respectively allocate a plurality of nodes to a customer queue and a service queue to obtain a plurality of customer nodes in the customer queue and a plurality of service nodes in the service queue; number the plurality of customer nodes and the plurality of service nodes to obtain a plurality of customer node numbers of the plurality of customer nodes and a plurality of service node numbers of the plurality of service nodes; and based on the plurality of customer node numbers and the plurality of service node numbers, send a connectivity detection instruction to the plurality of customer nodes and the plurality of service nodes, so that the plurality of customer nodes and the plurality of service nodes are in a state to be detected.

[0090] In some embodiments, the first detection unit 420 is configured to: based on the plurality of customer node numbers and the plurality of service node numbers, pair the plurality of customer nodes and the plurality of service nodes to obtain a plurality of first node combinations after pairing; perform a first network detection on the plurality of first node combinations to determine the network connectivity between the customer node and the service node in each first node combination, and determine the first node combination with normal network connectivity as a target node combination with a target state of the node state.

[0091] In some embodiments, the second detection unit 430 is configured to: when the number of the plurality of target nodes is greater than or equal to a preset number, determine the queues of the plurality of target nodes; based on the queues of the plurality of target nodes, determine a plurality of first nodes that are network-paired with the plurality of target nodes; pair the plurality of first nodes and the plurality of target nodes to obtain a plurality of second node combinations after pairing; perform a second network detection on the plurality of second node combinations to determine the network connectivity between the target node and the first node in each second node combination; when the network connectivity between the first node and the target node is normal, determine that the node state of the first node is the target state; and when the network connectivity between the first node and the target node is abnormal, determine that the node state of the first node is the abnormal state.

[0092] In some embodiments, the second detection unit 430 is configured to: if the queue of the target node is the customer queue, determine a first node that is network-paired with the target node from the service queue.

[0093] In some embodiments, the second detection unit 430 is configured to: if the queue of the target node is the service queue, determine a first node that is network-paired with the target node from the customer queue.

[0094] In some embodiments, the second detection unit 430 is configured to: after determining the node states of the plurality of first nodes according to the detection results of the second network detection of the plurality of second node combinations, based on the plurality of customer node numbers and the plurality of service node numbers, send a detection shutdown instruction to the plurality of customer nodes and the plurality of service nodes, so that the plurality of customer nodes and the plurality of service nodes are in a detection shutdown state.

[0095] It should be noted that when the node status detection device provided in the above embodiment performs node status detection, only the division of the above program modules is used for illustration. In actual applications, the above processing can be allocated to different program modules according to needs, that is, the internal structure of the node status detection device is divided into different program modules to complete all or part of the processing described above. In addition, the node status detection device provided in the above embodiment and the node status detection method embodiment provided in the present disclosure belong to the same concept. For the specific implementation process, please refer to the method embodiment, which will not be elaborated here.

[0096] Figure 5 FIG. is a schematic diagram of the hardware composition structure of the electronic device provided in the embodiment of the present disclosure. As Figure 5 shown, the electronic device 500 includes at least one processor 502; and a memory 501 communicatively connected to the at least one processor 502; wherein, the memory 501 stores instructions executable by the at least one processor 502, and the instructions are executed by the at least one processor 502 to implement the steps of the node status detection method provided in the embodiment of the present disclosure.

[0097] Optionally, the electronic device may specifically be the node status detection device of the embodiment of the present disclosure, and the electronic device can implement the corresponding processes implemented by the node status detection device in each method of the embodiment of the present disclosure. For the sake of brevity, it will not be elaborated here.

[0098] It can be understood that the electronic device further includes a communication interface 503. Each component in the electronic device is coupled together through a bus system 504. It can be understood that the bus system 504 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 504 further includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 5 all the various buses are labeled as the bus system 504.

[0099] It can be understood that the memory 501 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), direct rambus random access memory (DRRAM).The memory 501 described in the embodiments of the present disclosure is intended to include but not limited to these and any other suitable types of memories.

[0100] The method disclosed in the embodiments of the present disclosure above can be applied to or implemented by the processor 502. The processor 502 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 502 or the instructions in the form of software. The above-mentioned processor 502 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 502 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present disclosure, it can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the memory 501. The processor 502 reads the information in the memory 501 and combines its hardware to complete the steps of the foregoing method.

[0101] In an exemplary embodiment, the electronic device may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic components, and is used to execute the foregoing method.

[0102] The embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to execute the steps of the node state detection method described in the embodiments of the present disclosure when executed.

[0103] The embodiments of the present disclosure also provide a computer program product, including a computer program, and the computer program implements the steps of the node state detection method described in the embodiments of the present disclosure when executed by the processor.

[0104] Optionally, the computer-readable storage medium can be applied to the node state detection device in the embodiments of the present disclosure, and the computer instructions cause the computer to execute the corresponding processes implemented by the node state detection device in each method of the embodiments of the present disclosure. For the sake of brevity, it will not be elaborated here.

[0105] In several embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.

[0106] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0107] In addition, each functional unit in the embodiments of the present disclosure can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.

[0108] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical disks and other various media that can store program codes.

[0109] Alternatively, if the above-mentioned integrated units of the present disclosure are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present disclosure essentially or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in various embodiments of the present disclosure. And the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical disks and other various media that can store program codes.

[0110] As described above, it is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims described above.

Claims

1. A node state detection method, characterized in that, Including: Group the multiple nodes to obtain multiple customer nodes and multiple service nodes in a state to be detected; Pair the multiple customer nodes and the multiple service nodes, and perform a first network detection on the multiple first node combinations after pairing to determine multiple target node combinations with the node state being the target state; When the number of multiple target nodes in the multiple target node combinations is greater than or equal to the preset number, perform a second network detection on the multiple second node combinations after pairing the multiple target nodes and the multiple first nodes, and determine the node states of the multiple first nodes according to the detection results of the second network detection of the multiple second node combinations, where the multiple first nodes are the nodes other than the multiple target nodes among the multiple customer nodes and the multiple service nodes.

2. The method according to claim 1, characterized in that, The grouping process of the multiple nodes to obtain multiple customer nodes and multiple service nodes in a state to be detected includes: Allocate the multiple nodes to a customer queue and a service queue respectively to obtain multiple customer nodes in the customer queue and multiple service nodes in the service queue; Number the multiple customer nodes and the multiple service nodes to obtain multiple customer node numbers of the multiple customer nodes and multiple service node numbers of the multiple service nodes; Based on the multiple customer node numbers and the multiple service node numbers, send a connectivity detection instruction to the multiple customer nodes and the multiple service nodes to make the multiple customer nodes and the multiple service nodes in a state to be detected.

3. The method according to claim 2, wherein The pairing of the multiple customer nodes and the multiple service nodes, and the first network detection of the multiple first node combinations after pairing to determine multiple target node combinations with the node state being the target state includes: Based on the multiple customer node numbers and the multiple service node numbers, pair the multiple customer nodes and the multiple service nodes to obtain multiple first node combinations after pairing; Perform a first network detection on the multiple first node combinations to determine the network connectivity between the customer node and the service node in each first node combination, and determine the first node combination with normal network connectivity as the target node combination with the node state being the target state.

4. The method according to claim 1, wherein When the number of multiple target nodes in the multiple target node combinations is greater than or equal to the preset number, perform a second network detection on the multiple second node combinations after pairing the multiple target nodes and the multiple first nodes, and determine the node states of the multiple first nodes according to the detection results of the second network detection of the multiple second node combinations includes: When the number of the multiple target nodes is greater than or equal to the preset number, determine the queues of the multiple target nodes; Based on the queues of the multiple target nodes, determine multiple first nodes that are network-paired with the multiple target nodes; Pair the multiple first nodes and the multiple target nodes to obtain multiple second node combinations after pairing; Perform a second network detection on the multiple second node combinations to determine the network connectivity between the target node and the first node in each second node combination; When the network connection between the first node and the target node is normal, determine the node state of the first node as the target state; When the network connection between the first node and the target node is abnormal, determine the node state of the first node as the abnormal state.

5. The method according to claim 4, wherein The determining of the multiple first nodes for network pairing with the multiple target nodes based on the queues of the multiple target nodes includes: If the queue of the target node is a customer queue, determine the first node for network pairing with the target node from the service queue.

6. The method according to claim 4, wherein The determining of the multiple first nodes for network pairing with the multiple target nodes based on the queues of the multiple target nodes includes: If the queue of the target node is a service queue, determine the first node for network pairing with the target node from the customer queue.

7. The method according to claim 2, characterized in that After determining the node states of the multiple first nodes according to the detection results of the second network detection of the multiple second node combinations, the method includes: Based on the multiple customer node numbers and the multiple service node numbers, send a detection shutdown instruction to the multiple customer nodes and the multiple service nodes, so that the multiple customer nodes and the multiple service nodes are in the detection shutdown state.

8. A node status detection device, characterized in that, The device includes: A grouping unit, configured to perform grouping processing on multiple nodes to obtain multiple customer nodes and multiple service nodes in a to-be-detected state; A first detection unit, configured to pair the multiple customer nodes and the multiple service nodes, and perform first network detection on the paired multiple first node combinations to determine multiple target node combinations with the node state being the target state; A second detection unit, configured to, when the number of multiple target nodes in the multiple target node combinations is greater than or equal to a preset number, perform second network detection on the multiple second node combinations after pairing the multiple target nodes and the multiple first nodes, and determine the node states of the multiple first nodes according to the detection results of the second network detection of the multiple second node combinations, where the multiple first nodes are the nodes other than the multiple target nodes among the multiple customer nodes and the multiple service nodes.

9. An electronic device, characterized in that, Includes: A processor and a memory for storing a computer program that can run on the processor, wherein, when the processor is used to run the computer program, it executes the node state detection method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instruction is used to cause the computer to execute the node state detection method according to any one of claims 1-7.

Citation Information

Cited By

  • Cluster system-oriented performance detection method and device, electronic equipment and medium

    CN121509282A

  • Performance detection method and device for cluster system, electronic equipment and medium

    CN121509282B