Cluster management method, target cluster system, electronic equipment and storage medium
By configuring the RDMA network card for storage nodes in the Redis cluster and using the RDMA handshake protocol, data is directly transmitted between storage nodes, solving the problems of transmission delay and low efficiency during data migration, and achieving efficient data migration and real-time monitoring.
Patent Information
- Application Number
- CN202510357068.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-01
AI Technical Summary
In Redis cluster, data migration efficiency is affected due to frequent switching between kernel and user states, which affects data migration efficiency.
Configure an RDMA network card for storage nodes in the Redis cluster, use the RDMA handshake protocol to establish a network connection, directly transmit data between storage nodes, and monitor transmission progress through the cluster management nodes, and adjust data migration strategies.
It improves the data migration efficiency between storage nodes, reduces the invalid overhead during transmission, and realizes real-time monitoring and policy adjustment of the data migration process.
Smart Images

Figure CN120238562A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of cluster management, and particularly to a cluster management method, a target cluster system, an electronic device, and a storage medium. Background Art
[0002] At present, with the booming development of big data and cloud computing, the amount of data has increased explosively. As a solution to cope with massive data and high-concurrency requests, in a Redis cluster, each Redis node depends on network sockets for communication, and data transmission requires multiple switches between the kernel mode and the user mode, which will introduce additional overhead and significantly increase the transmission delay. For example, during data migration, a large amount of data needs to be transmitted between different storage nodes, and this inefficient transmission method will cause a significant extension of the data migration time. Summary of the Invention
[0003] This application provides a cluster management method, a target cluster system, an electronic device, and a storage medium, which can improve the data migration efficiency between each storage node. The technical solution is as follows:
[0004] According to one aspect of this application, a cluster management method is provided. The method is applied to a target cluster system, and the target cluster system includes a cluster management node and multiple storage nodes. Each storage node is configured with an RDMA network card. The method includes:
[0005] When there is a data migration requirement between a source node and a target node among the multiple storage nodes, the cluster management node is used to control the source node and the target node to establish a network connection based on the RDMA handshake protocol;
[0006] The source node is used to send target migration data to the memory of the target node based on the RDMA network card;
[0007] The target node is used to receive the target migration data sent by the source node based on the RDMA network card;
[0008] During the transmission of the target migration data, the cluster management node is used to receive the amount of transmitted data and the remaining data of the target migration data fed back by the source node and the target node;
[0009] The cluster management node is used to adjust the data migration strategy between the source node and the target node based on the amount of transmitted data and the remaining data.
[0010] According to another aspect of the present application, a target cluster system is provided. The target cluster system includes a cluster management node and multiple storage nodes, and each storage node is configured with an RDMA network card:
[0011] The cluster management node is used to control the source node and the target node to establish a network connection based on the RDMA handshake protocol when there is a data migration requirement between the source node and the target node among the multiple storage nodes;
[0012] The source node is used to send target migration data to the memory of the target node based on the RDMA network card;
[0013] The target node is used to receive the target migration data sent by the source node based on the RDMA network card;
[0014] The cluster management node is further used to receive the transmitted data volume and the remaining data volume of the target migration data fed back by the source node and the target node during the transmission of the target migration data;
[0015] The cluster management node is further used to adjust the target migration strategy between the source node and the target node based on the transmitted data volume and the remaining data volume.
[0016] According to one aspect of the present application, an electronic device is provided, including: a processor and a memory storing a program. The program includes instructions that, when executed by the processor, cause the processor to execute the cluster management method as described above.
[0017] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided. The computer instructions are used to cause the computer to execute the cluster management method as described above.
[0018] According to another aspect of the present application, a computer program product is provided. The computer program product includes computer instructions stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the above-mentioned cluster management method.
[0019] The beneficial effects brought by the technical solution provided by the embodiments of the present application at least include:
[0020] By configuring RDMA network cards for each storage node in the target cluster management system, when data transfer or data migration is performed between storage nodes, the source node can directly transfer the target migration data to the memory of the target node through the RDMA network card without the participation of the operating systems of the source node and the target node, and correspondingly, there is no need to switch between the user mode and the kernel mode, which can improve the data migration efficiency between storage nodes. Moreover, a transmission progress monitoring and feedback function is also added, so that information such as the amount of transferred data and the remaining data during the data migration process can be fed back to the cluster management node, enabling the cluster management node to adjust the data migration strategy in a timely manner according to the feedback information and ensuring the data migration efficiency between storage nodes. Brief Description of the Drawings
[0021] In the following description of exemplary embodiments with reference to the drawings, more details, features, and advantages of the present application are disclosed. In the drawings:
[0022] Figure 1 A flowchart of a cluster management method according to an exemplary embodiment of the present application is shown;
[0023] Figure 2 A flowchart of another cluster management method according to an exemplary embodiment of the present application is shown;
[0024] Figure 3 A flowchart of another cluster management method according to an exemplary embodiment of the present application is shown;
[0025] Figure 4 A schematic structural diagram of a target cluster system provided by an embodiment of the present application;
[0026] Figure 5 A block diagram of the structure of an exemplary electronic device capable of implementing the embodiments of the present application is shown. Detailed Description of the Embodiments
[0027] The embodiments of the present application will be described in more detail below with reference to the drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.
[0028] It should be understood that the steps recited in the method embodiments of the present application can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this regard.
[0029] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first", "second", etc. mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. It should be noted that the modifications of "one" and "multiple" mentioned in this application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more". The names of the messages or information exchanged between multiple devices in the embodiments of this application are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0030] The solution of the present invention will be described below with reference to the accompanying drawings. The technical solution provided by the embodiments of the present invention will be described in detail through specific embodiments and their application scenarios.
[0031] Please refer to Figure 1 , which shows a flowchart of a cluster management method according to an exemplary embodiment of the present application. This method will be described by taking its application to a target cluster system as an example. As Figure 1 shown, the method includes:
[0032] Step 101, when there is a data migration requirement between a source node and a target node among multiple storage nodes, the cluster management node controls the source node and the target node to establish a network connection based on the RDMA handshake protocol.
[0033] Among them, the target cluster system includes a cluster management node (CM, Cluster Manager) and multiple storage nodes. The cluster management node, as the core of the target cluster system, is responsible for the full life cycle management of the target cluster system; in terms of hardware, high-performance servers are adopted, equipped with multi-core high-speed CPUs, large-capacity memories, and high-speed storage devices; at the software level, multiple key modules are integrated: a node management module, a configuration management module, and a fault detection and recovery module. Among them, the node management module is used to handle the joining, exiting, and status monitoring of storage nodes; when a new storage node joins, authentication is performed, and a unique node ID is assigned, such as using an algorithm that combines a timestamp and a random number to generate a unique node ID (Identity Document), ensuring the uniqueness and traceability of the node ID. The configuration management module is used to maintain cluster configuration information, including node topology, data distribution rules, etc., which are stored in a binary file combined with a memory cache to ensure efficient reading and writing. The fault detection and recovery module is used to send probe messages to each node regularly through the RDMA heartbeat detection mechanism. If no response is received within a specified time (such as 2 seconds), the node is determined to be faulty, and the recovery process is quickly started, such as restoring data from a backup node and reallocating data storage. The storage node is a Redis node, which is a data storage and processing unit, equipped with an RDMA network card to achieve RDMA-based data interaction. On the basis of traditional Redis, the data storage structure and communication module are optimized, and an RDMA data transmission interface and a load information reporting function are added.
[0034] Traditional storage nodes rely on network sockets for communication. When data is transmitted, multiple kernel-mode to user-mode switches are required, resulting in high transmission latency and low bandwidth utilization. In the embodiment of the present application, by configuring an RDMA network card for each storage node, the Remote Direct Memory Access (RDMA) technology allows network nodes to directly access each other's memory without going through multiple kernel-mode to user-mode switches. When the storage nodes communicate and transfer data through the RDMA network card, the invalid overhead in the data transmission process can be reduced, and the data transmission efficiency can be improved.
[0035] On the basis of configuring an RDMA network card in the storage node, when there is a data migration requirement or a data transmission requirement between the source node and the target node among multiple storage nodes, the cluster management node first notifies the source node and the target node to perform data migration, so that the source node and the target node establish a network connection based on the RDMA handshake protocol, and then the source node and the target node can perform data transmission based on the network connection.
[0036] Step 102, the source node sends the target migration data to the memory of the target node through the RDMA network card.
[0037] Among them, the source node is the storage node from which data is migrated out. The source node uses the zero-copy and direct memory access characteristics of RDMA to directly transfer data to the target node. Before the transfer, the source node and the target node have established a reliable connection through the RDMA handshake protocol and negotiated transfer parameters such as data block size, transfer rate limit, etc. Correspondingly, during the transfer process, the source node directly sends the target migration data to the memory of the target node based on the RDMA network card. During the data sending process, the operating systems of the source node and the target node do not need to participate, and correspondingly, there is no switching between the user state and the kernel state, which can improve the data transfer efficiency.
[0038] Optionally, a pipelined transfer method is also adopted during the transfer, that is, after the source node sends the current data block, it directly sends the next data block without waiting for confirmation, improving the transfer efficiency. At the same time, the CRC64 check algorithm is used to ensure data integrity. If the target node fails the check, it requests retransmission through the RDMA fast retransmission mechanism.
[0039] Step 103, the target node receives the target migration data sent by the source node based on the RDMA network card.
[0040] The target node is the storage node to which data is migrated in. When the source node and the target node perform data transfer, the source node can directly send the target migration data to be transferred to the memory of the target node through the RDMA network card. Correspondingly, the memory of the target node receives the target migration data sent by the source node based on the RDMA network card.
[0041] Step 104, during the process of transferring the target migration data, the cluster management node receives the amount of transferred data and the remaining data of the target migration data fed back by the source node and the target node.
[0042] On the basis of optimizing the network communication between each storage node, a transmission progress monitoring and feedback function is also added. In a possible implementation manner, during the process of transferring the target migration data, the source node and the target node will update the amount of transferred data and the remaining data of the target migration data in real time, and feedback the progress information to the cluster management node every preset time period. Correspondingly, the cluster management node can receive the amount of transferred data and the remaining data of the target migration data fed back by the source node and the target node.
[0043] Exemplarily, the preset time period can be 500 milliseconds.
[0044] Optionally, this transmission progress monitoring and feedback function is implemented by embedding corresponding monitoring fields in the RDMA transmission protocol.
[0045] Step 105, the cluster management node adjusts the data migration strategy between the source node and the target node based on the amount of transferred data and the remaining data.
[0046] After the cluster management node receives the amount of data already transmitted and the remaining amount of data, the cluster management node calculates the transmission rate of the target migration data based on this feedback information. In the case where the transmission rate is lower than the rate threshold, the cluster management node can adjust the data migration strategy between the source node and the target node, and this data migration strategy refers to the data migration path between the source node and the target node.
[0047] In summary, the embodiment of the present application provides a cluster management method: by configuring RDMA network cards for each storage node in the target cluster management system, when data is transmitted or migrated between storage nodes, the source node can directly transmit the target migration data to the memory of the target node through the RDMA network card without the participation of the operating systems of the source node and the target node, and correspondingly, there is no need to switch between the user mode and the kernel mode, which can improve the data transmission efficiency; moreover, a transmission progress monitoring and feedback function is also added, so that information such as the amount of data already transmitted and the remaining amount of data during the data migration process can be fed back to the cluster management node, enabling the cluster management node to adjust the data migration strategy in a timely manner according to the feedback information and ensuring the data migration efficiency between storage nodes.
[0048] In addition to the cluster management node and multiple storage nodes, the target system cluster also includes a load balancing node (load balancer) for evenly distributing access requests to each storage node to achieve load balancing of the target system cluster and improve the processing capacity of the target system cluster under high concurrency.
[0049] Please refer to Figure 2 , which shows a flowchart of another cluster management method according to an exemplary embodiment of the present application. This method is described by taking its application to the target cluster system as an example. As Figure 2 shown, this method includes:
[0050] Step 201, each storage node reports multiple key load indicators to the load balancing node based on the RDMA network card.
[0051] Among them, the target system cluster includes a cluster management node, a load balancing node, and multiple storage nodes. The load balancing node is a component for realizing efficient request distribution by RDMA, adopting a distributed architecture and consisting of multiple load balancing sub-nodes. Each sub-node real-time collects the RDMA transmission status and key load indicators of each storage node, and uses intelligent algorithms to reasonably route client requests to the storage node with lighter load, realizing load balancing of the target system cluster and improving the processing capacity of the target system cluster under high concurrency.
[0052] Exemplarily, multiple key load metrics include data transfer rate, bandwidth utilization, receive and / or transmit queue length, CPU usage, and memory usage.
[0053] In a possible implementation, each storage node can collect various key load metrics in real time through the RDMA network card driver and the operating system interface, including RDMA data transfer rate, bandwidth utilization, receive and / or transmit queue length, CPU usage, memory usage, etc. Through the RDMA network, each storage node reports these key load metrics to the load balancing node in a compressed binary format at regular intervals (e.g., 100 milliseconds), and the corresponding load balancing node receives multiple key load metrics reported by each storage node.
[0054] Step 202: Determine the current load value of each storage node based on multiple key load metrics by the load balancing node.
[0055] The load balancing node is configured with a load evaluation model and algorithm. After receiving the key load metrics reported by each storage node, it can evaluate the current load value of each storage node based on the load evaluation model and algorithm.
[0056] Specifically, the load evaluation model and algorithm determine the current load value based on the key load metrics and metric weights. Exemplarily, the algorithm formula for the current load value can be as shown in formula (1):
[0057] L = wR·R + wB·B + wQr·Qr + wQs·Qs + wC·C + wM·M (1)
[0058] Where, L represents the current load value, R represents the data transfer rate, B represents the bandwidth utilization, Qr represents the receive queue length, Qs represents the transmit queue length, C represents the CPU usage, M represents the memory usage; wR represents the metric weight of the data transfer rate, wB represents the metric weight of the bandwidth utilization, wQr represents the metric weight of the receive queue length, wQs represents the metric weight of the transmit queue length, wC represents the metric weight of the CPU usage, and wM represents the metric weight of the memory usage.
[0059] Based on formula (1), when the load balancing node calculates the current load value of each storage node, it first obtains the metric weight of each key load metric among the multiple key load metrics, and then substitutes the metric weight and the key load metrics into formula (1), and can determine the current load value of each storage node based on the metric weight and the key load metrics.
[0060] Optionally, each key load metric has an initial metric weight value. Exemplarily, the initial weight values of each key load metric can be: wR = 0.3, wB = 0.2, wQr = 0.1, wQs = 0.1, wC = 0.2, wM = 0.3.
[0061] In other possible scenarios, the load balancing node is also set with a weight adjustment algorithm, which is used to dynamically adjust the metric weights of each key load metric according to the current access scenario. For example, in a high-concurrency read scenario, the weight adjustment algorithm appropriately increases the weights of data transfer rate and bandwidth utilization; in a write-intensive scenario, it increases the weights of CPU usage and memory usage. In a possible implementation, the load balancing node first determines the current target access scenario, and then determines the metric weights of each key load metric according to the target access scenario.
[0062] Optionally, the cluster manager can pre-set the business request characteristics and / or metric change trend characteristics of different target access scenarios, as well as the metric weights corresponding to each target access scenario, so that in the actual application process, the load balancing node determines the current target access scenario according to the current business request characteristics and / or metric change trend characteristics, and obtains the metric weights associated with the target access scenario for the evaluation of the current load values of each storage node.
[0063] Step 203, when the load balancing node receives a target access request, based on the current load value, the load balancing node determines a target routing node from multiple storage nodes and forwards the target access request to the target routing node.
[0064] After the load balancing node evaluates the current load values of each storage node, it can perform dynamic request allocation based on the current load values. For example, when the current load value of a certain storage node is large, a new access request can be routed to a storage node with a lower load. In a possible implementation, when the load balancing node receives a target access request, based on the current load value, it determines a target routing node from multiple storage nodes and forwards the target access request to the target routing node.
[0065] Among them, the target access request is an access request from the client to the storage node, which can be a read request, that is, reading data from the storage node, or a write request, that is, writing data to the storage node.
[0066] When the load balancing node forwards the target access request to the target routing node, it can use the RDMA hardware acceleration function, such as the hash function of the RDMA network card, to quickly forward the target access request to the target routing node and reduce the software processing overhead.
[0067] Step 204, in the case that there are storage nodes among multiple storage nodes whose current load value is greater than a preset threshold, the load balancing node determines a load balancing strategy based on the current load values, historical load values, and service request types of the multiple storage nodes. The load balancing strategy includes the target node and source node for data migration, as well as a request rerouting rule.
[0068] If the current load of a certain storage node is relatively high, in order to avoid abnormal subsequent access to this storage node caused by uneven load, the load balancing node will also formulate a load balancing strategy according to the current load conditions of each storage node. In a possible implementation manner, in the case that there are storage nodes among multiple storage nodes whose current load value is greater than a preset threshold, the load balancing node determines that the storage nodes with a current load value greater than the preset threshold have an excessive load and starts the load balancing strategy formulation process. The load balancing node will comprehensively consider factors such as the current load, historical load, processing capacity, and service request type of each storage node to formulate a detailed load balancing strategy, which at least includes the target node and source node for data migration, as well as a request rerouting rule.
[0069] Specifically, the load balancing strategy can be: routing some read requests to storage nodes with lighter load and stronger read performance, and reasonably allocating write requests according to the data consistency requirements and storage capabilities of the storage nodes; it also includes the source node and target node for data migration, the rule for request rerouting, etc.
[0070] Exemplarily, the preset threshold can be 0.8, that is, after the current load value is greater than 0.8, it is determined that the load of this storage node is excessive and the load balancing strategy formulation process needs to be started.
[0071] Step 205, the load balancing node pushes the load balancing strategy to the cluster management node.
[0072] After the load balancing node determines the load balancing strategy, it can push the load balancing strategy to the cluster management node.
[0073] Step 206, the cluster management node controls the target node and source node to execute the load balancing strategy.
[0074] After receiving the load balancing policy, the cluster management node can coordinate the target node and the source node to perform the load balancing operation. For data migration, after determining the source node and the target node, the cluster management node guides both parties to establish an RDMA connection and perform data transmission. During the transmission process, the cluster management node monitors the migration progress in real time. If it is found that the bandwidth utilization rate of a certain migration link is low, the migration process is optimized by adjusting the transmission parameters or reallocating the data migration path. For request rerouting, the cluster management node notifies the relevant nodes to update the request routing table to ensure that new requests can be distributed according to the load balancing policy. At the same time, the cluster management node updates the cluster configuration information and synchronizes the latest configuration information to all nodes through RDMA broadcast to ensure that the configurations of all storage nodes are consistent. During the data migration process, the steps shown in steps 101 to 105 in the above embodiments can be executed, and this embodiment will not be elaborated here.
[0075] In this embodiment, each storage node reports key load indicators to the load balancing node in real time, so that the load balancing node can determine the current load value of each storage node based on the key load indicators, and then perform request routing according to the current load value, avoiding routing access requests to storage nodes with high loads, and load balancing of each storage node can be achieved; moreover, after the load balancing node discovers a storage node with high load, it will immediately formulate a new load balancing policy according to relevant load information and request information to timely adjust the load of each storage node.
[0076] As the business develops, the number of storage nodes in the target system cluster will increase or decrease according to the amount of business storage data. Taking the addition of storage nodes as an example, this embodiment describes the operations that the cluster management node needs to perform after adding storage nodes.
[0077] Please refer to Figure 3 , which shows a flowchart of another cluster management method according to an exemplary embodiment of the present application. This method is described by taking its application to the target cluster system as an example. As Figure 3 shown, this method includes:
[0078] Step 301, in the case of adding a node to the target cluster system, the cluster management node determines a data migration policy based on the current loads and data access hotness of multiple storage nodes, and the data migration policy includes the target node and the source node of the data migration.
[0079] The main processes for adding a node to the target cluster system include: authentication and ID assignment, configuration file generation and distribution, connection establishment and security authentication, and data migration and initialization.
[0080] Among them, in the authentication and ID assignment phase, when a new node adds a storage node to the target cluster system, the new node sends a join request containing information such as hardware configuration, software version, RDMA network card information, and the signature of the node creator to the cluster management node. After receiving the request, the cluster management node first verifies the legality of the signature to confirm that the source of the new node is reliable. Subsequently, a unique node ID is generated using an algorithm based on the combination of timestamp and random number. For example, the current timestamp is accurate to milliseconds and concatenated with a 128-bit random number, and then a unique node ID is generated through the hash algorithm to ensure the uniqueness and traceability of the node ID.
[0081] Configuration file generation and distribution phase: The cluster management node generates an initial configuration file containing RDMA communication parameters (such as the port number of the RDMA connection, communication key, etc.), cluster topology information (IP addresses of other nodes, node IDs, etc.), and a preliminary data migration plan according to the configuration information of the current cluster. This configuration file is compressed in binary format and sent to the new node through the reliable transport protocol of RDMA to ensure the efficiency and stability of the transmission.
[0082] Connection establishment and security authentication phase: The cluster management node notifies other storage nodes of the new node's joining through RDMA broadcast. When the new node establishes an RDMA connection with the existing storage nodes, the Diffie-Hellman key exchange algorithm is adopted, and both parties securely negotiate a shared key in an insecure network environment for encryption and decryption of subsequent communications. At the same time, in order to prevent man-in-the-middle attacks, digital certificates are introduced for identity authentication during the key exchange process to ensure the authenticity of the identities of both communication parties.
[0083] Data migration and initialization phase; The cluster management node determines the data migration set and target nodes according to the current load value, data distribution, and hot data situation of the target cluster system, using the consistent hashing combined with the hot data recognition algorithm. During the data migration process, the source node and the target node establish a connection through an RDMA handshake. The source node utilizes the zero-copy and pipelined transmission characteristics of RDMA to transfer data to the target node. The CRC64 checksum and fast retransmission mechanism are used during the transmission to ensure the reliability of the data. After the migration is completed, the new node performs an integrity check on the received data. After ensuring that the data is correct, it sends a confirmation message to the cluster management node to complete the initialization process.
[0084] In a possible implementation, when adding a node to the target cluster system, the cluster management node determines the data migration strategy based on the current load and data access popularity of multiple storage nodes to ensure balanced data distribution and minimize the impact of hot data migration on the business. The data migration strategy includes the target node and the source node of the data migration.
[0085] In an exemplary example, a consistent hashing ring is introduced, and storage nodes and data keys are mapped onto the hashing ring. Assume the hashing ring is an integer ring from 0 to 2^32. The hash value of node Ni is hash(Ni), and the hash value of data key K is hash(K). The data key K is stored on the first node Ni in the clockwise direction that is greater than or equal to hash(K). When a new node Nnew is added, the hashing ring is recalculated to determine the affected data set, and data with low access popularity is preferentially migrated. Hot data is identified by recording the access frequency. For example, within a certain time window (such as 10 minutes), data with the number of accesses exceeding a threshold (such as 1000 times) is determined as hot data.
[0086] Optionally, in the case of adding a new node, the target nodes of the data migration strategy include the newly added node.
[0087] Step 302, the cluster management node controls the target node and the source node to execute the data migration strategy.
[0088] After determining the data migration strategy, the cluster management node can control the target node and the source node to execute the data migration strategy. The process of executing the data migration strategy between the target node and the source node may include steps 303 to 307.
[0089] Optionally, the cluster management node also periodically performs fault detection on each storage node. Specifically, the cluster management node sends probe requests to multiple storage nodes based on the RDMA heartbeat detection mechanism and judges node faults based on the response results of the probe requests. If the response from the target node is not received within a specified time (for example, 2 seconds), the cluster management node immediately conducts a secondary confirmation through other storage nodes. If it is confirmed that the target node has failed, the cluster management node quickly notifies the load balancing node to reroute the requests of the target node to the backup node or other available nodes. At the same time, the cluster management node records the relevant information of the faulty node (target node), including the fault occurrence time, the status of the last heartbeat, etc., for subsequent analysis of the fault cause.
[0090] Optionally, after determining that the target node has failed, the cluster management node also coordinates other storage nodes to recover the data of the target node from the backup node. During the recovery process, a fast data transfer method based on RDMA is adopted to ensure that the data can be recovered as soon as possible. After the recovery is completed, the cluster management node formulates a new data migration strategy according to the current load values and data distribution of each storage node in the cluster. For example, by analyzing the load conditions and data storage amounts of each storage node, the data recovered from the target node is reasonably dispersed to other storage nodes with lighter loads to ensure data integrity and consistency, while avoiding new load imbalance problems.
[0091] After the failed node (target node) is repaired, the target node sends a rejoin request to the cluster management node, and the cluster management node conducts a comprehensive health check on the target node, including hardware status, software version, data integrity, etc. After passing the check, the target node is re-incorporated into the cluster management according to the new node joining process to ensure the stability and reliability of the target cluster system.
[0092] Step 303, in the case of a data migration requirement between the source node and the target node among multiple storage nodes, the cluster management node controls the source node and the target node to establish a network connection based on the RDMA handshake protocol;
[0093] Step 304, the source node sends the target migration data to the memory of the target node based on the RDMA network card;
[0094] Step 305, the target node receives the target migration data sent by the source node based on the RDMA network card;
[0095] Step 306, during the process of transmitting the target migration data, the cluster management node receives the amount of transmitted data and the remaining data of the target migration data fed back by the source node and the target node;
[0096] Step 307, the cluster management node adjusts the data migration strategy between the source node and the target node based on the amount of transmitted data and the remaining data.
[0097] The implementation manners of Steps 303 to 307 can refer to the above embodiments, and are not elaborated herein.
[0098] In this embodiment, when a new node joins, the initialization of the new node is completed through authentication, configuration file distribution, secure connection establishment, and data migration guidance; during fault handling, a fast reroute request is made, data is restored from the backup node, and the migration strategy is re-formulated to ensure data integrity and consistency.
[0099] Please refer to Figure 4 , which is a schematic structural diagram of a target cluster system provided by an embodiment of the present application. Exemplarily, as Figure 4 shown, the target cluster system 400 includes a cluster management node 401, multiple storage nodes 402, and a load balancing node 403, and each storage node is configured with an RDMA network card.
[0100] The cluster management node 401 is used to control the source node and the target node to establish a network connection based on the RDMA handshake protocol in the case of a data migration requirement between the source node and the target node among the multiple storage nodes;
[0101] The source node is used to send target migration data to the memory of the target node based on the RDMA network card;
[0102] The target node is used to receive the target migration data sent by the source node based on the RDMA network card;
[0103] The cluster management node 401 is further used to receive the amount of transmitted data and the remaining data of the target migration data fed back by the source node and the target node during the transmission of the target migration data;
[0104] The cluster management node 401 is further used to adjust the target migration strategy between the source node and the target node based on the amount of transmitted data and the remaining data.
[0105] Optionally, the multiple storage nodes 402 are used to report multiple key load metrics to the load balancing node based on the RDMA network card;
[0106] The load balancing node 403 is used to determine the current load value of each storage node based on the multiple key load metrics;
[0107] The load balancing node 403 is further used to, when the load balancing node receives a target access request, determine a target routing node from the multiple storage nodes based on the current load value, and forward the target access request to the target routing node.
[0108] Optionally, the load balancing node 403 is further used to: obtain the metric weight of each key load metric among the multiple key load metrics; determine the current load value of each storage node based on the metric weight and each key load metric.
[0109] Optionally, the load balancing node 403 is further used to: determine the current target access scenario; determine the metric weight of each key load metric based on the target access scenario.
[0110] Optionally, when there is a storage node among the multiple storage nodes whose current load value is greater than a preset threshold, the load balancing node 403 is used to determine a load balancing strategy based on the current load value, historical load value, and service request type of the multiple storage nodes, where the load balancing strategy includes the target node and the source node for data migration, and a request rerouting rule; the load balancing node 403 is further used to push the load balancing strategy to the cluster management node; the cluster management node 401 is further used to control the target node and the source node to execute the load balancing strategy.
[0111] Optionally, the cluster management node 401 is further configured to: when adding a node to the target cluster system, determine a data migration policy based on the current loads and data access heat of the multiple storage nodes, where the data migration policy includes the target node and the source node for data migration; control the target node and the source node to execute the data migration policy.
[0112] Optionally, the cluster management node 401 is further configured to: send a detection request to multiple storage nodes based on the RDMA heartbeat detection mechanism; perform node failure determination based on the response result of the detection request.
[0113] It should be noted that the cluster management node, load balancing node, and storage node in the above embodiments can all be implemented as an electronic device. An exemplary embodiment of the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program that can be executed by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the cluster management method according to the embodiments of the present application.
[0114] An exemplary embodiment of the present application further provides a non-transitory computer-readable storage medium storing a computer program, where the computer program is used to cause a computer to execute the cluster management method according to the embodiments of the present application when executed by a processor of the computer.
[0115] An exemplary embodiment of the present application further provides a computer program product, including a computer program, where the computer program is used to cause a computer to execute the cluster management method according to the embodiments of the present application when executed by a processor of the computer.
[0116] Reference Figure 5 , the structural block diagram of an electronic device 500 that can be used as a server or client of the present application will now be described. It is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0117] As Figure 5As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to computer programs stored in the ROM 502 or computer programs loaded from the storage unit 508 into the RAM 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. The I / O interface 505 is also connected to the bus 504.
[0118] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. The input unit 506 can be any type of device that can input information into the electronic device 500. The input unit 506 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 507 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 508 can include but is not limited to a magnetic disk, an optical disk. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0119] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above. For example, in some embodiments, Figure 1 、 Figure 2 、 Figure 3 The method shown can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. In some embodiments, the computing unit 501 can be configured to execute Figure 1 、 Figure 2 、 Figure 3 The method shown in any other appropriate manner (for example, by means of firmware).
[0120] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0121] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0122] As used in the present application, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., a disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal for providing machine instructions and / or data to a programmable processor.
[0123] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0124] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0125] The computer system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The client - server relationship is generated by computer programs running on the respective computers and having a client - server relationship with each other.
Claims
1. A cluster management method, characterized in that: The method is applied to a target cluster system, the target cluster system includes a cluster management node and multiple storage nodes, each of the storage nodes is configured with an RDMA network card, and the method includes: When there is a data migration demand between a source node and a target node among the multiple storage nodes, the cluster management node controls the source node and the target node to establish a network connection based on the RDMA handshake protocol; Sending target migration data to the memory of the target node through the source node based on the RDMA network card; Receiving the target migration data sent by the source node through the target node based on the RDMA network card; In the process of transmitting the target migration data, receiving, through the cluster management node, the amount of transmitted data and the amount of remaining data of the target migration data fed back by the source node and the target node; The data migration strategy between the source node and the target node is adjusted by the cluster management node based on the amount of transmitted data and the amount of remaining data.
2. The method according to claim 1, characterized in that The target cluster system also includes a load balancing node, and the method further includes: Reporting multiple key load indicators to the load balancing node through each storage node based on the RDMA network card; Determining, by the load balancing node, a current load value of each of the storage nodes based on the multiple key load indicators; When the load balancing node receives a target access request, the load balancing node determines a target routing node from the multiple storage nodes based on the current load value, and forwards the target access request to the target routing node.
3. The method according to claim 2, characterized in that Determining the current load value of each storage node based on the multiple key load indicators through the load balancing node includes: Acquire, through the load balancing node, an indicator weight of each of the multiple key load indicators; The load balancing node determines the current load value of each storage node based on the indicator weight and each key load indicator.
4. The method according to claim 3, characterized in that The obtaining, through the load balancing node, the indicator weight of each of the multiple key load indicators includes: Determine the current target access scenario through the load balancing node; The load balancing node determines the indicator weight of each of the key load indicators based on the target access scenario.
5. The method according to claim 2, characterized in that: The method further comprises: In the case that there is a storage node whose current load value is greater than a preset threshold value among the multiple storage nodes, determining a load balancing strategy through the load balancing node based on the current load values, historical load values and service request types of the multiple storage nodes, wherein the load balancing strategy includes the target node and the source node of data migration, and a request rerouting rule; Pushing the load balancing strategy to the cluster management node through the load balancing node; The target node and the source node are controlled by the cluster management node to execute the load balancing strategy.
6. The method according to claim 2, characterized in that The method further comprises: In the case of adding a new node to the target cluster system, determining the data migration strategy based on the current load and data access heat of the multiple storage nodes by the cluster management node, the data migration strategy including the target node and the source node of the data migration; The target node and the source node are controlled by the cluster management node to execute the data migration strategy.
7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Sending a detection request to the plurality of storage nodes based on the RDMA heartbeat detection mechanism by the cluster management node; The cluster management node performs node failure judgment based on the response result of the detection request.
8. A target cluster system, characterized in that: The target cluster system includes a cluster management node and multiple storage nodes, each of which is configured with an RDMA network card: The cluster management node is used to control the source node and the target node in the multiple storage nodes to establish a network connection based on the RDMA handshake protocol when there is a data migration demand between the source node and the target node; The source node is used to send target migration data to the memory of the target node based on the RDMA network card; The target node is used to receive the target migration data sent by the source node based on the RDMA network card; The cluster management node is further configured to receive, during the process of transmitting the target migration data, the amount of transmitted data and the amount of remaining data of the target migration data fed back by the source node and the target node; The cluster management node is further configured to adjust a target migration strategy between the source node and the target node based on the amount of transmitted data and the amount of remaining data.
9. An electronic device, comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to execute the cluster management method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the cluster management method according to any one of claims 1 to 7.
Citation Information
Cited By
Data processing method of storage system and electronic equipment
CN120872262A