Master-slave operation control method, device and equipment and readable storage medium
By conducting business processing capabilities test and real-time monitoring of each node in a distributed cluster, the master-slave role is dynamically adjusted, and the problem of master-slave role not being adjusted according to performance is solved, and the effect of efficient utilization of cluster resources and improving overall business performance is achieved.
Patent Information
- Application Number
- CN202311541368.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-11-17
AI Technical Summary
During the operation of a distributed cluster, the role of the master and slave is not adjusted according to performance, resulting in high-performance devices being unable to fully utilize their performance advantages and causing waste of cluster resources.
By conducting business processing capability tests on each node when creating a distributed cluster, the management area is divided, so that each management area includes at least one host and one slave, and the master-slave switches according to the results of the real-time service processing capability test to ensure that the host is not lagging behind the host.
It realizes dynamic adjustment of master-slave roles based on business performance, fully utilizes the performance advantages of each node in the distributed cluster, and improves overall business performance.
Smart Images

Figure CN120075229A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed clusters, and particularly to a master-slave operation control method, device, equipment and readable storage medium. Background Art
[0002] A distributed cluster refers to a system composed of multiple computer nodes. The nodes communicate and cooperate with each other through a network to jointly complete a task or provide a service. Distributed clusters have advantages such as high availability, high performance, and scalability, and are therefore widely used in scenarios such as large-scale computing, storage, and processing.
[0003] To manage a distributed cluster well, one or more nodes are selected as the master node(s) when the distributed cluster is established, and the other nodes are slave nodes. Usually, the node with the smallest network address is default selected as the master node, or the master node is elected by voting among the nodes. During the operation of the distributed cluster, if the master node fails, a new master node is re-elected, and the services of the old master node are switched to the new master node.
[0004] During the operation of a distributed cluster, the data processing capacity and storage capacity of each node device will continuously change as the business progresses. If the roles of the master and slave nodes are not adjusted according to performance, it will cause high-performance devices to not fully utilize their performance advantages, resulting in waste of cluster resources.
[0005] Providing a distributed cluster master-slave implementation solution based on business performance to give full play to the performance advantages of the distributed cluster is a technical problem that those skilled in the art need to solve. Summary of the Invention
[0006] The purpose of the present invention is to provide a master-slave operation control method, device, equipment and readable storage medium for implementing a distributed cluster master-slave allocation scheme based on business performance, so as to give full play to the performance advantages of each node in the distributed cluster and improve the overall business performance of the distributed cluster.
[0007] To solve the above technical problem, the present invention provides a master-slave operation control method, including:
[0008] When creating a distributed cluster, according to the service types that each node of the distributed cluster needs to execute, perform a service processing capacity test on each node to obtain the initial service processing capacity test results of each node;
[0009] According to the initial service processing capacity test results, select multiple master nodes among each node and divide the management areas where each master node is located, so that each management area includes at least one master node and one slave node, and the initial service processing capacity test results of the nodes in the same management area are similar;
[0010] During the operation of the distributed cluster, monitor the service processing capabilities of the nodes to obtain the real-time service processing capability test results of the nodes.
[0011] Determine the target management area for which the master-slave switch operation is to be performed according to the real-time service processing capability test results.
[0012] Perform the master-slave switch operation on the target management area to ensure that the master in each management area is not the host with poor performance in the management area.
[0013] In some embodiments, the selecting multiple hosts from the nodes according to the initial service processing capability test results and dividing the management areas where the hosts are located, so that each management area includes at least one host and one slave, and the initial service processing capability test results of the nodes in the same management area are similar, includes:
[0014] Randomly select multiple nodes from the nodes as the initial clustering centroid nodes.
[0015] Assign values according to the initial service processing capability test results, calculate the clustering radius between the remaining nodes and the initial clustering centroid nodes, and divide the remaining nodes into the initial management areas where the initial clustering centroid nodes with close radius distances are located.
[0016] Loop to perform re-selecting the clustering centroid nodes in each management area, assigning values according to the initial service processing capability test results, calculating the clustering radius between the remaining nodes and the clustering centroid nodes, and dividing the remaining nodes into the management areas where the clustering centroid nodes with close radius distances are located, until each management area includes at least one host and one slave, and the initial service processing capability test results of the nodes in the same management area are similar.
[0017] In some embodiments, the determining the target management area for which the master-slave switch operation is to be performed according to the real-time service processing capability test results includes:
[0018] If the real-time service processing capability test result of the host in the management area is lower than that of the slaves in the management area that reach the first quantity threshold, and there is a high-performance slave among the slaves whose real-time service processing capability test result is within the first quantity threshold and the remaining storage space is greater than the second threshold, and there are auxiliary slaves that reach the third quantity threshold among the slaves, then determine the management area as the target management area.
[0019] In some embodiments, performing the master-slave switch operation on the target management area includes:
[0020] Transferring the host service data of the host to the auxiliary slave through the fiber optic switch, and recording the increment of the host service data processed by the auxiliary slave;
[0021] After transferring all the host service data of the host to the auxiliary slave, transferring the management authority and the monitoring authority of the host to the new host in the management area;
[0022] The new host sends control instructions to each slave, and determines that the permission switch is completed after receiving the response instructions from all the slaves in the management area;
[0023] After the new host determines that the permission switch is completed, it receives the host service data fed back by each auxiliary slave and the increment of the host service data processed by each auxiliary slave, and continues to process the host service data from the latest execution time point of the host service data.
[0024] In some embodiments, the distributed cluster is a distributed storage cluster;
[0025] Performing a service processing capacity test on each node according to the service types required to be executed by each node of the distributed cluster, and obtaining the initial service processing capacity test results of each node, including:
[0026] After each node is powered on, performing read and write processing of a fourth quantity threshold of service data on each node;
[0027] Using at least one of the throughput, average number of operations per second, and latency time when the node processes service data as the initial service processing capacity test result of the node;
[0028] Monitoring the service processing capacity of each node, and obtaining the real-time service processing capacity test results of each node, including:
[0029] Monitoring at least one of the throughput, average number of operations per second, and latency time when the node processes actual service data, and obtaining the real-time service processing capacity test results of each node.
[0030] In some embodiments, it further includes:
[0031] For the distributed cluster, sequentially allocate service data with priorities from high to low and / or service data with data volumes from large to small in the order of the real-time service processing capacity test results of the management area from high to low;
[0032] If all the service data of the distributed cluster have been allocated to the management area, end the cluster-level allocation optimization of the service data for the current batch.
[0033] If there is still remaining service data after allocating all the management areas, continue to allocate the remaining service data with the highest to lowest priorities and / or the remaining service data with the largest to smallest data volumes in the order of the test results of the real-time service processing capabilities of the management areas from high to low until all the service data of the distributed cluster are allocated to the management areas.
[0034] For the management areas, allocate them to the nodes with the test results of the real-time service processing capabilities from high to low in the order of the priorities of the service data from high to low and / or the data volumes of the service data from large to small for data processing and disk writing, and select the nodes for backing up the processed service data in the order of the test results of the real-time service processing capabilities from high to low and the remaining storage spaces from large to small.
[0035] In some embodiments, each of the nodes is interconnected through a gigabit switch to transmit control data.
[0036] Each of the nodes is interconnected through a fiber optic switch to transmit service data.
[0037] In some embodiments, selecting multiple hosts from each of the nodes according to the initial service processing capability test results and dividing the management areas where each host is located, so that each management area includes at least one host and one slave, and the initial service processing capability test results of the nodes in the same management area are similar, includes:
[0038] Randomly select multiple nodes from each of the nodes as initial clustering centroid nodes.
[0039] Assign values according to the initial service processing capability test results, calculate the clustering radii between the remaining nodes and the initial clustering centroid nodes, and divide the remaining nodes into the initial management areas where the initial clustering centroid nodes with close radius distances are located.
[0040] Loop to execute reselecting clustering centroid nodes in each management area, assigning values according to the initial service processing capability test results, calculating the clustering radii between the remaining nodes and the clustering centroid nodes, and dividing the remaining nodes into the management areas where the clustering centroid nodes with close radius distances are located until each management area includes at least one host and one slave, and the initial service processing capability test results of the nodes in the same management area are similar.
[0041] Determining a target management area for performing a master-slave switch operation according to the real-time service processing capacity test result includes:
[0042] If the real-time service processing capacity test result of the master in the management area is lower than that of the slaves in the management area that reach a first quantity threshold, and there is a high-performance slave among the slaves whose real-time service processing capacity test result is within the first quantity threshold and whose remaining storage space is greater than a second threshold, and there are auxiliary slaves that reach a third quantity threshold among the slaves, then determine the management area as the target management area;
[0043] Performing a master-slave switch operation on the target management area to ensure that the master in each management area is not a performance-backward host within the management area includes:
[0044] Transfer the host service data of the host to the auxiliary slave through a fiber optic switch, and record the increment of the auxiliary slave processing the host service data;
[0045] After transferring all the host service data of the host to the auxiliary slave, transfer the management authority and the monitoring authority of the host to the new host in the management area;
[0046] The new host sends control instructions to each slave, and determines that the permission switch is completed after receiving the response instructions from all the slaves in the management area;
[0047] After the new host determines that the permission switch is completed, receive the host service data fed back by each auxiliary slave and the increment of each auxiliary slave processing the host service data, and continue to process the host service data from the latest execution time point of the host service data.
[0048] To solve the above technical problems, the present invention also provides a master-slave operation control device, including:
[0049] A test unit, configured to, when creating a distributed cluster, perform a service processing capacity test on each node according to the service types required to be executed by each node of the distributed cluster, and obtain the initial service processing capacity test results of each node;
[0050] A partitioning unit, configured to select multiple hosts from each node according to the initial service processing capacity test results and partition the management areas where each host is located, so that each management area includes at least one host and one slave, and the initial service processing capacity test results of the nodes within the same management area are similar;
[0051] A monitoring unit, configured to monitor the service processing capabilities of the nodes during the operation of the distributed cluster, and obtain the real-time service processing capability test results of the nodes;
[0052] A determination unit, configured to determine a target management area for performing the master-slave switch operation according to the real-time service processing capability test results;
[0053] A control unit, configured to perform a master-slave switch operation on the target management area to ensure that the master in each management area is not a host with poor performance in the management area.
[0054] To solve the above technical problems, the present invention further provides a master-slave operation control device, including:
[0055] A memory, configured to store a computer program;
[0056] A processor, configured to execute the computer program, and when the computer program is executed by the processor, the steps of the master-slave operation control method as described in any one of the above are implemented.
[0057] To solve the above technical problems, the present invention further provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the master-slave operation control method as described in any one of the above are implemented.
[0058] The master-slave operation control method provided by the present invention performs a service processing capability test on each node when creating a distributed cluster, divides management areas according to the test results, so that at least one master and one slave are included in one management area and the initial service processing capability test results of each node are similar; during the operation of the distributed cluster, monitors the service processing capabilities of the nodes to obtain the real-time service processing capability test results of the nodes, determines the target management area and performs a master-slave switch operation on the target management area, so that the master in each management area is not a host with poor performance in the management area, realizing a master-slave allocation scheme for a distributed cluster based on service performance, which can give full play to the service performance of the nodes in the distributed cluster to improve the overall service performance of the distributed cluster.
[0059] The master-slave operation control method provided by the present invention uses a clustering algorithm to divide management areas, ensuring that nodes with similar initial service processing capability test results are divided into the same management area, and then selects the one with the strongest initial service capability as the master, laying a good foundation for optimizing the master-slave operation performance of the distributed cluster.
[0060] The master-slave operation control method provided by the present invention selects an auxiliary slave with better performance to temporarily process the service data of the master when the performance of the master in the management area lags behind, and then switches the management authority and monitoring authority of the master to the high-performance slave in the management area, so as to keep the master in the management area as a high-performance master and achieve zero-latency switching between the master and the slave.
[0061] The present invention also provides a master-slave operation control device, equipment and readable storage medium, which have the above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solutions of the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0063] Figure 1 It is an architecture diagram of a distributed cluster provided by an embodiment of the present invention;
[0064] Figure 2 It is a flowchart of a master-slave operation control method provided by an embodiment of the present invention;
[0065] Figure 3 It is a schematic diagram of dividing a management area based on a clustering algorithm provided by an embodiment of the present invention;
[0066] Figure 4 It is a schematic structural diagram of a master-slave operation control device provided by an embodiment of the present invention;
[0067] Figure 5 It is a schematic structural diagram of a master-slave operation control device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] The core of the present invention is to provide a master-slave operation control method, device, equipment and readable storage medium, which are used to implement a master-slave allocation scheme for a distributed cluster based on service performance, so as to give full play to the performance advantages of each node in the distributed cluster and improve the overall service performance of the distributed cluster.
[0069] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0070] The following describes Embodiment 1 of the present invention.
[0071] Figure 1 It is an architecture diagram of a distributed cluster provided by an embodiment of the present invention.
[0072] For ease of understanding, first, the system architecture applicable to the present invention is introduced. The specific implementation provided by the embodiments of the present invention can be applied to any type of distributed cluster, such as a distributed storage cluster, a distributed computing cluster, a distributed training cluster, etc. A distributed cluster is a system composed of multiple computer nodes (hereinafter referred to as nodes). During operation, the nodes communicate and cooperate with each other through a network to jointly complete a task or provide a service.
[0073] As Figure 1 shown, in the distributed cluster provided by the embodiments of the present invention, each node can be interconnected through a switch respectively. When the number of switch network ports is insufficient, multiple switches can be cascaded to achieve the interconnection between nodes. A distributed cluster may include dozens or even hundreds of nodes, and each node can act as a host or a slave.
[0074] To facilitate the decoupling of the control function and the service function of the distributed cluster, in the distributed cluster provided by the embodiments of the present invention, a switch for transmitting control data and a switch for transmitting service data can be respectively set. As Figure 1 shown, the switch for transmitting control data serves as a control link switching unit and is not used for transmitting service data. Therefore, it does not need to require a high transmission strategy and can be composed of multiple S5700S-52P-LI-AC 48-port gigabit switches 102. This switch has 48 gigabit network ports, which can meet the requirements for data interaction between multiple storage devices, and can be extended to a maximum of 48 devices. At the same time, the power consumption of the device during normal operation is less than 50W, meeting the requirements for building a green and energy-saving cluster.
[0075] The switch for transmitting service data is used to realize the transfer of service data between nodes and has certain requirements for the transmission rate. Then, multiple BR-310-B-0008 24-port fiber switches can be used. This switch has 24 fiber interfaces, which can support SAN storage devices to perform high-speed data transmission through an external FC card. The transmission rate can reach 8Gb / s, and it supports cascading of multiple fiber switches to form a data transmission and interaction unit.
[0076] In the distributed cluster provided by the embodiments of the present invention, each node is divided into different management areas, and a master and a slave are determined in each management area to respectively manage the operation of the master and slave in each management area. Compared with the traditional solution of selecting one or more masters in a distributed cluster to manage the entire distributed cluster, it is more adaptable to the management scenario of a large-scale distributed cluster and can give full play to the node performance.
[0077] Based on the above architecture, the master-slave operation control method provided by the embodiments of the present invention will be described below with reference to the accompanying drawings.
[0078] The following describes Embodiment 2 of the present invention.
[0079] Figure 2 It is a flowchart of a master-slave operation control method provided by the embodiments of the present invention.
[0080] As Figure 2 shown, the master-slave operation control method provided by the embodiments of the present invention includes:
[0081] S201: When creating a distributed cluster, according to the service types required to be executed by each node of the distributed cluster, perform a service processing capacity test on each node to obtain the initial service processing capacity test results of each node.
[0082] S202: According to the initial service processing capacity test results, select multiple masters from each node and divide the management areas where each master is located, so that each management area includes at least one master and one slave, and the initial service processing capacity test results of the nodes in the same management area are similar.
[0083] S203: During the operation of the distributed cluster, monitor the service processing capacity of each node to obtain the real-time service processing capacity test results of each node.
[0084] S204: Determine the target management area for which the master-slave switch operation is to be performed according to the real-time service processing capacity test results.
[0085] S205: Perform a master-slave switch operation on the target management area to ensure that the master in each management area is not the host with poor performance in the non-management area.
[0086] In specific implementation, the master-slave operation control method provided by the embodiments of the present invention can be executed based on an upper computer outside the distributed cluster. In other embodiments, the master-slave operation control method provided by the embodiments of the present invention can be deployed in all nodes of the distributed cluster, and any node can perform the service processing capacity monitoring task and cooperate to complete the master-slave switch operation after becoming the master.
[0087] In the master-slave operation control method provided by the embodiments of the present invention, each node can be interconnected through a gigabit switch to transmit control data; in addition, each node can be interconnected through a fiber optic switch to transmit service data. For the specific implementation, please refer to the description of Embodiment 1 of the present invention.
[0088] For S201 and S202, since a large number of nodes are included in the distributed cluster, randomly allocating the master and slave machines may cause poor business processing, and may also cause many nodes to apply to become the master, resulting in communication anomalies and system downtime. Therefore, in the master-slave operation control method provided by the embodiments of the present invention, after the cluster establishment instruction is issued, the upper computer immediately evaluates the business processing capabilities of all nodes in the distributed cluster, so as to divide the management area and select the master machine according to the evaluation results.
[0089] It should be noted that, according to the service types that each node of the distributed cluster needs to execute, the business processing capabilities of each node are tested, which is a test for the common service types to be executed by the distributed cluster. For example, if the distributed cluster is a distributed storage cluster, only the read and write capabilities of each node need to be tested. If the distributed cluster is a distributed computing cluster, it may also include testing the data calculation of each node. By making the nodes perform data preprocessing when the nodes are powered on, and counting information such as the time for all nodes to complete the service data and whether it is accurate, the initial business processing capability test results of each node are assigned values.
[0090] Then, in the master-slave operation control method provided by the embodiments of the present invention, if the distributed cluster is a distributed storage cluster, S201: According to the service types that each node of the distributed cluster needs to execute, the business processing capabilities of each node are tested, and the initial business processing capability test results of each node are obtained, which may include:
[0091] After each node is powered on, read and write processing of service data with a fourth quantity threshold (such as 64GB) is performed on each node;
[0092] Using at least one of the throughput, average number of operations per second, and latency time when the node processes the service data as the initial business processing capability test result of the node.
[0093] For example, each node can be made to execute data preprocessing when the device is powered on for a certain amount of service data (such as 64GB of service data), count the time required for each device to process this set of data respectively, and whether the data processing is accurate. Finally, an accurate score is obtained. For a distributed storage cluster, the scoring method can be set such that the higher the throughput when processing the service data, the higher the score, the more the average number of operations per second, the higher the score, and the smaller the latency time, the higher the score. The initial business processing capability test results of each node are characterized by the scores.
[0094] For S202, the management areas are divided according to the initial service processing capacity test results, and the hosts of each management area are selected. In order to reduce the latency of the master-slave switchover during the operation of the distributed cluster, in the embodiments of the present invention, the nodes with similar initial service processing capacity test results are divided into the same management area. At the same time, the minimum number of nodes in each management area is determined (for example, at least 20 nodes are allocated to each management area) to ensure normal data processing capacity. The nodes in each management area can be increased or decreased.
[0095] When dividing, all the nodes in the distributed cluster can be sorted according to the initial service processing capacity test results first, and then divided into multiple management areas according to the minimum number of nodes in each management area. The node with the best initial service processing capacity test result in each management area is used as the host. After determining the host, the host sends a delayed system startup signal to the other nodes in the management area to ensure that other nodes will not start and compete for the host after the host starts. The startup signal is transmitted through the control link interaction unit. The delayed signal sent in each management area has a check bit, and only the slaves in this management area can correctly identify this check bit. Because during the signal sending process through the switch, there will be delayed signals sent to other management areas. This check bit identification method can ensure that the master-slave selection processes between different management areas will not interfere with each other.
[0096] After creating the distributed cluster, service data is allocated to each management area according to the service processing situation of the distributed cluster. The sum or average value of the initial service processing capacity test results (assigned scores) of all the nodes in the management area is used as the initial service processing capacity test result (assigned score) of the management area as a whole. Important and large-volume data services, such as financial data, medical data, etc., can be allocated to the management areas with strong initial service processing capacity test results (high assigned scores). Secondary data such as databases and log records are allocated to the management areas with weak initial service processing capacity test results (low assigned scores) to ensure that each management area of the distributed cluster can exert its maximum performance. And in the management area, service data can also be allocated according to the initial service processing capacity test results (assigned scores) of each node. Based on this, the master-slave operation control method provided by the embodiments of the present invention may further include:
[0097] For the distributed cluster, according to the real-time service processing capacity test results of the management areas from high to low, service data with priorities from high to low and / or service data with data volumes from large to small are sequentially allocated;
[0098] If all the service data of the distributed cluster has been allocated to the management areas, the cluster-level allocation optimization of the service data in the current batch is ended;
[0099] If there is still remaining service data after all management areas are allocated, continue to allocate the remaining service data with priorities from high to low and / or the remaining service data with data volumes from large to small in the order from high to low of the real-time service processing capacity test results of the management areas until all the service data of the distributed cluster are allocated to the management areas;
[0100] For the management areas, allocate them to the nodes with real-time service processing capacity test results from high to low in the order from high to low of the priorities of the service data and / or in the order from large to small of the data volumes of the service data for data processing and disk writing, and select nodes for backup of the processed service data in the order from high to low of the real-time service processing capacity test results and from large to small of the remaining storage spaces.
[0101] Among them, the real-time service processing capacity test result can be the initial service processing capacity test result measured when creating the distributed cluster and / or the real-time service processing capacity test result during the operation of the distributed cluster. The embodiment of the present invention provides a two-layer optimization algorithm for service data allocation. When allocating service data each time, first divide the service data into multiple copies (Service 1, Service 2...), and allocate the service data to the management areas with scores from high to low in the order of the importance of the service data and / or the service data volume. If the service data can be allocated completely after one round of allocation, the cluster-level allocation optimization ends. If there is still unallocated service data, re-allocate the service data to the management areas with scores from high to low.
[0102] On the basis of the first-layer cluster-level allocation optimization, the second-layer optimization is carried out within each management area. According to the deviation of the real-time service processing capacity test results (scores) of each node in the management area, preferentially let the nodes with high scores process and write the service data, and transmit the processed service data to the nodes with low scores but large remaining storage spaces through the fiber optic switch for data backup. After the data processing is completed, the next round of service processing can be carried out, and the processed data are all backed up in the background, so as to ensure that the data processing speed in each management area reaches the optimum.
[0103] For S203, during the operation of the distributed cluster, each node is processing service data, so the service processing capacity of each node can be directly monitored. For example, for a distributed storage cluster, monitor the read and write performance of each node. For example, for a distributed computing cluster, test the computing performance of each node. Then in the master-slave machine operation control method provided by the embodiment of the present invention, if the distributed cluster is a distributed storage cluster, monitoring the service processing capacity of each node in S203 to obtain the real-time service processing capacity test results of each node may include: monitoring at least one of the throughput, average number of operations per second, and latency time when the node processes actual service data to obtain the real-time service processing capacity test results of each node.
[0104] For S204 and S205, based on the real-time service processing ability test results of each node monitored in real time, it is determined whether the master-slave switch within the management area needs to be performed to ensure that high-performance nodes always serve as the master in each management area.
[0105] If the real-time service processing ability test result of the master in the management area is located in the latter preset percentage of this management area, it is considered that the performance of this master is lagging behind. It is necessary to select a slave with a better real-time service processing ability test result in this management area as the new master. At this time, this management area can be considered as the target management area, and the master-slave switch operation needs to be executed.
[0106] When performing the master-slave switch operation on the target management area, the service executed on the master can be paused first, and the operation of switching the management permission and monitoring permission of the master to the new master can be carried out. However, this will cause the master service to be paused. To achieve zero-latency master-slave switching, the service data related to the control permission and management permission of the master can be copied to the new master first and then the switching of the control permission and management permission can be executed.
[0107] The master-slave operation control method provided by the embodiments of the present invention performs a service processing ability test on each node when creating a distributed cluster, divides the management area according to the test results, so that at least one master and one slave are included in one management area and the initial service processing ability test results of each node are similar; during the operation of the distributed cluster, monitors the service processing ability of each node to obtain the real-time service processing ability test results of each node, determines the target management area and performs the master-slave switch operation on the target management area, so that the master in each management area is not the master with lagging performance in this management area, realizing a master-slave allocation scheme for a distributed cluster based on service performance, which can give full play to the service performance of the nodes in the distributed cluster to improve the overall service performance of the distributed cluster.
[0108] The following describes Embodiment 3 of the present invention.
[0109] Figure 3 It is a schematic diagram of dividing the management area based on the clustering algorithm provided by the embodiments of the present invention.
[0110] On the basis of the above embodiments, the embodiments of the present invention provide another scheme for dividing the management area. In the master-slave operation control method provided by the embodiments of the present invention, S202: According to the initial service processing ability test results, select multiple masters among each node and divide the management area where each master is located, so that each management area includes at least one master and one slave, and the initial service processing ability test results of each node in the same management area are similar, which may include:
[0111] Randomly select multiple nodes from each node as the initial clustering centroid nodes;
[0112] Assign values based on the initial business processing ability test results, calculate the clustering radius between the remaining nodes and the initial clustering centroid nodes, and divide the remaining nodes into the initial management areas where the initial clustering centroid nodes with closer radius distances are located;
[0113] Loop to reselect the clustering centroid nodes within each management area, assign values based on the initial business processing ability test results, calculate the clustering radius between the remaining nodes and the clustering centroid nodes, and divide the remaining nodes into the management areas where the clustering centroid nodes with closer radius distances are located until each management area includes at least one host and one slave, and the initial business processing ability test results of the nodes within the same management area are similar.
[0114] In a specific implementation, after issuing the cluster establishment instruction, immediately evaluate the business processing capabilities of all nodes in the distributed cluster through the upper computer and assign values to the initial business processing ability test results of each node. As Figure 3 shown, mark the assigned values of each node on the coordinate axis, perform clustering analysis on the assigned values, and the upper computer executes the number of clusters according to the data processing volume. The number of clusters is related to the data volume and the total number of nodes. For example, in the second embodiment of the present invention, the minimum number of nodes in each management area (such as at least 20 nodes are allocated in each management area), then the number of clusters k = total number of nodes / 20, so as to ensure that there are 20 nodes in each class after clustering to ensure normal data processing capabilities. As the data volume increases, the number of nodes in each class can fluctuate around 20.
[0115] According to the K-means clustering algorithm, first randomly select (the number of clusters) k devices as the initial clustering centroid nodes, and then calculate the clustering radius between the assigned value of each node and the assigned value of the initial clustering centroid nodes. The closer the radius distance is to which initial clustering centroid node indicates that the device belongs to the assigned value interval where the initial clustering centroid node is located. Assuming that four assigned value intervals (assigned value interval 1, assigned value interval 2, assigned value interval 3, assigned value interval 4) are set, then after sequential calculations, four initial clustering centroid nodes and their affiliated nodes are obtained. After recalculating the clustering centroid nodes within this assigned value interval, new clustering centroid nodes are obtained, and then the clustering radius calculation and node division are performed again. This is executed sequentially until the assigned values of all nodes in each assigned value interval are equivalent, and then the clustering is completed.
[0116] This method of dividing management areas and determining the host 301 based on the clustering algorithm can ensure that the business processing capabilities of the nodes in a single management area are balanced. The node at the centroid is selected as the host 301 of this management area, and the other devices in this management area are slaves 302.
[0117] Based on the above embodiments, the master-slave operation control method provided by the embodiments of the present invention uses a clustering algorithm to divide the management area, ensuring that nodes with similar initial business processing capacity test results are divided into the same management area. Then, the node with the strongest initial business capacity is selected as the master, laying a good foundation for optimizing the operation performance of the master-slave in the distributed cluster.
[0118] The following describes Embodiment 4 of the present invention.
[0119] Based on the above embodiments, the embodiments of the present invention further illustrate the master-slave switching scheme.
[0120] To achieve zero-delay switching of the master-slave, in the master-slave operation control method provided by the embodiments of the present invention, S204: determining the target management area to perform the master-slave switching operation according to the real-time business processing capacity test results may include:
[0121] If the real-time business processing capacity test result of the master in the management area is lower than that of the slaves exceeding the first quantity threshold in the management area, and there is a high-performance slave among the slaves whose real-time business processing capacity test result is within the first quantity threshold and the remaining storage space is greater than the second threshold, and there are auxiliary slaves reaching the third quantity threshold among the slaves, then the management area is determined as the target management area.
[0122] In specific implementation, the first quantity threshold may be 30%, that is, when the real-time business processing capacity test result of the master is lower than the top 30% of the management area where it is located, it is considered that the master has a performance lag.
[0123] If the master-slave switch is directly performed, it will cause the host service to pause, bringing potential safety hazards to this management area and a bad experience to users. Therefore, it is also necessary to determine whether there are auxiliary slaves with redundant capabilities in the management area to share the service data of the host during the master-slave switching process. To avoid affecting the original services of the auxiliary slaves, multiple auxiliary slaves need to be determined, and the remaining storage space and computing performance of the auxiliary slaves should meet the requirement of being able to share the service data of the host. The third quantity threshold may be three.
[0124] In addition, the selected new master should have better performance than the old master and have enough space to store the service data of the old master. Then, the new master should meet the condition that the real-time business processing capacity test result is within the first quantity threshold and the remaining storage space is greater than the second threshold, such as a slave whose real-time business processing capacity test result is within the top 30% of the management area where it is located and has sufficient remaining space.
[0125] After meeting the above conditions, the management area is considered as the target management area with the condition of master-slave switch. If only the performance of the master is backward in the management area, but there is no auxiliary slave that can assist the master to complete zero-delay switching, it can be considered that the management area does not meet the master-slave switch condition, and wait for the appearance of a qualified auxiliary slave before performing the master-slave switch.
[0126] Based on this, S205: Performing the master-slave switch operation on the target management area may include:
[0127] Transfer the host service data of the host to the auxiliary slave through the fiber optic switch, and record the increment of the auxiliary slave processing the host service data;
[0128] After transferring all the host service data of the host to the auxiliary slave, transfer the management authority and monitoring authority of the host to the new host in the management area;
[0129] The new host sends control instructions to each slave, and determines that the permission switch is completed after receiving the response instructions from all slaves in the management area;
[0130] After the new host determines that the permission switch is completed, receive the host service data fed back by each auxiliary slave and the increment of each auxiliary slave processing the host service data, and continue to process the host service data from the latest execution time point of the host service data.
[0131] In a specific implementation, through the fiber optic switch, the host service data can be quickly transmitted to the auxiliary slave. The auxiliary slave starts to execute the host service externally and records the incremental content generated after receiving the host service data. After determining that all the service data of the host has been transferred to the auxiliary slave, then perform the task of transferring the management authority and monitoring authority of the host to the new host in the management area. Specifically, it can further include copying the control data of the host to the new host first, and then performing the switch of the management authority and monitoring authority. Then, as described in the second embodiment of the present invention, the new host sends control instructions carrying the specific identifier of the management area where it is located to each slave to inform other slaves of the information of the host switch, and determines that the permission switch is completed only after receiving the response instructions from all slaves. If the response instructions from all slaves are not received after the timeout, it is considered that the master-slave switch fails. After the new host determines that the permission switch is completed, it obtains the host service data from each auxiliary slave and the increment generated by each auxiliary slave processing the host service data during the master-slave switch, and continues to process the host data from the latest execution time point of the host service data. Thus, for the entire management area and for the users of this part of the service data, a zero-delay and imperceptible master-slave switch is achieved.
[0132] The master-slave operation control method provided by the embodiment of the present invention selects an auxiliary slave with better performance to temporarily process the service data of the master when the performance of the master in the management area lags behind, and then switches the management authority and monitoring authority of the master to the high-performance slave in the management area, so as to keep the master in the management area as a high-performance master and achieve zero-latency switching between the master and the slave.
[0133] The following describes Embodiment 5 of the present invention.
[0134] Based on the above embodiments, the embodiment of the present invention provides another master-slave operation control method.
[0135] For the specific implementation manner of S201, please refer to the description of the above embodiments.
[0136] In the embodiment of the present invention, S202: According to the initial service processing capacity test results, select multiple hosts from each node and divide the management areas where each host is located, so that each management area includes at least one host and one slave, and the initial service processing capacity test results of the nodes in the same management area are similar. It may include:
[0137] Randomly select multiple nodes from each node as the initial clustering centroid nodes;
[0138] Assign values according to the initial service processing capacity test results, calculate the clustering radius between the remaining nodes and the initial clustering centroid nodes, so as to divide the remaining nodes into the initial management areas where the initial clustering centroid nodes with close radius distances are located;
[0139] Loop to execute re-selecting the clustering centroid nodes in each management area, and assign values according to the initial service processing capacity test results, calculate the clustering radius between the remaining nodes and the clustering centroid nodes, so as to divide the remaining nodes into the management areas where the clustering centroid nodes with close radius distances are located, until each management area includes at least one host and one slave, and the initial service processing capacity test results of the nodes in the same management area are similar.
[0140] For the specific implementation manner of S203, please refer to the description of the above embodiments.
[0141] S204: Determine the target management area for which the master-slave switching operation is to be performed according to the real-time service processing capacity test results. It may include:
[0142] If the real-time service processing capacity test result of the host in the management area is lower than that of the slaves in the management area that reach the first quantity threshold, and there is a high-performance slave among the slaves whose real-time service processing capacity test result is within the first quantity threshold and the remaining storage space is greater than the second threshold, and there are auxiliary slaves among the slaves that reach the third quantity threshold, then determine the management area as the target management area.
[0143] S205: Perform the master-slave switch operation on the target management area to ensure that the master in each management area is not a host with poor performance in the non-management area, which may include:
[0144] Transfer the host service data of the host to the auxiliary slave through the fiber optic switch, and record the increment of the host service data processed by the auxiliary slave;
[0145] After transferring all the host service data of the host to the auxiliary slave, transfer the management authority and monitoring authority of the host to the new host in the management area;
[0146] The new host sends control instructions to each slave, and determines that the permission switch is completed after receiving the response instructions from all slaves in the management area;
[0147] After the new host determines that the permission switch is completed, receive the host service data fed back by each auxiliary slave and the increment of the host service data processed by each auxiliary slave, and continue to process the host service data from the latest execution time point of the host service data.
[0148] The master-slave operation control method provided by the present invention tests the service processing capabilities of each node when creating a distributed cluster, divides the management area according to the test results, so that at least one host and one slave are included in one management area and the initial service processing capability test results of each node are similar; during the operation of the distributed cluster, monitor the service processing capabilities of each node to obtain the real-time service processing capability test results of each node, so as to determine the target management area and perform the master-slave switch operation on the target management area, so that the master in each management area is not a host with poor performance in this management area, realizing a distributed cluster master-slave allocation scheme based on service performance, which can give full play to the service performance of the nodes in the distributed cluster to improve the overall service performance of the distributed cluster. The clustering algorithm is also used to divide the management area to ensure that nodes with similar initial service processing capability test results are divided into the same management area, and then the one with the strongest initial service ability is selected as the master, laying a good foundation for optimizing the master-slave operation performance of the distributed cluster. Also, when the performance of the master in the management area is backward, select an auxiliary slave with better performance to temporarily process the service data of the master, and then switch the management authority and monitoring authority of the master to the high-performance slave in the management area to achieve keeping the master in the management area as a high-performance host and zero-latency switching of the master and slave.
[0149] The above details each embodiment corresponding to the master-slave operation control method. On this basis, the present invention also discloses a master-slave operation control device, equipment and readable storage medium corresponding to the above method.
[0150] The following describes Embodiment 6 of the present invention.
[0151] Figure 4 The figure is a schematic structural diagram of a master-slave machine operation control device provided by an embodiment of the present invention.
[0152] As Figure 4 shown, the master-slave machine operation control device provided by the embodiment of the present invention includes:
[0153] A test unit 401, configured to, when creating a distributed cluster, perform a service processing capacity test on each node according to the service types required to be executed by each node of the distributed cluster, and obtain an initial service processing capacity test result of each node;
[0154] A partitioning unit 402, configured to select multiple hosts from each node and partition the management areas where each host is located according to the initial service processing capacity test result, so that each management area includes at least one host and one slave, and the initial service processing capacity test results of the nodes in the same management area are similar;
[0155] A monitoring unit 403, configured to monitor the service processing capacity of each node during the operation of the distributed cluster, and obtain a real-time service processing capacity test result of each node;
[0156] A determining unit 404, configured to determine a target management area for performing a master-slave machine switching operation according to the real-time service processing capacity test result;
[0157] A control unit 405, configured to perform a master-slave machine switching operation on the target management area to ensure that the host in each management area is a host with poor performance in the non-management area.
[0158] In some embodiments, the partitioning unit 402 selects multiple hosts from each node and partitions the management areas where each host is located according to the initial service processing capacity test result, so that each management area includes at least one host and one slave, and the initial service processing capacity test results of the nodes in the same management area are similar, including:
[0159] Randomly select multiple nodes from each node as initial clustering centroid nodes;
[0160] Assign values according to the initial service processing capacity test result, calculate the clustering radius between the remaining nodes and the initial clustering centroid nodes, and partition the remaining nodes into the initial management areas where the initial clustering centroid nodes with close radius distances are located;
[0161] Loop and execute to reselect the clustering centroid nodes in each management area, and use the initial business processing capacity test results as assignments to calculate the clustering radius between the remaining nodes and the clustering centroid nodes, so as to divide the remaining nodes into the management area where the clustering centroid nodes with closer radius distances are located, until each management area includes at least one host and one slave, and the initial business processing capacity test results of the nodes in the same management area are similar.
[0162] In some embodiments, the determining unit 404 determines the target management area for performing the master-slave switch operation according to the real-time business processing capacity test results, including:
[0163] If the real-time business processing capacity test result of the host in the management area is lower than that of the slaves exceeding the first quantity threshold in the management area, and there is a high-performance slave among the slaves whose real-time business processing capacity test result is within the first quantity threshold and the remaining storage space is greater than the second threshold, and there are auxiliary slaves reaching the third quantity threshold among the slaves, then determine the management area as the target management area.
[0164] In some embodiments, the control unit 405 performs the master-slave switch operation on the target management area, including:
[0165] Transfer the host service data of the host to the auxiliary slave through the fiber optic switch, and record the increment of the auxiliary slave processing the host service data;
[0166] After transferring all the host service data of the host to the auxiliary slave, transfer the management authority and the monitoring authority of the host to the new host in the management area;
[0167] The new host sends control instructions to each slave, and determines that the permission switch is completed after receiving the response instructions of all the slaves in the management area;
[0168] After the new host determines that the permission switch is completed, receive the host service data fed back by each auxiliary slave and the increment of each auxiliary slave processing the host service data, and continue to process the host service data from the latest execution time point of the host service data.
[0169] In some embodiments, the distributed cluster is a distributed storage cluster;
[0170] The testing unit 401 performs business processing capacity tests on each node according to the business types required to be executed by each node of the distributed cluster, and obtains the initial business processing capacity test results of each node, including:
[0171] After each node is powered on, perform read and write processing on the business data of the fourth quantity threshold for each node;
[0172] Taking at least one of throughput, average number of operations per second, and latency time when a node processes service data as the initial service processing capacity test result of the node;
[0173] The monitoring unit 403 monitors the service processing capabilities of each node, and obtains the real-time service processing capacity test results of each node, including:
[0174] Monitoring at least one of throughput, average number of operations per second, and latency time when a node processes actual service data, and obtaining the real-time service processing capacity test results of each node.
[0175] In some embodiments, the master-slave operation control device provided by the embodiments of the present invention may further include:
[0176] A first distribution unit, configured to, for a distributed cluster, sequentially distribute service data with priorities from high to low and / or service data with data volumes from large to small to management areas in descending order of the real-time service processing capacity test results of the management areas; if all service data of the distributed cluster has been distributed to the management areas, end the cluster-level distribution optimization of the service data in the current batch; if there is still remaining service data after all management areas have been distributed, continue to sequentially distribute the remaining service data with priorities from high to low and / or the remaining service data with data volumes from large to small to the management areas in descending order of the real-time service processing capacity test results of the management areas until all service data of the distributed cluster has been distributed to the management areas;
[0177] A second distribution unit, configured to, for a management area, sequentially distribute the service data to nodes with real-time service processing capacity test results from high to low for data processing and disk writing in descending order of the priority of the service data and / or in descending order of the data volume of the service data, and select nodes for backup of the processed service data in descending order of the real-time service processing capacity test results and in descending order of the remaining storage space.
[0178] In some embodiments, each node is interconnected through a gigabit switch to transmit control data;
[0179] Each node is interconnected through a fiber optic switch to transmit service data.
[0180] In some embodiments, the partitioning unit 402 selects multiple hosts from each node according to the initial service processing capacity test results, and partitions the management areas where each host is located, so that each management area includes at least one host and one slave, and the initial service processing capacity test results of the nodes in the same management area are similar, including:
[0181] Randomly selecting multiple nodes from each node as initial clustering centroid nodes;
[0182] Assign values based on the initial business processing capacity test results, calculate the clustering radius between the remaining nodes and the initial clustering centroid nodes, and divide the remaining nodes into the initial management areas where the initial clustering centroid nodes with closer radius distances are located;
[0183] Loop to execute reselecting clustering centroid nodes in each management area, assign values based on the initial business processing capacity test results, calculate the clustering radius between the remaining nodes and the clustering centroid nodes, and divide the remaining nodes into the management areas where the clustering centroid nodes with closer radius distances are located until each management area includes at least one host and one slave, and the initial business processing capacity test results of the nodes in the same management area are similar;
[0184] The determination unit 404 determines the target management area to perform the master-slave switch operation according to the real-time business processing capacity test results, including:
[0185] If the real-time business processing capacity test result of the host in the management area is lower than that of the slaves exceeding the first quantity threshold in the management area, and there is a high-performance slave among the slaves whose real-time business processing capacity test result is within the first quantity threshold and the remaining storage space is greater than the second threshold, and there are auxiliary slaves reaching the third quantity threshold among the slaves, then determine the management area as the target management area;
[0186] The control unit 405 performs the master-slave switch operation on the target management area to ensure that the host in each management area is not a host with poor performance in the non-management area, including:
[0187] Transfer the host business data of the host to the auxiliary slave through the fiber optic switch, and record the increment of the auxiliary slave processing the host business data;
[0188] After transferring all the host business data of the host to the auxiliary slave, transfer the management authority and the monitoring authority of the host to the new host in the management area;
[0189] The new host sends control instructions to each slave, and determines that the permission switch is completed after receiving the response instructions from all the slaves in the management area;
[0190] After the new host determines that the permission switch is completed, receive the host business data fed back by each auxiliary slave and the increment of each auxiliary slave processing the host business data, and continue to process the host business data from the latest execution time point of the host business data.
[0191] Since the embodiments in the device part correspond to the embodiments in the method part, for the embodiments in the device part, please refer to the description of the embodiments in the method part, which will not be elaborated here for the time being.
[0192] The following describes Embodiment 7 of the present invention.
[0193] Figure 5 The structure diagram of a master-slave machine operation control device provided by an embodiment of the present invention.
[0194] As Figure 5 shown, the master-slave machine operation control device provided by an embodiment of the present invention includes:
[0195] A memory 510 for storing a computer program 511;
[0196] A processor 520 for executing the computer program 511, and when the computer program 511 is executed by the processor 520, the steps of the master-slave machine operation control method described in any one of the above embodiments are implemented.
[0197] Among them, the processor 520 may include one or more processing cores, such as a 3-core processor, an 8-core processor, etc. The processor 520 may be implemented in at least one hardware form of a digital signal processor DSP (Digital Signal Processing), a field programmable gate array FPGA (Field-Programmable Gate Array), and a programmable logic array PLA (Programmable Logic Array). The processor 520 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the central processing unit CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 520 may be integrated with a graphics processing unit GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 520 may further include an artificial intelligence AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0198] The memory 510 may include one or more readable storage media, which may be non-transitory. The memory 510 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 510 is at least used to store the following computer program 511. After the computer program 511 is loaded and executed by the processor 520, it can implement the relevant steps in the master-slave operation control method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 510 may also include an operating system 512, data 513, etc., and the storage method may be transient storage or permanent storage. Among them, the operating system 512 may be Windows. The data 513 may include, but is not limited to, the data involved in the above method.
[0199] In some embodiments, the master-slave operation control device may further include a display screen 530, a power supply 540, a communication interface 550, an input / output interface 560, a sensor 570, and a communication bus 580.
[0200] Those skilled in the art can understand that Figure 5 the structure shown in does not constitute a limitation on the master-slave operation control device, and may include more or fewer components than shown in the figure.
[0201] The master-slave operation control device provided by the embodiment of the present invention includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the master-slave operation control method as described above, and the effect is the same.
[0202] Next, Embodiment VIII of the present invention will be described.
[0203] It should be noted that the device and equipment embodiments described above are only illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other may be through some interfaces, and the indirect couplings or communication connections of devices or modules may be electrical, mechanical or other forms. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0204] In addition, in each embodiment of the present invention, each functional module can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0205] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods described in each embodiment of the present invention.
[0206] Therefore, an embodiment of the present invention further provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the master-slave machine operation control method as described above.
[0207] The readable storage medium may include: various media such as USB flash drives, mobile hard disks, read-only memory ROM (Read-Only Memory), random access memory RAM (Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0208] The computer program included in the readable storage medium provided in this embodiment can implement the steps of the master-slave machine operation control method as described above when executed by a processor, and the effect is the same as above.
[0209] The above has introduced in detail a master-slave machine operation control method, device, equipment, and readable storage medium provided by the present invention. The embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices, equipment, and readable storage mediums disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description of the method part. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
[0210] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
Claims
1. A master-slave machine operation control method, characterized in that, it includes: When creating a distributed cluster, according to the service types required to be executed by each node of the distributed cluster, perform service processing capacity tests on each of the nodes to obtain the initial service processing capacity test results of each of the nodes; According to the initial service processing capacity test results, select multiple hosts from each of the nodes and divide the management areas where each of the hosts is located, so that each of the management areas includes at least one of the hosts and one slave machine, and the initial service processing capacity test results of each of the nodes in the same management area are similar; During the operation of the distributed cluster, monitor the service processing capacity of each of the nodes to obtain the real-time service processing capacity test results of each of the nodes; Determine the target management area to perform the master-slave machine switching operation according to the real-time service processing capacity test results; Perform the master-slave machine switching operation on the target management area to ensure that the host in each of the management areas is not the performance-backward host within the management area.
2. The master-slave machine operation control method according to claim 1, characterized in that, the step of selecting multiple hosts from each of the nodes and dividing the management areas where each of the hosts is located according to the initial service processing capacity test results, so that each of the management areas includes at least one of the hosts and one slave machine, and the initial service processing capacity test results of each of the nodes in the same management area are similar, includes: Randomly select multiple of the nodes from each of the nodes as initial clustering centroid nodes; Assign values according to the initial service processing capacity test results, calculate the clustering radius between the remaining nodes and the initial clustering centroid nodes, and divide the remaining nodes into the initial management areas where the initial clustering centroid nodes with closer radius distances are located; Loop to execute the step of reselecting clustering centroid nodes in each of the management areas, and assign values according to the initial service processing capacity test results, calculate the clustering radius between the remaining nodes and the clustering centroid nodes, and divide the remaining nodes into the management areas where the clustering centroid nodes with closer radius distances are located, until each of the management areas includes at least one of the hosts and one of the slave machines, and the initial service processing capacity test results of each of the nodes in the same management area are similar.
3. The master-slave machine operation control method according to claim 1, characterized in that, the step of determining the target management area to perform the master-slave machine switching operation according to the real-time service processing capacity test results, includes: If the real-time service processing capacity test result of the host in the management area is lower than that of the first number threshold of slave machines in the management area, and there is a high-performance slave machine among the slave machines whose real-time service processing capacity test result is within the first number threshold and the remaining storage space is greater than the second threshold, and there are auxiliary slave machines reaching the third number threshold among the slave machines, then determine the management area as the target management area.
4. The master-slave machine operation control method according to claim 3, characterized in that, Performing the master-slave switch operation on the target management area includes: Transferring the host service data of the host to the auxiliary slave through a fiber optic switch, and recording the increment of the host service data processed by the auxiliary slave; After transferring all the host service data of the host to the auxiliary slave, transferring the management authority and the monitoring authority of the host to the new host in the management area; The new host sends control instructions to each slave, and determines that the permission switch is completed after receiving the response instructions from all the slaves in the management area; After determining that the permission switch is completed, the new host receives the host service data fed back by each auxiliary slave and the increment of the host service data processed by each auxiliary slave, and continues to process the host service data from the latest execution time point of the host service data.
5. The master-slave operation control method according to claim 1, characterized in that, The distributed cluster is a distributed storage cluster; Performing a service processing capacity test on each node according to the service types required to be executed by each node of the distributed cluster to obtain the initial service processing capacity test results of each node, including: After each node is powered on, performing read and write processing of a fourth quantity threshold of service data on each node; Using at least one of the throughput, average number of operations per second, and latency time when the node processes service data as the initial service processing capacity test result of the node; Monitoring the service processing capacity of each node to obtain the real-time service processing capacity test results of each node, including: Monitoring at least one of the throughput, average number of operations per second, and latency time when the node processes actual service data to obtain the real-time service processing capacity test results of each node.
6. The master-slave operation control method according to claim 1, characterized in that, It further includes: For the distributed cluster, sequentially allocating service data with decreasing priority and / or service data with decreasing data volume in the order from high to low of the real-time service processing capacity test results of the management area; If all the service data of the distributed cluster has been allocated to the management area, end the cluster-level allocation optimization of the service data in the current batch; If there is still remaining service data after allocating all the management areas, continue to sequentially allocate the remaining service data with decreasing priority and / or the remaining service data with decreasing data volume in the order from high to low of the real-time service processing capacity test results of the management area until all the service data of the distributed cluster is allocated to the management area; For the management area, sequentially allocate to the nodes with decreasing real-time service processing capacity test results for data processing and disk writing according to the order of decreasing priority of the service data and / or the order of decreasing data volume of the service data, and select the nodes for backup of the processed service data in the order of decreasing real-time service processing capacity test results and decreasing remaining storage space.
7. The master-slave operation control method according to claim 1, characterized in that, each of the nodes is interconnected through a gigabit switch to transmit control data; each of the nodes is interconnected through a fiber optic switch to transmit service data.
8. The master-slave operation control method according to claim 1, characterized in that, selecting multiple hosts from each of the nodes and dividing the management areas where the hosts are located according to the initial service processing capacity test results, so that each of the management areas includes at least one of the hosts and one slave, and the initial service processing capacity test results of the nodes in the same management area are similar, including: randomly selecting a plurality of the nodes from each of the nodes as initial clustering centroid nodes; assigning values based on the initial service processing capacity test results, calculating the clustering radius between the remaining nodes and the initial clustering centroid nodes, and dividing the remaining nodes into the initial management areas where the initial clustering centroid nodes with close radius distances are located; repeatedly execute the steps of reselecting clustering centroid nodes in each of the management areas, assigning values based on the initial service processing capacity test results, calculating the clustering radius between the remaining nodes and the clustering centroid nodes, and dividing the remaining nodes into the management areas where the clustering centroid nodes with close radius distances are located, until each of the management areas includes at least one of the hosts and one of the slaves, and the initial service processing capacity test results of the nodes in the same management area are similar; determining the target management area for performing the master-slave switching operation according to the real-time service processing capacity test results, including: if the real-time service processing capacity test result of the host in the management area is lower than that of the slaves in the management area that reach the first quantity threshold, and there is a high-performance slave among the slaves whose real-time service processing capacity test result is within the first quantity threshold and the remaining storage space is greater than the second threshold, and there are auxiliary slaves in the slaves that reach the third quantity threshold, then determine the management area as the target management area; performing a master-slave switching operation on the target management area to ensure that the host in each management area is not the host with backward performance in the management area, including: transferring the host service data of the host to the auxiliary slave through a fiber optic switch, and recording the increment of the host service data processed by the auxiliary slave; after transferring all the host service data of the host to the auxiliary slave, transferring the management authority and the monitoring authority of the host to the new host in the management area; the new host sends control instructions to each of the slaves, and determines that the permission switching is completed after receiving the response instructions from all the slaves in the management area; after the new host determines that the permission switching is completed, receiving the host service data fed back by each of the auxiliary slaves and the increment of the host service data processed by each of the auxiliary slaves, and continuing to process the host service data from the latest execution time point of the host service data.
9. A master-slave machine operation control device, characterized in that, it includes: A test unit, which is used to test the service processing capabilities of each node according to the service types required by each node of the distributed cluster when creating the distributed cluster, and obtain the initial service processing capability test results of each node; A partitioning unit, which is used to select multiple hosts from each node and partition the management areas where each host is located according to the initial service processing capability test results, so that each management area includes at least one host and one slave, and the initial service processing capability test results of each node in the same management area are similar; A monitoring unit, which is used to monitor the service processing capabilities of each node during the operation of the distributed cluster, and obtain the real-time service processing capability test results of each node; A determination unit, which is used to determine the target management area to perform the master-slave switch operation according to the real-time service processing capability test results; A control unit, which is used to perform the master-slave switch operation on the target management area to ensure that the host in each management area is not the host with poor performance in the management area.
10. A master-slave machine operation control device, characterized in that, it includes: A memory, which is used to store computer programs; A processor, which is used to execute the computer program. When the computer program is executed by the processor, it realizes the steps of the master-slave machine operation control method according to any one of claims 1 to 8.
11. A readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, it realizes the steps of the master-slave machine operation control method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Packet scheduling method and system for multi-antenna radio communication system
CN101034923A
Distributed ad-hoc network intelligent garbage can clearing system and method
CN111422530A
Monitoring device and method based on multiple unmanned aerial vehicles
CN113970931A
Cluster deployment method and device, equipment, and computer readable storage medium
CN114070739A
Task scheduling method and device under cloud platform, medium, equipment and program product
CN114791855A