Data balancing method and related equipment
Through intelligent scheduling services and multi-balancer configuration, the node status is monitored in real time, and local balance operations are identified and triggered, which solves the problem of unbalanced data block distribution in the HDFS system, and efficient and flexible data balance is achieved, improving system stability and resource utilization.
Patent Information
- Application Number
- CN202510203725.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-04
AI Technical Summary
The uneven distribution of data blocks between nodes in the existing HDFS system leads to waste of storage resources and system stability issues. The existing tools rely on a single balancer instance for global equalization, with high bandwidth consumption and low flexibility.
Introduce intelligent scheduling services, by monitoring the internal node status of the computer room in real time, identifying the unbalanced computer room, triggering the local balancer to perform internal balance operations, reducing data transmission across computer rooms, and adopting multi-balancer configuration to improve flexibility and efficiency.
It reduces operation and maintenance complexity, reduces network consumption across computer rooms, improves system stability and flexibility, ensures even data distribution, and improves the overall operating performance of the system.
Smart Images

Figure CN120256089A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of distributed storage technologies, and more particularly to a data balancing method, apparatus, server, computer-readable storage medium, and computer program product. Background Art
[0002] Data in HDFS (Hadoop Distributed File System) is stored in the form of blocks, and each data block will be replicated multiple times. As time goes by, the amount of data stored in the system will continue to grow, and the distribution of data blocks among nodes may become unbalanced. Some nodes may store too much data, while other nodes may be idle, resulting in waste of storage resources. In order to improve the utilization rate of storage and ensure that data is evenly distributed across all nodes, data balancing is crucial in the HDFS deployment with a dual-data center or multi-data center architecture.
[0003] Currently, there are some tools for HDFS data balancing, such as Cloudera Manager, Tectonic, etc. These tools rely on a single Balancer instance for global balancing, consuming a large amount of bandwidth and having a large network latency. Moreover, the balancing flexibility is relatively low through manual management and scheduling strategies via the interface. Summary of the Invention
[0004] In view of the above-mentioned defects or deficiencies in the prior art, it is desirable to provide a data balancing method, apparatus, server, computer-readable storage medium, and computer program product, which can significantly reduce bandwidth consumption while achieving efficient and flexible load balancing control.
[0005] In a first aspect, an embodiment of the present application provides a data balancing method, including:
[0006] Monitoring the node status in each data center, evaluating the node load distribution within the data center, and generating a load evaluation result;
[0007] Identifying, based on the load evaluation result, the data center with unbalanced node loads as the target data center;
[0008] Triggering the balancer within the target data center to perform local balancing operations on the internal nodes of the data center.
[0009] In one embodiment, it further includes:
[0010] Obtaining the historical load evaluation results of each data center;
[0011] Predict the time of load imbalance for each computer room according to the historical load evaluation results, and use each computer room as the target computer room to obtain the predicted time corresponding to each target computer room;
[0012] Determine the startup time of the balancer according to the predicted time;
[0013] When the startup time is reached, execute the step of triggering the balancer in the target computer room to perform local balancing operations on the internal nodes of the affiliated computer room.
[0014] In one embodiment, it further includes:
[0015] Identify data blocks whose access frequency exceeds the threshold as hot data;
[0016] Generate data copies of the hot data;
[0017] Trigger the balancer in the computer room where the hot data belongs to perform local balancing operations on the hot data and the corresponding data copies.
[0018] In one embodiment, it further includes:
[0019] If the node status shows that the node is abnormal, obtain the load evaluation results to evaluate the cross-computer room balancing requirements, and generate a requirements evaluation result;
[0020] If the requirements evaluation result meets the cross-computer room balancing standard, perform cross-computer room data migration on the nodes showing abnormalities.
[0021] In one embodiment, the monitoring of the node status in each computer room includes:
[0022] Monitor the load status and network status of each node in the computer room as the node status.
[0023] In one embodiment, if there are two or more target computer rooms, triggering the balancer in the target computer room to perform local balancing operations on the internal nodes of the affiliated computer room includes:
[0024] Parallelly trigger the balancers in two or more target computer rooms at the same time to perform local balancing operations on the internal nodes of the computer rooms to which each balancer belongs.
[0025] In a second aspect, an embodiment of the present application provides a data balancing device, including:
[0026] A node monitoring unit, configured to monitor the node status in each computer room, evaluate the node load distribution in the computer room, and generate a load evaluation result;
[0027] A load evaluation unit, configured to identify, according to the load evaluation result, a computer room with unbalanced node loads as a target computer room;
[0028] An equilibrium execution unit, configured to trigger a balancer in the target computer room to perform a local balancing operation on internal nodes of the computer room to which it belongs.
[0029] In a third aspect, an embodiment of the present application provides a server, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method described in the embodiment of the present application are implemented.
[0030] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in the embodiment of the present application are implemented.
[0031] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method described in the embodiment of the present application are implemented.
[0032] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent:
[0034] Figure 1 A flowchart showing the data balancing method provided by the embodiment of the present application;
[0035] Figure 2 An architecture diagram showing an adaptive balancing control method provided by the embodiment of the present application;
[0036] Figure 3 A schematic diagram showing a fine control method for cross-computer room balancing provided by the embodiment of the present application;
[0037] Figure 4 A schematic block diagram showing an exemplary structure of the data balancing device provided by the embodiment of the present application;
[0038] Figure 5 A schematic diagram showing the structure of a computer system of a server suitable for implementing the embodiment of the present application. DETAILED DESCRIPTION
[0039] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the convenience of description, only the parts related to the invention are shown in the drawings.
[0040] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments. Although the embodiments of the present application provide the method operation instruction steps as shown in the following embodiments or drawings, more or fewer operation instruction steps may be included in the method based on routine or non-creative labor. In the steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided in the embodiments of the present application. When the method is actually processed or executed by the device, it can be executed in the method order shown in the embodiments or drawings or executed in parallel.
[0041] Under the disaster recovery architecture of a distributed file system, generally two or more computer room architectures are set up, relying on a single balancer instance for global balancing. However, cross-computer room data transmission will bring relatively high network latency and bandwidth consumption. Especially when performing data balancing, cross-computer room data scheduling may cause unnecessary network overhead. In addition, HDFS by default only allows one balancer service to run simultaneously, which limits flexible computer room-level balancing strategies. In order to achieve efficient data balancing within different computer rooms, avoid frequent cross-computer room operations, and at the same time utilize the intelligent scheduling service to dynamically control the process of data balancing, the present application proposes a data balancing method. Please refer to Figure 1 , Figure 1 which shows the schematic flowchart of the data balancing method provided by an embodiment of the present application. As Figure 1 shown, the method includes:
[0042] S101. Monitor the node status in each computer room, evaluate the node load distribution in the computer room, and generate a load evaluation result;
[0043] The existing HDFS balancer is usually started manually, lacking dynamic scheduling and automated operation. This solution introduces an intelligent scheduling service, which monitors the node status inside the computer room in real time, determines the node load distribution by evaluating the node status, determines the current load situation, automatically triggers the balancing operation when the node load is unbalanced, and dynamically starts or stops the balancer, thereby realizing intelligent balanced scheduling of data. On the one hand, it reduces manual intervention and reduces the complexity of operation and maintenance; on the other hand, by real-time monitoring of the node status, it can timely discover the imbalance of node load, and adjust the data distribution of the cluster in time according to the real-time load, ensuring that the system can be load balanced at any time, and improving the stability and reliability of the overall operation of the system; in addition, based on real-time monitoring data, more flexible load adjustment strategies can be adopted according to different conditions of node load. For example, the balancing strategy can be dynamically adjusted according to the load difference between computer rooms, thereby improving the flexibility of data balancing management.
[0044] Among them, the monitoring of the node status in the computer room can be achieved by integrating Hadoop's own monitoring system or using external monitoring tools (such as Prometheus); in addition, the specific monitoring indicators of node status monitoring are not limited in this embodiment, including but not limited to the node's load status (CPU, storage, network usage, etc.) and network status (data access mode, etc.).
[0045] According to the node status obtained by monitoring, the node load distribution in the computer room is evaluated, and the high-load nodes and low-load nodes in the computer room are identified, which can help determine whether there is a load imbalance in the computer room. Among them, the evaluation indicators are set according to the monitoring items of the node status. For example, by analyzing the usage of CPU, memory and disk I / O, a high load threshold is set (for example, CPU usage exceeds 80%, memory exceeds 75%). When the node's indicators exceed this threshold, the node is considered to be heavily loaded; a low load threshold can also be further set (for example, CPU usage is less than 15%, memory is less than 20%). When the node's indicators are lower than this threshold, the node is considered to be a low-load node and is not effectively used.
[0046] S102, according to the load evaluation result, identifying a computer room with unbalanced node load as a target computer room;
[0047] According to the load evaluation results, identify whether the status of each computer room meets the preset conditions for node load imbalance. If it meets, determine that the computer room is identified as the target computer room with node load imbalance. Among them, the preset conditions for node load imbalance are not limited. For example, some nodes are overloaded while other node resources are idle (load is too low). The manifestations of node overload include, for example, too high CPU usage rate of the node, too large memory occupancy, excessive disk I / O or network bandwidth usage. The manifestations of node resource idleness include, for example, low usage rates of resources such as CPU, memory, and disk of the node, or even almost idle.
[0048] S103. Trigger the balancer in the target computer room to perform local balancing operations on the internal nodes of the affiliated computer room.
[0049] For the identified target computer room with node load imbalance, in this method, trigger the built-in balancer of the target computer room to perform local balancing operations on the internal nodes of the computer room to which the balancer belongs.
[0050] In traditional HDFS deployments, HDFS only allows a single balancer instance to run. Therefore, the granularity of the balancing operation is large and it is impossible to perform local control on specific computer rooms or nodes. In this case of cross-computer room deployment, this may lead to a large amount of cross-computer room data transmission, increasing the network load. In this method, an independent balancer instance runs in each computer room. Only trigger the balancer in the target computer room, and through the balancer in the computer room, perform local balancing operations on the internal nodes of the affiliated computer room. The balancing operation is only carried out within this computer room. By restricting the operation scope of each balancer, unnecessary cross-computer room data transmission is greatly reduced, thereby reducing the cross-computer room network consumption.
[0051] Among them, the implementation method of controlling the balancer in the target computer room to only perform local balancing operations on the internal nodes of the affiliated computer room is not limited. For example, it can be achieved by configuring the dfs.hosts.include and dfs.hosts.exclude files of HDFS in each computer room, and restricting each instance to only perform balancing operations on the internal nodes of the computer room through the include / exclude configuration files of HDFS. Of course, it can also be achieved through other configuration methods, such as customizing the balancer, adjusting the balancer parameters, etc., which will not be elaborated here.
[0052] Dynamic monitoring and real-time balancing trigger for each computer room node, dynamically starting and stopping the balancer according to the actual situation. Through intelligent automatic scheduling trigger, it can reduce user operations while ensuring more reasonable data distribution. In addition, the load balancing strategies of balancers in different computer rooms can be configured accordingly according to different computer room states. For example, it can automatically match corresponding strategies according to the real-time load evaluation results to execute local balancing operations and reallocate resources. The configuration management of various load balancing strategies in the multi-balancer configuration is more flexible and diverse than that of traditional single balancers, thus realizing more flexible balancing management.
[0053] It should be noted that the number of target computer rooms identified in this embodiment is not limited, and it can be one, or two or more. If the number of target computer rooms is two or more, since the balancer only processes nodes within a specified range, to accelerate the overall balancing speed, the balancers in the target computer rooms are triggered to perform local balancing operations on the internal nodes of their respective computer rooms. Specifically, it can be: triggering the balancers in two or more target computer rooms simultaneously in parallel to perform local balancing operations on the internal nodes of the computer rooms to which each balancer belongs. This way of parallel balancing in multiple computer rooms accelerates the overall balancing speed, ensures that the data in the system is quickly distributed to the nodes with lighter loads, and greatly improves the efficiency and flexibility of the balancing operation. Of course, it is also possible not to perform the parallel balancing operation of multiple balancers. This is not limited in this embodiment and can be set according to the actual application scenario, and will not be elaborated here.
[0054] It should be noted that although the operations of the method of the present invention are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result.
[0055] Based on the above introduction, the method provided in this embodiment introduces an intelligent scheduling service. By real-time monitoring the node status inside the computer room, determining the node load distribution by evaluating the node status, determining the current load situation, automatically triggering the balancing operation when the node load is unbalanced, and dynamically starting or stopping the balancer, thus realizing the intelligent balanced scheduling of data. On the one hand, it reduces manual intervention and the complexity of operation and maintenance. On the other hand, by real-time monitoring the node status, it can timely detect the unbalanced situation of node load, and can timely adjust the data distribution of the cluster according to the real-time load, improving the stability and reliability of the overall system operation. In addition, in this method, each computer room runs an independent balancer instance, and only the balancer in the target computer room with unbalanced node load is triggered. The balancer in the computer room performs local balancing operations on the internal nodes of its respective computer room, and the balancing operation is only carried out within this computer room, greatly reducing unnecessary cross-computer-room data transmission, thereby reducing the cross-computer-room network consumption.
[0056] For better understanding, in this embodiment, an implementation process of a data balancing method under a specific monitoring metric is introduced. In this embodiment, the monitoring metrics of node status include: CPU usage rate, memory usage rate, disk I / O usage rate, and network bandwidth usage rate as an example.
[0057] Suppose there are 6 nodes (node 1 to node 6) in computer room A, and the node status data obtained by monitoring is as shown in Table 1 below:
[0058]
[0059]
[0060] Table 1
[0061] Based on the above node status data, evaluating the node load distribution in the computer room, the following load evaluation results can be obtained: The CPU, memory, disk I / O, and network bandwidth usage rates of node 1 and node 4 are relatively high. Especially for node 4, the CPU usage rate (85%), memory usage rate (90%), and disk I / O usage rate (75%) are close to the critical value. The load of node 4 is too heavy, which may affect its performance and the response of other nodes.
[0062] The loads of node 2 and node 3 are relatively low, and almost all metrics are at a low level, which may lead to resource waste.
[0063] According to the above load evaluation results, identify the computer rooms with unbalanced node loads. Here, only the load evaluation results of computer room 1 are used to determine whether the load is balanced. The specific process is as follows:
[0064] The load data of the 6 nodes in computer room A is significantly unbalanced. Node 4 has an overload risk: The high load of node 4 (especially the memory usage rate of 90%) makes it a performance bottleneck of the entire computer room, and there is a risk of system crash or response delay. At the same time, node 3 has resource waste: The load of node 3 is low, and there are a large number of idle resources, which may lead to resource waste, especially the disk I / O usage rate is only 30%. Therefore, computer room A is identified as the target computer room with unbalanced node loads.
[0065] Trigger the balancer built into computer room A. This balancer can select a suitable strategy (such as resource-aware balancing, minimum load algorithm, etc.) according to the real-time load evaluation results to perform local balancing operations and reallocate resources.
[0066] Furthermore, based on the above embodiment, an adaptive balancing control method is proposed in this embodiment, as Figure 2The figure shows the architecture diagram of an adaptive equalization control method. Besides the daily node status monitoring and balancer triggering, by predicting future load fluctuations, equalization operations are carried out in advance to avoid potential bottleneck problems in the future.
[0067] Besides the above steps S101 to S103, the following steps can be further executed:
[0068] S104. Obtain the historical load evaluation results of each computer room;
[0069] Obtain the load evaluation results of each computer room within a certain historical period. The specific historical time range can be set according to factors such as the actual data analysis ability and the available resource volume.
[0070] S105. Predict the time of load imbalance for each computer room based on the historical load evaluation results, and take each computer room as the target computer room to obtain the corresponding prediction time for each target computer room;
[0071] Obtain the historical load evaluation of each computer room. By analyzing the historical load evaluation, it can help identify the resource requirements and performance performance of the system under different time periods and different business loads, so as to predict the time when the computer room becomes the target computer room with load imbalance as the prediction time. For example, during certain specific periods (such as daytime working hours, holidays, etc.), the system's access volume usually increases sharply. By analyzing historical data, the scheduling service can predict these peak periods as the prediction time, so as to prepare the corresponding load balancing measures in advance. Another example is that some services may be triggered regularly, while other services may become sudden due to certain events or external requests. By analyzing the pattern of historical request frequencies, the system can identify when there may be a surge in requests as the prediction time and allocate resources in advance.
[0072] S106. Determine the start time of the balancer according to the prediction time; when the start time is reached, execute S103.
[0073] Determine the appropriate start time according to the predicted future possible load imbalance time, so as to ensure that the equalization operation is started in advance before the prediction time. The determination of the start time should avoid unnecessary resource consumption caused by starting too early, and also avoid missing the best opportunity to adjust the load caused by starting too late, which may further lead to system performance degradation or failures. For the specific determination method of the start time, it can be through training a load prediction model to predict the advance time, setting a trigger threshold close to the load imbalance, etc., which will not be elaborated here. When the start time is reached, execute S103 to trigger the balancer in the corresponding target computer room and start the data equalization operation in advance.
[0074] In addition to real-time data balancing under real-time node status monitoring, the method provided in this embodiment proposes a method for automatically starting balancing operations based on predicted load imbalance. By analyzing historical data access patterns and load fluctuations, the scheduling service can actively identify potential load imbalances based on data-driven predictions and pre-start load balancing operations before problems occur. This not only improves resource utilization efficiency, but also significantly reduces the risk of performance degradation and system crashes caused by load imbalance.
[0075] In a distributed system, hot data refers to frequently requested data or resources. This data usually becomes the center of access for a specific reason (for example, hot news, inventory information of a specific product, real-time transaction data, etc.). Hot data may be concentrated on a certain node rather than evenly distributed on all nodes. Although many existing load balancing solutions can distribute loads based on data copies, they are usually not specially optimized for hot data. The main goal of these solutions is to ensure that data copies are evenly distributed in the cluster to avoid excessive occupation of certain node resources. However, when the system does not optimize hot data, the access volume may be concentrated on a certain storage node or replica, causing the load of the node to rise sharply, while other nodes are relatively idle. This not only increases the latency of the system, but may also lead to slower response time and even the risk of system crash, which in turn affects the stability and high availability of the entire cluster.
[0076] In order to avoid excessive load caused by frequently accessed data, this embodiment proposes a data balancing method for hot data.
[0077] In addition to the above steps S101 to S103, the following steps may be further performed:
[0078] S107, identifying data blocks whose access frequency exceeds a threshold value as hot data;
[0079] Combined with HDFS data access records, data access patterns are analyzed to automatically identify frequently accessed data blocks as hot data. The value setting of the access frequency threshold can be determined according to the specific application scenario and data object, and is not limited here.
[0080] S108, generating a data copy of the hotspot data;
[0081] For the identified hotspot data, more copies are generated in the computer room where the hotspot node is located to ensure that the copies of the hotspot data are evenly distributed to multiple nodes within the computer room, reducing the load pressure on a single node.
[0082] S109: Trigger the balancer in the computer room where the hotspot data belongs to perform a local balancing operation on the hotspot data and the corresponding data replica.
[0083] For hot data and additional data copies, use a balancer for local balancing to ensure that data blocks do not concentrate on certain nodes.
[0084] Based on the above introduction, the method provided in this embodiment actively identifies hot data, creates additional copies near the hot nodes, and triggers the balancer in the computer room to perform local balancing operations on the hot data and its copies, which can alleviate the problem of excessive access load of hot data on a single node or in the computer room and improve the performance and availability of the overall system.
[0085] The above embodiments introduce several data balancing methods in the computer room to avoid the increase in network load and latency caused by excessive cross-computer-room data flow. In order to ensure data balance between computer rooms while reducing unnecessary network overhead, on the basis of the above embodiments, a fine control method for cross-computer-room balancing is proposed in this embodiment, such as Figure 3 shown in the schematic diagram of the fine control method for cross-computer-room balancing.
[0086] In addition to the above steps S101 to S103, the following steps can be further executed:
[0087] S110. If the node status shows that the node is abnormal, obtain the load evaluation result to evaluate the cross-computer-room balancing requirement, and generate a requirement evaluation result;
[0088] By default, the balancers in each computer room only perform data balancing locally to avoid cross-computer-room data transmission and reduce the network bandwidth pressure. When it is detected that a node in a certain computer room is abnormal, such as node failure or extreme data imbalance, etc., the load evaluation result is further obtained to evaluate the cross-computer-room balancing requirement. According to the network status and load evaluation, it is evaluated whether it will have a serious adverse impact on the overall stable operation of the system to minimize cross-computer-room data transmission. Specifically, the cross-computer-room balancing requirement can be evaluated by evaluating aspects such as the network bandwidth between computer rooms, the latency of cross-computer-room data transmission, and the stability and reliability of the network connection between computer rooms. This embodiment does not limit this, and only the above is used as an example for introduction. Other evaluation methods can refer to the introduction of this embodiment and will not be elaborated here.
[0089] S111. If the requirement evaluation result meets the cross-computer-room balancing standard, perform cross-computer-room data migration on the node with abnormal display.
[0090] If the requirement evaluation result meets the preset cross-computer-room balancing standard, perform cross-computer-room data migration on the node with abnormal display.
[0091] This embodiment proposes a data balancing method with internal balancing in the computer room taking precedence and cross-computer room balancing being restricted. By strictly controlling cross-computer room balancing operations and only allowing cross-computer room data balancing under special circumstances, unnecessary network overhead is greatly reduced while ensuring overall system data balance.
[0092] In the above embodiments, several automated data balancing solutions are proposed. It should be noted that in addition to the above automated means, in order to expand and enrich the means of data balancing and meet different usage requirements, an API interface for data balancing scheduling can be further configured to facilitate the administrator to manually or automatically control the start and stop of balancer instances, achieving more flexible management. The specific settings of the manual control API can refer to the implementation of related technologies and will not be elaborated here.
[0093] Further referring to Figure 4 , which shows an exemplary structural block diagram of a data balancing device according to an embodiment of the present application, mainly including:
[0094] A node monitoring unit 101 for monitoring the status of nodes in each computer room, evaluating the node load distribution in the computer room, and generating a load evaluation result;
[0095] A load evaluation unit 102 for identifying the computer room with unbalanced node loads as the target computer room according to the load evaluation result;
[0096] An equilibrium execution unit 103 for triggering the balancer in the target computer room to perform local balancing operations on the internal nodes of the affiliated computer room.
[0097] It should be understood that the various units described in the above device correspond to the respective steps in the method described with reference to Figure 1 . Therefore, the operations and features described above for the method also apply to the device and the units included therein, and will not be elaborated here. The device can be pre-implemented in the browser or other secure applications of the server, or can be loaded into the browser or its secure application of the server by means of downloading, etc. The corresponding units in the device can cooperate with the units in the server to implement the solution of the embodiment of the present application.
[0098] For the several units mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described units can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided into being embodied by multiple units.
[0099] It should be noted that for the details not disclosed in the data balancing device of the embodiment of the present application, please refer to the details disclosed in the above embodiments of the present application and will not be elaborated here.
[0100] Refer to the following Figure 5 , Figure 5 which shows a schematic structural diagram of a computer system of a server suitable for implementing the embodiments of the present application.
[0101] As Figure 5 shown, the computer system includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation instructions of the system are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0102] The following components are connected to the I / O interface 505; an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as required. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as required, so that the computer program read from it can be installed into the storage section 508 as required.
[0103] Specifically, according to the embodiments of the present application, the process described above with reference to the flowchart Figure 1 can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowchart. In such an embodiment, the computer program includes program codes for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above functions defined in the system of the present application are executed.
[0104] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operation instructions of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the foregoing module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two connected blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that executes the specified functions or operation instructions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0106] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. Among them, the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0107] As another aspect, this application also provides a computer-readable storage medium. The computer-readable storage medium can be included in the server described in the above embodiments, or can exist alone without being assembled into the server. The above computer-readable storage medium stores one or more programs, and when the above programs are executed by one or more processors, they are used to implement the data balancing method described in this application.
[0108] The above description is only a preferred embodiment of this application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of disclosure involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the foregoing disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A data balancing method, characterized in that, It includes: Monitor the node status in each computer room, evaluate the node load distribution in the computer room, and generate a load evaluation result; According to the load evaluation result, identify the computer rooms with unbalanced node loads as target computer rooms; Trigger the balancer in the target computer room to perform local balancing operations on the internal nodes of the affiliated computer room.
2. The method according to claim 1, characterized in that It also includes: Obtain the historical load evaluation results of each computer room; Predict the time of load imbalance for each computer room based on the historical load evaluation results, and regard each computer room as the target computer room to obtain the corresponding prediction time for each target computer room; Determine the startup time of the balancer according to the prediction time; When the startup time is reached, execute the step of triggering the balancer in the target computer room to perform local balancing operations on the internal nodes of the affiliated computer room.
3. The method according to claim 1, characterized in that, It also includes: Identify the data blocks whose access frequency exceeds the threshold as hot data; Generate data copies of the hot data; Trigger the balancer in the computer room where the hot data belongs to perform local balancing operations on the hot data and the corresponding data copies.
4. The method according to claim 1, characterized in that, It also includes: If the node status shows node anomalies, obtain the load evaluation result to evaluate the cross-computer room balancing requirements and generate a requirements evaluation result; If the requirements evaluation result meets the cross-computer room balancing standard, perform cross-computer room data migration on the nodes showing anomalies.
5. The method according to claim 1, wherein The monitoring of the node status in each computer room includes: Monitor the load status and network status of each node in the computer room as the node status.
6. The method according to claim 1, characterized in that, If there are two or more target computer rooms, triggering the balancer in the target computer room to perform local balancing operations on the internal nodes of the affiliated computer room includes: Parallelly trigger the balancers in two or more target computer rooms at the same time to perform local balancing operations on the internal nodes of the computer rooms to which each balancer belongs.
7. A data balancing device, characterized in that It includes: A node monitoring unit for monitoring the node status in each computer room, evaluating the node load distribution in the computer room, and generating a load evaluation result; A load evaluation unit for identifying the computer rooms with unbalanced node loads as target computer rooms according to the load evaluation result; An equilibrium execution unit for triggering the balancer in the target computer room to perform local balancing operations on the internal nodes of the affiliated computer room.
8. A server, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the program, it implements the steps of the method described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 6.