Data storage method and device, related equipment and computer readable storage medium
By acquiring multiple adjacent storage nodes in a distributed storage system, load detection and virtual node creation, the problem of low storage data adjustment efficiency in the existing technology is solved, and load balancing and system stability are improved.
Patent Information
- Application Number
- CN202411865385.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-06
AI Technical Summary
When the number of physical servers changes, the existing distributed storage system needs to adjust the storage location of all stored data, which reduces the efficiency of system update and iteration.
By acquiring multiple storage nodes, each node is adjacent to two nodes, performing load detection and statistics, determining load exception nodes, creating virtual nodes, and adjusting the storage data of load exception nodes through data migration to achieve load balancing.
It realizes full utilization of storage resources, improves resource utilization and system stability, avoids failures caused by excessive load, and simplifies the expansion of storage capacity.
Smart Images

Figure CN119937914A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data storage, and in particular to a data storage method, apparatus, related equipment and computer-readable storage medium. Background Art
[0002] With the development of distributed technology, there are more and more scenarios where distributed systems are used to execute business. In order to support the operation of distributed systems, data is often stored in a distributed storage manner to meet the distributed system's rapid call for stored data and improve the efficiency of executing business. At present, related technologies often store corresponding storage data according to the distribution of physical servers. Although this can achieve the purpose of distributed storage, when the number of physical servers changes, the storage location of all stored data in the distributed system needs to be adjusted, which reduces the efficiency of distributed system updates and iterations. Summary of the invention
[0003] In order to solve the technical problems existing in the related art, the embodiments of the present application provide a data storage method, apparatus, related equipment and computer-readable storage medium.
[0004] To achieve the above purpose, the technical solution of the embodiment of the present application is implemented as follows:
[0005] In one aspect, an embodiment of the present application provides a data storage method, the method comprising:
[0006] Acquire multiple storage nodes, wherein each storage node is adjacent to two storage nodes;
[0007] Performing load detection on the multiple storage nodes to obtain load data of each storage node, and performing statistics on the load data of each storage node to obtain load parameters of the multiple storage nodes;
[0008] Based on the load parameter and the load data, determining a load-abnormal node in the storage node, and creating a virtual node based on the load-abnormal node;
[0009] By performing data migration on the abnormal load node and the virtual node, the storage data of the abnormal load node is adjusted.
[0010] On the other hand, an embodiment of the present application provides a data storage device, including:
[0011] An acquisition module, used for acquiring a plurality of storage nodes, wherein each storage node is adjacent to two storage nodes;
[0012] A load detection module, used to perform load detection on the multiple storage nodes to obtain load data of each storage node, and to perform statistics on the load data of each storage node to obtain load parameters of the multiple storage nodes;
[0013] A node determination module, configured to determine a load-abnormal node in the storage node based on the load parameter and the load data, and to create a virtual node based on the load-abnormal node;
[0014] The data migration module is used to adjust the storage data of the load-abnormal node by migrating data between the load-abnormal node and the virtual node.
[0015] On the other hand, an embodiment of the present application further provides an electronic device, comprising: a processor and a memory for storing a computer program that can be run on the processor, wherein the processor is used to execute the steps in the above method when running the computer program.
[0016] On the other hand, an embodiment of the present application further provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method are implemented.
[0017] On the other hand, an embodiment of the present application further provides a computer program product, including a computer program, which implements the steps in the above method when executed by a processor.
[0018] The data storage method, apparatus, related equipment and computer-readable storage medium provided in the embodiments of the present application obtain multiple storage nodes, wherein each storage node is adjacent to two storage nodes. By storing the storage data in multiple different storage nodes, the storage capacity can be easily expanded by adding storage nodes to meet the demand for rapid growth in data volume. Afterwards, load detection can be performed on the multiple storage nodes to obtain the load data of each storage node, and the load data of each storage node can be counted to obtain the load parameters of the multiple storage nodes. Based on the load parameters and the load data, the load-abnormal nodes in the storage nodes are determined. By counting the current load conditions of each storage node, the nodes in the storage nodes with abnormal loads can be accurately determined. Finally, based on the load-abnormal nodes, virtual nodes are created. By performing data migration on the load-abnormal nodes and the virtual nodes, the storage data of the load-abnormal nodes are adjusted. By creating virtual nodes, the data stored in the nodes with abnormal loads can be adjusted to achieve load balancing between the nodes, which can ensure that all storage resources are fully utilized and improve resource utilization. At the same time, it can avoid storage node failures due to excessive loads and improve the stability of the storage system. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A schematic diagram of a data storage method according to an embodiment of the present application;
[0020] Figure 2 It is a schematic diagram of the principle of data storage provided by an embodiment of the present application;
[0021] Figure 3 A schematic diagram of the structure of a data storage device according to an embodiment of the present application;
[0022] Figure 4 Schematic diagram of the hardware structure of the embodiment of the present application. DETAILED DESCRIPTION
[0023] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0025] As distributed computing becomes increasingly mature, numerous distributed products are being applied in more and more scenarios. A typical scenario is that distributed products help users achieve decentralized storage and access to massive amounts of data, supporting high system availability. The core principle is a specific data distribution algorithm. Especially in the era of cloud computing, cloud-based distributed systems and products are hot areas of current technological development. The quality of a distributed algorithm directly affects the read and write efficiency and risk response capabilities of the distributed system, and is therefore extremely important.
[0026] Assume that 100 data (assuming they are 100 integers from 0 to 99) need to be stored, and there are 5 physical servers. The 100 data need to be stored on the 5 physical servers, marked as 0 / 1 / 2 / 3 / 4, a total of 5 physical nodes. The technology adopted is mainly the basic hash algorithm. By taking the modulus, the data is taken modulo 5, and the obtained data (0 / 1 / 2 / 3 / 4) is the identifier of the target physical node where the data should be stored. Then the 100 data are stored in these five target physical nodes. Taking the modulus through the hash algorithm is relatively simple, which is conducive to uniform distribution of data. However, when dealing with the addition of new physical nodes (if the number of nodes changes from 5 to 6), all data need to recalculate the storage location and re-hash, which brings huge transmission and network pressure, and can easily cause system crashes.
[0027] Based on this, an embodiment of the present application proposes a data storage method. In various embodiments of the present application, multiple storage nodes are obtained, wherein each storage node is adjacent to two storage nodes. By storing the storage data in multiple different storage nodes, the storage capacity can be easily expanded by adding storage nodes to meet the demand for rapid growth in data volume. Afterwards, load detection can be performed on the multiple storage nodes to obtain the load data of each storage node, and the load data of each storage node is statistically analyzed to obtain the load parameters of the multiple storage nodes. Based on the load parameters and the load data, the load-abnormal nodes in the storage nodes are determined. By counting the current load conditions of each storage node, the nodes in the storage nodes with abnormal loads can be accurately determined. Finally, based on the load-abnormal nodes, virtual nodes are created. By performing data migration on the load-abnormal nodes and the virtual nodes, the storage data of the load-abnormal nodes is adjusted. By creating virtual nodes, the data stored in the nodes with abnormal loads can be adjusted to achieve load balancing between the nodes. This can ensure that all storage resources are fully utilized, thereby improving resource utilization. At the same time, it can avoid storage node failures due to excessive loads, thereby improving the stability of the storage system.
[0028] The present application provides a data storage method. Figure 1 The data storage method of the present invention is shown in FIG. Figure 1 ;like Figure 1 As shown, the method includes:
[0029] Step 101: Acquire multiple storage nodes, wherein each storage node is adjacent to two storage nodes.
[0030] In actual implementation, each storage node may correspond to a physical server. For example, when there are ten physical servers for storing data in a distributed system, there may be ten corresponding storage nodes, and the data stored in each storage node is actually stored in the physical server corresponding to the storage node.
[0031] In actual implementation, each storage node may be adjacent to two other storage nodes. Specifically, the storage nodes may be arranged in an order of a circular ring with the ends connected. For example, the storage nodes include storage node A, storage node B, storage node C, and storage node D, wherein the positions of the storage nodes in the circular ring are storage node A, storage node B, storage node C, storage node D, and storage node A in a clockwise order of the circular ring. At this time, storage node A is adjacent to storage node D and storage node B, storage node B is adjacent to storage node A and storage node C, storage node C is adjacent to storage node B and storage node D, and storage node D is adjacent to storage node C and storage node A.
[0032] It should be noted that in order to ensure the stability of the system after adding new nodes and avoid errors in data transmission, after setting up the ring and the storage nodes on the ring, it is necessary to limit the storage nodes to migrate the stored data in a specific order. For example, the storage node can only migrate data to the next storage node in the clockwise direction. Continuing with the above example, although storage node A is adjacent to storage node D and storage node B, the storage data in storage node A can only be migrated to storage node B, and the storage data in storage node A cannot be migrated to storage node D. Of course, the storage node can also only migrate data to the next storage node in the counterclockwise direction. Continuing with the above example, the storage data in storage node A can only be migrated to storage node D, and the storage data in storage node A cannot be migrated to storage node B.
[0033] Step 102: Perform load detection on multiple storage nodes to obtain load data of each storage node, and perform statistics on the load data of each storage node to obtain load parameters of multiple storage nodes.
[0034] In actual implementation, since the amount of data stored in each storage node is different, the load of each storage node is different. This will result in the resources in the storage nodes with less load may not be fully utilized, while the resources in the storage nodes with more load are overloaded, resulting in a decrease in overall resource utilization; overloaded resources may cause slower processing speed and increased response time, thereby affecting the performance of the entire system; occupied resources that are overloaded for a long time may cause failures due to exceeding their processing capacity, affecting the stability of the system. Based on this, when storing data, it is necessary to ensure the load balance of each storage node.
[0035] In some embodiments, the load detection of multiple storage nodes in step 102 to obtain the load data of each storage node can be achieved through the following technical solution: counting the amount of storage data included in each storage node and the processor utilization, number of cores and memory utilization corresponding to each storage node; based on the processor utilization, number of cores and memory utilization, determining the current load of each storage node, and using the current load and amount of storage data as load data.
[0036] In actual implementation, a load detector can be deployed to detect the load included in each storage node in real time. When a fixed information transmission cycle is reached, or when a storage node with a seriously excessive load is detected, the load information of the storage node obtained by the load detector is transmitted to the controller so that the controller can perform subsequent operations.
[0037] It should be noted that the fixed information transmission period can be a value set based on experience, and can be a time period such as 1 day, 1 hour, etc. The specific information transmission period can be set according to actual conditions and is not specifically limited here.
[0038] In actual implementation, the storage data volume of each storage node, as well as the processor utilization, core number and memory utilization of each storage node can be detected first. Then, the current load of the storage node can be determined through the processor utilization, core number and memory utilization of each storage node. Finally, the current load and storage data volume of the storage node can be used as load data.
[0039] As an example, the current load of the storage node may be determined by first determining the average processor utilization rate using the following formula (1):
[0040]
[0041] In formula (1), P is the average processor utilization, A is the processor utilization, and N is the number of processor cores.
[0042] In actual implementation, after determining the average processor utilization, the average processor utilization and memory utilization of the storage node can be compared, and the larger value can be used as the current load of the storage node. For example, if the average processor utilization of storage node A is 20% and the memory utilization of storage node A is 40%, then the current load of storage node A is 40%.
[0043] Through the above manner, the load data that can reflect the current operating status of each storage node can be accurately determined, thereby improving the accuracy of subsequent adjustment of the storage data of each storage node.
[0044] In some embodiments, the load data includes the amount of stored data. The statistics of the load data of each storage node in step 102 to obtain the load parameters of multiple storage nodes can be achieved by the following technical solution: averaging the amount of stored data of each storage node to obtain the average value of the amount of stored data; based on the average value, determining the standard deviation of the data storage amount of each storage node; dividing the standard deviation and the average value to obtain the coefficient of variation, and using the coefficient of variation and the average value as load parameters.
[0045] In actual implementation, the controller may store the data storage capacity in the received load data of each storage node into a queue, wherein each queue element included in the queue is the data storage capacity of each storage node. Then, the average value of the queue may be determined by the following formula (2):
[0046]
[0047] In formula (2), μ is the average value of the queue (i.e., the average value of the data storage capacity), N is the number of elements included in the queue (i.e., the number of data storage capacities of the storage nodes), and M is i is the value of the i-th data storage capacity.
[0048] By using the above method, the average value of the data storage volume can be determined, and then the standard deviation of the data storage volume can be determined by the following formula (3):
[0049]
[0050] In formula (3), α is the standard deviation of data storage, M i is the value of the i-th data storage capacity, μ is the average data storage capacity, and N is the number of data storage capacities of the storage nodes.
[0051] Finally, the coefficient of variation can be determined by the following formula (4):
[0052]
[0053] In formula (4), CV is the coefficient of variation, α is the standard deviation of the data storage amount, and μ is the average value of the data storage amount.
[0054] Through the above method, the coefficient of variation that reflects whether the current load is abnormal and the severity of the abnormality can be accurately determined, thereby improving the accuracy of subsequent adjustment of storage data of each storage node through the coefficient of variation.
[0055] Step 103: Based on the load parameters and the load data, determine the abnormally loaded nodes in the storage nodes, and create virtual nodes based on the abnormally loaded nodes.
[0056] In some embodiments, the load parameters include a coefficient of variation and an average value, and the load data include a storage data volume. The load parameters and load data in step 103 can be used to determine the load abnormal node in the storage node by the following technical solution: when the coefficient of variation is greater than the coefficient of variation threshold, determine the first difference between the storage data volume of each storage node and the average value; and select the storage node with the largest absolute value of the first difference as the load abnormal node.
[0057] In actual implementation, it is necessary to first determine the relationship between the coefficient of variation and the coefficient of variation threshold. If the coefficient of variation is less than the coefficient of variation threshold, it indicates that the load processing of each storage node is currently balanced, and there is no need to migrate the storage data included in each storage node. If the coefficient of variation is greater than the coefficient of variation threshold, it indicates that the load of each storage node is in an uneven state, and there are nodes with abnormal load. It is necessary to determine the nodes with abnormal load in the storage nodes, and then adjust the storage data in each storage node.
[0058] As an example, when the coefficient of variation is greater than the coefficient of variation threshold, the first difference value of each storage node can be determined by the following formula (5):
[0059] Δ=M i -μ (5)
[0060] In formula (5), Δ is the first difference, M i is the value of the i-th data storage capacity, and μ is the average value of the data storage capacity.
[0061] After determining the first difference of each storage node, the storage node with the largest absolute value of the first difference can be used as the abnormal load node. For example, if the first difference A is 10 and the first difference B is -12, the storage node corresponding to the first difference B can be used as the abnormal load node.
[0062] In some embodiments, the creation of a virtual node based on the abnormal load node in step 103 may be implemented through steps 1031 to 1033 as shown below:
[0063] In step 1031, a positive or negative value of a first difference of a node with abnormal load is determined;
[0064] In practical applications, the absolute value of the first difference represents the difference between the load of the abnormal load node and the average load, and the positive or negative value of the first difference represents whether the load of the abnormal load node is greater than the average value or less than the average value.
[0065] In step 1032, if the first difference of the abnormal load node is a positive value, a first virtual node is created between the abnormal load node and an adjacent node of the abnormal load node based on the coefficient of variation, wherein the first virtual node is a virtual node corresponding to a storage node other than the abnormal load node;
[0066] In actual implementation, if the first difference is a positive value, it indicates that the load of the abnormal load node exceeds the average load, and therefore the storage data of the abnormal load node needs to be migrated to another storage node.
[0067] In actual implementation, since the data in the storage node can only migrate to the storage node in a fixed direction (clockwise or counterclockwise), if storage node A is an abnormally loaded node, the direction of data migration is clockwise, and the adjacent node in the clockwise direction of storage node A is storage node B, then a first virtual node can be created between storage node A and storage node B, wherein the first virtual node is the virtual node except the virtual node corresponding to storage node A. It should be noted that the virtual node is the virtual node of the server corresponding to the entity, for example, virtual node A corresponds to server B. If the input is migrated to virtual node A at this time, it means that the data is stored in server B.
[0068] In some embodiments, the creation of a virtual node between the load abnormality node and the adjacent nodes of the load abnormality node based on the coefficient of variation in step 1032 can be achieved through the following technical solution: dividing the coefficient of variation by the quantity threshold to obtain the number of virtual nodes; creating a first virtual node of the number of nodes between the load abnormality node and the adjacent nodes of the load abnormality node.
[0069] In actual implementation, since the coefficient of variation also represents the degree of load balancing, if the coefficient of variation is too large, it represents that the current load is particularly unbalanced. In this case, multiple first virtual nodes need to be created to migrate storage data.
[0070] As an example, the number of first virtual nodes to be created can be determined by the value of the coefficient of variation. Specifically, the coefficient of variation can be divided by the quantity threshold to obtain a value that is the number of first virtual nodes. For example, if the coefficient of variation is 4 and the quantity threshold is 2, then the number of virtual nodes is 2. It should be noted that the value of the first virtual node can only be an integer. If the value obtained by dividing the coefficient of variation by the quantity threshold is not an integer, it can be rounded up as the number of first virtual nodes. For example, if the coefficient of variation is 3 and the quantity threshold is 2, then the number of first virtual nodes is 2.
[0071] It should be noted that the number of first virtual nodes obtained here can be the number of other storage nodes except the load abnormality node. Continuing with the above example, if the number of first virtual nodes is 1, the storage node corresponding to the first virtual node is storage node A, and in addition to storage node A, it also includes storage node B, storage node C and storage node D, then the first virtual node is created between storage node A and storage node B, and the created first virtual node is three nodes, corresponding to storage node B, storage node C and storage node D respectively.
[0072] In step 1033, if the first difference value of the load abnormality node is a negative value, a second virtual node is created between other storage nodes except the load abnormality node based on the coefficient of variation, wherein the second virtual node is a virtual node corresponding to the load abnormality node.
[0073] In actual implementation, if the first difference of the load-abnormal node is a negative value, it indicates that the load of the load-abnormal node is lower than the average load. At this time, it can be determined whether the load of other storage nodes exceeds the load threshold. If no storage node's load exceeds the load threshold, subsequent data migration will not be performed. If there is a node whose load exceeds the load threshold, it is necessary to migrate the storage data of other storage nodes to the load-abnormal node. Since the data in the storage node can only be migrated to the storage node in a fixed direction (clockwise or counterclockwise), if storage node A is a load-abnormal node and the direction of data migration is clockwise, a second virtual node corresponding to storage node A can be created in the clockwise direction of other storage nodes. A virtual node is created so that the storage data of storage node D can be migrated to the virtual node. It should be noted that in order to ensure that all storage nodes can migrate the storage data to storage node A, a second virtual node can be created between each storage node except storage node A. For example, the storage nodes are storage node A, storage node B, storage node C, storage node D and storage node A in clockwise order. Storage node A is an abnormal storage node. At this time, a second virtual node can be created between storage node A and storage node D, a second virtual node can be created between storage node B and storage node C, and a second virtual node can be created between storage node C and storage node D.
[0074] As an example, the number of second virtual nodes to be created can be determined by the value of the coefficient of variation. Specifically, the coefficient of variation can be divided by the quantity threshold to obtain a value that is the number of second virtual nodes. For example, if the coefficient of variation is 4 and the quantity threshold is 2, then the number of virtual nodes is 2. It should be noted that the value of the second virtual node can only be an integer. If the value obtained by dividing the coefficient of variation by the quantity threshold is not an integer, it can be rounded up as the number of second virtual nodes. For example, if the coefficient of variation is 3 and the quantity threshold is 2, then the number of first virtual nodes is 2.
[0075] It should be noted that the number of second virtual nodes obtained here can be the number of second virtual nodes created between other storage nodes except the load abnormality node. Continuing with the above example, if the number of first virtual nodes is 1, the storage node corresponding to the first virtual node is storage node A, and in addition to storage node A, it also includes storage node B, storage node C and storage node D, then a second virtual node is created between storage node B and storage node C, a second virtual node is created between storage node C and storage node D, and a second virtual node is created between storage node D and storage node A.
[0076] Step 104: Adjust the storage data of the abnormally loaded node by performing data migration on the abnormally loaded node and the virtual node.
[0077] In some embodiments, the data migration of the load abnormality node and the virtual node in step 104 can be implemented through the following technical solution: if the first difference value of the load abnormality node is a positive value, the data stored in the load abnormality node is migrated to the virtual node; if the first difference value of the load abnormality node is a negative value, the data stored in the storage nodes other than the load abnormality node is migrated to the virtual node.
[0078] In actual implementation, if the first difference value of the load abnormality node is a positive value, it indicates that the storage data of the load abnormality node needs to be migrated out. At this time, the storage data of the load abnormality node can be migrated to the virtual node. Correspondingly, if the first difference value of the load abnormality node is a negative value, it indicates that the storage data of other abnormal nodes need to be migrated to the load abnormality node. Therefore, the storage nodes of other storage nodes can be migrated to the virtual node corresponding to the load abnormality node first.
[0079] The present application is described below in conjunction with application examples.
[0080] See also Figure 2 , Figure 2 It is a schematic diagram of the principle of data storage provided in an embodiment of the present application.
[0081] exist Figure 2 In the example, it is assumed that the number of known physical nodes (i.e., the storage nodes mentioned above) is n, and then 32 Modulo, get the unique identifier of each node (such as IP address and other conventional units to identify the machine), such as Figure 2 The three entity nodes obtained are P1, P2 and P3 respectively.
[0082] In step 201, the detector detects the amount of data M (unit: pieces) mapped on each node (i.e. the amount of stored data mentioned above), and the load LP (i.e. the load data mentioned above) (unit: percentage) of the server corresponding to each entity node;
[0083] The load LP calculation formula is: LP = max (CPU utilization / number of cores, memory utilization).
[0084] Among them, CPU utilization / number of cores is used to calculate the average single-core CPU load.
[0085] The detector module regularly transmits data to the controller. The interval period is set to T (T can be determined according to experience. The overall principle is to set a long period when the cluster load is small, such as 1 day; set a short period when the cluster data storage is large or the load is high, such as 1 hour). This period setting function is open to users through the graphical interface or interface. The time setting range is (1h, 1d), and the step size is 1h.
[0086] In step 202: The controller receives the values of each node M (i.e., the above-mentioned stored data volume) and L (i.e., the above-mentioned current load) transmitted by the detector, inputs them into the built-in intelligent scheduler module (abbreviated as scheduler), and starts intelligent calculation. The process is as follows:
[0087] The scheduler stores the data volumes on all current nodes (i.e., the above-mentioned storage nodes) into a queue Q{M1, M2,..., Mn} (Mn refers to the data quantity on the nth node) for comparative analysis, and obtains the coefficient of variation CV of the queue Q (i.e., CV = standard deviation / average value), the maximum distance to the average MDTA (i.e., max distance to average, that is, the point farthest from the average value), and the corresponding node P (i.e., the above-mentioned load abnormal node).
[0088] If CV>10% (i.e., the above-mentioned coefficient of variation threshold), it indicates that the data in this queue has strong variability, that is, it represents that the data distribution on n nodes is uneven and needs to be handled specifically. At this time, compare MDTA and the mean Mean:
[0089] If DMTA>Mean, it means that the load on a certain node (i.e., the above-mentioned load abnormal node) far exceeds the average pressure of the cluster. At this time, if the load of node P (i.e., the above-mentioned load abnormal node) >= ML (i.e., the above-mentioned load threshold), it means that the pressure on this node is relatively large. If a large amount of new data continues to be added, there may be a hot spot problem, which may cause a failure. At this time, it is necessary to transfer the data of this node P to other nodes and determine the instruction Comm{nP, S}.
[0090] If DMTA<Mean, it means that the load on a certain node is much smaller than the average pressure, and this node is the node with the smallest pressure.
[0091] If there is a situation where the load value of a node other than node P exceeds the load threshold ML, it means that this node P needs to bear the data of other nodes, and determine the instruction Comm{P, S};
[0092] If the load values of all nodes do not exceed the load threshold ML, it means that the overall state of the cluster is good and no handling is required.
[0093] It should be noted that, in the instructions Comm{P, S} and Comm{nP, S}, P represents a virtual node newly added to the P node, nP represents a virtual node created for nodes other than P, and S represents the number of virtual nodes created.
[0094] The calculation logic of S is as follows: the range of CV values may be > 100%, which indicates severe variation. More virtual nodes need to be created for the target node. Define S = 1 + CV / 100%, that is, CV is rounded upward. Example: If CV = 1.3, then S = 2; if CV = 0.8, then S = 1).
[0095] If CV <= 10%, no action is taken.
[0096] In step 203, the built-in scheduler feeds back the final instruction to the controller, and the controller carries out specific virtual node creation actions. For Comm{P, S} and Comm{nP, S}, the specific creation rules are as follows:
[0097] 1. If the instruction is Comm{P, S} (with P node as Figure 2 ), then S virtual nodes corresponding to the P2 nodes are created in the counterclockwise direction of each node except the P2 node (such as Figure 2 Two virtual nodes P21 and P22 corresponding to storage node P2 are created, and the value of S is 1 at this time).
[0098] 2. If the instruction is Comm{nP, S}, hash each non-P2 corresponding virtual node S at equal distances between the P2 node and the first node in the counterclockwise direction of the P2 node (such as Figure 2 A virtual node P11 corresponding to storage node P1 is created, and a virtual node P31 corresponding to storage node P3 is created. At this time, the value of S is 1).
[0099] The above process is then repeated continuously to ensure that the data of the N cluster nodes remains relatively uniform and that all physical servers remain in a relatively healthy state.
[0100] The above method can fully consider the load and data distribution of each physical node in the distributed system, and use this information to determine whether and how to efficiently create virtual nodes, thereby solving the problem of uneven data distribution. Under this method, through multiple rounds of cycles, the load of each node can be more and more balanced, thereby improving the overall system throughput and ensuring the best performance.
[0101] In order to implement the data storage method on the first client side of the embodiment of the present application, the embodiment of the present application also provides a data storage device, Figure 3Schematic diagram of the structure of the data storage device according to the embodiment of the present application. Figure 3 As shown, the data storage device comprises:
[0102] An acquisition module 31 is used to acquire multiple storage nodes, wherein each storage node is adjacent to two storage nodes;
[0103] A load detection module 32, configured to perform load detection on the plurality of storage nodes to obtain load data of each storage node, and to perform statistics on the load data of each storage node to obtain load parameters of the plurality of storage nodes;
[0104] A node determination module 33, configured to determine a load-abnormal node in the storage node based on the load parameter and the load data, and create a virtual node based on the load-abnormal node;
[0105] The data migration module 34 is used to adjust the storage data of the abnormal load node by performing data migration on the abnormal load node and the virtual node.
[0106] In one embodiment, the load detection module 32 is also used to count the amount of storage data included in each storage node and the processor utilization, core number and memory utilization corresponding to each storage node; based on the processor utilization, the core number and the memory utilization, the current load of each storage node is determined, and the current load and the amount of storage data are used as the load data.
[0107] In one embodiment, the load detection module 32 is also used to average the amount of storage data of each storage node to obtain an average value of the amount of storage data; based on the average value, determine the standard deviation of the data storage amount of each storage node; divide the standard deviation and the average value to obtain a coefficient of variation, and use the coefficient of variation and the average value as the load parameters.
[0108] In one embodiment, the node determination module 33 is also used to determine a first difference between the amount of stored data of each storage node and an average value of the amount of stored data when the coefficient of variation is greater than a coefficient of variation threshold; and to select a storage node having the largest absolute value of the first difference as the load abnormality node.
[0109] In one embodiment, the node determination module 33 is also used to create a first virtual node between the load abnormality node and an adjacent node of the load abnormality node based on the coefficient of variation if the first difference value of the load abnormality node is a positive value, wherein the first virtual node is a virtual node corresponding to a storage node other than the load abnormality node; if the first difference value of the load abnormality node is a negative value, then based on the coefficient of variation, create a second virtual node between other storage nodes other than the load abnormality node, wherein the second virtual node is a virtual node corresponding to the load abnormality node.
[0110] In one embodiment, the node determination module 33 is also used to divide the coefficient of variation by a quantity threshold to obtain the number of nodes of the virtual node; and to create the first virtual node of the number of nodes between the load abnormality node and the adjacent nodes of the load abnormality node.
[0111] In one embodiment, the data migration module 34 is also used to migrate the data stored in the load abnormality node to the virtual node if the first difference value of the load abnormality node is a positive value; if the first difference value of the load abnormality node is a negative value, migrate the data stored in the storage nodes other than the load abnormality node to the virtual node.
[0112] It should be noted that: the data storage device provided in the above embodiment only uses the division of the above-mentioned program modules as an example when performing data storage. In actual applications, the above-mentioned processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above.
[0113] It should be noted that: the data storage device provided in the above embodiment only uses the division of the above-mentioned program modules as an example when performing data storage. In actual applications, the above-mentioned processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above.
[0114] Based on the hardware implementation of the above program modules, and in order to implement the data storage method of the embodiment of the present application, the embodiment of the present application further provides a first client, Figure 4 The hardware structure diagram of the embodiment of the present application is as follows: Figure 4 As shown, the first client 40 includes:
[0115] A first communication interface 41, capable of exchanging information with other devices;
[0116] The first processor 42 is connected to the first communication interface 41 to implement information interaction with other devices and is used to execute the data storage method provided above when running a computer program, and the computer program is stored in the first memory 43.
[0117] Specifically, the first processor 42 is used to obtain a plurality of storage nodes, wherein each storage node is adjacent to two storage nodes, perform load detection on the plurality of storage nodes to obtain load data of each storage node, perform statistics on the load data of each storage node to obtain load parameters of the plurality of storage nodes, determine a load-abnormal node among the storage nodes based on the load parameters and the load data, create a virtual node based on the load-abnormal node, and adjust the storage data of the load-abnormal node by performing data migration on the load-abnormal node and the virtual node;
[0118] The first communication interface 41 is used to feed back the adjusted status of each storage node to the user.
[0119] In one embodiment, the first processor 42 is also used to count the amount of storage data included in each storage node and the processor utilization, core number and memory utilization corresponding to each storage node; based on the processor utilization, the core number and the memory utilization, the current load of each storage node is determined, and the current load and the amount of storage data are used as the load data.
[0120] In one embodiment, the first processor 42 is further used to average the amount of storage data of each storage node to obtain an average value of the amount of storage data; based on the average value, determine the standard deviation of the data storage amount of each storage node; divide the standard deviation and the average value to obtain a coefficient of variation, and use the coefficient of variation and the average value as the load parameters.
[0121] In one embodiment, the first processor 42 is further used to determine a first difference between the amount of stored data of each storage node and the average value when the coefficient of variation is greater than a coefficient of variation threshold; and to use the storage node with the largest absolute value of the first difference as the load abnormality node.
[0122] In one embodiment, the first processor 42 is further used to create a first virtual node between the load abnormality node and an adjacent node of the load abnormality node based on the coefficient of variation if the first difference value of the load abnormality node is a positive value, wherein the first virtual node is a virtual node corresponding to a storage node other than the load abnormality node; if the first difference value of the load abnormality node is a negative value, then create a second virtual node between other storage nodes other than the load abnormality node based on the coefficient of variation, wherein the second virtual node is a virtual node corresponding to the load abnormality node.
[0123] In one embodiment, the first processor 42 is further used to divide the coefficient of variation by a quantity threshold to obtain the number of nodes of the first virtual node; and to create the first virtual nodes of the number of nodes between the load abnormality node and the adjacent nodes of the load abnormality node.
[0124] In one embodiment, the first processor 42 is further used to migrate the data stored in the load abnormality node to the virtual node if the first difference value of the load abnormality node is a positive value; if the first difference value of the load abnormality node is a negative value, migrate the data stored in the storage nodes other than the load abnormality node to the virtual node.
[0125] It should be noted that the specific processing process of the first communication interface 41 and the first processor 42 can be understood by referring to the above data storage method.
[0126] Of course, in actual application, the various components in the first client 40 are coupled together through the first bus system 44. It can be understood that the first bus system 44 is used to realize the connection and communication between these components. In addition to the data bus, the first bus system 44 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 4 In the figure, various buses are labeled as a first bus system 44 .
[0127] The first memory 43 in the embodiment of the present application is used to store various types of data to support the operation of the first client 40. Examples of such data include: any computer program used to operate on the first client 40.
[0128] The data storage method disclosed in the above embodiment of the present application can be applied to the first processor 42, or implemented by the first processor 42. The first processor 42 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above data storage method can be completed by the hardware integrated logic circuit or software instructions in the first processor 42. The above-mentioned first processor 42 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The first processor 42 can implement or execute the data storage method, steps and logic block diagram disclosed in the embodiment of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the data storage method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the first memory 43. The first processor 42 reads the information in the first memory 43 and completes the steps of the aforementioned data storage method in combination with its hardware.
[0129] In an exemplary embodiment, the first client 40 may be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), field programmable gate array (FPGA), general processor, controller, microcontroller (MCU), microprocessor, or other electronic components to execute the aforementioned data storage method.
[0130] In an exemplary embodiment, the embodiment of the present application further provides an electronic device, including a processor and a memory for storing a computer program that can be run on the processor, wherein the processor is used to execute the steps of any of the above methods when running the computer program.
[0131] The embodiment of the present application further provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, for example, including a memory 503 storing a computer program, and the computer program can be executed by a processor 502 of an electronic device 500 to complete the steps of the aforementioned method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM.
[0132] An embodiment of the present application also provides a computer program product, including a computer program, which implements the steps of any of the above methods when executed by a processor.
[0133] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. The term "and / or" herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the term "one or more" herein represents any combination of at least two of any one or more of a plurality of items. For example, including one or more of A, B, and C can represent including any one or at least two or more elements selected from the set consisting of A, B, and C.
[0134] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0135] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.
Claims
1. A data storage method, characterized in that: The method comprises: Acquire multiple storage nodes, wherein each storage node is adjacent to two storage nodes; Performing load detection on the multiple storage nodes to obtain load data of each storage node, and performing statistics on the load data of each storage node to obtain load parameters of the multiple storage nodes; Based on the load parameter and the load data, determining a load-abnormal node in the storage node, and creating a virtual node based on the load-abnormal node; By performing data migration on the abnormal load node and the virtual node, the storage data of the abnormal load node is adjusted.
2. The method according to claim 1, characterized in that The performing load detection on the plurality of storage nodes to obtain load data of each storage node includes: Count the amount of storage data included in each storage node and the processor utilization, core number and memory utilization corresponding to each storage node; Based on the processor utilization, the number of cores and the memory utilization, the current load of each storage node is determined, and the current load and the amount of stored data are used as the load data.
3. The method according to claim 1, characterized in that: The load data includes the amount of stored data; The counting of the load data of each storage node to obtain load parameters of multiple storage nodes includes: The amount of stored data of each storage node is averaged to obtain the average value of the amount of stored data; Based on the average value, determine a standard deviation of the data storage amount of each storage node; The standard deviation and the average value are divided to obtain a coefficient of variation, and the coefficient of variation and the average value are used as the load parameters.
4. The method according to claim 1, characterized in that: The load parameters include a coefficient of variation and an average value of the amount of stored data, and the load data includes the amount of stored data; The determining, based on the load parameter and the load data, a load abnormality node in the storage node includes: When the coefficient of variation is greater than a coefficient of variation threshold, determining a first difference between the amount of stored data of each storage node and an average value of the amount of stored data; The storage node with the largest absolute value of the first difference is used as the abnormal load node.
5. The method according to claim 4, characterized in that The step of creating a virtual node based on the abnormal load node includes: If the first difference of the load abnormality node is a positive value, based on the coefficient of variation, a first virtual node is created between the load abnormality node and an adjacent node of the load abnormality node, wherein the first virtual node is a virtual node corresponding to a storage node other than the load abnormality node; If the first difference of the load abnormality node is a negative value, a second virtual node is created between other storage nodes except the load abnormality node based on the variation coefficient, wherein the second virtual node is a virtual node corresponding to the load abnormality node.
6. The method according to claim 5, characterized in that The step of creating a first virtual node between the abnormal load node and an adjacent node of the abnormal load node includes: Dividing the coefficient of variation by a quantity threshold to obtain the number of nodes of the first virtual node; The first virtual nodes of the number of nodes are created between the abnormal load node and the adjacent nodes of the abnormal load node.
7. The method according to claim 4, characterized in that The step of performing data migration on the abnormally loaded node and the virtual node includes: If the first difference of the abnormal load node is a positive value, migrating the data stored in the abnormal load node to the virtual node; If the first difference value of the abnormal load node is a negative value, the data stored in the storage nodes other than the abnormal load node are migrated to the virtual node.
8. A data storage device, characterized in that: include: An acquisition module, used for acquiring a plurality of storage nodes, wherein each storage node is adjacent to two storage nodes; A load detection module, used to perform load detection on the multiple storage nodes to obtain load data of each storage node, and to perform statistics on the load data of each storage node to obtain load parameters of the multiple storage nodes; A node determination module, configured to determine a load-abnormal node in the storage node based on the load parameter and the load data, and to create a virtual node based on the load-abnormal node; The data migration module is used to adjust the storage data of the load-abnormal node by migrating data between the load-abnormal node and the virtual node.
9. An electronic device, characterized in that: include: A processor and a memory for storing a computer program that can be executed on the processor, wherein: The processor is used to execute the steps of the method according to any one of claims 1 to 7 when running a computer program.
10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.