Method for optimizing data distribution and terminal
By prioritizing data allocation based on the memory level of nodes within the cluster, the problem of uneven server resource allocation is solved, achieving balanced use of node resources and improving data processing efficiency and overall balance.
Patent Information
- Application Number
- CN202410567489.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-09
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-05-09
AI Technical Summary
In existing technologies, server resource allocation is crude and cannot truly reflect the actual situation of the application, resulting in uneven data distribution and affecting overall data processing efficiency.
By periodically obtaining the memory usage of each node in the cluster, new messages are preferentially allocated to the node with the largest remaining memory. Once the memory usage of all nodes is equal, messages are distributed evenly using a round-robin method. A new central node is added for message distribution to improve efficiency.
It achieves balanced use of node resources within the cluster, improves the efficiency and balance of data processing, ensures balanced load on each node, and reasonably balances resource usage.
Smart Images

Figure CN118519761B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a data distribution optimization method and terminal. BACKGROUND
[0002] In current Internet systems, the scheduling of various tasks, data distribution and other operations are controllable, and can be distributed or scheduled according to various conditions and rules. Usually, the distribution rules are triggered according to certain conditions, such as the used conditions of server resources, CPU, memory, IO resources, or the remaining conditions of resources, to distribute data.
[0003] Since multiple applications are deployed on current servers, the resource conditions of the servers cannot truly reflect the actual conditions of the applications. Even if only one type or one application is deployed on a server, the actual remaining resources of the server are not the same concept as the actual available remaining resources of the application. Therefore, the data distribution method according to the server resources can only be considered as a rough data distribution method. SUMMARY
[0004] The technical problem to be solved by the present application is to provide a data distribution optimization method and terminal, which can reasonably balance the resource usage conditions of each node in the cluster and improve the efficiency and balance of overall data processing.
[0005] To solve the above technical problems, the technical scheme adopted by the present application is:
[0006] A data distribution optimization method, comprising the steps of:
[0007] S1, periodically acquiring the memory water level usage conditions of each node in the cluster;
[0008] S2, when the remaining water levels of each node are not uniform, preferentially distributing the newly acquired messages to the node with the largest remaining water level;
[0009] S3, repeating step S2 until the remaining water levels of each node are consistent, and then distributing the newly acquired messages to each node in a round-robin manner.
[0010] To solve the above technical problems, another technical scheme adopted by the present application is:
[0011] A data distribution optimization terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:
[0012] S1, periodically acquiring the memory water level usage conditions of each node in the cluster;
[0013] S2, when there is a remaining water level of each node is uneven, the newly acquired message is preferentially distributed to the node with the largest remaining water level;
[0014] S3, repeating step S2 until the remaining water level of each node is consistent, and then adopting a polling manner to distribute the newly acquired message to each node.
[0015] The present application has the beneficial effect of providing an optimized data distribution method and terminal, by obtaining the water level occupation of each node in the cluster as the judgment condition for data distribution, if one of the nodes in the cluster has more remaining water level, the data is preferentially distributed to this node, otherwise, when the remaining water level of all nodes tends to be consistent, the data is evenly distributed, by this way, the real resource usage of each node in the cluster can be greatly guaranteed, the load balance of each node can be maintained, the resource usage of each node can be reasonably balanced, and the overall data processing efficiency and balance can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 A flow chart of the optimized data distribution method of the embodiment of the present application;
[0017] Figure 2 A structure schematic diagram of the optimized data distribution terminal of the embodiment of the present application.
[0018] Label explanation:
[0019] 1. An optimized data distribution terminal, 2. a memory, 3. a processor. DETAILED DESCRIPTION
[0020] To explain the technical content, the achieved purposes and effects of the present application in detail, the following will be explained in combination with the embodiments and the drawings.
[0021] Please refer to Figure 1 An optimized data distribution method, comprising the steps of:
[0022] S1, periodically obtaining the memory water level usage of each node in the cluster;
[0023] S2, when there is a remaining water level of each node is uneven, the newly acquired message is preferentially distributed to the node with the largest remaining water level;
[0024] S3, repeating step S2 until the remaining water level of each node is consistent, and then adopting a polling manner to distribute the newly acquired message to each node.
[0025] From the above description, the beneficial effects of the present application are that: provide a kind of optimization method of data distribution, by obtaining the water level occupation condition of each node in cluster as the judgment condition of data distribution, if the residual water level of a node is more in the multiple nodes of a cluster, then data is preferentially distributed to this node, otherwise, when the residual water level of all nodes tends to be consistent, then the data is distributed in the way of equal division, by this kind of way, the real resource use condition of each node in cluster can be guaranteed to a great extent, so that each node maintains the balance of load, reasonably balances the resource use condition of each node, improves the efficiency and balance of overall data processing.
[0026] Further, the step S1 further includes before:
[0027] S0, a central node is newly built in the cluster, for message distribution.
[0028] From the above description, the newly added central node is specially used to distribute messages to each node of the cluster, to improve the data distribution efficiency.
[0029] Further, the step S1 is specifically:
[0030] The central node periodically obtains the memory water level use condition of each node in the cluster, and calculates the residual water level of each node according to the memory water level use condition of each node, and stores it in local memory.
[0031] From the above description, the calculated residual water level of each node is stored in local memory, so that the central node can be called at any time, to further improve the efficiency of subsequent data distribution.
[0032] Further, the step S2 is specifically:
[0033] S21, the central node periodically obtains the residual water level of each node stored in the local memory, and judges whether the difference of the residual water level of each node exceeds a preset threshold value, if yes, enter step S22, otherwise, adopt polling mode to distribute new message to each node;
[0034] S22, determine the proportion of the residual water level of each node, and distribute new message to the corresponding node according to the proportion, until the difference of the residual water level of each node is less than the preset threshold value, then adopt polling mode to distribute new message to each node.
[0035] From the above description, it can be seen that whether the residual water levels of the nodes are inconsistent can be determined by comparing the difference between the residual water levels of the nodes with the preset threshold, so as to avoid the problem of low allocation efficiency caused by the fact that slight difference is determined as inconsistent residual water level and enters the subsequent preferential allocation step; meanwhile, the message is allocated according to the proportion between the residual water levels of the nodes when the residual water levels are inconsistent, so as to achieve optimal allocation.
[0036] Further, the step S3 comprises a step of:
[0037] S4, returning to step S1 when the message speed of a certain node is too slow.
[0038] From the above description, when the message speed of a certain node is too slow, it is likely that the memory water level usage of the node has reached the peak value, so the step S1 is returned to re-enter the optimal allocation mode according to the residual water level, so as to ensure that the water levels of the nodes are consistent again in time, and the resource usage of each node is reasonably balanced, so as to improve the efficiency and balance of the overall data processing.
[0039] Please refer to Figure 2 An optimization terminal for data allocation, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements the following steps when executing the computer program:
[0040] S1, periodically acquiring the memory water level usage of each node in the cluster;
[0041] S2, when the residual water levels of the nodes are inconsistent, preferentially allocating the newly acquired message to the node with the largest residual water level;
[0042] S3, repeating step S2 until the residual water levels of the nodes are consistent, and then allocating the newly acquired message to each node in a polling manner.
[0043] From the above description, the beneficial effects of the present application are that: based on the same technical concept, in combination with the above-mentioned optimization method for data allocation, an optimization terminal for data allocation is provided, which acquires the water level occupation of each node in the cluster as a judgment condition for data allocation, if the residual water level of a certain node is more among multiple nodes in a cluster, the data is preferentially allocated to this node, otherwise, when the residual water levels of all nodes tend to be consistent, the data is allocated in an equal manner, through this manner, the real resource usage of each node in the cluster can be greatly ensured, each node can maintain load balance, the resource usage of each node can be reasonably balanced, and the efficiency and balance of the overall data processing can be improved.
[0044] Further, the step S1 further comprises the following step before:
[0045] S0, a center node is newly built in the cluster for message distribution.
[0046] As described above, the newly-built center node is specially used for distributing messages to each node in the cluster, thereby improving the data distribution efficiency.
[0047] Further, the step S1 specifically comprises the following step:
[0048] The center node periodically acquires the memory water level usage of each node in the cluster, and calculates the residual water level of each node according to the memory water level usage of each node, and stores it in the local memory.
[0049] As described above, the calculated residual water level of each node is stored in the local memory, so that the center node can be called at any time, thereby further improving the efficiency of subsequent data distribution.
[0050] Further, the step S2 specifically comprises the following step:
[0051] S21, the center node periodically acquires the residual water level of each node stored in the local memory, and judges whether the difference of the residual water level of each node exceeds a preset threshold, if yes, it enters step S22, otherwise it distributes new messages to each node in a polling manner;
[0052] S22, the proportion of the residual water level of each node is determined, and new messages are distributed to the corresponding nodes according to the proportion, until the difference of the residual water level of each node is less than the preset threshold, and then new messages are distributed to each node in a polling manner.
[0053] As described above, the difference between the residual water level of each node and the preset threshold can be compared to determine whether the residual water level of each node is inconsistent, thereby avoiding the problem that slight difference is determined as inconsistent residual water level and enters the subsequent preferential distribution step, causing low distribution efficiency; at the same time, when the residual water level is inconsistent, the messages are distributed according to the proportion between the residual water levels of each node, thereby achieving optimal distribution.
[0054] Further, the step S3 further comprises the following step:
[0055] S4, when the message speed of a certain node is too slow, return to step S1.
[0056] From the above description, when the message speed of a certain node is too slow, it is likely that the memory water level usage of the node has reached the peak, so it returns to step S1 to re-enter the optimal allocation mode according to the remaining water level, to ensure that the water levels of the nodes are timely consistent again, to reasonably balance the resource usage of each node, and to improve the efficiency and balance of the overall data processing.
[0057] The application provides an optimization method and terminal for data distribution, which is mainly applied to the scene of task scheduling and data distribution.
[0058] Please refer to Figure 1 The embodiment one of the application is:
[0059] An optimization method for data distribution, as shown in the figure, comprises the following steps: Figure 1
[0060] S1, periodically acquire the memory water level usage of each node in the cluster.
[0061] S2, when the remaining water levels of the nodes are uneven, preferentially distribute the newly acquired message to the node with the largest remaining water level.
[0062] S3, repeat step S2 until the remaining water levels of the nodes are consistent, and then distribute the newly acquired message to the nodes in a polling manner.
[0063] In this embodiment, the water level occupation of each node in the cluster is acquired as a judgment condition for data distribution, if the remaining water level of a certain node is more in the multiple nodes in a cluster, the data is preferentially distributed to the node, otherwise, when the remaining water levels of all nodes are consistent, the data is distributed in an equal manner, through this manner, the real resource usage of each node in the cluster can be greatly ensured, the load balance of each node can be maintained, the resource usage of each node can be reasonably balanced, and the efficiency and balance of the overall data processing can be improved.
[0064] The embodiment two of the application is:
[0065] An optimization method for data distribution, on the basis of the above-mentioned embodiment one, before step S1, the embodiment further comprises:
[0066] S0, a center node is newly established in the cluster, and is used for message distribution.
[0067] In this embodiment, the center node is newly added and is specially used for distributing the message to each node in the cluster, so that the data distribution efficiency is improved.
[0068] Step S1 is specifically:
[0069] The central node periodically acquires the memory water level usage of each node in the cluster, and calculates the residual water level of each node according to the memory water level usage of each node, and stores it in the local memory. That is, the calculated residual water level of each node is stored in the local memory, so that the central node can be called at any time, further improving the efficiency of subsequent data allocation.
[0070] In this embodiment, step S2 is specifically:
[0071] S21, the central node periodically acquires the residual water level of each node stored in the local memory, and judges whether the difference of the residual water level of each node exceeds a preset threshold, if yes, go to step S22, otherwise, distribute new messages to each node in a polling manner;
[0072] S22, determine the proportion of the residual water level of each node, and distribute new messages to the corresponding node according to the proportion, until the difference of the residual water level of each node is less than the preset threshold, and then distribute new messages to each node in a polling manner.
[0073] That is, the difference between the residual water level of each node and the preset threshold can be compared to determine whether the residual water level of each node is inconsistent, avoiding the problem that slight difference is determined as inconsistent residual water level and entering the subsequent priority allocation step to cause low allocation efficiency. At the same time, it is also limited to distribute messages according to the proportion between the residual water levels of each node when the residual water levels are uneven, to achieve optimal distribution.
[0074] Taking a rabbitmq (message middleware) cluster as an example, the management console can monitor the water level usage of each node in the cluster in real time. Assuming that there are three nodes in a rabbitmq cluster, and the highest water level (i.e. the maximum resource occupied memory) of each node is different. The highest memory water level of node a is 30G, the highest memory water level of node b is 40G, and the highest memory water level of node c is 50G. The total physical memory of the server where each node is located is 100G, and only one node is deployed on a server.
[0075] Assuming that at a certain moment, node a has used 10G of memory, node b has used 30G of memory, and node c has used 35G of memory, then at this moment, the residual water level of node a is 20G, the residual water level of node b is 10G, and the residual water level of node c is 15G. These residual water levels are periodically stored in the local memory of rabbitmq.
[0076] In a RabbitMQ cluster, a central node is pre-built to distribute all messages. The central node periodically retrieves the remaining water level of nodes a, b, and c from its local memory. This interval can be the same as the time taken by the management console to retrieve and store the remaining water level of nodes a, b, and c in its local memory, or it can be different; no limitation is made here.
[0077] When new messages arrive and require allocation, the central node retrieves the information stored in its local memory and allocates the data based on the remaining water level of each node.
[0078] If the remaining water level of nodes a, b, and c is obtained, and the difference between any two nodes is greater than 0.5G (a preset threshold), then the remaining water level of the nodes in the RabbitMQ cluster is considered uneven. In this case, the ratio of the remaining water levels of the three nodes is 4:2:3. The central node can then distribute new messages to the corresponding nodes according to this 4:2:3 ratio. Alternatively, in this embodiment, another distribution ratio can be used. For example, assuming the ratio of the remaining water levels of the three nodes is 3:2:1, 50% of the new messages can be allocated to node a first, and the remaining 50% can be distributed to nodes b and c in a 2:1 ratio. The method of distributing new messages according to the remaining water level ratio is not limited to the two methods mentioned above and can be designed according to the actual situation.
[0079] In addition, in this embodiment, step S3 is followed by the following step:
[0080] S4. When a node experiences slow message delivery, return to step S1.
[0081] When a node experiences slow message delivery, it is likely that the memory usage of that node has reached its peak. Therefore, the process returns to step S1 and re-enters the optimal allocation method based on the remaining memory level. This ensures that the memory levels of each node become consistent again in a timely manner, achieving a reasonable balance in the resource usage of each node and improving the overall efficiency and balance of data processing.
[0082] Please refer to Figure 2 Embodiment four of the present invention is as follows:
[0083] An optimized data allocation terminal 1 includes a memory 2, a processor 3, and a computer program stored on the memory 2 and executable on the processor 3. When the processor 3 executes the computer program, it completes the steps of the optimized data allocation method in any of the embodiments 1 to 3 described above.
[0084] In summary, the data allocation optimization method and terminal provided by this invention can greatly ensure the actual resource usage of each node within the cluster, maintain load balance among nodes, reasonably balance the resource usage of each node, and improve the overall efficiency and balance of data processing.
[0085] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent modifications made based on the content of the present invention specification and drawings, or direct or indirect applications in related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. An optimization method for data allocation, characterized in that, Including the following steps: S1. Periodically obtain the memory usage of each node in the cluster; S2. When there is an uneven remaining water level among the nodes, the newly acquired messages are preferentially allocated to the node with the largest remaining water level. S3. Repeat step S2 until the remaining water level of each node is consistent, and then use a polling method to distribute the newly acquired messages to each node. The procedure preceding step S1 also includes: S0. Create a new central node in the cluster for message distribution; Step S1 specifically involves: The central node periodically obtains the memory usage of each node in the cluster, calculates the remaining memory level of each node based on the memory usage of each node, and stores it in local memory. Step S2 specifically involves: S21. The central node periodically obtains the remaining water level of each node stored in the local memory, and determines whether the difference between the remaining water levels of each node exceeds a preset threshold. If so, proceed to step S22; otherwise, use a polling method to distribute new messages to each node. S22. Determine the proportion of the remaining water level of each node, and allocate new messages to the corresponding nodes according to the proportion until the difference of the remaining water level of each node is less than the preset threshold, and then use a round-robin method to allocate new messages to each node.
2. The data allocation optimization method according to claim 1, characterized in that, Step S3 is followed by the following steps: S4. When a node experiences slow message delivery, return to step S1.
3. An optimized terminal for data allocation, characterized in that, Includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, performs the following steps: S1. Periodically obtain the memory usage of each node in the cluster; S2. When there is an uneven remaining water level among the nodes, the newly acquired messages are preferentially allocated to the node with the largest remaining water level. S3. Repeat step S2 until the remaining water level of each node is consistent, and then use a polling method to distribute the newly acquired messages to each node. The procedure preceding step S1 also includes: S0. Create a new central node in the cluster for message distribution; Step S1 specifically involves: The central node periodically obtains the memory usage of each node in the cluster, calculates the remaining memory level of each node based on the memory usage of each node, and stores it in local memory. Step S2 specifically involves: S21. The central node periodically obtains the remaining water level of each node stored in the local memory, and determines whether the difference between the remaining water levels of each node exceeds a preset threshold. If so, proceed to step S22; otherwise, use a polling method to distribute new messages to each node. S22. Determine the proportion of the remaining water level of each node, and allocate new messages to the corresponding nodes according to the proportion until the difference of the remaining water level of each node is less than the preset threshold, and then use a round-robin method to allocate new messages to each node.
4. The optimized data allocation terminal according to claim 3, characterized in that, Step S3 is followed by the following steps: S4. When a node experiences slow message delivery, return to step S1.
Citation Information
Patent Citations
HDFS-based data equalization optimization method, system terminal and storage medium
CN110928836A
Task allocation method, system, equipment and medium
CN113806045A