Distributed data processing method
By establishing a multi-regional node network in the distributed data processing system, and monitoring and dynamically adjusting task allocation priorities in real time, the problem of data processing interruption caused by node failures or load fluctuations in existing technologies is solved, achieving continuous and efficient data processing, and improving the system's reliability and resource utilization.
Patent Information
- Application Number
- CN202411399141.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2044-10-09
AI Technical Summary
Existing distributed data processing systems suffer from problems such as low resource utilization, unreasonable task allocation, unbalanced load, untimely and improper handling of node failures, and unresolved resource issues. These unresolved issues include inadequate node status monitoring in existing technologies. These are common technical problems in existing data processing systems, including data processing interruptions or low efficiency, improper resource allocation, and other unresolved issues. Furthermore, existing technologies often exhibit passive handling of node failures or load fluctuations, lacking effective monitoring and early warning mechanisms, leading to data processing interruptions or significant efficiency drops.
By establishing a distributed node network, the area covered by the data processing system is divided into multiple processing zones, and different types of servers, including physical servers, virtual servers, and cloud servers, are deployed in each processing zone. These servers are interconnected through high-speed Ethernet, fiber optic networks, and dedicated networks. The system monitors the operating status parameters of the nodes in real time, including CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization. It calculates a comprehensive load assessment coefficient, dynamically adjusts task allocation priorities and node technologies, and achieves effective management and optimization of node load.
It achieves data processing continuity and system stability during node failures or load fluctuations, improves resource utilization and processing efficiency, ensures high efficiency and reliability of data processing, and optimizes system operating efficiency and response speed through real-time monitoring and dynamic adjustment.
Smart Images

Figure CN119336500B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a distributed data processing method. Background Technology
[0002] Data processing is generally divided into centralized data processing and distributed data processing. The biggest problem with centralized data processing is its slow processing speed, while also placing high demands on computer performance. The significant improvements in personal computer performance and their widespread use have made it possible to distribute processing power across all computers on a network. Distributed computing is the opposite of centralized computing; in distributed computing, data can be distributed over a large area.
[0003] Existing data processing systems generally suffer from several problems: First, resource utilization is low, with some nodes operating at low load for extended periods while others are overloaded, resulting in low processing efficiency. Second, task allocation lacks intelligence, failing to allocate tasks reasonably based on factors such as task priority, node load status, and data characteristics, leading to delays in processing important tasks and impacting overall system performance. Third, handling node failures or load fluctuations is reactive, lacking effective monitoring and early warning mechanisms. Once a problem occurs, it is difficult to adjust in a timely manner, potentially leading to data processing interruptions or a significant drop in efficiency. Summary of the Invention
[0004] The purpose of this invention is to provide a distributed data processing method to solve the problems mentioned in the background above.
[0005] The objective of this invention can be achieved through the following technical solution: a distributed data processing method, comprising the following steps:
[0006] Step 1: Establish a distributed node network: Divide the area covered by the data processing system into several processing regions, and set up several servers within each processing region. Deploy the distributed nodes on the corresponding servers in each processing region.
[0007] Step 2, Node Status Monitoring: Real-time monitoring of the running status of each node on each server in each distributed processing area to obtain the running status parameters of each node on each server in each distributed processing area;
[0008] Step 3: Node load analysis: Perform comprehensive calculation and analysis on the CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization of each node on each server in each distributed processing area to obtain the comprehensive load evaluation coefficient of each node on each server in each distributed processing area.
[0009] Step 4: Node load status determination: Obtain the comprehensive load evaluation coefficient classification threshold of each node on each server in each distributed processing area from the system database, and compare and analyze the comprehensive load evaluation coefficient of each node on each server in each distributed processing area with the comprehensive load evaluation coefficient classification threshold of each node on each server in each distributed processing area.
[0010] Step 5: Data processing priority determination: Extract the features of the data to be processed, combine the features of the data to be processed with the load status of each node on each server in each distributed processing area into a mapping table, and assign a task priority to each combination. At the same time, a decision rule for determining the task allocation priority is formulated.
[0011] Step 6: Data processing task allocation: Assign high-priority tasks to the low-load nodes on the corresponding servers in each of the distributed processing areas. While ensuring that high-priority tasks are processed first, assign medium-priority tasks to the nodes on the corresponding servers in each of the distributed processing areas. Finally, schedule low-priority tasks for processing during system idle time.
[0012] Step 7: Distributed processing node optimization: Real-time monitoring and acquisition of the comprehensive load evaluation coefficient of each node on each server in each distributed processing area. The monitoring period is divided into several monitoring periods. When a node on each server in each distributed processing area frequently experiences delays or excessive load, a load anomaly signal is generated, and the task allocation weight of that node is adjusted.
[0013] Preferably, in step one, the nodes on each server in each distributed processing area are interconnected via high-speed Ethernet, fiber optic networks, and dedicated networks.
[0014] Preferably, in step two, the running status parameters of each node on each server in each distributed processing area include CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization.
[0015] Preferably, in step three, the specific method for comprehensively calculating and analyzing the CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization of each node on each server in each distributed processing area is as follows:
[0016] To obtain the mean CPU utilization of each node on each server in each distributed processing region from time t1 to time t2, sum the m CPU utilization values and divide by m. Then, calculate the variance of the CPU utilization of each node on each server in each distributed processing region using the variance calculation formula. Next, calculate the square root of the variance of the CPU utilization of each node on each server in each distributed processing region to obtain the standard deviation of the CPU utilization of each node on each server in each distributed processing region. Finally, subtract the CPU utilization of each node on each server in each distributed processing region from the mean CPU utilization and divide the result by the standard deviation of the CPU utilization to obtain the standard value of the CPU utilization of each node on each server in each distributed processing region, denoted as CN. i i represents the number of each node, i = 1, 2, ..., n, and n represents the total number of node numbers;
[0017] Similarly, calculate the standard values for memory utilization, network bandwidth utilization, and disk space utilization for each node on each server in each distributed processing region, denoted as MN. i NN i and DN i ;
[0018] According to the formula Calculate the comprehensive evaluation value of each node on each server in each distributed processing region. a1, a2, a3, and a4 represent the weighting factors corresponding to the set CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization, respectively.
[0019] According to the formula Calculate the CPU utilization index value of each node on each server in each distributed processing region. Where α represents the set attenuation factor, 0 < α < 1. Similarly, the exponential values corresponding to the memory utilization, network bandwidth utilization, and disk space utilization of each node on each server in each distributed processing region are calculated, and denoted as follows: and
[0020] By combining the standard values of CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization of each node on each server in each distributed processing region with the comprehensive evaluation value of each node on each server in each distributed processing region, the comprehensive load evaluation coefficient of each node on each server in each distributed processing region is obtained.
[0021] Preferably, the comprehensive load evaluation coefficient for each node on each server in each distributed processing region is calculated and analyzed in the following way:
[0022] According to the formula Calculate the comprehensive load evaluation coefficient γ for each node on each server in each distributed processing region, where β represents the set adjustment factor, 0 < β < 1, and w1, w2, w3, and w4 represent the weight factors corresponding to the set standard values of CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization, respectively.
[0023] Preferably, in step four, the comprehensive load evaluation coefficient of each node on each server in each distributed processing region is compared and analyzed with the classification threshold of the comprehensive load evaluation coefficient of each node on each server in each distributed processing region. The specific comparison and analysis method is as follows:
[0024] The comprehensive load evaluation coefficient classification thresholds for each node on each server in each distributed processing region are obtained from the system database and denoted as k1 and k2, respectively, where 0 < k1 < k2 < 1.
[0025] If γ < k1, it means that the node has enough resources to handle new tasks, and it is recorded as a low-load node.
[0026] If k1 < γ < k2, it means that the resource usage of this node is relatively reasonable, and it is recorded as a medium load node.
[0027] If γ > k2, it indicates that the node's resources are close to saturation, and it is recorded as a high-load node.
[0028] Preferably, in step five, the data characteristics include data type, data size, data timeliness requirements, and data processing complexity.
[0029] Preferably, in step seven, by obtaining the number of nodes on each server in each distributed processing area that exhibit load anomaly signals during a monitoring period, denoted as g, the number of nodes g on each server in each distributed processing area during a monitoring period is compared with the number of nodes n on each server in each distributed processing area. If the calculation result is greater than 0.8, it indicates that the load margin of each node on each server in each distributed processing area during the monitoring period is insufficient, and a node expansion optimization signal is generated. The processing system increases the number of nodes on each server in each distributed processing area during the monitoring period by 50% of the current number of nodes. After the new node is added to the system, the load balancer automatically adjusts the load balancing to evenly distribute the data processing tasks to each node. If the calculation result is less than 0.4, it indicates that the load margin of each node on each server in each distributed processing area during the monitoring period is sufficient, and a node contraction optimization signal is generated. The processing system reduces the number of nodes on each server in each distributed processing area during the monitoring period by 20% of the current number of nodes.
[0030] The beneficial effects of this invention are:
[0031] This invention establishes a distributed node network, dividing the area covered by the data processing system into multiple processing zones. Different types of servers, including physical servers, virtual servers, and cloud servers, are deployed within each zone. These servers are interconnected via high-speed Ethernet, fiber optic networks, and dedicated networks. This ensures that even if a network in one zone fails or suffers a natural disaster, nodes in other zones can still function normally, guaranteeing the continuity of data processing and improving system reliability and fault tolerance. Furthermore, by monitoring the real-time operating status parameters of nodes on each server in each processing zone, including CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization, node failures or performance issues can be detected and addressed promptly, further optimizing system efficiency and stability.
[0032] This invention comprehensively calculates and analyzes the CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization of each node on each server in a distributed processing area to obtain the comprehensive load evaluation coefficient for each node. Further, by calculating the mean, variance, and standard deviation of these resource utilization rates for each node, the standard values of each resource utilization rate for each node can be calculated. Then, based on these standard values and weighting factors, the index values and comprehensive evaluation values of each resource utilization rate for each node are calculated. Finally, by comprehensively calculating the standard values, comprehensive evaluation values, and adjustment factors of each node's resource utilization rates, the comprehensive load evaluation coefficient for each node can be obtained. This coefficient helps the system accurately identify the load status of each node, classifying nodes into low-load, medium-load, and high-load nodes. Low-load nodes indicate that the node has sufficient resources to process tasks and the system performance is good; medium-load nodes indicate that the node's resource utilization is relatively reasonable but performance changes need to be monitored; and high-load nodes indicate that the node's resources are close to saturation, system performance may be affected, and measures need to be taken to reduce the load, thereby achieving effective management and optimization of node operating load.
[0033] This invention generates a mapping table by comprehensively considering the data type, size, timeliness requirements, and processing complexity of the data to be processed, and combining this with the real-time load status of each node on each server in each processing area. Based on this table, tasks are prioritized. The system pays particular attention to tasks with high timeliness, giving them higher priority regardless of their data characteristics or node load status. Alternatively, it prioritizes high-complexity data when processing under low load and adjusts priorities when under high load. Simultaneously, by collecting real-time data on task processing time and node load changes, the system dynamically optimizes the mapping table and decision rules to ensure effective resource utilization, achieve more efficient distributed data processing, guarantee data processing quality and speed, and improve the overall system's operating efficiency and response speed. Attached Figure Description
[0034] The invention will now be further described with reference to the accompanying drawings.
[0035] Figure 1 This is a flowchart of the distributed data processing method of the present invention. Detailed Implementation
[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0037] Please see Figure 1 As shown, this invention is a distributed data processing method, comprising the following steps:
[0038] Step 1: Establish a distributed node network: Divide the area covered by the data processing system into several processing regions, and set up several servers within each processing region. Deploy distributed nodes on the corresponding servers in each processing region. The nodes are interconnected through high-speed Ethernet, fiber optic networks, and dedicated networks. By deploying distributed nodes in different geographical locations, if the network in one region fails or suffers a natural disaster, nodes in other regions can still work normally, ensuring the continuity of data processing while effectively improving the reliability and fault tolerance of the system.
[0039] It should be noted that servers include physical servers, virtual servers, and cloud servers. Depending on actual needs and budget, the most suitable server type can be selected. By setting up different types of servers in each processing area, physical servers offer higher performance and stability, while virtual servers and cloud servers offer greater flexibility and scalability. Using multiple different types of servers can improve data processing efficiency to some extent.
[0040] Step 2: Node Status Monitoring: Real-time monitoring of the running status of each node on each server in each distributed processing area is performed to obtain the running status parameters of each node on each server in each distributed processing area. The running status parameters of each node on each server in each distributed processing area include CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization.
[0041] It should be noted that by monitoring the running status of each node on each server in each distributed processing area in real time, node failures or performance problems can be detected in a timely manner, and corresponding measures can be taken to deal with them, such as reassigning tasks and repairing faulty nodes.
[0042] In one specific embodiment, this invention establishes a distributed node network, dividing the area covered by the data processing system into multiple processing zones. Different types of servers, including physical servers, virtual servers, and cloud servers, are deployed within each processing zone. These servers are interconnected via high-speed Ethernet, fiber optic networks, and dedicated networks. This ensures that even if a network in one region fails or suffers a natural disaster, nodes in other regions can still function normally, guaranteeing the continuity of data processing and improving the system's reliability and fault tolerance. Furthermore, by monitoring the real-time operating status parameters of nodes on each server in each processing zone, including CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization, node failures or performance issues can be detected and addressed promptly, thereby further optimizing the system's operating efficiency and stability.
[0043] Step 3: Node Load Analysis: Extract the values of CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization from the running status parameters of each node on each server in each distributed processing region. This yields the CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization of each node on each server in each distributed processing region. By comprehensively calculating and analyzing the CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization of each node on each server in each distributed processing region, a comprehensive load evaluation coefficient for each node on each server in each distributed processing region is obtained.
[0044] Furthermore, the comprehensive load evaluation coefficient for each node on each server in each distributed processing region is calculated and analyzed as follows:
[0045] To obtain the mean CPU utilization of each node on each server in each distributed processing region from time t1 to time t2, sum the m CPU utilization values and divide by m. Then, calculate the variance of the CPU utilization of each node on each server in each distributed processing region using the variance calculation formula. Next, calculate the square root of the variance of the CPU utilization of each node on each server in each distributed processing region to obtain the standard deviation of the CPU utilization of each node on each server in each distributed processing region. Finally, subtract the CPU utilization of each node on each server in each distributed processing region from the mean CPU utilization and divide the result by the standard deviation of the CPU utilization to obtain the standard value of the CPU utilization of each node on each server in each distributed processing region, denoted as CN. i i represents the number of each node, i = 1, 2, ..., n, and n represents the total number of node numbers;
[0046] Similarly, calculate the standard values for memory utilization, network bandwidth utilization, and disk space utilization for each node on each server in each distributed processing region, denoted as MN. i NN i and DN i ;
[0047] It should be noted that the calculation of the CPU utilization variance of each node on each server in each distributed processing region is obtained by calculating the difference between the CPU utilization mean at each time point, squaring the difference, summing the squared results of the differences at each time point, and dividing by m. The variances of memory utilization, network bandwidth utilization, and disk space utilization of each node on each server in each distributed processing region are calculated in the same way.
[0048] According to the formula Calculate the comprehensive evaluation value of each node on each server in each distributed processing region. a1, a2, a3, and a4 represent the weighting factors corresponding to the set CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization, respectively.
[0049] According to the formula Calculate the CPU utilization index value of each node on each server in each distributed processing region. Where α represents the set attenuation factor, 0 < α < 1. Similarly, the exponential values corresponding to the memory utilization, network bandwidth utilization, and disk space utilization of each node on each server in each distributed processing region are calculated, and denoted as follows: and
[0050] By combining the standard values of CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization of each node on each server in each distributed processing region with the comprehensive evaluation value of each node on each server in each distributed processing region, the comprehensive load evaluation coefficient of each node on each server in each distributed processing region is obtained. The specific calculation and analysis method is as follows:
[0051] According to the formula Calculate the comprehensive load evaluation coefficient γ for each node on each server in each distributed processing region, where β represents the set adjustment factor, 0 < β < 1, and w1, w2, w3, and w4 represent the weight factors corresponding to the set standard values of CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization, respectively.
[0052] Step 4: Node load status determination: Obtain the comprehensive load evaluation coefficient classification threshold of each node on each server in each distributed processing area from the system database, and denot it as k1 and k2 respectively, where 0 < k1 < k2 < 1;
[0053] It should be noted that the classification threshold for the comprehensive load evaluation coefficient of each node on each server in each distributed processing area is set by analyzing historical load data to determine some typical load ranges, and then setting thresholds for low load, medium load and high load based on the typical load ranges, which are represented by k1 and k2 respectively.
[0054] The comprehensive load evaluation coefficient of each node on each server in each distributed processing region is compared and analyzed with the classification threshold of the comprehensive load evaluation coefficient of each node on each server in each distributed processing region. The specific comparison and analysis method is as follows:
[0055] If γ < k1, it means that the node has enough resources to handle new tasks, the system performance is good, and the task processing time is short. It is recorded as a low-load node.
[0056] If k1 < γ < k2, it means that the resource utilization of this node is relatively reasonable, but as the load increases, attention may need to be paid to the performance changes, and the task processing time may be slightly extended. It is recorded as a medium load node.
[0057] If γ > k2, it means that the node's resources are close to saturation, the system performance may be affected, the task processing time may be significantly extended, and measures need to be taken to reduce the load. It is recorded as a high-load node.
[0058] In a specific embodiment, this invention analyzes parameters such as CPU utilization, memory utilization, network bandwidth utilization, and disk space utilization of each server node in a distributed processing area to calculate the comprehensive load evaluation coefficient of each node, thereby assessing the load status of each node. Specifically, firstly, the mean and variance of each parameter of each node within a specified time period are calculated. Then, the standard deviation and standardized value are calculated based on the variance. A comprehensive evaluation is performed based on the set weighting factors to finally obtain the comprehensive load evaluation coefficient. Next, low, medium, and high load classification thresholds are set based on historical load data, and the comprehensive load evaluation coefficient of each node is compared with them to determine the load status of each node. This determines whether resources are sufficient and performance is good (low load node), resources are used reasonably and performance changes need to be noted (medium load node), or resources are close to saturation and affecting performance (high load node), thereby achieving effective determination and optimization of the node load status in the distributed processing system.
[0059] Step 5: Data Processing Priority Determination: Extract the characteristics of the data to be processed. The data characteristics include data type, data size, data timeliness requirements, and data processing complexity. The data type classification specifically includes structured data, semi-structured data, and unstructured data. Combine the data characteristics to be processed with the load status of each node on each server in each distributed processing area into a mapping table, and assign a task priority to each combination. At the same time, a decision rule for determining the task allocation priority is formulated.
[0060] The data size is classified as follows: less than 10MB is considered small data, between 10MB and 100MB is considered medium data, and more than 100MB is considered large data.
[0061] The specific classification of data timeliness requirements is as follows: data that needs to be completed within 30 seconds is considered high timeliness, data that needs to be completed within 30 minutes is considered medium timeliness, and data with no completion time requirement is considered low timeliness.
[0062] The data processing complexity is classified as follows: simple processing is defined as data filtering and statistics that only require data filtering and statistics; medium processing is defined as data involving complex algorithms but not requiring a lot of computation; and complex processing is defined as data involving training of simulated learning models.
[0063] The data characteristics to be processed and the load status of each node on each server in each distributed processing area are combined into a mapping table. The specific method of the mapping table is shown in the table below.
[0064]
[0065] In addition, decision-making rules for determining task allocation priorities have been established, including but not limited to:
[0066] 1. When data has high timeliness requirements, regardless of data characteristics and node load status, give it a higher priority in task allocation;
[0067] 2. When processing highly complex data, if the node is under low load, give it high priority; if the node is under high load, reduce its priority.
[0068] It should be noted that as the system's operation and data processing status change, the mapping table and decision rules are dynamically adjusted and optimized in real time. By collecting actual data on the effects of task allocation, such as task processing time and node load changes, if it is found that the priority of task allocation for certain combinations is unreasonable, the mapping table can be adjusted. In addition, based on changes in business needs, the importance of data features and the impact of node load status can be reassessed, and the weights and priority judgment criteria of decision rules can be adjusted. By establishing mapping tables or decision rules, the priority of task allocation can be determined more systematically based on different data features and node load status, thereby achieving more efficient distributed data processing.
[0069] In one specific embodiment, this invention constructs a mapping table by comprehensively considering the data type, size, timeliness requirements, and processing complexity of the data to be processed, as well as the current load status of each server and node in the distributed processing environment. Based on this, a task allocation priority is assigned to each data processing task. Specifically, when data has high timeliness requirements, the task will be assigned a higher priority regardless of other conditions. For data with high processing complexity, the task will be assigned a high priority when the node is under low load. Conversely, if the node is under high load, the task priority will be reduced. In addition, by continuously collecting data such as actual task processing time and node load changes, the system can dynamically adjust the mapping table and decision rules to ensure the rationality and effectiveness of the allocation strategy. This enables efficient and orderly processing of data of different types and sizes in a distributed environment, meeting the data timeliness and processing complexity requirements under different business needs, and significantly improving the overall efficiency and performance of data processing.
[0070] Step 6: Data processing task allocation: Assign high-priority tasks to the low-load nodes on the corresponding servers in each distributed processing area. While ensuring that high-priority tasks are processed first, assign medium-priority tasks to the nodes on the corresponding servers in each distributed processing area. Finally, schedule low-priority tasks for processing during system idle time, such as at night or during other low-load node idle periods.
[0071] Step 7: Distributed processing node optimization: Real-time monitoring and acquisition of the comprehensive load evaluation coefficient of each node on each server in each distributed processing area. The monitoring period is divided into several monitoring periods. When a node on each server in each distributed processing area frequently experiences delays or excessive load, a load anomaly signal is generated, and the task allocation weight of that node is adjusted.
[0072] The system obtains the number of nodes on each server in each distributed processing region that exhibit load anomaly signals during a monitoring period, denoted as g. It then compares g with the number of nodes n on each server in the distributed processing region during the monitoring period. If the result is greater than 0.8, it indicates insufficient load margin for each node on each server in the distributed processing region during the monitoring period, generating a node expansion optimization signal. The processing system increases the number of nodes on each server in the distributed processing region during the monitoring period by 50%. After the new node joins the system, the load balancer automatically adjusts the load balancing, evenly distributing data processing tasks across all nodes. If the result is less than 0.4, it indicates sufficient load margin for each node on each server in the distributed processing region during the monitoring period, generating a node contraction optimization signal. The processing system reduces the number of nodes on each server in the distributed processing region during the monitoring period by 20%.
[0073] In one specific embodiment, this invention utilizes refined task allocation and node optimization strategies to efficiently utilize server resources within the distributed processing area, thereby ensuring the priority processing of core tasks and the stable operation of the overall system. First, the system assigns high-priority tasks to low-load nodes on each server for immediate processing. Then, medium-priority tasks are allocated to other nodes to minimize latency. Finally, low-priority tasks are scheduled for processing during system idle periods, such as at night or when node load is low. During distributed processing, the system also monitors the overall load of each node in real time. If frequent latency or overload is detected on a node, the load pressure is alleviated by adjusting task allocation weights. Elsewhere, the system counts the number of nodes exhibiting abnormal load signals within a monitoring period and compares this number with the total number of nodes. If the ratio exceeds 80%, it indicates a relatively high load, and the system triggers a node expansion optimization signal, increasing the number of existing nodes by 50% to enhance processing capacity. The load balancer automatically adjusts to ensure even task distribution. Conversely, if the ratio is below 40%, it indicates a relatively light load, and the system triggers a node contraction optimization signal, reducing the number of current nodes by 20% to utilize existing resources more efficiently. By dynamically adjusting the number of nodes and intelligently allocating tasks, the system can effectively handle the processing needs of tasks with different priorities and continuously optimize resource utilization during operation, ensuring the stability and efficiency of the overall system performance.
[0074] The above description is merely an example and illustration of the structure of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the structure of the invention or exceed the scope defined in the claims, all of which should fall within the protection scope of the present invention.
Claims
1. A distributed data processing method, comprising the following steps: Step one, establishing a distributed node network; Step two, node state monitoring; Step three, node running load analysis: comprehensive calculation and analysis of CPU usage, memory usage, network bandwidth occupancy and disk space usage of each node on each server in each processing area of the distributed system, to obtain the comprehensive load evaluation coefficient of each node on each server in each processing area of the distributed system; The specific way of comprehensive calculation and analysis of CPU usage, memory usage, network bandwidth occupancy and disk space usage of each node on each server in each processing area of the distributed system is as follows: The CPU usage rate of each node on each server in each processing area of the distribution is obtained by summing the CPU usage rate of each node on each server in each processing area of the distribution from the time point t1 to the time point t2, dividing by m, and then taking the average. The CPU usage rate variance of each node on each server in each processing area of the distribution is calculated according to the CPU usage rate of each node on each server in each processing area of the distribution. The square root of the CPU usage rate variance of each node on each server in each processing area of the distribution is calculated, and the CPU usage rate standard deviation of each node on each server in each processing area of the distribution is obtained. Finally, the CPU usage rate of each node on each server in each processing area of the distribution is subtracted from the CPU usage rate average, and the result is divided by the CPU usage rate standard deviation to obtain the CPU usage rate standard value of each node on each server in each processing area of the distribution, denoted as CN i , where i represents the number of each node. Similarly, the standard values corresponding to the memory usage, network bandwidth occupation and disk space usage of each node on each server in each processing area of the distributed system are calculated, and are denoted as MN i , NN i and DN i , respectively. According to the formula The comprehensive evaluation value corresponding to each node on each server in each processing area of the distribution is calculated a1, a2, a3, a4 respectively represent the weight factors corresponding to the set CPU usage, memory usage, network bandwidth occupancy and disk space usage. According to the formula The CPU usage index value of each node on each server in each processing area of the distribution is calculated , where α represents a set attenuation factor, 0 < α < 1, and the memory usage, network bandwidth occupancy, and disk space usage of each node on each server in each processing area of the distribution are calculated in the same way, and are respectively denoted as , , and ; The comprehensive load evaluation coefficient of each node on each server in each processing area of the distributed system is calculated and analyzed as follows: Step four, node load state determination: obtain the comprehensive load evaluation coefficient classification threshold of each node on each server in each processing area of the distributed system from the system database, and compare and analyze the comprehensive load evaluation coefficient of each node on each server in each processing area of the distributed system with the comprehensive load evaluation coefficient classification threshold of each node on each server in each processing area of the distributed system; According to the formula The comprehensive load evaluation coefficient of each node on each server in each processing area is calculated Wherein, β represents a set adjustment factor, w1, w2, w3, w4 respectively represent weight factors corresponding to the set CPU usage index value, the memory usage index value, the network bandwidth occupation rate index value and the disk space usage rate index value. Step five, data processing priority determination: extract the data features to be processed, combine the data features to be processed with the load state of each node on each server in each processing area of the distributed system into a mapping table, and assign a task allocation priority to each combination, while formulating a decision rule for determining the task allocation priority; Step six, data processing task allocation: allocate high-priority tasks to low-load nodes on each server in each processing area of the distributed system, allocate medium-priority tasks to other nodes on each server in each processing area of the distributed system under the premise of ensuring high-priority task processing, and finally arrange low-priority tasks to be processed during node idle time; Step seven, distributed processing node optimization: real-time monitoring of the comprehensive load evaluation coefficient of each node on each server in each processing area of the distributed system, dividing the monitoring period into several monitoring periods, and generating a load abnormal signal when a node in each processing area of the distributed system has a delay or a high load, and adjusting the task allocation weight of the node. The specific process of step one is to divide the regions covered by the data processing system into several processing areas, and set several servers in each processing area, and deploy each node in each processing area on the corresponding server, wherein each node in each processing area of the distributed system is connected to each other through high-speed Ethernet, optical fiber network and special network.
2. The distributed data processing method of claim 1, wherein: 3. The distributed data processing method of claim 1, wherein: The specific process of step two is: real-time monitoring of the running state of each node on the corresponding server in each processing area of the distributed system to obtain the running state parameters of each node on the corresponding server in each processing area of the distributed system, wherein the running state parameters include CPU usage, memory usage, network bandwidth occupation rate and disk space usage.
4. The distributed data processing method of claim 1, wherein: In step four, the comprehensive load evaluation coefficient of each node on the corresponding server in each processing area of the distributed system and the comprehensive load evaluation coefficient classification threshold of each node on the corresponding server in each processing area of the distributed system are compared and analyzed, and the specific comparison and analysis method is as follows: The comprehensive load evaluation coefficient classification threshold of each node on the corresponding server in each processing area of the distributed system is obtained from the system database, and is denoted as k1 and k2, wherein 0 < k1 < k2 < 1. If <k1, then it is denoted that the node has spare resources to handle new tasks, and it is recorded as a low-load node; If k1 < If k2 < k2, it means that the node's resource usage is reasonable, and it is recorded as a medium load node; If >k2, it means that the node resource is close to saturation, and it is recorded as a high-load node.
5. The distributed data processing method of claim 1, wherein: In step five, the data characteristics include data type, data size, data timeliness requirement and data processing complexity.
6. The distributed data processing method of claim 1, wherein: In step seven, the number of nodes that generate load abnormal signals in each processing area of the distributed system within a monitoring period is obtained and denoted as g. The number of nodes that generate load abnormal signals in each processing area of the distributed system within a monitoring period g and the number of nodes n in each processing area of the distributed system are compared. If the calculation result is greater than 0.8, it indicates that the load margin of each node in each processing area of the distributed system within the monitoring period is insufficient, a node expansion optimization signal is generated, the number of nodes in each processing area of the distributed system within the monitoring period is increased by 50% of the current number of nodes, and after the new nodes are added to the system, the load balancer automatically adjusts the load balancing and uniformly distributes the data processing tasks to each node. If the calculation result is less than 0.4, it indicates that the load margin of each node in each processing area of the distributed system within the monitoring period is sufficient, a node contraction optimization signal is generated, and the number of nodes in each processing area of the distributed system within the monitoring period is reduced by 20% of the current number of nodes.
Citation Information
Patent Citations
Dynamic load balancing method for distributed file system under cloud environment
CN108200156A
Method and system for cluster dynamic balance expansion
CN117950858A