Distributed data processing method and system based on cloud computing

By constructing a multi-dimensional node feature vector and a dynamic weight allocation model, combined with a graph matching algorithm and a task priority strategy, the problems of node overload and high communication costs in distributed data processing systems are solved, and efficient and energy-saving task scheduling and system stability are achieved.

CN120602486AInactive Publication Date: 2025-09-05NANJING ZERO 323 TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510902063.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-09-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In traditional distributed data processing systems, static scheduling strategies cannot be dynamically adjusted according to the actual node status, resulting in overload or resource waste of some nodes. They also fail to fully consider the time and economic costs of cross-node communication, affecting the overall execution efficiency of the system.

Method used

By obtaining the heterogeneous resource parameters of computing nodes, constructing a multi-dimensional node feature vector, using a dynamic weight distribution model and graph matching algorithm to optimize data sharding, setting task priorities and adopting cross-availability zone multi-copy storage and erasure code encoding storage, real-time monitoring of load and cost, triggering the migration of differential state snapshots, and realizing dynamic re-sharding.

Benefits of technology

It improves the flexibility and execution efficiency of task scheduling, reduces communication costs and delays, enhances the robustness and stability of the system, and ensures the high availability and disaster recovery capabilities of tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602486A_ABST
    Figure CN120602486A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed data processing method and system based on cloud computing, and relates to the technical field of distributed data processing.The method comprises the steps that heterogeneous resource parameters of computing nodes are obtained through a cloud platform interface, and the heterogeneous resource parameters comprise computing power indexes, network topology distances, real-time energy consumption efficiency and cloud service pricing data; constructing a multi-dimensional node feature vector; according to the multi-dimensional node feature vector, a dynamic weight distribution model is adopted to calculate a fragment weight value of each node, the weight distribution model introduces a normalization coefficient of energy consumption efficiency and cloud service pricing, and to-be-processed data is divided into data fragments in direct proportion to the fragment weight values; generating an inter-node communication cost matrix based on the network topology distance and cloud service pricing data, and distributing the data fragments to target computing nodes by adopting a graph matching algorithm to minimize the total communication cost; and setting task priorities, and screening key tasks and non-key tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distributed data processing, and in particular to a distributed data processing method and system based on cloud computing. Background Art

[0002] Distributed data processing technology involves distributing and storing data across multiple computing nodes over a network, leveraging the computing power of these nodes to collaboratively complete data processing tasks. Therefore, utilizing advanced technologies to improve the intelligence and security of distributed data processing has become a pressing issue.

[0003] In the field of distributed data processing, traditional static scheduling strategies often allocate tasks according to fixed rules or ratios and cannot be dynamically adjusted according to the actual node status, which can easily cause some nodes to be overloaded or resources to be wasted. In addition, many distributed systems do not fully consider the time and economic costs of cross-node communication when scheduling tasks, resulting in excessively high total communication costs and affecting the overall execution efficiency of the system. At the same time, in a cloud computing environment, if resource recycling or node failure occurs, the ongoing critical tasks will be interrupted, affecting business continuity. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a distributed data processing method based on cloud computing to solve the problem that traditional static scheduling strategies often allocate tasks according to fixed rules or proportions, cannot be dynamically adjusted according to the actual node status, and easily cause some nodes to be overloaded or resources to be wasted. In addition, many distributed systems do not fully consider the time and economic costs of cross-node communication when scheduling tasks, resulting in excessively high total communication costs, which affects the overall execution efficiency of the system.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a distributed data processing method based on cloud computing, comprising:

[0008] Obtaining heterogeneous resource parameters of computing nodes through the cloud platform interface, including computing power indicators, network topology distance, real-time energy efficiency, and cloud service pricing data, and constructing a multi-dimensional node feature vector;

[0009] According to the multi-dimensional node feature vector, a dynamic weight distribution model is used to calculate the shard weight value of each node. The weight distribution model introduces a normalization coefficient of energy efficiency and cloud service pricing to divide the data to be processed into data shards proportional to the shard weight value;

[0010] Based on the network topology distance and cloud service pricing data, a communication cost matrix between nodes is generated, and a graph matching algorithm is used to distribute the data shards to target computing nodes to minimize the total communication cost;

[0011] Set task priorities, filter critical tasks and non-critical tasks, use cross-availability zone multi-copy storage for critical tasks, use erasure coding for non-critical tasks, and trigger the migration of differential state snapshots to private cloud nodes when cloud platform resource recovery signals are detected;

[0012] Monitor the actual load and cost consumption of each node, use feedback to adjust the calculation coefficient of the sharding weight value, and dynamically re-shard tasks that exceed the execution threshold time.

[0013] As a preferred solution of the distributed data processing method based on cloud computing of the present invention, the construction of a multi-dimensional node feature vector includes:

[0014] Obtain the heterogeneous resource parameters of the computing node through the cloud platform interface, where the computing power indicator is expressed as the number of floating-point operations completed per unit time FLOPS, denoted as f i ;

[0015] The network topology distance is expressed as the network delay between nodes, denoted as d ij , where i and j represent two nodes respectively;

[0016] Real-time energy efficiency is expressed as the energy consumption value per unit time, denoted as e i ;

[0017] Cloud service pricing data is expressed as cost per unit time, denoted as c i ;

[0018] Normalize the above four parameters to obtain the standardized eigenvalues;

[0019] Combine the normalized results into a multidimensional node feature vector V according to the weights i .

[0020] As a preferred solution of the distributed data processing method based on cloud computing of the present invention, the construction and sharding process of the dynamic weight distribution model includes the following steps:

[0021] Based on the multidimensional node feature vector V i Construct a comprehensive scoring function;

[0022] The comprehensive score S of all nodes i Perform standardization processing to obtain the standardized comprehensive score of node i

[0023] Based on the standardized comprehensive score Calculate the shard weight value W of each node i ;

[0024] Assume that the total amount of data to be processed is T, then according to the shard weight W of each node i , divide the data into several subsets and assign them to corresponding nodes.

[0025] As a preferred solution of the distributed data processing method based on cloud computing of the present invention, the communication cost matrix generation and graph matching optimization process includes:

[0026] First, define the time and economic cost incurred when node i sends a unit of data to node j;

[0027] The communication cost consists of two parts: network transmission time cost and receiving end operation cost;

[0028] Calculate the element M in the communication cost matrix ij ;

[0029] For all node pairs (i, j), generate the communication cost matrix M;

[0030] If node i is not allowed to send data to itself, then M ii Set to infinity or a maximum value to prohibit self-loops. N is the total number of available computing nodes.

[0031] Consider each data shard as a task node to form a task set, and the set of available computing nodes as a candidate node set. Construct a bipartite graph G = (U ∪ V, E), where U represents the set of task nodes, i.e., data shards, and V represents the set of candidate nodes, i.e., computing nodes.

[0032] Each edge (u, v) in the edge set E corresponds to the cost of task u∈U assigned to node v∈V, which is M ij ;

[0033] The Kuhn-Munkres algorithm is used for optimal matching. The Kuhn-Munkres algorithm is used to find a way to allocate tasks to nodes so as to minimize the total communication cost under the premise of a given cost matrix m×n.

[0034] The input is the communication cost matrix M, and the output is the optimal matching solution;

[0035] If the number of tasks m is greater than the number of nodes n, multiple rounds of scheduling or the introduction of load balancing strategies are required;

[0036] If the number of nodes n is greater than the number of tasks m, some nodes will not be selected to avoid resource waste.

[0037] As a preferred solution of the distributed data processing method based on cloud computing of the present invention, the task priority screening and storage strategy implementation process includes:

[0038] Set the task priority threshold P th , obtain the task priority value P according to the task type or user-specified information i ;

[0039] If the condition P is met i >P th , then mark the task as a critical task, otherwise mark it as a non-critical task;

[0040] For data marked as mission-critical, a cross-availability zone multi-copy storage strategy is adopted:

[0041] Copy the original task data into multiple completely identical copies, and the number of copies is recorded as R k ;

[0042] Each replica is stored independently in a different availability zone;

[0043] The availability zone is an area in the cloud platform that has physical isolation capabilities;

[0044] When a failure occurs in one availability zone, other replicas can still ensure that tasks continue to execute;

[0045] For data marked as non-critical, an erasure code storage strategy is used;

[0046] Use (n,k) erasure code to encode task data;

[0047] Where n represents the total number of segments after encoding, and k represents the number of data segments into which the original data is divided;

[0048] Generate nk check fragments using Reed-Solomon coding;

[0049] All n fragments are stored in different nodes in the cluster;

[0050] Only any k fragments are needed to restore the original data;

[0051] Real-time monitoring of whether the cloud platform sends a resource recovery signal;

[0052] If a resource recycling signal is detected, the differential state snapshot mechanism is triggered, including:

[0053] Record the execution status and memory snapshot of the current task;

[0054] A differential snapshot contains only the parts that have changed since the previous snapshot;

[0055] Compress and package incremental snapshots;

[0056] Migrate to the private cloud node, and the standby node loads the snapshot and restores the task context;

[0057] Allows the task to continue from the interruption point to avoid interruption losses.

[0058] As a preferred solution of the distributed data processing method based on cloud computing of the present invention, the load and cost monitoring feedback includes the following steps:

[0059] During the task execution process, the actual operating status parameters of each node are collected, including: the number of tasks currently being processed by the node, the maximum number of concurrent tasks of the node, the node's elapsed running time, and the node's unit time cost;

[0060] Based on the actual operating status parameters, calculate the current load value L of the node i ;

[0061] At the same time, the cumulative cost consumption TC of the node is calculated i ;

[0062] Set the system load threshold L th and cost budget ratio threshold δ;

[0063] The feedback adjustment mechanism is triggered when any of the following conditions are met:

[0064] L i >L th , that is, the node load is too high;

[0065] TC i >δ·TC total , that is, the cost of this node exceeds the set proportion of the overall budget;

[0066] After triggering, recalculate the node's shard weight value W' i ;

[0067] Updated weight value W′ i Used for the next round of data sharding and scheduling decisions;

[0068] The feedback mechanism is executed periodically to ensure dynamic balance of cluster resources and controllable costs.

[0069] As a preferred solution of the distributed data processing method based on cloud computing of the present invention, the dynamic resharding mechanism includes the following steps:

[0070] Set the task execution time threshold T th , preset by the user or the system;

[0071] Monitor the execution time of each task;

[0072] When the current execution time of a task is T exec >T th , it is determined to be a timed task;

[0073] To start dynamic resharding for timed tasks, follow these steps:

[0074] Analyze the data dependency graph of the task and identify a set of subtasks that can be executed independently;

[0075] Each subtask has independent input data and output results, records the dependency order between subtasks, and generates a subtask scheduling topology graph;

[0076] Obtain status information of all currently available nodes, including current load, unit time cost, and network topology distance;

[0077] Based on the state information, calculate the comprehensive score Q of each candidate node j ;

[0078] Adopt the minimum priority queue strategy to assign subtasks to the node with the lowest score in sequence;

[0079] Pack the selected subtask data and send it to the target node;

[0080] The original node retains the main control logic and coordinates the synchronization and result aggregation between subtasks;

[0081] Update the global task schedule, record the allocation of subtasks, and generate a migration log, including the task ID, subtask number, source node number, target node number, migration timestamp, and data size.

[0082] In a second aspect, the present invention provides a distributed data processing system based on cloud computing, comprising:

[0083] Resource assessment module, task allocation module, communication optimization module, state migration module and load control module;

[0084] The resource evaluation module is used to obtain the heterogeneous resource parameters of the computing nodes through the cloud platform interface, including computing power indicators, network topology distance, energy efficiency and cloud service pricing, and normalize the parameters and combine them into a multi-dimensional node feature vector;

[0085] The task allocation module is used to calculate the shard weight value based on the comprehensive score and standardized score of each node, and divide the data to be processed into multiple subsets according to the weight ratio and allocate them to the corresponding nodes respectively;

[0086] The communication optimization module is used to construct a communication cost matrix based on network topology distance and unit time cost, and use a graph matching algorithm to optimally match data shards with target nodes to minimize the total communication cost;

[0087] The state migration module is used to set task priority thresholds and screen critical tasks from non-critical tasks. It uses cross-availability zone multi-copy storage for critical tasks and erasure coding for non-critical tasks. When a resource recycling signal is detected, a differential state snapshot mechanism is triggered to migrate the task state to a private cloud node to ensure execution continuity.

[0088] The load control module is used to monitor the actual load and cost consumption of each node in real time. When the node load or cost exceeds the preset threshold, it provides feedback to adjust the calculation coefficient of the sharding weight value and dynamically re-shards the timed-out tasks.

[0089] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the distributed data processing method based on cloud computing as described in the first aspect of the present invention is implemented.

[0090] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the distributed data processing method based on cloud computing as described in the first aspect of the present invention.

[0091] The beneficial effects of the present invention are as follows: through the dynamic weight distribution model and the sharding mechanism, the personalized and flexible task scheduling is achieved, and the problems of uneven resource utilization and rigid scheduling in the traditional static scheduling method are overcome, thereby effectively improving the execution efficiency and reducing the operating cost, and achieving the effect of efficient, energy-saving and economical collaborative computing effect; through the communication cost modeling and graph matching optimization mechanism, the optimal selection of the communication path is achieved, which significantly reduces the task waiting time and the resource waste caused by cross-node communication, improves the system's throughput and response speed, and achieves the effect of reducing communication delay and reducing overall cost; through the task priority screening and state migration control mechanism, the high availability and elastic recovery capability of task execution are achieved, and the problem of task interruption due to resource recovery or failure in the cloud computing environment is solved, thereby enhancing the robustness and stability of the system, and achieving the effect of ensuring task integrity and improving the system's disaster recovery capability. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0093] Figure 1 This is a flow chart of the distributed data processing method based on cloud computing in Example 1.

[0094] Figure 2 Schematic diagram of a distributed data processing system based on cloud computing in Example 1. DETAILED DESCRIPTION

[0095] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0096] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0097] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0098] Example, see Figure 1 and Figure 2 , is an embodiment of the present invention, which provides a distributed data processing method based on cloud computing, comprising the following steps:

[0099] S1. Obtain the heterogeneous resource parameters of computing nodes through the cloud platform interface. The heterogeneous resource parameters include computing power indicators, network topology distance, real-time energy efficiency and cloud service pricing data, and construct a multi-dimensional node feature vector.

[0100] Furthermore, constructing a multi-dimensional node feature vector includes:

[0101] Obtain the heterogeneous resource parameters of the computing node through the cloud platform interface, where the computing power indicator is expressed as the number of floating-point operations completed per unit time FLOPS, denoted as f i ;

[0102] The network topology distance is expressed as the network delay between nodes, denoted as dij , where i and j represent two nodes respectively;

[0103] Real-time energy efficiency is expressed as the energy consumption value per unit time, denoted as e i ;

[0104] Cloud service pricing data is expressed as cost per unit time, denoted as c i ;

[0105] Normalize the above four parameters to obtain the standardized eigenvalues:

[0106]

[0107] Combine the normalized results into a multidimensional node feature vector V according to the weights i , the expression is:

[0108] V i =[w1·f i ,w2·e i ,w3·c i ];

[0109] Among them, w1, w2, and w3 represent the weight coefficients corresponding to computing power, energy efficiency, and cost, respectively, which are configured by the system administrator or dynamically adjusted based on historical task execution;

[0110] It should be noted that by defining the node computing power indicator as the number of floating-point operations completed per unit time, and normalizing it in combination with network latency, energy efficiency and cloud service pricing, a multi-dimensional node feature vector of unified dimension is constructed, so that heterogeneous resources can be quantitatively evaluated at the same scale. This method not only improves the accuracy of resource perception, but also enhances the adaptability of task scheduling strategies to different types of nodes, thereby effectively improving the overall resource utilization of the system and task execution efficiency.

[0111] S2. Based on the multi-dimensional node feature vector, a dynamic weight allocation model is used to calculate the shard weight value of each node. The weight allocation model introduces normalization coefficients of energy efficiency and cloud service pricing to divide the data to be processed into data shards proportional to the shard weight value;

[0112] Furthermore, the construction and sharding process of the dynamic weight distribution model includes the following steps:

[0113] Based on the multidimensional node feature vector V i Construct a comprehensive scoring function, which is expressed as:

[0114] S i =w a ·f i +w b ·ei -w c c i ;

[0115] Among them, S i is the comprehensive score of node i, w a is the weighted coefficient of computing power dimension, w b is the weighted coefficient of energy efficiency dimension, w c is the weighting coefficient of the cost dimension;

[0116] The comprehensive score S of all nodes i Perform standardization processing to obtain the standardized comprehensive score of node i

[0117] Based on the standardized comprehensive score Calculate the shard weight value W of each node i , the expression is:

[0118]

[0119] Among them, N is the total number of nodes participating in task scheduling, represents the sum of the standardized scores of all nodes, W i Assign weights to node i’s tasks in the entire cluster;

[0120] Assume that the total amount of data to be processed is T, then according to the shard weight W of each node i , divide the data into several subsets and assign them to corresponding nodes respectively. The expression is:

[0121] T i =T·W i ;

[0122] Among them, T is the total amount of data to be processed, T i The amount of data that should be allocated to node i;

[0123] It should be noted that a comprehensive scoring function encompassing computing power, energy efficiency, and cost is constructed based on multidimensional feature vectors, and a weighting coefficient is introduced for dynamic adjustment, enabling personalized calculation of task allocation weights. Sharding weights are determined by calculating the ratio of each node's standardized score, ensuring that data partitioning is proportional to node capabilities, avoiding resource waste and uneven load. This method enhances the flexibility and intelligence of task scheduling, enabling the system to adaptively optimize the balance between performance and cost based on different business objectives.

[0124] S3. Generate an inter-node communication cost matrix based on network topology distance and cloud service pricing data, and use a graph matching algorithm to distribute data shards to target computing nodes to minimize the total communication cost.

[0125] Furthermore, the communication cost matrix generation and graph matching optimization process includes:

[0126] First, define the time and economic cost incurred when node i sends a unit of data to node j;

[0127] The communication cost consists of two parts: the network transmission time cost and the receiving end operation cost;

[0128] Calculate the element M in the communication cost matrix ij , the expression is:

[0129] M ij =d ij ·r+p·C j ;

[0130] Among them, M ij is the cost of sending unit data from node i to node j, d ij is the network delay between node i and node j, r is the unit data transmission rate, p is the size of the transmitted data, C j is the unit time cost of the receiver node j;

[0131] For all node pairs (i, j), generate the communication cost matrix M, the matrix form is:

[0132]

[0133] If node i is not allowed to send data to itself, then M ii Set to infinity or a maximum value to prohibit self-loops. N is the total number of available computing nodes.

[0134] Consider each data shard as a task node to form a task set, and the set of available computing nodes as a candidate node set. Construct a bipartite graph G = (U ∪ V, E), where U represents the set of task nodes, i.e., data shards, and V represents the set of candidate nodes, i.e., computing nodes.

[0135] Each edge (u, v) in the edge set E corresponds to the cost of task u∈U assigned to node v∈V, which is M ij ;

[0136] The Kuhn-Munkres algorithm is used for optimal matching. The Kuhn-Munkres algorithm is used to find the allocation method of tasks to nodes to minimize the total communication cost under the premise of a given cost matrix m×n;

[0137] The input is the communication cost matrix M, and the output is the optimal matching solution;

[0138] If the number of tasks m is greater than the number of nodes n, multiple rounds of scheduling or the introduction of load balancing strategies are required;

[0139] If the number of nodes n is greater than the number of tasks m, some nodes will not be selected to avoid resource waste;

[0140] It should be noted that by establishing a communication cost model that includes network transmission time costs and receiving end operating costs, and using the Kuhn-Munkres algorithm for graph matching optimization, the global optimal allocation between tasks and nodes is achieved. While ensuring the rationality of task allocation, the mechanism effectively reduces the delay and cost caused by cross-node communication. It is particularly suitable for complex scheduling scenarios where the number of nodes is inconsistent with the number of tasks, and significantly improves the system's communication efficiency and resource scheduling capabilities.

[0141] S4. Set task priorities, filter critical tasks and non-critical tasks, use cross-availability zone multi-copy storage for critical tasks, use erasure coding for non-critical tasks, and trigger the migration of differential state snapshots to private cloud nodes when a cloud platform resource recovery signal is detected;

[0142] Furthermore, the task priority screening and storage policy implementation process includes:

[0143] Set the task priority threshold P th , obtain the task priority value P according to the task type or user-specified information i ;

[0144] If the condition P is met i >P th , then mark the task as a critical task, otherwise mark it as a non-critical task;

[0145] For data marked as mission-critical, a cross-availability zone multi-copy storage strategy is adopted:

[0146] Copy the original task data into multiple completely identical copies, and the number of copies is recorded as R k ;

[0147] Each replica is stored independently in a different availability zone;

[0148] Availability Zones are physically isolated areas within a cloud platform.

[0149] When a failure occurs in one availability zone, other replicas can still ensure that tasks continue to execute;

[0150] For data marked as non-critical, an erasure code storage strategy is used;

[0151] Use (n,k) erasure code to encode task data;

[0152] Where n represents the total number of segments after encoding, and k represents the number of data segments into which the original data is divided;

[0153] Generate nk check fragments using Reed-Solomon coding;

[0154] All n fragments are stored in different nodes in the cluster;

[0155] Only any k fragments are needed to restore the original data;

[0156] Real-time monitoring of whether the cloud platform sends a resource recovery signal;

[0157] If a resource recycling signal is detected, the differential state snapshot mechanism is triggered, including:

[0158] Record the execution status and memory snapshot of the current task;

[0159] A differential snapshot contains only the parts that have changed since the previous snapshot;

[0160] Compress and package incremental snapshots;

[0161] Migrate to the private cloud node, and the standby node loads the snapshot and restores the task context;

[0162] Allow the task to continue from the interruption point to avoid interruption losses;

[0163] It should be noted that by setting task priority thresholds to distinguish between critical tasks and non-critical tasks, and adopting cross-availability zone multi-copy storage and erasure code storage strategies respectively, a reasonable configuration of task data between reliability and economy is achieved. When the cloud platform resource recovery signal is detected, the differential state snapshot mechanism is triggered, and only the changed part of the task state is migrated, which greatly reduces the migration overhead and ensures the continuity of task execution and disaster recovery capabilities. The method significantly improves the system robustness and task recovery efficiency, and is particularly suitable for cloud computing environments in high concurrency and high availability scenarios.

[0164] S5. Monitor the actual load and cost consumption of each node, provide feedback to adjust the calculation coefficient of the sharding weight value, and dynamically re-shard tasks that exceed the execution threshold time;

[0165] Furthermore, load and cost monitoring feedback includes the following steps:

[0166] During the task execution process, the actual operating status parameters of each node are collected, including: the number of tasks currently being processed by the node, the maximum number of concurrent tasks of the node, the node's elapsed running time, and the node's unit time cost;

[0167] Based on the actual operating status parameters, calculate the current load value L of the node i , the expression is:

[0168]

[0169] Among them, L i Indicates the current load ratio of the node. is the number of tasks currently being processed by node i, is the maximum number of concurrent tasks for node i;

[0170] At the same time, the cumulative cost consumption TC of the node is calculated i , the expression is:

[0171] TC i =t i ·C i ;

[0172] Among them, TC i Indicates the total cost consumed by the node since the start of the task;

[0173] Set the system load threshold L th and cost budget ratio threshold δ;

[0174] The feedback adjustment mechanism is triggered when any of the following conditions are met:

[0175] L i >L th , that is, the node load is too high;

[0176] TC i >δ·TC total , that is, the cost of this node exceeds the set proportion of the overall budget;

[0177] After triggering, recalculate the node's shard weight value W' i , the expression is:

[0178] W′ i =W i ·(1-α·(L i -L th ));

[0179] Among them, W i is the original shard weight, α is the attenuation coefficient, which is used to control the impact of overload on the weight. If W′ i <0, it is set to 0, indicating that no new tasks will be assigned to the node;

[0180] Updated weight value W′ i Used for the next round of data sharding and scheduling decisions;

[0181] The feedback mechanism is executed periodically to ensure dynamic balance of cluster resources and controllable costs;

[0182] The dynamic resharding mechanism includes the following steps:

[0183] Set the task execution time threshold T th , preset by the user or the system;

[0184] Monitor the execution time of each task;

[0185] When the current execution time of a task is T exec >T th , it is determined to be a timed task;

[0186] To start dynamic resharding for timed tasks, follow these steps:

[0187] Analyze the data dependency graph of the task and identify a set of subtasks that can be executed independently;

[0188] Each subtask has independent input data and output results, records the dependency order between subtasks, and generates a subtask scheduling topology graph;

[0189] Obtain status information of all currently available nodes, including current load, unit time cost, and network topology distance;

[0190] Based on the state information, calculate the comprehensive score Q of each candidate node j , the expression is:

[0191] Q j =w d ·d ij +w c ·C j +w l ·L j ;

[0192] Among them, w d , w c , w l are the weighted coefficients of network delay, cost, and load, respectively. j The smaller the value, the more suitable the node is for migration.

[0193] Adopt the minimum priority queue strategy to assign subtasks to the node with the lowest score in sequence;

[0194] Pack the selected subtask data and send it to the target node;

[0195] The original node retains the main control logic and coordinates the synchronization and result aggregation between subtasks;

[0196] Update the global task schedule, record the allocation of subtasks, and generate a migration log, including the task ID, subtask number, source node number, target node number, migration timestamp, and data size;

[0197] It should be noted that through real-time monitoring of node load and cost consumption, combined with dynamic adjustment of sharding weight values ​​based on preset thresholds, closed-loop feedback control of system resources is achieved to prevent node overload or resource idleness, and maintain cluster load balance. At the same time, a dynamic re-sharding mechanism is initiated for timed tasks, and sub-task decomposition and candidate node scoring strategies are used to reschedule tasks to better nodes for execution, thereby improving task response speed and system elastic scheduling capabilities. The mechanism enhances the system's ability to respond to abnormal situations and improves the stability and timeliness of task execution.

[0198] This embodiment also provides a distributed data processing system based on cloud computing, including:

[0199] Resource assessment module, task allocation module, communication optimization module, state migration module and load control module;

[0200] The resource evaluation module is used to obtain the heterogeneous resource parameters of computing nodes through the cloud platform interface, including computing power indicators, network topology distance, energy efficiency and cloud service pricing, and normalize the parameters and combine them into a multi-dimensional node feature vector;

[0201] The task allocation module is used to calculate the shard weight value based on the comprehensive score and standardized score of each node, and divide the data to be processed into multiple subsets according to the weight ratio and assign them to the corresponding nodes respectively;

[0202] The communication optimization module is used to construct a communication cost matrix based on the network topology distance and unit time cost, and use a graph matching algorithm to optimally match data shards with target nodes to minimize the total communication cost;

[0203] The state migration module is used to set task priority thresholds and screen critical tasks from non-critical tasks. It uses cross-availability zone multi-replica storage for critical tasks and erasure coding for non-critical tasks. When a resource recycling signal is detected, the differential state snapshot mechanism is triggered to migrate the task state to the private cloud node to ensure execution continuity.

[0204] The load control module is used to monitor the actual load and cost consumption of each node in real time. When the node load or cost exceeds the preset threshold, it will feedback and adjust the calculation coefficient of the sharding weight value, and dynamically re-shard the execution timed-out tasks.

[0205] This embodiment also provides a computer device suitable for a distributed data processing method based on cloud computing, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the distributed data processing method based on cloud computing proposed in the above embodiment.

[0206] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.

[0207] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the distributed data processing method based on cloud computing proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0208] In summary, the present invention realizes the personalization and flexibility of task scheduling through a dynamic weight distribution model and a sharding mechanism, overcomes the problems of uneven resource utilization and rigid scheduling in traditional static scheduling methods, thereby effectively improving execution efficiency, reducing operating costs, and achieving efficient, energy-saving, and economical collaborative computing effects. Through communication cost modeling and graph matching optimization mechanism, the optimal selection of communication paths is achieved, which significantly reduces task waiting time and resource waste caused by cross-node communication, improves the system's throughput and response speed, and achieves the effect of reducing communication delay and reducing overall costs. Through task priority screening and state migration control mechanism, high availability and elastic recovery capability of task execution are achieved, and the problem of task interruption due to resource recovery or failure in cloud computing environment is solved, thereby enhancing the robustness and stability of the system, and achieving the effect of ensuring task integrity and improving system disaster recovery capability.

[0209] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A distributed data processing method based on cloud computing, characterized by: include: Obtaining heterogeneous resource parameters of computing nodes through the cloud platform interface, including computing power indicators, network topology distance, real-time energy efficiency, and cloud service pricing data, and constructing a multi-dimensional node feature vector; According to the multi-dimensional node feature vector, a dynamic weight distribution model is used to calculate the shard weight value of each node. The weight distribution model introduces a normalization coefficient of energy efficiency and cloud service pricing to divide the data to be processed into data shards proportional to the shard weight value; Based on the network topology distance and cloud service pricing data, a communication cost matrix between nodes is generated, and a graph matching algorithm is used to distribute the data shards to target computing nodes to minimize the total communication cost; Set task priorities, filter critical tasks and non-critical tasks, use cross-availability zone multi-copy storage for critical tasks, use erasure coding for non-critical tasks, and trigger the migration of differential state snapshots to private cloud nodes when cloud platform resource recovery signals are detected; Monitor the actual load and cost consumption of each node, use feedback to adjust the calculation coefficient of the sharding weight value, and dynamically re-shard tasks that exceed the execution threshold time.

2. The distributed data processing method based on cloud computing according to claim 1, wherein: The constructing of a multi-dimensional node feature vector comprises: Obtain the heterogeneous resource parameters of the computing node through the cloud platform interface, where the computing power indicator is expressed as the number of floating-point operations completed per unit time FLOPS, denoted as f i ; The network topology distance is expressed as the network delay between nodes, denoted as d ij , where i and j represent two nodes respectively; Real-time energy efficiency is expressed as the energy consumption value per unit time, denoted as e i ; Cloud service pricing data is expressed as cost per unit time, denoted as c i ; Normalize the above four parameters to obtain the standardized eigenvalues Combine the normalized results into a multidimensional node feature vector V according to the weights i .

3. The distributed data processing method based on cloud computing according to claim 2, wherein: The construction and sharding process of the dynamic weight distribution model includes the following steps: Based on the multidimensional node feature vector V i Construct comprehensive scoring function S i ; The comprehensive score S of all nodes i Perform standardization processing to obtain the standardized comprehensive score of node i Based on the standardized comprehensive score Calculate the shard weight value W of each node i ; Assume that the total amount of data to be processed is T, then according to the shard weight W of each node i , divide the data into several subsets and assign them to corresponding nodes.

4. The distributed data processing method based on cloud computing according to claim 3, wherein: The communication cost matrix generation and graph matching optimization process includes: First, define the time and economic cost incurred when node i sends a unit of data to node j; The communication cost consists of two parts: network transmission time cost and receiving end operation cost; Calculate the element M in the communication cost matrix ij ; For all node pairs (i, j), generate the communication cost matrix M; If node i is not allowed to send data to itself, then M ii Set to infinity or a maximum value to prohibit self-loops. N is the total number of available computing nodes. Consider each data shard as a task node to form a task set, and the set of available computing nodes as a candidate node set. Construct a bipartite graph G = (U ∪ V, E), where U represents the set of task nodes, i.e., data shards, and V represents the set of candidate nodes, i.e., computing nodes. Each edge (u, v) in the edge set E corresponds to the cost of task u∈U assigned to node v∈V, which is M ij ; The Kuhn-Munkres algorithm is used for optimal matching. The Kuhn-Munkres algorithm is used to find a way to allocate tasks to nodes so as to minimize the total communication cost under the premise of a given cost matrix m×n. The input is the communication cost matrix M, and the output is the optimal matching solution; If the number of tasks m is greater than the number of nodes n, multiple rounds of scheduling or the introduction of load balancing strategies are required; If the number of nodes n is greater than the number of tasks m, some nodes will not be selected to avoid resource waste.

5. The distributed data processing method based on cloud computing according to claim 4, characterized in that: The task priority screening and storage strategy implementation process includes: Set the task priority threshold P th , obtain the task priority value P according to the task type or user-specified information i ; If the condition P is met i >P th , then mark the task as a critical task, otherwise mark it as a non-critical task; For data marked as mission-critical, a cross-availability zone multi-copy storage strategy is used: Copy the original task data into multiple completely identical copies, and the number of copies is recorded as R k ; Each replica is stored independently in a different availability zone; The availability zone is an area in the cloud platform that has physical isolation capabilities; When a failure occurs in one availability zone, other replicas can still ensure that tasks continue to execute; For data marked as non-critical, an erasure code storage strategy is used; Use (n,k) erasure code to encode task data; Where n represents the total number of segments after encoding, and k represents the number of data segments into which the original data is divided; Generate nk check fragments using Reed-Solomon coding; All n fragments are stored in different nodes in the cluster; Only any k fragments are needed to restore the original data; Real-time monitoring of whether the cloud platform sends a resource recovery signal; If a resource recycling signal is detected, the differential state snapshot mechanism is triggered, including: Record the execution status and memory snapshot of the current task; A differential snapshot contains only the parts that have changed since the previous snapshot; Compress and package incremental snapshots; Migrate to the private cloud node, and the standby node loads the snapshot and restores the task context; Allows the task to continue from the interruption point to avoid interruption losses.

6. The distributed data processing method based on cloud computing according to claim 5, wherein: The load and cost monitoring feedback includes the following steps: During the task execution process, the actual operating status parameters of each node are collected, including: the number of tasks currently being processed by the node, the maximum number of concurrent tasks of the node, the node's elapsed running time, and the node's unit time cost; Based on the actual operating status parameters, calculate the current load value L of the node i ; At the same time, the cumulative cost consumption TC of the node is calculated i ; Set the system load threshold L th and cost budget ratio threshold δ; The feedback adjustment mechanism is triggered when any of the following conditions are met: L i >L th , that is, the node load is too high; TC i >δ·TC total , that is, the cost of this node exceeds the set proportion of the overall budget; After triggering, recalculate the node's shard weight value W' i ; Updated weight value W′ i Used for the next round of data sharding and scheduling decisions; The feedback mechanism is executed periodically to ensure dynamic balance of cluster resources and controllable costs.

7. The distributed data processing method based on cloud computing according to claim 6, characterized in that: The dynamic resharding mechanism includes the following steps: Set the task execution time threshold T th , preset by the user or the system; Monitor the execution time of each task; When the current execution time of a task is T exec >T th , it is determined to be a timed task; To start dynamic resharding for timed tasks, follow these steps: Analyze the data dependency graph of the task and identify a set of subtasks that can be executed independently; Each subtask has independent input data and output results, records the dependency order between subtasks, and generates a subtask scheduling topology graph; Obtain status information of all currently available nodes, including current load, unit time cost, and network topology distance; Based on the state information, calculate the comprehensive score Q of each candidate node j ; Adopt the minimum priority queue strategy to assign subtasks to the node with the lowest score in sequence; Pack the selected subtask data and send it to the target node; The original node retains the main control logic and coordinates the synchronization and result aggregation between subtasks; Update the global task schedule, record the allocation of subtasks, and generate a migration log, including the task ID, subtask number, source node number, target node number, migration timestamp, and data size.

8. A distributed data processing system based on cloud computing, based on the distributed data processing method based on cloud computing according to any one of claims 1 to 7, characterized in that: include: Resource assessment module, task allocation module, communication optimization module, state migration module and load control module; The resource evaluation module is used to obtain the heterogeneous resource parameters of the computing nodes through the cloud platform interface, including computing power indicators, network topology distance, energy efficiency and cloud service pricing, and normalize the parameters and combine them into a multi-dimensional node feature vector; The task allocation module is used to calculate the shard weight value based on the comprehensive score and standardized score of each node, and divide the data to be processed into multiple subsets according to the weight ratio and allocate them to the corresponding nodes respectively; The communication optimization module is used to construct a communication cost matrix based on network topology distance and unit time cost, and use a graph matching algorithm to optimally match data shards with target nodes to minimize the total communication cost; The state migration module is used to set task priority thresholds and screen critical tasks from non-critical tasks. It uses cross-availability zone multi-copy storage for critical tasks and erasure coding for non-critical tasks. When a resource recycling signal is detected, a differential state snapshot mechanism is triggered to migrate the task state to a private cloud node to ensure execution continuity. The load control module is used to monitor the actual load and cost consumption of each node in real time. When the node load or cost exceeds the preset threshold, it provides feedback to adjust the calculation coefficient of the sharding weight value and dynamically re-shards the timed-out tasks.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the distributed data processing method based on cloud computing according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the distributed data processing method based on cloud computing according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Parallel optimization method and system for high performance computers

    CN122346356A