Strong-adaptation distributed data distribution method supporting dynamic expansion

Through deep reinforcement learning technology, the distributed task allocation is optimized, which solves the problem that traditional strategies are difficult to dynamically adjust in complex dynamic environments, and efficient and accurate data distribution and node resource utilization are achieved, enhancing the system's adaptability and stability.

CN119960991AActive Publication Date: 2025-05-09CHENGDU HAIQING TECH CO LTD

Patent Information

Application Number
CN202510048436.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-09
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

Traditional distributed data distribution strategies are difficult to dynamically adjust when facing complex dynamic environments, resulting in uneven task distribution, node overload or resource waste, and lack the ability to quickly adapt to dynamic changes of nodes.

Method used

Deep reinforcement learning technology is used to optimize distributed task allocation. By initializing the deep reinforcement learning model, periodically detecting node status, dynamically monitoring task characteristics and node load, and designing reward mechanisms, so that the model outputs allocation strategies based on historical experience and current status, and supports dynamic expansion and exit of nodes.

Benefits of technology

It improves the efficiency and accuracy of data distribution, enhances the system's ability to adapt to complex dynamic environments, supports flexible expansion and exit of nodes, and ensures that the system maintains efficient and stable operation during expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960991A_ABST
    Figure CN119960991A_ABST
Patent Text Reader

Abstract

The invention discloses a highly-adaptive distributed data distribution method supporting dynamic extension, which realizes intelligent distribution and efficient processing of data tasks by introducing a task distribution mechanism based on deep reinforcement learning and combining dynamic states of nodes in a distributed system. Specifically, the method comprises the following steps: firstly, collecting real-time state information, including load, bandwidth, delay and the like, of each node in a system, and classifying data tasks according to types and priorities; thirdly, modeling node states and task features through a deep reinforcement learning model, and outputting an optimal task allocation strategy; in the task execution process, the node state is monitored in real time, and the model is dynamically updated through task feedback so as to continuously optimize the distribution strategy. According to the method, distributed task allocation is optimized through a deep reinforcement learning technology, the efficiency and accuracy of data distribution are improved, and meanwhile, the adaptability of the system to a complex dynamic environment is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distributed computing and data processing technology, and in particular to a highly adaptive distributed data distribution strategy that supports dynamic expansion, which aims to optimize the data transmission efficiency and node resource utilization in a distributed system through an intelligent task allocation and resource scheduling mechanism, and is suitable for efficient processing of complex tasks in a large-scale distributed network environment. Background Art

[0002] With the rapid development of information technology, distributed systems have gradually become the core infrastructure for processing large-scale data tasks and have been widely used in cloud computing, edge computing, and the Internet of Things. Such systems can effectively improve computing power and data processing efficiency by completing tasks through the collaboration of multiple nodes. However, with the continuous growth of data scale and complexity, the challenges faced by distributed systems have become increasingly severe. Especially in an environment with huge data transmission volume and dynamic changes in node status, how to achieve efficient data distribution and task allocation has become the focus and difficulty of research.

[0003] Traditional distributed data distribution strategies usually rely on static rules or simple scheduling algorithms, such as load-balancing based polling allocation or random allocation. Although these methods can achieve basic functions under the conditions of fixed nodes and stable task requirements, they show obvious limitations in complex environments. When the node load or network conditions change, static allocation strategies are often difficult to adjust the allocation plan in time, which can easily lead to uneven task distribution, overload of certain nodes, or waste of resources. In addition, these methods lack flexibility when facing diverse task types (such as real-time streaming data and batch data tasks), and it is difficult to dynamically optimize the allocation strategy according to task characteristics and priorities.

[0004] As the scale of distributed systems expands, the dynamic joining and exiting of nodes has also become an important factor affecting the stability of the system. In traditional distributed systems, the joining of new nodes usually requires manual registration and configuration, which prolongs the time it takes for the system to adapt to changes. Node exit may cause task interruption and even affect the normal operation of the entire system. Existing methods lack the ability to quickly adapt to dynamic changes in nodes, especially in scenarios with large task volumes and high real-time requirements. This deficiency will significantly reduce the efficiency and stability of the system.

[0005] In recent years, with the rise of artificial intelligence technology, especially the widespread application of deep learning and reinforcement learning, more and more research has begun to explore intelligent distributed task allocation methods. The allocation model based on deep reinforcement learning can dynamically perceive the characteristics of tasks and node status, and adjust the allocation decision according to the real-time environment. This type of method optimizes itself by continuously accumulating system operation data, which can theoretically greatly improve the adaptability and efficiency of distributed systems. However, the existing task allocation schemes based on deep reinforcement learning still have certain limitations, such as insufficient real-time performance, poor model adaptability, and limited support for complex task scenarios. In addition, these methods have not yet formed a unified and effective solution for the dynamic expansion of nodes and optimal resource utilization.

[0006] In summary, in order to address the shortcomings of traditional distributed data distribution strategies and give full play to the advantages of artificial intelligence technology in task allocation, it is urgent to propose a new data distribution strategy. Summary of the invention

[0007] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a highly adaptive distributed data distribution method that supports dynamic expansion and optimizes distributed task allocation through deep reinforcement learning technology, which can improve the efficiency and accuracy of data distribution and enhance the system's adaptability to complex dynamic environments.

[0008] The object of the present invention is to achieve the following technical solution: a strongly adaptive distributed data distribution method supporting dynamic expansion, comprising the following steps:

[0009] S1. Initialize all nodes of the distributed system, configure the basic architecture of the deep reinforcement learning model, and establish the communication protocol between nodes; configure hardware and network parameters for all distributed nodes to ensure interoperability between nodes; build a deep reinforcement learning model, define input states, output actions, and reward mechanisms;

[0010] S2, receiving task requests from the distributed system, assigning priorities to each task after parsing the task type, and converting task features into the input format of the deep reinforcement learning model; classifying tasks according to their type and priority; converting task features into a data format that can be processed by the deep reinforcement learning model;

[0011] S3. Periodically detect the operating status of all nodes, record the real-time load information of the nodes, and remove nodes with abnormal status or overload; dynamically monitor the CPU utilization, memory usage, and network bandwidth information of the nodes, and block nodes that are operating abnormally or exceed the load threshold;

[0012] S4. Input the task characteristics and node status into the deep reinforcement learning model, design the reward mechanism, make the model output the allocation strategy according to historical experience and current status, and perform reinforcement learning training on the model according to the state-action-reward data;

[0013] S5. According to the allocation decision output by the trained model, the task is transferred to the target node, and the lock-free queue technology is used to optimize the task scheduling;

[0014] S6. Collect task execution results, update the experience pool of the deep reinforcement learning model accordingly, record the action-state-reward data of task assignments, and retrain the model regularly;

[0015] S7, dynamically handle the joining or exit of nodes, and dynamically adjust the resource configuration of nodes according to the system task load; register and synchronize the status of newly joined nodes in real time; reallocate tasks that are not completed due to node exit, and dynamically optimize resource allocation;

[0016] S8. Regularly evaluate the performance of distributed systems and adjust deep reinforcement learning model parameters based on the evaluation results to form best practices for dynamic task allocation.

[0017] The beneficial effects of the present invention are as follows: the distribution strategy of the present invention should be able to dynamically perceive the state changes of nodes and tasks, optimize the distributed task allocation through deep reinforcement learning technology, improve the efficiency and accuracy of data distribution, and enhance the system's adaptability to complex dynamic environments. The present invention also supports the flexible expansion and exit of nodes to ensure that the system can still maintain efficient and stable operation during expansion. The present invention can provide strong technical support for large-scale distributed systems and promote their further application and development in the fields of cloud computing, the Internet of Things, and edge computing. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A flowchart of a highly adaptive distributed data distribution strategy supporting dynamic expansion according to the present invention;

[0019] Figure 2 Schematic diagram of the distributed system architecture in step S1 in an embodiment of the present invention;

[0020] Figure 3 This is a structural diagram of the deep reinforcement learning task allocation model used in step S4 in an embodiment of the present invention. DETAILED DESCRIPTION

[0021] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0022] like Figure 1 As shown, a strongly adaptive distributed data distribution method supporting dynamic expansion of the present invention comprises the following steps:

[0023] S1. Initialize all nodes of the distributed system, configure the basic architecture of the deep reinforcement learning (DQN) model, and establish a communication protocol between nodes; the distributed system includes a central controller and multiple distributed nodes, where the central controller is a high-performance resource management computer equipped with a multi-core CPU, large-capacity memory, and high-performance storage devices, and runs task decomposition and resource management programs. The distributed nodes are multiple sub-computers, each of which is equipped with a multi-core CPU, GPU accelerator, 32GB or more memory, and a high-speed network interface. The central controller can communicate with all distributed nodes through the TCP / IP protocol, and assign tasks to each node through the central controller. The CPU in each node collects the operating status information, and uploads the status information to the central controller, which is then processed by the multi-core CPU of the central controller. Its network structure is as follows: Figure 2 Configure hardware parameters and network parameters for all distributed nodes to ensure interoperability between nodes; build a deep reinforcement learning model, define input states, output actions, and reward mechanisms to optimize task allocation strategies; details are as follows:

[0024] S11. Initialize the system nodes and configure the hardware parameters for all nodes in the distributed system, including but not limited to CPU performance C i 、Memory capacity M i and network bandwidth B i ; Where i represents the i-th node; configure the communication protocol between nodes to ensure low latency and high reliability of data transmission, use the Transmission Control Protocol (TCP) or User Datagram Protocol (UDP) to achieve communication between nodes and ensure the stability of the system;

[0025] S12. Define the input state vector S of the deep Q network (DQN), including task feature information (such as task size T size , Task priority T priority ) and node status information (such as node load ratio L i , Remaining bandwidth ratio B i / B max , the remaining computing power is C i / C max ); define action set A = {a1, a2, ..., a N}, a i represents the action of assigning tasks to the i-th node, N is the total number of nodes; the reward function R(s,a) is designed as follows:

[0026] R(s,a)=α·E-β·L-γ·D

[0027] Among them, E represents the task completion rate, L represents the node load balancing index, D represents the communication delay of task allocation, and α, β, and γ are weight coefficients used to balance the importance of different goals;

[0028] S13. Determine the unique identifier (ID) of the distributed node and establish a routing table between nodes to describe the logical topological relationship between nodes; configure an efficient data transmission interface and define the data packet format, including the task identifier, target node ID, task type and data content, to ensure the accuracy and completeness of distributed task allocation.

[0029] S14. Based on the CPU performance and network bandwidth of the node, define the initial load threshold L for each node th,i :

[0030] L th,i =δ c ·C i +δ b B i

[0031] Among them, δ c and δ b The weight coefficient of the node's CPU performance and network bandwidth ensures the rationality of load distribution;

[0032] S15. Configure a distributed monitoring module in the system to regularly collect the running status information of the nodes and upload it to the central controller through a unified data collection interface. Determine the status collection frequency and node status update cycle to balance the monitoring overhead and system real-time performance.

[0033] S2, receiving task requests from the distributed system, assigning priorities to each task after parsing the task type, and converting task features into the input format of the deep reinforcement learning model; classifying tasks according to task type (real-time data stream, batch tasks, etc.) and priority; converting task features (such as task size, completion time limit) into a data format that can be processed by the deep reinforcement learning model. The details are as follows:

[0034] S21. After receiving the task request, the system analyzes the basic attributes of the task, including the task type, task size, and computing requirements. Task types are divided into real-time streaming tasks, batch tasks, and delay-tolerant tasks, which are marked as high, medium, and low priorities, respectively. The priority of real-time streaming tasks is marked as 1, batch tasks are marked as 2, and delay-tolerant tasks are marked as 3. The task size T size Calculate the required T in bytes compute It is quantified in floating point operations (FLOPs) to describe the computational intensity required to execute a task.

[0035] S22. Based on the parsed task attributes, construct a task feature vector T = [T size ,T compute ,T priority ], where T size is the data size of the task, T compute To calculate the demand, T priority is the task priority; the task feature vector is used to represent the key features of the task to support subsequent allocation decisions.

[0036] S23. Standardize the task feature vector to ensure that the value range of all features is consistent; the standardization formula is as follows:

[0037]

[0038] Among them, T min and T max They represent the minimum and maximum values ​​of the feature in the historical task data respectively; the standardized task feature vector T′ is mapped to the interval [0,1] to facilitate deep learning model processing and optimized calculation.

[0039] S24, the standardized task feature vector T′ system current node state N state Combined to form a complete input state vector S = [T ′, N state ]; where the node status information includes the node load ratio L i , Remaining bandwidth ratio B i / B max And the remaining computing power is C i / C msx , to reflect the current resource usage of each node in the system.

[0040] S25. By integrating the task features and node states, the input state vector S for the deep reinforcement learning model is finally generated. The input state vector not only contains the characteristic description of the task, but also combines the resource status of the current node, providing a comprehensive information basis for the model and supporting the output of efficient task allocation decisions.

[0041] S3. Periodically detect the operating status of all nodes, record the real-time load information of the nodes, and remove nodes with abnormal status or overload; dynamically monitor the CPU utilization, memory usage, network bandwidth information and other status information of the nodes, and shield the nodes with abnormal operation or exceeding the load threshold to ensure the reliability of task allocation. The details are as follows:

[0042] S31, periodically collect the operating status information of all nodes through the distributed monitoring module, and the monitoring content includes the CPU utilization U of the i-th node CPU,i 、Memory usage U MEM,i, Network bandwidth usage U BW,i And the task queue length Q i ; Through the predefined sampling period Δt, it ensures that the status information can be updated in time and reflect the real-time operation of the system.

[0043] S32. Calculate the node's comprehensive load index L' based on the collected status information i :

[0044] L′ i =δ·U CPU,i +ε·U MEM,i +∈·U BW,i

[0045] Among them, δ, ε and ∈ are load weight parameters, which are used to adjust the contribution ratio of different resource types to the comprehensive load. Comprehensive load L′ i The higher the value, the more serious the current resource usage of the node.

[0046] S33, calculate the comprehensive load index L' i The load threshold L set by the system th For comparison: If L′ i >L th , the node is determined to be overloaded and marked as abnormal; otherwise, the node is marked as normal; nodes in normal state will continue to be included in the candidate range for task allocation.

[0047] S34: For abnormal nodes, remove them from the candidate list for task allocation and record abnormal status information E i , so as to facilitate subsequent system alarm or troubleshooting. At the same time, by reallocating unfinished tasks to other normal nodes, task processing interruptions caused by abnormal nodes can be avoided.

[0048] S35, upload the updated node operation status information to the central controller for unified storage and synchronization, ensuring that all modules in the distributed system can share the latest node status in real time. Through the synchronization mechanism, other modules can make dynamic adjustments based on the latest status information to provide support for subsequent task allocation.

[0049] S4. Input the task characteristics and node status into the deep reinforcement learning model, design a reward mechanism, and make the model output the allocation strategy according to the historical experience and the current status to train the DQN model; the reward mechanism includes positive rewards such as high task completion efficiency, load balancing, and high system resource utilization, and the model is reinforced according to the state-action-reward data. The details are as follows:

[0050] S41, collect the characteristic vector of the task T′=[T size ,T compute,T priority ] and the node's state vector N state =[L i ,C i / C max ,B i / B max ], and combine them into the input state vector S of the deep reinforcement learning model =

[0051] [T′,N state ]; where T size is the task data volume, T compute is the amount of computing resources required for the task, T priority is the task priority, L i , C i / C max , B i / B max They represent the node's load ratio, remaining computing power ratio, and remaining bandwidth ratio respectively.

[0052] S42. Construct a reinforcement learning model based on deep reinforcement learning (DQN), where the state vector S is used as the input of the model, and the action set A = {a1, a2, ..., a N} is the output, indicating the selection decision of task allocation to a certain node; Figure 3 As shown in the figure, the DQN network includes three hidden layers, where the first hidden layer is a fully connected layer containing 128 nodes and a ReLU layer, the second hidden layer is a fully connected layer containing 64 nodes and a ReLU layer, and the third hidden layer is a fully connected layer containing N nodes. i Corresponds to assigning the task to the i-th node.

[0053] S43. The design goal of the reward function R(s,a) is to balance task completion efficiency, load balancing, and delay minimization. The specific formula is as follows:

[0054]

[0055] Among them, T complete is the task completion time, L i is the load ratio of the target node; D is the communication delay of task transmission, α,

[0056] β and γ are weight parameters, which represent the degree of concern for efficiency, load balancing, and latency, respectively.

[0057] S44, input the input state vector S into the deep Q network model, and obtain each action a through the forward calculation of the model iThe Q value of the task is selected, and the action with the largest Q value is selected as the allocation decision of the current task; the experience pool of the model is updated using the task execution results and the reward function, and the state-action-reward data (S, a, R) is recorded; based on the experience pool data, the model parameters are optimized through back propagation to ensure that the model gradually learns the optimal task allocation strategy;

[0058] S45. During the training process, the model input and experience pool are updated in real time based on new tasks and node status to ensure the timeliness and adaptability of the strategy. After the model is trained, the optimal allocation action a is selected for the current task. i * , assign the task to the corresponding target node.

[0059] S5. Based on the allocation decision output by the trained model, the task is transferred to the target node, and the lock-free queue technology is used to optimize task scheduling, avoid thread conflicts, and improve execution efficiency. The model decision is used to select the optimal target node; based on the lock-free queue technology, the efficient transmission and concurrent processing of tasks are guaranteed.

[0060] S5 is as follows:

[0061] S51, assign action a according to the output of the deep reinforcement learning model i * , assign the task to the target node N i ; Wherein, the target node N i The highest allocation decision score (Q value) is obtained through model calculation. The system prioritizes nodes with sufficient resources and low latency based on the priority of the task and the status information of the node to ensure efficient execution of the task.

[0062] S52, after assigning the task, the system transmits the task data to the target node N through the network communication module i ; During the transmission process, the optimal routing algorithm is used to select the transmission path to minimize network delay D trans At the same time, the traffic monitoring module detects the transmission status in real time to ensure the integrity and accuracy of the task data.

[0063] S53. In the target node, a lock-free queue is used to schedule and manage tasks. The lock-free queue design uses atomic operations to ensure concurrency safety in a multi-threaded environment and effectively avoid conflicts and delays caused by thread competition. Specifically, tasks are scheduled and managed according to priority T. priority Sort by tasks, with high priority tasks entering the head of the queue for processing.

[0064] S54, after receiving the task data, the target node calculates the task data according to the task characteristics (such as the computing requirements Tcompute ) and its own resource conditions (such as remaining computing capacity C i ) allocates corresponding computing resources and storage resources; the node selects the appropriate execution mode according to the task type: for real-time tasks, the quick response mode is adopted; for batch tasks, the batch processing mode is adopted; during task execution, the resource usage is monitored in real time; during task execution, the resource usage is monitored in real time to avoid resource overload.

[0065] S55. During the task execution process, the target node regularly generates progress feedback information of the task execution, including the current amount of processed data, the estimated completion time T remain The feedback information is uploaded to the central controller through the communication module for subsequent task monitoring and scheduling optimization.

[0066] S56. After the task is completed, the target node uploads the final execution result to the central controller and records the completion time T complete and resource usage R used ; If an error occurs during task execution, the node will generate an error log and notify the central controller, triggering the reallocation mechanism to reallocate the unfinished tasks to other candidate nodes.

[0067] S6. Collect task execution results, update the experience pool of the deep reinforcement learning model accordingly, record the action-state-reward data of task assignment, and retrain the model regularly; collect data such as task completion time and node status feedback; retrain the DQN model regularly and continuously optimize the task assignment strategy. The details are as follows:

[0068] S61. After the task is completed, the target node sends the task execution result and completion time T complete , resource usage R used The system also records whether the task is completed successfully or whether there are any abnormal situations (such as timeout, interruption), providing basic data for subsequent model adjustment and optimization.

[0069] S62, compare the task execution result with the corresponding input state vector S = [T′, N state ], assign action a i , the reward value R(s,a) is stored in the experience pool of the deep reinforcement learning model to form state-action-reward data (S,a,R);

[0070] S63. The system manages the data in the experience pool. When the storage capacity of the experience pool reaches the upper limit, the system adopts a priority elimination strategy to remove older or low-value data entries to ensure that new data can be updated to the experience pool in a timely manner. Prioritize the retention of samples that are valuable for model learning, such as data records of highly complex tasks or abnormal tasks, to enhance the adaptability of the model.

[0071] S64. Regularly extract batch samples ((S, a, R) from the experience pool to retrain the deep reinforcement learning model; optimize the model parameters through the back propagation algorithm and update the model's state-action mapping strategy;

[0072] S65. After retraining is completed, the new model strategy is deployed to the system, and the effectiveness of the strategy is verified through real-time monitoring, including the efficiency of task allocation, the utilization of system resources, and the load balancing of nodes; if the new strategy causes performance degradation, the system will roll back to the previous version of the strategy to ensure the stability and efficiency of the distributed system.

[0073] S66. Regularly evaluate the performance of the deep reinforcement learning model, using task completion rate, average completion time T avg , node load balancing coefficient B balance Quantify the model effect; if the evaluation results show that the model does not meet the expected performance standards, adjust the weight parameters α, β, γ or other hyperparameters of the reward function to further optimize model training.

[0074] S7, dynamically handle the joining or exit of nodes, and dynamically adjust the resource configuration of nodes according to the system task load; register and synchronize the status of newly joined nodes in real time; reallocate tasks that are not completed due to node exit, and dynamically optimize resource allocation; the details are as follows:

[0075] S71, when the new node T avg When joining a distributed system, the system assigns a unique identifier (ID) to it through a registration mechanism and records its hardware parameters (such as computing power C new 、Memory capacity M new 、Network bandwidth B new ) and initial state (such as load ratio ); Subsequently, the system synchronizes the routing information of the new node with other nodes and updates the global node topology to ensure that the new node can participate in task allocation in a timely manner.

[0076] S72. For the node N that joins new,The system regularly collects its status information through the monitoring module, including CPU utilization, memory occupancy and network bandwidth utilization. The collected status information is uploaded to the ,central controller through the communication module as the basis for subsequent ,allocation decisions and load balancing. At the same time, the addition of new nodes will trigger ,the update of the global node status table to ensure that all nodes can share the ,latest status information.

[0077] S73. To adapt to the addition of new nodes, the system recalculates the task allocation strategy and includes the new nodes in the task candidate allocation range; through real-time decision-making of the deep reinforcement learning model, new nodes are preferentially assigned tasks that are suitable for their resource characteristics, such as allocating low computing demand tasks to new nodes with fewer resources, thereby avoiding resource waste.

[0078] S74, when node N exit When exiting due to a failure or planned maintenance, the system detects an abnormal state (such as communication interruption or resource unavailability) through the anomaly detection mechanism and marks it as "exit state"; then, the system stops assigning new tasks to the exiting node and transfers its unfinished tasks to other normal nodes N alt ,The transfer prioritizes nodes with lower load and sufficient resources to minimize the ,interruption time of task execution.

[0079] S75. For the exit node, the system records its exit time T exit , the amount of completed task data, and the cause of the failure (such as hardware failure or network interruption). These records are stored in the system log for subsequent system maintenance and fault analysis.

[0080] S76, after a node joins or leaves, the system re-evaluates the comprehensive load L of each node i , update the global load balancing strategy; by adjusting the load threshold L th,i ,Dynamically optimize task allocation to ensure efficient resource utilization and balanced task distribution when the number of nodes changes;

[0081] S77. After the node status changes, the system evaluates the expansion performance through an observation period; the evaluation indicators include task completion rate, node resource utilization, average task delay and load balancing coefficient; if it is found that the performance does not meet the expected goals, the system will adjust the task allocation strategy or optimize the model parameters to further improve the adaptability of dynamic expansion.

[0082] S8. Regularly evaluate the performance of distributed systems and adjust deep reinforcement learning model parameters based on the evaluation results to form best practices for dynamic task allocation and further improve system efficiency; optimize model parameters based on system operation data (throughput, latency, task success rate, etc.); improve the global performance of task allocation to ensure that the system adapts to dynamically changing needs.

[0083] S8 is as follows:

[0084] S81. The system continuously collects key performance indicator data during operation, including task completion rate R complete , average completion time T avg , node resource utilization U i And the load balancing factor B balance ; Among them, the task completion rate is defined as the ratio of the actual number of completed tasks to the total number of tasks, the average completion time represents the average time taken for a task to be completed from receipt to completion, the node resource utilization rate is the average occupancy ratio of the computing power, memory and bandwidth of each node, and the load balancing coefficient reflects the uniformity of the load between the nodes in the system.

[0085] S82. Calculate the system's comprehensive performance index P based on the collected performance data. sys , the specific formula is:

[0086] P sys =α·R complete -β·T avg -γ·(1-B balance )

[0087] Among them, α, β, and γ are weight coefficients, which respectively represent the importance of task completion rate, task delay, and load balancing. By evaluating the comprehensive performance indicators, it is judged whether the current operating status of the system meets the expected goals;

[0088] S83. If the evaluation results show that the system performance is lower than the target value, the system adjusts the parameters of the deep reinforcement learning model according to the actual operation; for example, by optimizing the weight parameters α, β, and γ in the reward function, the model's attention to specific performance goals (such as latency or load balancing) is enhanced. At the same time, the weights of the input features can be adjusted or the model architecture can be redefined according to the latest task distribution and node status to improve the prediction accuracy and allocation efficiency of the model.

[0089] S84. Apply the optimized strategy to the test task set to verify its improvement effect on system performance; record the performance data differences before and after optimization, including task delay, resource utilization, and task completion rate; if the test results show significant performance improvement, deploy the new strategy to the actual environment; otherwise, continue to adjust the model or strategy parameters until the expected goals are met;

[0090] S85. Based on the optimized model and strategy, the system makes global adjustments to the current task allocation scheme and resource allocation mechanism to ensure the stability of system performance in a dynamically changing environment. At the same time, the parameter adjustments involved in the evaluation and optimization process are recorded in the system log to provide a basis for subsequent improvement and maintenance.

[0091] S86. Combine historical performance data to analyze the long-term operating trends of the system and identify possible performance bottlenecks or room for improvement. For example, by analyzing the distribution characteristics of tasks and changes in node resources in different time periods, predict possible load peaks or resource shortages in the future, and take countermeasures in advance, such as adding nodes or optimizing communication protocols.

[0092] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A strongly adaptive distributed data distribution method supporting dynamic expansion, characterized in that: The following steps are involved: S1. Initialize all nodes of the distributed system, configure the basic architecture of the deep reinforcement learning model, and establish the communication protocol between nodes; configure hardware and network parameters for all distributed nodes to ensure interoperability between nodes; build a deep reinforcement learning model, define input states, output actions, and reward mechanisms; S2, receiving task requests from the distributed system, assigning priorities to each task after parsing the task type, and converting task features into the input format of the deep reinforcement learning model; classifying tasks according to their type and priority; converting task features into a data format that can be processed by the deep reinforcement learning model; S3. Periodically detect the operating status of all nodes, record the real-time load information of the nodes, and remove nodes with abnormal status or overload; dynamically monitor the CPU utilization, memory usage, and network bandwidth information of the nodes, and block nodes that are operating abnormally or exceed the load threshold; S4. Input the task characteristics and node status into the deep reinforcement learning model, design the reward mechanism, make the model output the allocation strategy according to historical experience and current status, and perform reinforcement learning training on the model according to the state-action-reward data; S5. According to the allocation decision output by the trained model, the task is transferred to the target node, and the lock-free queue technology is used to optimize the task scheduling; S6. Collect task execution results, update the experience pool of the deep reinforcement learning model accordingly, record the action-state-reward data of task assignments, and retrain the model regularly; S7, dynamically handle the joining or exit of nodes, and dynamically adjust the resource configuration of nodes according to the system task load; register and synchronize the status of newly joined nodes in real time; reallocate tasks that are not completed due to node exit, and dynamically optimize resource allocation; S8. Regularly evaluate the performance of distributed systems and adjust deep reinforcement learning model parameters based on the evaluation results to form best practices for dynamic task allocation.

2. A strongly adaptive distributed data distribution method supporting dynamic expansion according to claim 1, characterized in that: The step S1 is specifically as follows: S11, initialize the system nodes and configure the hardware parameters for all nodes in the distributed system, including CPU performance C i 、Memory capacity M i and network bandwidth B i ; Where i represents the i-th node; the communication between nodes is realized by using the transmission control protocol or the user datagram protocol; S12. Define the input state vector S of the deep Q network, including task feature information and node state information; define the action set A = {a1, a2, ..., a N }, a i represents the action of assigning tasks to the i-th node; the reward function R(s,a) is designed as follows: R(s,a)=α·E-β·L-γ·D Among them, E represents the task completion rate, L represents the node load balancing index, D represents the communication delay of task allocation, and α, β, and γ are weight coefficients used to balance the importance of different goals; S13, determine the unique identifier of the distributed node and establish a routing table between nodes; configure an efficient data transmission interface and define the data packet format, including the task identifier, target node ID, task type and data content; S14. Based on the CPU performance and network bandwidth of the node, define the initial load threshold L for each node th,i : L th,i =d c ·C i +d b ·B i Among them, δ c and δ b is the weight coefficient of the node's CPU performance and network bandwidth; S15. Configure a distributed monitoring module to regularly collect the operating status information of the nodes and upload it to the central controller through a unified data collection interface.

3. A strongly adaptive distributed data distribution method supporting dynamic expansion according to claim 1, characterized in that: The step S2 is specifically as follows: S21. After receiving the task request, the system analyzes the basic attributes of the task, including the task type, task size, and computing requirements. Task types are divided into real-time streaming tasks, batch tasks, and delay-tolerant tasks, which are marked as high, medium, and low priorities, respectively. The priority of real-time streaming tasks is marked as 1, batch tasks are marked as 2, and delay-tolerant tasks are marked as 3. The task size T size Calculate the required T in bytes compute Quantization is performed in units of floating-point operands; S22. Based on the parsed task attributes, construct a task feature vector T = [T size ,T compute ,T priority ], where T siz is the data size of the task, T compute To calculate the demand, T priority Prioritize the tasks; S23. Standardize the task feature vector to ensure that the value range of all features is consistent; the standardization formula is as follows: Among them, T min and T max Respectively represent the minimum and maximum values ​​of the feature in the historical task data; S24, the standardized task feature vector T′ system current node state N state Combined to form a complete input state vector S = [T ′, N state ]; where the node status information includes the node load ratio L i , Remaining bandwidth ratio B i / B mac And the remaining computing power is C i / C msx ; S25. By integrating task features and node states, the input state vector S for the deep reinforcement learning model is finally generated.

4. According to the method for distributing data with strong adaptability supporting dynamic expansion according to claim 1, the step S3 is specifically as follows: S31, periodically collect the operating status information of all nodes through the distributed monitoring module, and the monitoring content includes the CPU utilization U of the i-th node CPU,i 、Memory usage U MEM,i , Network bandwidth usage U BW,i And the task queue length Q i ; S32. Calculate the node's comprehensive load index L' based on the collected status information i : L′ i =δ·U CPU,i +ε·U MEM,i +∈·U BW,i in, δ, ε, and ∈ are load weight parameters; S33, calculate the comprehensive load index L' i The load threshold L set by the system th For comparison: If L′ i >L th , the node is determined to be overloaded and marked as abnormal; otherwise, the node is marked as normal; S34: For abnormal nodes, remove them from the candidate list for task allocation and record abnormal status information E i ; S35. Upload the updated node operation status information to the central controller for unified storage and synchronization, to ensure that all modules in the distributed system can share the latest node status in real time.

5. The method for distributing data with strong adaptability and supporting dynamic expansion according to claim 3 is characterized in that: The step S4 is specifically as follows: S41, collect the characteristic vector of the task T′=[T size ,T compute ,T priority ] and the node's state vector N state =[L i ,C i / C max ,B i / B max ], and combine them into the input state vector S = [T′, N state ]; S42. Construct a model based on deep reinforcement learning, where the state vector S is used as the input of the model, and the action set A = {a1, a2, ..., a N } is the output, which represents the selection decision of task assignment to a certain node. Each action a in the action set corresponds to assigning the task to the i-th node; S43. The design goal of the reward function R(s,a) is to balance task completion efficiency, load balancing, and delay minimization. The specific formula is as follows: Among them, T complete is the task completion time, L i is the load ratio of the target node; S44, input the input state vector S into the deep Q network model, and obtain each action a through the forward calculation of the model i The Q value of the task is selected, and the action with the largest Q value is selected as the allocation decision of the current task; the experience pool of the model is updated using the task execution results and the reward function, and the state-action-reward data (S, a, R) is recorded; based on the experience pool data, the model parameters are optimized through back propagation to ensure that the model gradually learns the optimal task allocation strategy; S45. During the training process, the model input and experience pool are updated in real time in combination with new tasks and node states; after the model is trained, the optimal allocation action a is selected for the current task. i * , assign the task to the corresponding target node.

6. A strongly adaptive distributed data distribution method supporting dynamic expansion according to claim 5, characterized in that: The step S5 is specifically as follows: S51, assign action a according to the output of the deep reinforcement learning model i * , assign the task to the target node N i ; S52, after assigning the task, the system transmits the task data to the target node N through the network communication module i ; The optimal routing algorithm is used to select the transmission path during the transmission process; S53. Inside the target node, a lock-free queue is used to schedule and manage tasks. Tasks are scheduled and managed according to their priority level T priority Sort by tasks with high priority, and place them at the head of the queue for processing; S54, after receiving the task data, the target node allocates corresponding computing resources and storage resources according to the task characteristics and its own resource conditions; selects the appropriate execution mode according to the task type: for real-time tasks, adopts the rapid response mode; for batch tasks, adopts the batch processing mode; monitors the resource usage in real time during the task execution; S55. During the task execution process, the target node regularly generates progress feedback information of the task execution, including the current amount of processed data, the estimated completion time T remain and resource usage status, and the feedback information is uploaded to the central controller through the communication module; S56. After the task is completed, the target node uploads the final execution result to the central controller and records the completion time T complete and resource usage R used ; If an error occurs during task execution, the node will generate an error log and notify the central controller, triggering the reallocation mechanism to reallocate the unfinished tasks to other candidate nodes.

7. A strongly adaptive distributed data distribution method supporting dynamic expansion according to claim 1, characterized in that: The step S6 is specifically as follows: S61. After the task is completed, the target node sends the task execution result and completion time T complete , resource usage R use The running status information is uploaded to the central controller; at the same time, it records whether the task is successfully completed or whether there are any abnormal conditions; S62, compare the task execution result with the corresponding input state vector S = [T', N state ], assign action a i , the reward value R(s,a) is stored in the experience pool of the deep reinforcement learning model to form state-action-reward data (S,a,R); S63, the system manages the data in the experience pool. When the storage capacity of the experience pool reaches the upper limit, a priority elimination strategy is adopted to remove older or low-value data entries. S64. Regularly extract batch samples ((S, a, R) from the experience pool to retrain the deep reinforcement learning model; optimize the model parameters through the back propagation algorithm and update the model's state-action mapping strategy; S65. After retraining is completed, the new model strategy is deployed to the system, and the effect of the strategy is verified through real-time monitoring. If the new strategy causes performance degradation, the system will roll back to the previous version of the strategy; S66. Regularly evaluate the performance of the deep reinforcement learning model. If the evaluation results show that the model does not meet the expected performance standards, adjust the parameters to further optimize the model training.

8. The method for distributing data with strong adaptability and supporting dynamic expansion according to claim 1, characterized in that: The step S7 is specifically as follows: S71, when the new node T avg When joining a distributed system, the system assigns it a unique identifier through a registration mechanism and records its hardware parameters and initial status; Afterwards, the system synchronizes the routing information of the new node with other nodes and updates the global node topology; S72. For the node N that joins new ,The system regularly collects its status information through the monitoring module, including CPU utilization, memory occupancy and network bandwidth utilization, and the collected status information is uploaded to the central controller through the communication module; ,At the same time, the addition of new nodes will trigger the update of the global node ,status table; S73, recalculate the task allocation strategy and include the new node in the task candidate allocation range; S74, when node N exit When a node needs to exit due to a failure or planned maintenance, the system detects that its status is abnormal through the anomaly detection mechanism and marks it as "exit status". Subsequently, the system stops assigning new tasks to the exiting node and transfers its unfinished tasks to other normal nodes N. alt ,The transfer prioritizes nodes with lower load and sufficient resources; S75. For the exit node, the system records its exit time T exit , the amount of completed task data and the cause of the failure; S76, after a node joins or leaves, the system re-evaluates the comprehensive load L of each node i , update the global load balancing strategy; by adjusting the load threshold L th,i ,dynamically optimize task allocation to ensure efficient resource utilization and balanced task distribution when the number of nodes changes; S77. After the node status changes, the system evaluates the expansion performance through an observation period; the evaluation indicators include task completion rate, node resource utilization, average task delay and load balancing coefficient; if it is found that the performance does not meet the expected goals, adjust the task allocation strategy or optimize the model parameters.

9. The method for distributing data with strong adaptability and supporting dynamic expansion according to claim 1, characterized in that: The step S8 is specifically as follows: S81. The system continuously collects key performance indicator data during operation, including task completion rate R complete , average completion time T avg , node resource utilization U i And the load balancing factor B balance ; S82. Calculate the system's comprehensive performance index P based on the collected performance data. sys , the specific formula is: P sys =α·R complete -β·T avg -γ·(1-B balance ) Among them, α, β, and γ are weight coefficients, which respectively represent the importance of task completion rate, task delay, and load balancing. By evaluating the comprehensive performance indicators, it is judged whether the current operating status of the system meets the expected goals; S83. If the evaluation result shows that the system performance is lower than the target value, the system adjusts the parameters of the deep reinforcement learning model according to the actual operation situation; S84. Apply the optimized strategy to the test task set to verify its improvement effect on system performance; record the performance data differences before and after optimization, including task delay, resource utilization, and task completion rate; if the test results show significant performance improvement, deploy the new strategy to the actual environment; otherwise, continue to adjust the model or strategy parameters until the expected goals are met; S85. Based on the optimized model and strategy, globally adjust the current task allocation plan and resource allocation mechanism; S86. Analyze the long-term operating trends of the system based on historical performance data.

Citation Information

Patent Citations

  • Calculation power distribution network online scheduling method based on dual deep reinforcement learning

    CN117834624A

  • Industrial network optimization system based on edge computing

    CN118474100A

  • Cloud computing platform real-time CPU load balancing method and system based on reinforcement learning

    CN118885291A

  • Dynamic load balancing method and system of AI model intelligent machine

    CN119201435A

  • Scheduling automation system application state management method

    CN119292745A

Cited By

  • Classification model training method based on big data distributed computing

    CN120234158A

  • A classification model training method based on big data distributed computing

    CN120234158B

  • Large language model distributed training method and system based on dynamic resource scheduling

    CN120278283A

  • Large language model distributed training method and system based on dynamic resource scheduling

    CN120278283B

  • Operation management method and system of server cluster

    CN120416256A