A cooperative data migration scheduling method based on topology awareness and deep reinforcement learning
By employing a collaborative data migration scheduling method combining topology awareness and deep reinforcement learning, the problems of uneven task allocation and resource conflicts in dynamic environments during data migration are solved, achieving an efficient and stable data migration process.
Patent Information
- Application Number
- CN202510975359.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-15
AI Technical Summary
Existing data migration methods struggle to adapt to dynamic network environments, leading to uneven task allocation, resource overruns, and path congestion, which negatively impact migration efficiency and service quality.
A collaborative data migration scheduling method based on topology awareness and deep reinforcement learning is adopted. Node and path representation vectors are generated through graph neural networks, and dynamic scheduling is performed by combining deep reinforcement learning agents to optimize data block selection, worker selection and transmission path. The path information is executed by a software-defined network controller.
It improved migration efficiency, ensured business stability and resource controllability, enhanced network adaptability and path optimization, and achieved the integrity and non-duplication of data migration.
Smart Images

Figure CN120469982B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data migration technology, and in particular to a collaborative data migration scheduling method based on topology awareness and deep reinforcement learning. Background Technology
[0002] In modern data center operations and large-scale data management, data migration serves as a core supporting technology for scenarios such as storage system upgrades, disaster recovery synchronization, and data distribution across computing clusters. It is widely applied in key areas such as cloud computing, distributed storage, and edge computing. Related technologies typically construct data migration systems through the synergy of multi-machine parallel operations, network path planning, and resource scheduling strategies. Specifically, this system covers the entire process from source data retrieval, data sharding, task allocation, network path selection, to target storage writing, including key aspects such as storage performance monitoring, network topology discovery, task scheduling algorithms, and resource usage threshold setting. With the expansion of data center scale and the increase in business complexity, traditional migration methods are gradually revealing limitations in dynamic environment adaptability, resource control precision, and path optimization capabilities, necessitating the introduction of more intelligent scheduling mechanisms to improve overall migration efficiency and system stability.
[0003] However, existing data migration methods often employ static task allocation strategies or path selection mechanisms based on single metrics, failing to fully integrate network topology and real-time resource status. This can lead to uneven task allocation, resource overruns, and path congestion. Furthermore, the lack of global awareness prevents multi-objective optimization, impacting migration efficiency and service quality. Specifically, traditional methods struggle to adapt to changes in network topology, fluctuations in worker load, and differences in storage response capabilities, resulting in decreased migration performance or system resource conflicts. Moreover, most systems lack a unified scheduling and resource guarantee framework; task allocation, path selection, and rate control modules operate independently, hindering collaborative optimization and further limiting the intelligence and practical effectiveness of migration systems. Summary of the Invention
[0004] This application aims to at least partially address one of the technical problems in the related art.
[0005] Therefore, the first objective of this application is to propose a collaborative data migration scheduling method based on topology awareness and deep reinforcement learning.
[0006] The second objective of this application is to propose a collaborative data migration scheduling device based on topology awareness and deep reinforcement learning.
[0007] The third objective of this application is to propose an electronic device.
[0008] The fourth objective of this application is to provide a computer-readable storage medium.
[0009] The fifth objective of this application is to provide a computer program product.
[0010] To achieve the above objectives, the first aspect of this application proposes a cooperative data migration scheduling method based on topology awareness and deep reinforcement learning, comprising:
[0011] In response to the start request of the data migration task, initialize the system configuration, including source data information, target storage information, worker cluster information and business protection policies;
[0012] The network topology discovery module constructs a weighted graph containing source servers, worker machines, target storage, and network devices, and calculates multiple candidate network paths based on the weighted graph.
[0013] The weighted graph is embedded using a graph neural network (GNN) processing module to generate node embedding vectors, and path representation vectors are generated by aggregating the path node sequences. Based on the node embedding vectors, path representation vectors, source data states, worker states, and target storage states, a system state vector is generated by fusing them together.
[0014] The system state vector is input into the deep reinforcement learning DRL agent module, and the DRL agent module outputs a composite action based on the system state vector. The composite action includes data block selection, worker selection, source-to-worker path index, worker-to-target path index, and migration rate suggestion.
[0015] When the path information in the composite action is determined, the path information is converted into flow rules by the network control module and then sent to the switch devices on the path for execution by the software-defined network controller.
[0016] Optionally, the network topology discovery module integrates with the network management system to obtain static topology information and dynamic link characteristics of network devices. The static topology information includes the physical connection relationships between nodes and the rated bandwidth of the links, and the dynamic link characteristics include the current load and fault status of the links.
[0017] Optionally, the graph neural network processing module employs a graph attention network, which uses an attention mechanism to weighted aggregate the features of nodes and links to generate low-dimensional embedding vectors to characterize the network topology.
[0018] Optionally, the path representation vector is generated by treating the path as a sequence of nodes and aggregating the embedding vectors of key nodes in the path. The aggregation process includes average pooling or sequence processing operations.
[0019] Optionally, the DRL agent module uses a proximal policy optimization algorithm for training and decision-making, which improves the stability of the learning process by limiting the policy update magnitude.
[0020] Optionally, the reward function of the DRL agent module includes at least one of the following:
[0021] The total amount of data successfully migrated within the period;
[0022] The number of data blocks that were newly migrated within the period;
[0023] Penalty value for exceeding the preset threshold for worker machine resource usage;
[0024] The penalty value for the selected network path causing link congestion or high packet loss;
[0025] A time step penalty value set to encourage tasks to be completed as quickly as possible.
[0026] Optionally, the penalty values in the reward function are positively correlated with the degree of resource overrun, and a non-linear penalty function is used to enhance the suppression effect on resource overrun.
[0027] Optionally, it also includes: during the migration process, continuously monitoring the migration status of each data block through the task management module, and when a data block migration failure is detected, adding the data block back to the list to be assigned, so as to ensure the integrity and non-duplication of the migration task.
[0028] To achieve the above objectives, a second aspect of this application proposes a cooperative data migration scheduling device based on topology awareness and deep reinforcement learning, comprising:
[0029] The data acquisition module is used to acquire source data information, target storage information, worker cluster information, and business protection strategies.
[0030] The network topology discovery module is used to construct a weighted graph containing source servers, worker machines, target storage, and network devices, and output candidate network paths;
[0031] The graph neural network (GNN) processing module is used to embed the weighted graph and generate node embedding vectors and path representation vectors.
[0032] The state fusion module is used to fuse source data state, worker state, target storage state and path representation vector to generate system state vector;
[0033] The Deep Reinforcement Learning (DRL) agent module is used to make decisions based on the system state vector and outputs a composite action including data block selection, worker selection, source-to-worker path index, worker-to-target path index, and migration rate suggestion.
[0034] The network control module is used to convert the path information in the composite action into flow rules, and to send them to the switch devices on the path for execution through the software-defined network controller;
[0035] The task management module is used to record and update the migration status of each data block to ensure the integrity and non-duplication of migration tasks.
[0036] To achieve the above objectives, a third aspect of this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0037] The memory stores computer-executed instructions;
[0038] The processor executes computer execution instructions stored in the memory to implement the method as described in any one of the first aspects.
[0039] To achieve the above objectives, a fourth aspect of this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the method as described in any one of the first aspects.
[0040] To achieve the above objectives, a fifth aspect of this application provides a computer program product that, when executed by a processor, implements the method described in any one of the first aspects.
[0041] The technical solutions provided by the embodiments of this application bring at least the following beneficial effects:
[0042] (1) Improve migration efficiency and shorten migration time: Through the dynamic scheduling of DRL agents, data blocks are allocated to the optimal worker machine and the best network path is selected (combining topology awareness and real-time status), thereby maximizing parallel processing capabilities and overall migration throughput, and significantly reducing the total time required for large-scale data migration.
[0043] (2) Ensure business stability and resource controllability: The system can dynamically adjust the intensity of migration tasks (such as transmission rate) through DRL according to the preset business protection strategy, and accurately control the consumption of local resources (CPU, memory, disk I / O, network bandwidth) of the migration activity, so as to ensure that the service quality of other normal business services on the work machine is not significantly affected.
[0044] (3) Enhance network adaptability and path optimization: Through the topology awareness capability assisted by GNN, the system can understand complex network structures and dynamically select and adjust data transmission paths in combination with real-time network status (such as available bandwidth, latency, and packet loss rate), effectively avoid congestion, make full use of available network resources, and adapt to the dynamically changing network environment.
[0045] (4) Achieve efficient collaboration and data integrity: Through the unified scheduling of the central DRL agent and the fine management of the data block status, multiple work machines can be effectively coordinated to work in parallel, avoiding task conflicts, duplication of work or task omissions, and ensuring the integrity and non-duplication of data migration.
[0046] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0047] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0048] Figure 1 A textual flowchart illustrating a collaborative data migration scheduling method based on topology awareness and deep reinforcement learning, provided as an embodiment of this application.
[0049] Figure 2 This is a simplified flowchart illustrating a collaborative data migration scheduling method based on topology awareness and deep reinforcement learning, provided as an embodiment of this application. Detailed Implementation
[0050] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0051] In modern data centers, as data volumes continue to grow, traditional data migration methods often suffer from problems when facing parallel migration across multiple machines. These problems include uneven task allocation, lack of topology awareness in path selection, uncontrollable resource usage, and scheduling strategies that cannot dynamically adapt to environmental changes. These issues lead to low migration efficiency, frequent resource conflicts, decreased service quality, and even duplication or omission of data migration tasks.
[0052] To address the aforementioned issues, embodiments of this application provide a collaborative data migration scheduling method based on topology awareness and deep reinforcement learning. Figure 1 and Figure 2 This is a flowchart illustrating the collaborative data migration scheduling method based on topology awareness and deep reinforcement learning provided in an embodiment of this application. Figure 1 and Figure 2 As shown, the method includes the following steps:
[0053] Step 101: In response to the start request of the data migration task, initialize the system configuration, including source data information, target storage information, worker cluster information, and business protection policies.
[0054] In one embodiment of this application, in step 101, after receiving the start request for the data migration task, the application first initializes and collects the required static configuration information, specifically including the following:
[0055] First, the system uses user input or preset scripts to determine the basic information of the source data, including but not limited to the location of the source data, such as the source server's IP address, storage volume ID or shared storage mount point information, the total data volume, directory structure, and file types. This information is used in the subsequent data logical sharding and migration scheduling process to ensure that data block allocation and status tracking have accurate basic information.
[0056] Secondly, the system acquires relevant information about the target storage, including its physical location, access protocol (such as NFS, CIFS, object storage API, etc.), authentication information required for access, and available capacity. Accurate identification and parameter configuration of the source and target storage provide support for subsequent data block transmission path selection and access strategies.
[0057] Furthermore, in this embodiment, the system also needs to initialize and configure the worker cluster information. The worker cluster includes several worker nodes that execute data migration tasks. The basic configuration parameters of each worker node need to be collected during the initialization phase, including but not limited to the number of CPU cores, available memory size, disk type and capacity, network card bandwidth and network interface configuration, etc. These static hardware resource parameters will serve as important inputs for the subsequent scheduling decisions of the DRL scheduling agent.
[0058] Furthermore, in this embodiment, the system also needs to load a service protection policy for each worker machine during the initialization phase. This service protection policy limits the upper limit of resources that migration tasks can occupy on each worker machine. For example, it can be set that the CPU load increase caused by the migration task does not exceed 30% of the worker machine's total CPU load, the network egress bandwidth occupied by the migration task does not exceed 60% of the worker machine's total network card bandwidth, and the disk I / O queue length does not exceed a preset threshold. This service protection policy provides resource usage boundaries for the DRL scheduling agent throughout the entire data migration process, preventing migration tasks from causing unacceptable performance impacts on other services running on the worker machines.
[0059] Through the above steps, this embodiment of the application completes the initial configuration and collection of source data, target storage, worker cluster, and resource protection strategies during the migration task startup phase, providing complete basic environmental parameters and constraints for subsequent topology discovery, path calculation, DRL scheduling, and dynamic resource control.
[0060] Step 102: Construct a weighted graph containing source server, worker machine, target storage and network devices through the network topology discovery module, and calculate multiple candidate network paths based on the weighted graph.
[0061] In another embodiment of this application, in step 102, after the system completes task initialization, it further automatically constructs a network topology model covering the source server, each worker machine, the target storage node and the key network devices between them through the network topology discovery module, so as to support path selection and scheduling optimization in subsequent data migration tasks.
[0062] Specifically, this application embodiment integrates with existing data center network management systems (such as software-defined networking (SDN) controllers) to acquire and maintain real-time information on the physical connectivity and dynamic status of links within the network. The static information extracted from the network management system by the network topology discovery module includes the physical connectivity between nodes, such as the connection between servers (source, worker, and target storage host machines) and top-of-rack (ToR) switches, the connection between ToR switches and aggregation / distribution switches, and the interconnection between aggregation / distribution switches and core switches. In typical data center environments, these layers constitute the basic structure of layered physical networks such as Leaf-Spine architectures or traditional three-layer network architectures.
[0063] In addition, the network topology discovery module also collects dynamic characteristics of the links, such as the current actual available bandwidth, end-to-end latency, packet loss rate, and current load status. This information is of great reference value for evaluating the merits of different transmission paths and can be used to guide the selection of subsequent data migration paths in real time.
[0064] In this embodiment, the network topology can be formally represented as a weighted graph. ,in This represents the set of all network nodes in the topology, including the source server, worker machines, target storage, and key network devices (such as switches and routers). This represents the set of links between nodes. The weight of a link can be a combination of key metrics such as link bandwidth, latency, packet loss rate, and current link utilization.
[0065] Based on this weighted graph, the system can calculate multiple candidate transmission paths from the source node to each worker machine, and from each worker machine to the target storage, using predefined shortest path or widest path algorithms. In practical applications, the path from the source to the worker machine is often not direct, but requires several hops via a combination of ToR, aggregation layer, and core layer switches for forwarding. For a typical example, if the source server S and a worker machine W1 are not in the same rack or under the same aggregation switch, there may be multiple possible paths, such as:
[0066] Path 1: S→ToR_S→Agg1→Core1→Agg_W1→ToR_W1→W1
[0067] Path 2: S→ToR_S→Agg2→Core2→Agg_W1→ToR_W1→W1
[0068] Similarly, there may be multiple physical or logical paths between the worker machine and the target storage. By comparing the differences of these candidate paths in terms of link bandwidth, current load, end-to-end latency, etc., the system can select several optimal or suboptimal paths (e.g., K candidate paths) to provide a feasible solution space for path selection for the subsequent deep reinforcement learning scheduling module.
[0069] Through the above steps, this embodiment of the application not only completes the modeling and visualization of network connection relationships in the network topology discovery stage, but also achieves accurate characterization of multi-path transmission in the complex multi-hop network environment inside the data center through real-time dynamic collection and weighted graph construction, thereby providing a reliable path selection basis for the scheduling and execution of subsequent data migration tasks.
[0070] Step 103: The weighted graph is embedded using the graph neural network (GNN) processing module to generate node embedding vectors, and path representation vectors are generated by aggregating the path node sequences. Based on the node embedding vectors, path representation vectors, source data states, worker states, and target storage states, system state vectors are fused together to generate system state vectors.
[0071] In one embodiment of this application, in step 103, after the system completes the construction of the network topology weighted graph and multi-path calculation, it further performs embedding processing on the weighted graph through the graph neural network (GNN) processing module to generate node embedding vectors and generate path representation vectors for candidate paths, and finally forms a complete system state vector that can be used by the deep reinforcement learning (DRL) scheduling module.
[0072] Specifically, in this embodiment, the graph neural network processing module uses a graph attention network (GAT) as its core architecture. The graph neural network processing module first receives the network topology weighted graph G=(V,E) output in step 102, where V represents the set of nodes in the network, including the source server, each worker node, the target storage node, and key network devices (such as rack-top ToR switches, aggregation layer switches, core switches, etc.), and E represents the physical or logical links between the aforementioned nodes. The link weights may include static and dynamic features.
[0073] In this embodiment of the application, the static features include at least the link's rated bandwidth and physical distance (which can be used to derive the basic transmission delay); the dynamic features include the link's current real-time load, known fault status, available bandwidth, and other information, all of which are provided and updated in real time by the network management system (such as an SDN controller).
[0074] When performing embedding computation, the graph attention network (GAT) uses an attention mechanism to weighted aggregate the features of each node's neighboring nodes and connected links, automatically learning the weights of the influence of different neighboring nodes on the current node. After multi-layer graph convolution and attention operations, GAT generates a low-dimensional node embedding vector for each key node in the network. In this embodiment, the embedding vector comprehensively reflects each node's relative position, connection density, local load, and overall environmental state within the network, thus providing richer input features than single static network information and suitable for intelligent scheduling.
[0075] To enable the DRL scheduling module to understand the complete path topology beyond simple link metrics during path selection, this embodiment further treats candidate paths as structures composed of node sequences. For the K candidate paths from the source server to each worker machine and the candidate paths from each worker machine to the target storage generated in step 102, i.e., for each path... (From source S to working machine) The j-th candidate path) and (from the working machine) The graph neural network processing module treats the path as a sequence of nodes and generates a path representation vector by aggregating the embedding vectors of key nodes in the path (the l-th candidate path to the target T).
[0076] Specifically, in this embodiment, the path representation vector can be generated in the following ways: On the one hand, mean pooling can be used to weighted average the embedding vectors of all nodes in the path to form the overall representation of the path; on the other hand, sequence modeling methods (such as small recurrent neural networks, RNNs) can be used to encode the path node sequence sequentially, capturing the node order information and potential contextual dependencies in the path, thereby generating a path representation vector containing more sequence features. Compared with simple static bandwidth or delay metrics, this path representation method can provide the DRL scheduling agent with a deeper understanding of the path topology and support more accurate dynamic path selection.
[0077] In addition, in order to form a complete input state that can be used by the deep reinforcement learning scheduling module, the embodiments of this application also incorporate the following multi-source state information:
[0078] Source data status: This includes a list of data blocks to be allocated, the size of each data block, its priority, and its current migration status (such as pending allocation, in transit, completed, failed, etc.). This information helps the DRL agent dynamically determine the priority of subsequent task allocation and data sharding scheme.
[0079] Worker status: This includes the current CPU utilization, memory utilization, disk I / O queue length and throughput, used bandwidth and remaining available bandwidth of each worker interface (source and destination), etc. It also records the data blocks that have been allocated but not yet completed and their migration progress information, as well as the gap between the current resource usage and the business protection threshold. This is used to dynamically adjust the task allocation intensity during decision-making to prevent resource overload.
[0080] This includes the target storage's current write load, available capacity, and potential access bottlenecks, to ensure that the write end does not become the performance bottleneck of the entire migration process.
[0081] Finally, the graph neural network processing module fuses the node embedding vector, path representation vector, and the aforementioned source data state, worker machine state, and target storage state to generate a complete system state vector. This state vector serves as the input to the DRL scheduling agent, helping the agent to comprehensively consider multi-dimensional factors such as data blocks, resource load, network topology, and path characteristics, and make dynamic scheduling decisions that conform to the globally optimal migration goal.
[0082] Through the steps described in the embodiments of this application, the system can achieve a deep integration of topology awareness and data migration scheduling when facing dynamically changing network and worker loads, significantly improving scheduling efficiency and resource usage security in large-scale data migration scenarios.
[0083] Step 104: Input the system state vector into the deep reinforcement learning DRL agent module. The DRL agent module outputs a composite action based on the system state vector. The composite action includes data block selection, worker selection, source-to-worker path index, worker-to-target path index, and migration rate suggestion.
[0084] In one embodiment of this application, in step 104, after generating the system state vector, the system... The input is fed into the Deep Reinforcement Learning (DRL) agent module, which then outputs a composite action based on the state vector to guide the execution of the data transfer task. .
[0085] The DRL agent module adopts an Actor-Critic structure, wherein:
[0086] The Actor network is used to generate the policy, and its input is the fusion system state vector generated in step 103. The output is a compound action. The probability distribution;
[0087] Critic networks are used to analyze the current state. The value is estimated to provide a value reference for policy optimization in Actor networks.
[0088] In this embodiment of the application, the structure of the Actor network includes:
[0089] Input layer: Receives the fused state vector described above. .
[0090] Hidden layers: Employ multi-layer fully connected neural networks (MLPs), for example, including 2-3 layers, each with 512 neurons, using the ReLU activation function;
[0091] Output layer: Generates compound actions The probability distributions of each component, specifically including:
[0092] : Data block selection: If there are multiple candidate data blocks, output the Softmax probability distribution to select a specific data block;
[0093] : Selecting the working machine and outputting the Softmax probability distribution for each candidate working machine;
[0094] The source path index is used to select a path from the K pre-calculated candidate paths using the Softmax output.
[0095] The working machine to target storage path index is also selected from pre-calculated candidate paths;
[0096] The migration rate is suggested, and the Softmax probability distribution is output in the form of discrete rate levels.
[0097] The Critic network is used for the value function of the current state. The estimation is performed, and the structure is similar to that of the Actor network, also including an input layer (receiver). ), multiple hidden layers (which may not share weights with the Actor), and an output layer (which outputs a single scalar as the state value).
[0098] In this embodiment, to guide the DRL agent to achieve a balance among multiple optimization objectives such as migration efficiency, resource protection, and network path load balancing, the DRL agent module employs a multi-objective reward function during training, formally defined as follows:
[0099]
[0100] in:
[0101] : Represents the total amount of data that all worker machines successfully pull from the source and write to the target storage within time step t, which is used as the main positive reward;
[0102] : Represents the number of newly migrated data blocks within time step t, as a stage completion reward;
[0103] If the resource usage (CPU, memory, network bandwidth, disk I / O) of any working machine exceeds the preset service protection threshold, a penalty will be imposed. The penalty magnitude is positively correlated with the degree of exceedance, and squared loss is used for calculation here.
[0104] If the selected path causes excessive link congestion (e.g., utilization exceeds 80%) or increased packet loss rate, a congestion penalty will be imposed.
[0105] A small negative reward is given for each decision-making step to encourage the task to be completed as quickly as possible;
[0106] This serves as a weighting factor for each reward item, used to flexibly balance the relative priorities among different optimization objectives such as data throughput, block completion, resource compliance, network load, and time efficiency.
[0107] Through the above design, the DRL agent in the embodiments of this application can be based on the input system state vector. Dynamically generate optimal composite actions in each scheduling cycle. This action includes: Data block selection: determining which data block to migrate in the current cycle; Worker selection: selecting a suitable worker to perform the migration task; Source-to-worker path selection: specifying which network path to choose from multiple candidate paths to transfer the data block; Worker-to-target storage path selection: similarly specifying the optimal push path from the candidate paths; Migration rate recommendation: recommending reasonable rate levels for the source-to-worker and worker-to-target transmissions to avoid resource conflicts and link congestion.
[0108] In summary, the embodiments of this application, by combining the Actor-Critic framework with the PPO algorithm, a complete multi-objective reward function design, and multi-source state information generated based on graph neural networks, effectively improve the scheduling intelligence, adaptability, and resource controllability of data migration scheduling tasks in complex data center environments.
[0109] Step 105: When the path information in the composite action is determined, the path information is converted into flow rules by the network control module and then sent to the switch devices on the path for execution by the software-defined network controller.
[0110] In one embodiment of this application, in step 105, when the DRL agent outputs a composite action... Once the path information is determined, the system uses the network control module to parse the path information into executable network flow rules, and then distributes them to the relevant switching devices on the selected path through the software-defined networking (SDN) controller to achieve precise control over the data transmission path.
[0111] Specifically, in this embodiment of the application, the trained DRL model is deployed within a central migration controller, which periodically performs the following operations: real-time collection of the current complete system state. This state includes: source data status (such as the list of data blocks to be migrated, data block size, priority, etc.); resource usage status of each worker machine (CPU, memory, disk I / O, network load, etc.); current load and available capacity of the target storage; node embedding vectors and path representation vectors generated by graph neural networks (GNN); and dynamic characteristics of the network link such as real-time bandwidth utilization, end-to-end latency, and packet loss rate.
[0112] The central migration controller will collect The input is fed into the DRL model for inference, and the DRL agent outputs the optimal composite action for the current time step. (i.e., data block allocation, worker selection, path specification, and rate suggestion).
[0113] In this embodiment, to achieve precise scheduling and isolation of the selected network path, the central migration controller calls the SDN controller interface through the network control module to convert the path information into explicit flow rules. The flow rules are based on the network topology path pre-calculated in steps 102 and 103, and include a complete node sequence and corresponding link information from the source to the worker machine and from the worker machine to the target storage.
[0114] Based on the path information, the SDN controller issues detailed flow table entries to the relevant switching devices (such as ToR, aggregation, core switches, etc.). Flow rules may include: matching conditions, such as source IP, destination IP, source port, destination port and protocol type (five-tuple); forwarding actions, such as forwarding the matched data flow to a specified port to ensure that the data is forwarded strictly along the specified path; priority or rate limiting parameters, used to match the recommended migration rate to avoid link overload.
[0115] By dynamically distributing SDN flow rules, this application embodiment can ensure that when data migration tasks pass through multi-hop networks within the data center, they are forwarded strictly according to the optimal path selected by the agent, avoiding packet detours, packet loss, or abnormal delays caused by network adaptive routing or flow table conflicts.
[0116] In this embodiment, the central migration controller continuously monitors the migration status of data blocks for which tasks have been assigned, records the transmission progress of each data block, and ensures that the same data block is not reassigned. If a data block transmission failure is detected during the migration process (e.g., link interruption, abnormal worker load, etc.), the system will automatically mark the data block as "pending allocation" and re-include it in subsequent scheduling cycles. The DRL agent will then reselect a suitable worker and transmission path for scheduling and execution based on the latest status.
[0117] Furthermore, in this embodiment, the central migration controller is equipped with a task management module for continuously tracking and managing the migration status of each data block during the migration task execution. The task management module monitors in real time: the transmission progress of each allocated data block; whether the data block currently in transmission has been completed, or whether it has failed or been abnormally interrupted.
[0118] Once a data block is detected to have failed during transmission (e.g., due to abnormal worker load, link interruption, write error, etc.), the task management module will immediately reset the status of the data block to "pending allocation" and add it back to the pending allocation list for the DRL agent in subsequent scheduling cycles to reschedule and allocate, and reselect the worker, path, and rate to retry the transmission.
[0119] Through this mechanism, the embodiments of this application guarantee: the integrity of the data block migration process, that is, all source data blocks can be successfully migrated to the target after multiple schedulings; the non-duplication of data block allocation, the system can avoid the same data block being repeatedly allocated or repeatedly pulled by multiple workers at the same time through globally unique data block status identifiers; and the self-recovery when encountering migration interruption or failure, automatic rescheduling can be achieved without manual intervention.
[0120] In summary, this application embodiment, by combining a central migration controller, an SDN controller, and a task management module, forms a self-closing, self-correcting end-to-end scheduling and path control system, which can effectively support the efficient and stable execution of data migration tasks in complex, multi-path, and dynamic load environments in large-scale data centers.
[0121] To implement the above embodiments, this application also proposes a cooperative data migration scheduling device based on topology awareness and deep reinforcement learning. The device includes:
[0122] The data acquisition module is used to acquire source data information, target storage information, worker cluster information, and business protection strategies.
[0123] The network topology discovery module is used to construct a weighted graph containing source servers, worker machines, target storage, and network devices, and output candidate network paths;
[0124] The graph neural network (GNN) processing module is used to embed the weighted graph and generate node embedding vectors and path representation vectors.
[0125] The state fusion module is used to fuse source data state, worker state, target storage state and path representation vector to generate system state vector;
[0126] The Deep Reinforcement Learning (DRL) agent module is used to make decisions based on the system state vector and outputs a composite action including data block selection, worker selection, source-to-worker path index, worker-to-target path index, and migration rate suggestion.
[0127] The network control module is used to convert the path information in the composite action into flow rules, and to send them to the switch devices on the path for execution through the software-defined network controller;
[0128] The task management module is used to record and update the migration status of each data block to ensure the integrity and non-duplication of migration tasks.
[0129] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0130] To implement the above embodiments, this application also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.
[0131] To implement the above embodiments, this application also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.
[0132] To implement the above embodiments, this application also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.
[0133] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0134] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.
[0135] This application is intended to provide an implementation scheme for users to selectively prevent the use or access to their personal information data. Specifically, this disclosure is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.
[0136] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0137] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0138] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0139] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0140] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0141] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0142] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0143] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
[0144] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.
[0145] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A collaborative data migration scheduling method based on topology awareness and deep reinforcement learning, characterized in that, Includes the following steps: In response to the start request of the data migration task, initialize the system configuration, including source data information, target storage information, worker cluster information and business protection policies; The network topology discovery module constructs a weighted graph containing source servers, worker machines, target storage, and network devices, and calculates multiple candidate network paths based on the weighted graph. The weighted graph is embedded using a graph neural network (GNN) processing module to generate node embedding vectors, and path representation vectors are generated by aggregating path node sequences. Based on the node embedding vectors, path representation vectors, source data states, worker machine states, and target storage states, a system state vector is generated by fusing them. During the embedding calculation, based on an attention mechanism, the features of each node's neighboring nodes and connected links are weighted and aggregated, and the weights of the influence of different neighboring nodes on the current node are automatically learned. After multi-layer graph convolution and attention operations, the graph neural network (GNN) processing module generates low-dimensional node embedding vectors for each key node in the network. The embedding vectors comprehensively reflect the relative position, connection density, local load, and overall state of the surrounding environment of each node in the entire network. The system state vector is input into the deep reinforcement learning DRL agent module, and the DRL agent module outputs a composite action based on the system state vector. The composite action includes data block selection, worker selection, source-to-worker path index, worker-to-target path index, and migration rate suggestion. When the path information in the composite action is determined, the path information is converted into flow rules by the network control module and then sent to the switch devices on the path for execution by the software-defined network controller.
2. The method according to claim 1, characterized in that, The network topology discovery module, through integration with the network management system, acquires static topology information and dynamic link characteristics of network devices. The static topology information includes the physical connection relationships between nodes and the rated bandwidth of the links, while the dynamic link characteristics include the current load and fault status of the links.
3. The method according to claim 2, characterized in that, The graph neural network (GNN) processing module employs a graph attention network, which uses an attention mechanism to weighted aggregate the features of nodes and links to generate low-dimensional embedding vectors to characterize the network topology.
4. The method according to claim 3, characterized in that, The path representation vector is generated by treating the path as a sequence of nodes and aggregating the embedding vectors of key nodes in the path. The aggregation process includes average pooling or sequence processing operations.
5. The method according to claim 4, characterized in that, The DRL agent module uses a proximal policy optimization algorithm for training and decision-making. The proximal policy optimization algorithm improves the stability of the learning process by limiting the policy update magnitude.
6. The method according to claim 5, characterized in that, The reward function of the DRL agent module includes at least one of the following: The total amount of data successfully migrated within the period; The number of data blocks that were newly migrated within the period; Penalty value for exceeding the preset threshold for worker machine resource usage; The penalty value for the selected network path causing link congestion or high packet loss; A time step penalty value set to encourage tasks to be completed as quickly as possible.
7. The method according to claim 6, characterized in that, The penalty values in the reward function are positively correlated with the degree of resource overrun, and a non-linear penalty function is used to enhance the suppression effect on resource overrun.
8. The method according to any one of claims 1-7, characterized in that, Also includes: During the migration process, the migration status of each data block is continuously monitored through the task management module. When a data block migration failure is detected, the data block is added back to the list to be assigned to ensure the integrity and non-duplication of the migration task.
9. A cooperative data migration scheduling device based on topology awareness and deep reinforcement learning, characterized in that, include: The data acquisition module is used to acquire source data information, target storage information, worker cluster information, and business protection strategies. The network topology discovery module is used to construct a weighted graph containing source servers, worker machines, target storage, and network devices, and output candidate network paths; The graph neural network (GNN) processing module is used to embed the weighted graph and generate node embedding vectors and path representation vectors. The state fusion module is used to fuse source data state, worker state, target storage state and path representation vector to generate system state vector; The Deep Reinforcement Learning (DRL) agent module is used to make decisions based on the system state vector and outputs a composite action including data block selection, worker selection, source-to-worker path index, worker-to-target path index, and migration rate suggestion. The network control module is used to convert the path information in the composite action into flow rules, and to send them to the switch devices on the path for execution through the software-defined network controller; The task management module is used to record and update the migration status of each data block to ensure the integrity and non-duplication of migration tasks. During the embedding computation, based on the attention mechanism, the features of each node's neighboring nodes and the links connected to them are weighted and aggregated. The weights of the influence of different neighboring nodes on the current node are automatically learned. After multi-layer graph convolution and attention operations, the graph neural network (GNN) processing module generates a low-dimensional node embedding vector for each key node in the network. The embedding vector comprehensively reflects the relative position, connection density, local load and overall state of the surrounding environment of each node in the entire network.
10. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Graph neural network-based computing power network workflow scheduling method and device
CN117978725A
Intelligent routing system and method based on knowledge definition network and graph reinforcement learning
CN118282918A