Distributed data migration method and system based on data gravitation mechanism
By constructing a multidimensional migration cost model through a gravity mechanism, and combining time constraints and resource compatibility, data migration decisions are optimized, solving the problem that existing technologies fail to comprehensively consider multidimensional factors, and achieving efficient and dynamic data migration and task scheduling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUZHOU UNIV
- Filing Date
- 2025-11-26
- Publication Date
- 2026-04-21
AI Technical Summary
Existing data migration methods fail to comprehensively consider multiple factors such as data scale, computing power, network distance and time constraints in geographically distributed environments. This results in a lack of theoretical guidance for migration decisions, repeated transmissions that increase total migration overhead, and static strategies that cannot predict resource unavailability, leading to task interruptions and resource mismatches.
A distributed data migration method based on the data gravity mechanism is adopted to construct a multi-dimensional migration cost model. Combining time constraints and resource compatibility, the optimal migration target is determined by gravity strength calculation, and multi-path transmission is used to optimize data migration decisions.
It achieves coordinated optimization of data migration and task scheduling, reduces unnecessary long-distance transmission, improves task execution efficiency, ensures task continuity, improves resource utilization, and reduces the overall data migration volume and task completion time.
Smart Images

Figure CN121900937A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of distributed scheduling technology, and in particular to a distributed data migration method and system based on a data gravity mechanism. Background Technology
[0002] Data migration is a key and complex technical means in geographically distributed scheduling tasks, aiming to solve the core contradiction of the mismatch between data storage location and computing task requirements. In the globalized cloud computing environment, data and computing resources are often dispersed in data centers in different geographical locations. Reasonable data migration strategies are crucial for reducing task execution latency, reducing network transmission overhead, and improving resource utilization. However, existing data migration methods face many limitations in geographically distributed environments. Traditional data migration strategies usually adopt bandwidth-based cost estimation models, which only consider network bandwidth and data size. This simplified model ignores several key factors in geographically distributed environments: (1) heterogeneity of computing resources, with significant differences in computing capabilities among different data centers; (2) the impact of network latency, with geographical distance and routing overhead causing additional transmission costs; (3) time window constraints, with data center availability changing dynamically over time; and (4) data-computing affinity, requiring data to be pre-migrated to areas with abundant computing resources.
[0003] Existing methods have the following main problems when dealing with geographically distributed data migration:
[0004] Traditional methods, relying solely on bandwidth or latency for cost-based decisions, fail to consider the synergistic effects of multiple factors such as data scale, computing power, and network distance. Experiments show that bandwidth-only approaches lead to longer data migration times for workloads. Furthermore, existing methods lack data-computation affinity modeling, making it difficult to quantify the attraction between data blocks and computing nodes, resulting in a lack of theoretical guidance for data migration decisions. For iterative workloads such as machine learning training, repeated cross-datacenter transfers increase total migration overhead. Ignoring spatiotemporal dynamics, data center resource availability is constrained by time windows, with some sites operating only during specific periods. Traditional static migration strategies cannot predict resource unavailability, leading to task interruptions and data re-migration. Finally, existing methods often treat data as a whole, lacking fine-grained management at the data block level and failing to develop differentiated migration strategies for data of different priorities and types. Therefore, there is an urgent need for a data migration optimization method that can comprehensively consider multiple factors such as data scale, computing power, network distance, and time constraints, quantify the affinity between data and computing resources, and support dynamic and adaptive migration. Summary of the Invention
[0005] In view of the aforementioned existing problems, the present invention is proposed.
[0006] Therefore, this invention provides a distributed data migration method based on the data gravity mechanism to address the problem of how to comprehensively consider multiple factors such as data scale, computing power, network distance, and time constraints, and quantify the affinity relationship between data and computing resources, supporting a dynamic and adaptive data migration optimization method.
[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0008] In a first aspect, the present invention provides a distributed data migration method based on a data gravity mechanism, which includes establishing a multi-dimensional migration cost model in a multi-node collaborative computing and cross-domain scheduling scenario;
[0009] We construct a physics-inspired affinity quantification model that maps data size, computing power, and network distance to a quality-distance structure.
[0010] Calculate the gravitational strength by combining time constraints and resource compatibility constraints;
[0011] By combining future resource availability and task iteration structure, a data placement strategy is obtained, and an optimal data migration decision is made based on the gravity strength.
[0012] As a preferred embodiment of the distributed data migration method based on the data gravity mechanism described in this invention, the establishment of the multidimensional migration cost model includes calculating the basic migration cost for the source node. To the target node Data migration, defining the basic migration cost ;
[0013] In a shared network environment, link contention affects effective bandwidth and transmission stability. The network contention effect can be modeled as a node-related phenomenon. The proportion of concurrent traffic load to its total egress bandwidth;
[0014] The basic migration cost is extended to a weighted cost to accommodate the needs of various service level agreements (SLAs) and multi-priority workloads, and importance is distinguished in the global scheduling strategy.
[0015] As a preferred embodiment of the distributed data migration method based on the data gravity mechanism described in this invention, the construction of the physically inspired affinity quantification model includes constructing a physically inspired affinity quantification model by analogy with the law of universal gravitation, mapping data scale, computing power and network distance to a comparable mass-distance structure.
[0016] Analogous to Newton's law of universal gravitation, the gravitational force between data and calculation can be defined as follows:
[0017]
[0018] in, Data quality is typically determined by data size or data importance. The quality of computation is defined by the overall available computing power of the target computing node. Network distance is represented by transmission cost and latency; The system-related gravitational constant is used to control the scale of the overall gravity; It represents the distance decay exponent, which adjusts the rate at which gravity decays with distance;
[0019] Decay Index Changing the rate at which gravity decays with distance influences migration decision preferences; decay exponent Adaptively select based on workload characteristics; for tasks with high data locality requirements, set... This causes gravity to weaken with distance, enhancing migration to nearby locations; for computationally intensive tasks, it sets... It employs linear decay, allowing execution on nodes with high computing power but located far apart.
[0020] As a preferred embodiment of the distributed data migration method based on the data gravity mechanism described in this invention, the mapping as a quality-distance structure includes, in order to accurately reflect the heterogeneity of modern diverse computing resources, the computing quality... Defined as a weighted combination of available capabilities across different resource dimensions; available capability parameters include: a weighted combination of remaining available CPU processing capacity, available GPU computing resources, available memory capacity, and available network processing capacity, with each resource weight dynamically adjusted based on workload characteristics;
[0021] For GPU-intensive tasks, give high weight to available GPU computing resources; for memory-intensive tasks, give dominant weight to available memory capacity.
[0022] network distance Constructed as a composite metric of geographical distance and transmission cost:
[0023]
[0024] in Indicates geographical distance. , It represents the balance coefficient, used to adjust the relative contribution of physical distance and network transmission characteristics. The composite distance metric reflects both physical separation and network transmission difficulty.
[0025] As a preferred embodiment of the distributed data migration method based on the data gravity mechanism described in this invention, the combination of time constraints and resource compatibility constraints includes integrating time constraints and resource compatibility to calculate the effective gravity strength;
[0026] Time window function, defines time availability function Feasibility of adjusting scheduling for different time periods:
[0027]
[0028] in, Represents a node At any moment Time availability function value, Indicates the current scheduling time; This indicates the start time of the active time window for node i; Represents a node The end time of the active time window; This represents the timeout penalty factor, with a value range of [value range missing]. ; This represents the exponential decay rate, controlling the rate at which the penalty intensity decreases; at scheduling time... When within the active window, the function value is 1, indicating full availability; when outside the window, an exponential penalty is applied based on the degree of excess; otherwise, the value is 0, indicating unavailability.
[0029] Define resource type compatibility functions This measures the degree of matching between task requirements and node capabilities.
[0030] As a preferred embodiment of the distributed data migration method based on the data gravity mechanism described in this invention, the calculation of gravity strength includes ensuring that data migration decisions simultaneously satisfy the triple constraints of gravitational attraction, time availability, and resource compatibility; integrating computational tasks At any moment For nodes The effective gravitational strength is expressed by the formula:
[0031]
[0032] in, Indicates task At any moment For nodes The effective gravitational strength; Indicates task With nodes The fundamental gravitational force between them was calculated using a physics-inspired gravitational model; Represents a node At any moment Time availability function value; Indicates task With nodes Resource compatibility function value;
[0033] Taking all factors into account, the complete scheduling gravity function is defined as follows:
[0034]
[0035] in, Indicates task At any moment For nodes The final scheduling gravity value; Represents the gravitational constants associated with the system, used to adjust the gravitational level; Indicates task The size of the data, i.e., the data quality; Represents a node At any moment The computing power, i.e., the quality of computing; Indicates task Current position of data to node Network distance; This represents the distance decay index.
[0036] As a preferred embodiment of the distributed data migration method based on the data gravity mechanism described in this invention, the step of making the optimal data migration decision based on gravity strength includes, based on a gravity model, sorting and selecting candidate computing nodes according to the attraction strength of data blocks to determine the optimal migration target node, specifically including:
[0037] Obtain the set of all data blocks to be migrated and the set of candidate nodes;
[0038] Calculate the gravitational strength value of each data block to be placed for each candidate computing node, and form a list of gravitational strength of data blocks;
[0039] Arrange the candidate nodes in descending order of gravitational strength to obtain a priority list;
[0040] Starting from the top of the priority list, each node is checked sequentially to determine if it has enough remaining storage capacity to accommodate the data block. If a node meets the capacity constraint, it is considered a candidate migration target; otherwise, the next node is checked.
[0041] If there are nodes that meet the capacity constraints, the node with the strongest attraction is selected as the migration target; if the capacity of all nodes is insufficient, the data sharding mechanism is triggered, and the data blocks are split and migrated to multiple nodes.
[0042] Calculate the optimal transmission path from the current location to the target node; consider network topology and bandwidth utilization, and use multi-path transmission to reduce contention; perform batch migration, prioritizing the transmission of data blocks with high gravitational intensity;
[0043] After the migration is complete, update the global data location index, node storage usage, and network traffic statistics to support the first-stage scheduling task of the next migration cycle.
[0044] Secondly, the present invention provides a distributed data migration system based on a data gravity mechanism, including a data migration cost modeling module, which establishes a multi-dimensional migration cost model that comprehensively considers bandwidth, latency, and network competition factors.
[0045] The gravity model building module constructs a data-computation affinity gravity model, mapping the scale of data, computing power, and network distance to mass and distance in the gravity model, quantifying the gravitational strength between data and computing resources; by ranking the gravitational strength of candidate nodes, the node with the strongest gravity is selected as the migration target, ensuring the efficiency of task scheduling and the optimization of resource utilization;
[0046] The data migration execution module is responsible for making data migration decisions, calculating the optimal transmission path from the current node to the target node, and performing multi-path transmission based on network topology and bandwidth utilization.
[0047] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the distributed data migration method based on the data gravity mechanism as described in the first aspect of the present invention.
[0048] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the distributed data migration method based on the data gravity mechanism as described in the first aspect of the present invention.
[0049] The beneficial effects of this invention are as follows: This invention proposes a data migration optimization method based on a gravity model, which can comprehensively consider multiple factors such as data scale, computing power, and network distance to achieve coordinated optimization of data migration and task scheduling. Compared with traditional methods that rely solely on bandwidth, this invention has a better ability to characterize data-computation affinity, effectively identifying and avoiding unnecessary long-distance data transmission and improving the efficiency of cross-data center task execution. This method also introduces a time window function and a predictive migration mechanism, enabling it to adapt to dynamically changing resource conditions and ensuring task continuity even in scenarios with significant time-varying characteristics, such as nighttime edge nodes. Regarding heterogeneous resource scheduling, this invention ensures that computing tasks match node capabilities through a compatibility function, avoiding performance loss due to resource mismatch. Furthermore, the method has adaptive parameter adjustment capabilities, dynamically optimizing key parameters according to workload characteristics. In iterative tasks, through cumulative gravity judgment and pre-migration strategies, it significantly reduces the overall data migration volume and task completion time. Attached Figure Description
[0050] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0051] Figure 1 This is a flowchart of a distributed data migration method based on the data gravity mechanism.
[0052] Figure 2 This is a schematic diagram of the data gravity field model for a distributed data migration method based on the data gravity mechanism.
[0053] Figure 3 This is a data center network topology diagram for a distributed data migration method based on the data gravity mechanism.
[0054] Figure 4 This is a comparative analysis of workloads for different data characteristics in distributed data migration methods based on data gravity mechanisms. Detailed Implementation
[0055] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0056] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0057] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0058] Reference Figures 1-4 As one embodiment of the present invention, this embodiment provides a distributed data migration method based on a data gravity mechanism, comprising the following steps:
[0059] S1: Establish a multi-dimensional migration cost model in multi-node collaborative computing and cross-domain scheduling scenarios.
[0060] Furthermore, in multi-node collaborative computing or cross-domain scheduling scenarios, data migration often constitutes a key part of the overall scheduling overhead.
[0061] Figure 4 This study demonstrates the performance of workload migration under different data characteristics, illustrating the significant impact of data size and access patterns on migration costs. To accurately characterize actual data transmission behavior under different network conditions, this phase establishes a multi-dimensional migration cost model that comprehensively considers factors such as bandwidth, latency, and network contention, providing a foundational quantitative input for subsequent "gravity" calculations.
[0062] The basic migration cost is calculated as follows: For data migration from source node j to target node i, the basic migration cost is defined as:
[0063]
[0064] in, Indicates the number of data blocks that need to be migrated; Let d be the size of the d-th data block; This represents the available network bandwidth between nodes j and i; This indicates the underlying network latency, such as round-trip time (RTT) or one-way transmission delay. Quantify the degree of network competition experienced by this link; Weighting coefficients are used to control the influence of network competition.
[0065] Network contention modeling reveals that in a shared network environment, link contention significantly impacts effective bandwidth and transmission stability. Therefore, the network contention effect is modeled... Model as nodes The proportion of concurrent traffic load to its total egress bandwidth:
[0066]
[0067] in, This represents the current flow from node j to node k; Let be the total outbound bandwidth of node j.
[0068] Weighted migration cost: To accommodate the needs of various service level agreements (SLAs) and multi-priority workloads, the basic migration cost is generalized to a weighted cost, so as to distinguish its importance in the global scheduling strategy.
[0069]
[0070] in, Indicates the number of priority categories; Priority Weighting coefficients; Indicates priority The data migration cost. Higher priority tasks are assigned greater weight to ensure their data migration receives priority.
[0071] It should be noted that a multi-dimensional migration cost model is established, comprehensively considering factors such as bandwidth, latency, and network contention. First, the basic migration cost is calculated based on information such as the size of the data block, network bandwidth between the source and target nodes, and network latency. By modeling actual data transmission behavior under different network conditions, necessary quantitative input is provided for subsequent gravity calculations.
[0072] S2: Construct a physics-inspired affinity quantification model that maps data size, computing power, and network distance to a quality-distance structure.
[0073] Furthermore, a physics-inspired gravity model is constructed to more intuitively characterize the attractive relationship between "data" and "computing resources" during data scheduling and task placement. This stage draws on Newton's law of universal gravitation from classical physics to build a physics-inspired affinity quantification model. This model maps data scale, computing power, and network distance into a comparable "mass-distance" structure, thereby providing interpretable and analyzable indicators for subsequent scheduling decisions.
[0074] The basic form of the gravitational model, analogous to Newton's law of universal gravitation:
[0075]
[0076] The attraction between data and computation is defined as:
[0077]
[0078] Where S represents data quality, usually determined by data size or importance; C represents computation quality, defined by the available computing power of the target computing node; D represents network distance, characterized by transmission cost and latency; G is the system-related gravitational constant used to control the scale of overall gravity; and k is the distance decay exponent, usually taken as 1 or 2, used to adjust the rate at which gravity decays with distance.
[0079] like Figure 2 As shown, the "Data Gravity Field Model" intuitively illustrates the gravitational relationship between data scale, computing power, and network distance.
[0080] The definition of computational quality, in order to accurately reflect the heterogeneity of modern diverse computing resources, defines computational quality as... Defined as a weighted combination of available capabilities across different resource dimensions:
[0081]
[0082] in, Indicates the remaining available CPU processing power; Indicates available GPU computing resources; Indicates the available memory capacity; This indicates available network processing capacity or throughput. Resource weights. , , , Adjust dynamically based on workload characteristics.
[0083] For example, for GPU-intensive tasks, It is given higher weight; for memory-intensive tasks, Dominant.
[0084] Network distance metrics are crucial because traditional physical distance is insufficient to describe communication costs in computing environments. Therefore, network distance... It is constructed as a composite metric of geographical distance and transmission cost, expressed by the following formula:
[0085]
[0086] in, Geographical distance, , This is a balancing factor used to adjust the relative contribution of physical distance to network transmission characteristics. This composite distance metric can simultaneously reflect both physical separation and network transmission difficulty. For example... Figure 3 As shown, Figure 3A typical data center network topology is presented to illustrate how communication paths and network distances are constructed between nodes in different geographical locations.
[0087] The choice of decay exponent k significantly alters the rate at which gravity decays with distance, thus influencing migration decision preferences. Therefore, the decay exponent k is adaptively selected based on workload characteristics. For tasks with high data locality requirements (such as iterative computation), setting k=2 allows gravity to decay rapidly with distance, reinforcing migration to the nearest node. For computationally intensive tasks, setting k=1 employs linear decay, allowing execution on computationally powerful but geographically distant nodes.
[0088] It should be noted that a data-computation affinity gravity model is constructed, mapping the scale of data, computing power, and network distance to mass and distance in the gravity model, thereby quantifying the gravitational strength between data and computing resources. The model adaptively adjusts according to the characteristics of different workloads to ensure the matching of computing tasks and resources, thereby improving the efficiency of data migration and the utilization rate of computing resources.
[0089] S3: Calculate the gravitational strength by combining time constraints and resource compatibility constraints.
[0090] Further, the gravitational strength calculation, in this stage, integrates time constraints and resource compatibility to calculate the effective gravitational strength.
[0091] Time window function, defines time availability function The feasibility of scheduling at different time periods can be adjusted, and the formula is expressed as:
[0092]
[0093] in, This represents the time availability function value of node i at time t, where t represents the current scheduling time. This indicates the start time of the active time window for node i. This indicates the end time of the active time window of node i. This is the timeout penalty factor, with a value range of (0,1). The exponential decay rate controls the rate at which the penalty intensity decays. At scheduling time... When within the active window, the function value is 1, indicating full availability; when outside the window, an exponential penalty is applied based on the degree of excess; otherwise, the value is 0, indicating unavailability.
[0094] Resource compatibility functions, defining resource type compatibility functions. Measure the degree of matching between task requirements and node capabilities:
[0095]
[0096] This represents the resource compatibility function value between task j and node i, where R represents the set of resource types, including CPU, GPU, memory, storage, bandwidth, etc. Indicates the number of resource types. This represents the available capacity of resource type r of node i. This indicates the amount of resource type r required by task j.
[0097] Effective gravitational strength: Integrating the above factors, calculate the effective gravitational force of task j on node i at time t:
[0098]
[0099] This formula ensures that data migration decisions simultaneously satisfy the triple constraints of gravitational attraction, time availability, and resource compatibility.
[0100] in, This represents the effective gravitational force exerted by task j on node i at time t. The fundamental gravitational force between task j and node i is calculated using the aforementioned physics-inspired gravity model. This represents the time availability function value of node i at time t. This represents the resource compatibility function value between task j and node i.
[0101] Finally, considering all factors, the complete scheduling gravity function is defined as follows:
[0102]
[0103] in, Let G represent the final scheduling gravity value of task j for node i at time t, and let G represent the system-related gravitational constant used to adjust the gravitational level. This indicates the data size of task j, i.e., data quality. This represents the computational capability of node i at time t, i.e., computational quality. The distance from the current location of task j to node i is represented by k, which represents the distance decay exponent and is usually 1 or 2.
[0104] It should be noted that the gravitational strength of each candidate node is calculated based on the gravity model, and the availability of the time window and resource compatibility are comprehensively considered to ultimately determine the target node for data migration. By ranking the gravitational strength of candidate nodes, the node with the strongest gravity is selected as the migration target, ensuring efficient task scheduling and optimal resource utilization.
[0105] S4: Combining future resource availability and task iteration structure, a data placement strategy is obtained, and the optimal data migration decision is made based on the gravity strength.
[0106] Furthermore, in the data migration decision-making phase, after completing migration cost modeling and gravity calculation, the aim of this phase is to make the optimal data migration decision based on gravity strength, and further combine future resource availability and task iteration structure to achieve a more intelligent, forward-looking and efficient data placement strategy.
[0107] This data migration decision algorithm, based on a gravity model, sorts and selects candidate computing nodes according to the attraction strength of data blocks, thereby determining the optimal migration target node. Its detailed execution flow is as follows:
[0108] (1) Initialization: Obtain the set of all data blocks to be migrated and candidate node set .
[0109] (2) Gravity calculation: For each data block to be placed, calculate its gravitational strength value to each candidate computing node to form a gravitational strength list of the data block.
[0110] (3) Node sorting: Arrange the candidate nodes in descending order of gravitational strength to obtain a priority list. .
[0111] (4) Capacity check: Starting from the top of the priority list, check each node sequentially to determine if it has enough remaining storage capacity to accommodate the data blocks. If a node meets the capacity constraint, it is considered a candidate migration target; otherwise, the process continues to check the next node.
[0112] (5) Migration decision: If there are nodes that meet the capacity constraints, select the node with the greatest attraction as the migration target. If the capacity of all nodes is insufficient, trigger the data sharding mechanism to split the data blocks and migrate them to multiple nodes.
[0113] (6) Migration Execution: Calculate the optimal transmission path from the current location to the target node. Considering network topology and bandwidth utilization, multi-path transmission is adopted to reduce contention. Migration is carried out in batches, prioritizing the transmission of data blocks with high gravitational intensity.
[0114] (7) Status update: After the migration is completed, the system updates the global data location index, node storage usage and network traffic statistics to support the next migration cycle or the next stage of scheduling tasks.
[0115] It should be noted that this function is responsible for executing data migration decisions, calculating the optimal transmission path from the current node to the target node, and performing multi-path transmission based on network topology and bandwidth utilization. It employs a batch migration strategy, prioritizing the transmission of data blocks with high gravitational pull, and updating the global data location index and node storage state to support subsequent scheduling tasks.
[0116] The following is an embodiment of the present invention, providing a method for optimizing the allocation of resources for collaborative construction across multiple trades. To verify the beneficial effects of the present invention, scientific demonstration is conducted through economic benefit calculations and simulation experiments. It can comprehensively consider multiple factors such as data scale, computing power, and network distance to achieve collaborative optimization of data migration and task scheduling.
[0117] Compared to traditional methods that rely solely on bandwidth, this invention offers superior data-computation affinity characterization, effectively identifying and avoiding unnecessary long-distance data transmission and improving the efficiency of cross-data center task execution. The method also introduces a time window function and a predictive migration mechanism, enabling it to adapt to dynamically changing resource conditions and ensuring task continuity even in scenarios with distinct time-varying characteristics, such as nighttime edge nodes. Regarding heterogeneous resource scheduling, this invention uses a compatibility function to ensure that computational tasks match node capabilities, avoiding performance losses due to resource mismatch. Furthermore, the method possesses adaptive parameter adjustment capabilities, dynamically optimizing key parameters based on workload characteristics. In iterative tasks, it significantly reduces the overall data migration volume and task completion time through cumulative gravity judgment and pre-migration strategies.
[0118] To verify the effectiveness of this invention, comparative experiments were conducted in a real-world test environment comprising 5 geographically distributed data centers and 25 computing nodes. Three representative workloads were selected for the experiments: WordCount (data-intensive), PageRank (iterative computation), and a machine learning pipeline (hybrid). The gravity model method of this invention was compared with traditional methods based solely on bandwidth. The experimental results are shown in Table 1 and... Figure 4 As shown. Figure 4 The above comparison results are presented in a visual format.
[0119] Table 1 Comparison of Data Migration Performance
[0120] workload method Data migration volume percentage Transmission time (seconds) WordCount Based on bandwidth only 8.2% 295 WordCount Gravity model 4.7% 156 PageRank Based on bandwidth only 12.5% 498 PageRank Gravity model 7.3% 287 ML Pipeline Based on bandwidth only 18.6% 695 ML Pipeline Gravity model 11.2% 423
[0121] The experimental data shows that for the WordCount workload, this invention reduces the data migration rate from 8.2% to 4.7% and the transmission time from 295 seconds to 156 seconds, a reduction of 47.1%; for PageRank iterative calculation, the transmission time is reduced from 498 seconds to 287 seconds, a reduction of 42.4%; and for the machine learning pipeline, the transmission time is reduced from 695 seconds to 423 seconds, a reduction of 39.1%.
[0122] Experimental results show that this method can serve as the core data migration component of the scheduling system, and it performs stably in various load types and distributed environments of different scales, demonstrating good scalability and versatility.
[0123] This embodiment also provides a distributed data migration system based on the data gravity mechanism, including: a data migration cost modeling module, which establishes a multi-dimensional migration cost model, comprehensively considering bandwidth, latency, and network competition factors.
[0124] The gravity model building module constructs a data-computation affinity gravity model, mapping the scale of data, computing power, and network distance to mass and distance in the gravity model, quantifying the gravitational strength between data and computing resources; by ranking the gravitational strength of candidate nodes, the node with the strongest gravity is selected as the migration target, ensuring efficient task scheduling and optimal resource utilization.
[0125] The data migration execution module is responsible for making data migration decisions, calculating the optimal transmission path from the current node to the target node, and performing multi-path transmission based on network topology and bandwidth utilization.
[0126] This embodiment also provides a computer device applicable to the distributed data migration method based on the data gravity mechanism, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the distributed data migration method based on the data gravity mechanism proposed in the above embodiment.
[0127] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0128] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the distributed data migration method based on the data gravity mechanism proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0129] In summary, this invention establishes a multi-dimensional migration cost model in multi-node collaborative computing and cross-domain scheduling scenarios; constructs a physics-inspired affinity quantification model to map data scale, computing power, and network distance into a quality-distance structure; calculates the gravitational strength by combining time constraints and resource compatibility constraints; and obtains a data placement strategy by combining future resource availability and task iteration structure, and makes the optimal data migration decision based on the gravitational strength.
[0130] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A distributed data migration method based on data gravity mechanism, characterized in that: This includes establishing a multi-dimensional migration cost model in multi-node collaborative computing and cross-domain scheduling scenarios; We construct a physics-inspired affinity quantification model that maps data size, computing power, and network distance to a quality-distance structure. Calculate the gravitational strength by combining time constraints and resource compatibility constraints; By combining future resource availability and task iteration structure, a data placement strategy is obtained, and an optimal data migration decision is made based on the gravity strength.
2. The distributed data migration method based on data gravity mechanism as described in claim 1, characterized in that: The establishment of the multidimensional migration cost model includes calculating the basic migration cost for the source node. To the target node Data migration, defining the basic migration cost ; In a shared network environment, link contention affects effective bandwidth and transmission stability. The network contention effect can be modeled as a node-related phenomenon. The proportion of concurrent traffic load to its total egress bandwidth; The basic migration cost is extended to a weighted cost to accommodate the needs of various service level agreements (SLAs) and multi-priority workloads, and importance is distinguished in the global scheduling strategy.
3. The distributed data migration method based on data gravity mechanism as described in claim 2, characterized in that: The construction of the physical-inspired affinity quantification model includes constructing a physical-inspired affinity quantification model based on the law of universal gravitation, mapping data scale, computing power and network distance to a comparable mass-distance structure; Analogous to Newton's law of universal gravitation, the gravitational force between data and calculation can be defined as follows: in, Data quality is typically determined by data size or data importance. The quality of computation is defined by the overall available computing power of the target computing node. Network distance is represented by transmission cost and latency; The system-related gravitational constant is used to control the scale of the overall gravity; It represents the distance decay exponent, which adjusts the rate at which gravity decays with distance; Decay Index Changing the rate at which gravity decays with distance influences migration decision preferences; decay exponent Adaptively select based on workload characteristics; for tasks with high data locality requirements, set... This causes gravity to weaken with distance, enhancing migration to nearby locations; for computationally intensive tasks, it sets... It employs linear decay, allowing execution on nodes with high computing power but located far apart.
4. The distributed data migration method based on data gravity mechanism as described in claim 3, characterized in that: The mapping, a quality-distance structure, includes, to reflect the heterogeneity of modern, diverse computing resources, the computational quality... Defined as a weighted combination of available capabilities across different resource dimensions; the available capabilities include: remaining available CPU processing capacity, available GPU computing resources, available memory capacity, and available network processing capacity, with the resource weight of each available capability dynamically adjusted based on workload characteristics; For GPU-intensive tasks, give high weight to available GPU computing resources; for memory-intensive tasks, give dominant weight to available memory capacity. network distance Constructed as a composite metric of geographical distance and transmission cost: in Indicates geographical distance. , It represents the balance coefficient, used to adjust the relative contribution of physical distance and network transmission characteristics. The composite distance metric reflects both physical separation and network transmission difficulty.
5. The distributed data migration method based on data gravity mechanism as described in claim 4, characterized in that: The combination of time constraints and resource compatibility constraints includes integrating time constraints and resource compatibility to calculate the effective gravitational strength; Time window function, defines time availability function Feasibility of adjusting scheduling for different time periods: in, Represents a node At any moment Time availability function value, Indicates the current scheduling time; This indicates the start time of the active time window for node i; Represents a node The end time of the active time window; This represents the timeout penalty factor, with a value range of [value range missing]. ; This represents the exponential decay rate, controlling the rate at which the penalty intensity decreases; at scheduling time... When within the active window, the function value is 1, indicating full availability; when outside the window, an exponential penalty is applied based on the degree of excess; otherwise, the value is 0, indicating unavailability. Define resource type compatibility functions This measures the degree of matching between task requirements and node capabilities.
6. The distributed data migration method based on data gravity mechanism as described in claim 5, characterized in that: The computational gravity strength includes ensuring that data migration decisions simultaneously satisfy the triple constraints of gravitational attraction, time availability, and resource compatibility; and integrating computational tasks. At any moment For nodes The effective gravitational strength is expressed by the formula: in, Indicates task At any moment For nodes The effective gravitational strength; Indicates task With nodes The fundamental gravitational force between them was calculated using a physics-inspired gravitational model; Represents a node At any moment Time availability function value; Indicates task With nodes Resource compatibility function value; Taking all factors into account, the complete scheduling gravity function is defined as follows: in, Indicates task At any moment For nodes The final scheduling gravity value; Represents the gravitational constants associated with the system, used to adjust the gravitational level; Indicates task The size of the data, i.e., the data quality; Represents a node At any moment The computing power, i.e., the quality of computing; Indicates task Data current position to node Network distance; This represents the distance decay index.
7. The distributed data migration method based on data gravity mechanism as described in claim 6, characterized in that: The optimal data migration decision based on gravitational strength includes, based on a gravity model, sorting and selecting candidate computing nodes according to the attraction strength of data blocks to determine the optimal migration target node, specifically including: Obtain the set of all data blocks to be migrated and the set of candidate nodes; Calculate the gravitational strength value of each data block to be placed for each candidate computing node, and form a list of gravitational strength of data blocks; Arrange the candidate nodes in descending order of gravitational strength to obtain a priority list; Starting from the top of the priority list, each node is checked sequentially to determine if it has enough remaining storage capacity to accommodate the data block. If a node meets the capacity constraint, it is considered a candidate migration target; otherwise, the next node is checked. If there are nodes that meet the capacity constraints, the node with the strongest attraction is selected as the migration target; if the capacity of all nodes is insufficient, the data sharding mechanism is triggered, and the data blocks are split and migrated to multiple nodes. Calculate the optimal transmission path from the current location to the target node; consider network topology and bandwidth utilization, and use multi-path transmission to reduce contention; perform batch migration, prioritizing the transmission of data blocks with high gravitational intensity; After the migration is complete, update the global data location index, node storage usage, and network traffic statistics to support the first-stage scheduling task of the next migration cycle.
8. A distributed data migration system based on a data gravity mechanism, comprising the distributed data migration method based on a data gravity mechanism as described in any one of claims 1 to 7, characterized in that: This includes a data migration cost modeling module, which establishes a multi-dimensional migration cost model that comprehensively considers factors such as bandwidth, latency, and network competition. The gravity model building module constructs a data-computation affinity gravity model, mapping the scale of data, computing power, and network distance to mass and distance in the gravity model, quantifying the gravitational strength between data and computing resources; by ranking the gravitational strength of candidate nodes, the node with the strongest gravity is selected as the migration target, ensuring the efficiency of task scheduling and the optimization of resource utilization; The data migration execution module is responsible for making data migration decisions, calculating the optimal transmission path from the current node to the target node, and performing multi-path transmission based on network topology and bandwidth utilization.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the distributed data migration method based on the data gravity mechanism as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the distributed data migration method based on the data gravity mechanism as described in any one of claims 1 to 7.