Data cross-region migration method and device, electronic equipment and storage medium

By optimizing cross-regional data migration through genetic algorithms and comprehensively considering multi-factor analysis, the high cost and low efficiency problems caused by single factors in existing technologies are solved, and efficient and economical data migration and resource utilization are achieved.

CN121349998APending Publication Date: 2026-01-16E SURFING VISION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511509788.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing cross-regional data migration strategies have limited considerations, resulting in low migration efficiency and high operating costs. They are difficult to formulate efficient and economical migration strategies for complex multi-regional storage environments and dynamic data access needs.

Method used

We employ multi-factor analysis based on genetic algorithms and an improved genetic algorithm, combined with divide-and-conquer and parallel evolution methods, to construct a mathematical model, optimize the scheduling process, and calculate the optimal data migration scheme.

Benefits of technology

It reduces the cost and impact of data storage and migration, improves resource utilization efficiency, enhances user experience continuity, solves the problem of requiring human intervention in hot storage areas, and achieves resource load balancing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121349998A_ABST
    Figure CN121349998A_ABST
Patent Text Reader

Abstract

The invention discloses a data cross-region migration method and device, electronic equipment and a storage medium. The technical problems that a current data cross-region migration strategy is single in consideration factor, low in migration efficiency and high in operation cost are solved. The method comprises the steps of obtaining to-be-migrated data and at least one target storage area of data cross-region migration; based on the to-be-migrated data, performing multi-factor analysis on each target storage area to obtain a multi-factor analysis result of each target storage area; on the basis of the multi-factor analysis result of each target storage area, optimal scheduling combining divide-and-conquer and parallel genetic evolution is carried out on the to-be-migrated data, and an optimal data migration scheme is obtained; and migrating the to-be-migrated data to at least one target storage area based on the optimal data migration scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of IT and software development technology, and in particular to a method, apparatus, electronic device, and storage medium for cross-regional data migration. Background Technology

[0002] In the construction of highly available distributed systems, to improve service availability and disaster recovery capabilities, reduce the impact of failures, balance storage resources, improve data access efficiency, and reduce storage costs, it is typically necessary to implement multi-AZ (Availability Zone) and unitized deployment of services. Here, AZ refers to geographically isolated, independent sets of resources within a region. Unitization refers to a self-contained set capable of completing all business operations. This set contains all the services required by all businesses, as well as the data allocated to that unit.

[0003] In the process of building multi-AZ and unitized systems, it is often necessary to migrate business data of existing users across regions and select the region to write business data of new users. However, current cross-regional data migration strategies often consider only one factor. Furthermore, most current migration scheduling algorithms use simple heuristics or fixed rules, lacking a robust mathematical model to accurately calculate the optimal migration scheme. Consequently, when facing complex multi-regional storage environments and dynamically changing data access needs, it is difficult to formulate efficient and economical data migration strategies, resulting in high operating costs and low resource utilization efficiency of the storage system. Therefore, there is an urgent need for a method that can comprehensively consider multiple factors and achieve efficient and low-cost cross-regional migration of storage data through scientific mathematical models and optimization algorithms. Summary of the Invention

[0004] This invention provides a method, apparatus, electronic device, and storage medium for cross-regional data migration, which solves or partially solves the technical problems of current cross-regional data migration strategies having limited consideration of factors, low migration efficiency, and high operating costs.

[0005] This invention provides a method for cross-regional data migration, the method comprising:

[0006] Obtain the data to be migrated, and at least one target storage area for the cross-regional data migration;

[0007] Based on the data to be migrated, a multi-factor analysis is performed on each of the target storage areas to obtain the multi-factor analysis results for each target storage area.

[0008] Based on the multi-factor analysis results of each target storage area, the data to be migrated is optimized by combining divide-and-conquer and parallel genetic evolution to obtain the optimal data migration scheme.

[0009] Based on the optimal data migration scheme, the data to be migrated is migrated to the at least one target storage area.

[0010] Optionally, the step of performing multi-factor analysis on each of the target storage areas based on the data to be migrated, and obtaining the multi-factor analysis results for each target storage area, includes:

[0011] For each target storage area, based on the data to be migrated, the target storage area is subjected to location matching degree analysis, storage resource utilization rate analysis, storage cost analysis, data access demand analysis, and network bandwidth analysis to obtain the location matching degree analysis results, storage resource utilization rate analysis results, storage cost analysis results, data access demand analysis results, and network bandwidth analysis results.

[0012] By combining the weighted penalty mechanism, and integrating the results of the location matching degree analysis, the storage resource utilization analysis, the storage cost analysis, the data access demand analysis, and the network bandwidth analysis by weighted summation, the multi-factor analysis results of the target storage area are obtained.

[0013] Optionally, the step of optimizing the data to be migrated by combining divide-and-conquer with parallel genetic evolution based on the multi-factor analysis results of each target storage area to obtain the optimal data migration scheme includes:

[0014] Extract all users whose data needs to be migrated across regions from the data to be migrated;

[0015] The users are grouped according to a preset grouping method to obtain multiple user groups.

[0016] Based on the multi-factor analysis results of each target storage area, parallel genetic optimization scheduling is performed on the multiple grouped users to obtain the optimal data migration scheme.

[0017] Optionally, the step of performing parallel genetic optimization scheduling on the multiple grouped users based on the multi-factor analysis results of each target storage area to obtain the optimal data migration scheme includes:

[0018] For each group of users, based on the multi-factor analysis results of each target storage area and combined with the load balancing degree of each target storage area, a fitness function for each user in the group under each target storage area is constructed;

[0019] With the goal of maximizing the fitness function, genetic optimization scheduling is performed on the group of users to obtain the optimal scheduling result for each user in the group, which is then used as the optimal migration sub-scheme for the user.

[0020] The optimal migration sub-schemes for each user in each of the aforementioned user groups are combined to form the global optimal data migration scheme.

[0021] Optionally, the optimal data migration scheme includes an optimal migration sub-scheme for all users who need to migrate data across regions; the step of migrating the data to be migrated to the at least one target storage area based on the optimal data migration scheme includes:

[0022] For each user who needs to migrate data across regions, retrieve the user's business data from the data to be migrated, and select a storage area to be migrated from the at least one target storage area based on the optimal migration sub-scheme;

[0023] The business data is written to the storage area to be migrated in batches and multiple times, and migration integrity verification is performed.

[0024] Once the migration integrity check is passed, the retrieval route update operation for the user is immediately executed, and after a preset time, the business data stored in the original data storage area is deleted.

[0025] Optionally, the execution process of the migration integrity verification includes:

[0026] The data volume of the business data is determined based on a preset data volume threshold, and the integrity of the data migration process of the business data is verified based on the data volume determination result to obtain the migration integrity verification result.

[0027] Optionally, the step of judging the data volume of the business data based on a preset data volume threshold, and performing integrity verification on the data migration process of the business data according to the data volume judgment result, to obtain the migration integrity verification result, includes:

[0028] Determine whether the amount of the business data is less than or equal to a preset data amount threshold;

[0029] If so, then after all the business data is written to the storage area to be migrated, an integrity check will be performed again;

[0030] If not, an integrity check will be performed after each batch of business data is written to the storage area to be migrated, and the next batch of data will be written only after the integrity check of the current batch of data passes.

[0031] When the integrity verification process is confirmed to be executed, the data volume before and after the migration of the currently successfully migrated business data is compared.

[0032] If the data volume before and after the migration is consistent, the current migration integrity verification is considered to have passed.

[0033] If the data volume before and after the migration is inconsistent, the current migration integrity check is deemed to have failed.

[0034] The present invention also provides a data cross-region migration device, comprising:

[0035] A data acquisition unit is used to acquire the data to be migrated, as well as at least one target storage area for cross-regional data migration.

[0036] A multi-factor analysis unit is used to perform multi-factor analysis on each of the target storage areas based on the data to be migrated, and to obtain the multi-factor analysis results for each of the target storage areas.

[0037] An optimization scheduling unit is used to perform optimization scheduling of the data to be migrated by combining divide-and-conquer and parallel genetic evolution based on the multi-factor analysis results of each target storage area, so as to obtain the optimal data migration scheme.

[0038] A data migration unit is used to migrate the data to be migrated to the at least one target storage area based on the optimal data migration scheme.

[0039] The present invention also provides an electronic device, the device comprising a processor and a memory:

[0040] The memory is used to store program code and transmit the program code to the processor;

[0041] The processor is configured to execute the data cross-region migration method as described above, according to instructions in the program code.

[0042] The present invention also provides a computer-readable storage medium for storing program code for performing the data cross-region migration method as described in any of the preceding claims.

[0043] As can be seen from the above technical solutions, the present invention has the following advantages:

[0044] This paper presents a method for cross-regional data migration. First, it obtains the data to be migrated and at least one target storage region for the cross-regional migration. Then, based on the data to be migrated, a multi-factor analysis is performed on each target storage region to obtain the multi-factor analysis results for each region. By comprehensively considering multi-factor analysis, the problem of requiring manual intervention in hot storage regions due to a single migration strategy can be solved. While ensuring data access needs are met, the impact of secondary manual intervention during the migration process on users is avoided, thereby reducing the impact and cost of data storage and migration, improving the resource utilization efficiency and economic benefits of the distributed cloud computing storage system, and enhancing user experience continuity. Then, based on the multi-factor analysis results of each target storage region, the data to be migrated is optimized and scheduled using a combination of divide-and-conquer and parallel genetic evolution to obtain the optimal data migration scheme. Finally, based on the optimal data migration scheme, the data to be migrated is migrated to at least one target storage region. Therefore, by using an improved genetic algorithm based on divide-and-conquer and parallel evolution to optimize the scheduling of data migration, the optimal data migration scheme can be obtained, achieving scientific and intelligent decision-making for data migration, significantly improving data migration efficiency, and making it suitable for different storage environments and data access needs. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a schematic diagram of the structure of a cross-regional data migration system;

[0047] Figure 2 A flowchart illustrating the steps of a method for cross-regional data migration;

[0048] Figure 3 This is a schematic diagram of a logical framework for multi-factor analysis;

[0049] Figure 4 This is a simplified schematic diagram of an optimized scheduling process;

[0050] Figure 5 This is a flowchart of a GA optimizer based on a genetic algorithm.

[0051] Figure 6 This is a schematic diagram of a logical framework for inter-region adjustments via a global coordinator.

[0052] Figure 7 This is a schematic diagram of a process for migrating business data across regions.

[0053] Figure 8 This is a schematic diagram illustrating the overall process of a method for cross-regional data migration.

[0054] Figure 9 This is a structural block diagram of a data migration device across regions. Detailed Implementation

[0055] This invention provides a method, apparatus, electronic device, and storage medium for cross-regional data migration, which solves or partially solves the technical problems of current cross-regional data migration strategies having limited considerations, low migration efficiency, and high operating costs.

[0056] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0057] As an example, in the process of building multi-AZ and unitized deployments, it is often necessary to migrate the business data of existing users across regions and select the region where the business data of new users is written. However, current cross-regional data migration strategies often consider only one factor. Specifically, they lack a comprehensive consideration of multiple factors such as user location, resource load, storage costs, network bandwidth costs, and data access requirements.

[0058] For example, some strategies focus solely on user location, migrating user data to the same storage region based on ownership, without considering that these regions may have high loads and resource contention. This leads to high data transmission latency and a poor user experience. Furthermore, secondary manual intervention during migration can cause short-term user unavailability, ultimately increasing overall costs instead of reducing them. Other strategies prioritize meeting data access needs, centrally migrating data to regions closer to users, but neglecting the higher storage costs and energy consumption in those regions. This prevents optimal resource allocation.

[0059] Furthermore, most current migration scheduling algorithms employ simple heuristics or fixed rules, lacking robust mathematical models to accurately calculate the optimal migration scheme. Consequently, when faced with complex multi-region storage environments and dynamically changing data access demands, it is difficult to formulate efficient and economical data migration strategies, resulting in high operating costs and low resource utilization of storage systems. Therefore, there is an urgent need for a method that comprehensively considers multiple factors and achieves efficient and low-cost cross-region migration of stored data through scientific mathematical models and optimization algorithms.

[0060] Therefore, one of the core inventive points of this invention is to provide a data cross-region migration method based on a genetic algorithm. Based on multi-factor analysis and an improved genetic algorithm, combined with a divide-and-conquer + parallel evolution approach, a precise mathematical model is constructed by comprehensively analyzing factors such as storage costs, network bandwidth costs, energy consumption, resource availability, and data access requirements in different regions. A corresponding migration optimization scheduling process is then designed based on the genetic algorithm to calculate the optimal data cross-region migration scheme.

[0061] By employing the technical solution of this invention, on the one hand, the problem of requiring manual intervention in hot storage areas caused by a single migration strategy can be solved. While ensuring that data access needs are met, the impact of secondary manual intervention during the migration process on users is avoided, thereby reducing the impact and cost of data storage and migration, improving the resource utilization efficiency and economic benefits of the distributed cloud computing storage system, and enhancing the continuity of user experience.

[0062] On the other hand, it can solve the problem of uneven resource storage load during the modular construction process aimed at data center disaster recovery and horizontal scaling. Based on data information collection, multi-factor intelligent analysis and optimized scheduling algorithms, it quantifies the data center weight of user data migration, formulates efficient and economical data migration strategies, solves the problem of large differences in the utilization rate of key resources (storage capacity) between different nodes, reduces operating costs and improves resource utilization.

[0063] Reference Figure 1 The diagram shows a schematic representation of a cross-regional data migration system provided by an embodiment of the present invention.

[0064] like Figure 1 As shown, a cross-regional data migration system based on a genetic algorithm mainly includes an application layer, a control layer, and a capability layer. The application layer includes various user application access terminals, such as mobile apps, browser-based systems (BSS), and mini-programs. Users can initiate user data migration requests through these application access terminals. These user data migration requests are then transmitted to the control layer.

[0065] The application scenario of the data cross-region migration method provided in this invention is mainly reflected in the control layer. The control layer mainly includes a data acquisition module, a multi-factor analysis module, an optimization scheduling algorithm module, and a migration execution module. The data acquisition module is used to collect indicators such as user location information, access frequency, and resource utilization of each storage area. The multi-factor analysis module is used to analyze and process the collected data to clarify the degree of influence of each factor on the data migration decision. The optimization scheduling algorithm module is used to solve the mathematical model based on a genetic algorithm combined with a divide-and-conquer + parallel evolution approach to find the optimal data migration scheme that meets the user's conditions. The migration execution module is used to execute the data cross-region migration operation according to the optimal data migration scheme obtained by the optimization scheduling algorithm module.

[0066] The capability layer connects with the control layer and mainly comprises a central data center and multiple unitized data centers. Each central data center and each unitized data center can contain business services and data storage areas. During data migration, the control layer, based on the user's data migration request, obtains the data to be migrated from the capability layer, collects core resources from each storage area, and, in conjunction with the control layer, realizes cross-regional data migration.

[0067] To enable those skilled in the art to better understand the technical solution of the present invention, the following is combined with... Figure 2 The main modules in the control layer of the data cross-region migration system mentioned in the foregoing embodiments are described in detail.

[0068] Reference Figure 2 The diagram illustrates a flowchart of a data cross-region migration method provided by an embodiment of the present invention, which may specifically include the following steps:

[0069] Step 201: Obtain the data to be migrated, and at least one target storage area for the cross-regional data migration;

[0070] The execution process of step 201 is mainly based on the data acquisition module. This module is primarily responsible for interacting with the big data platform and the automated operation and maintenance platform, collecting user information, access IPs, access log data from the big data platform, and storage area host information from the automated operation and maintenance platform via interfaces. The collected information includes, but is not limited to: user location information, access frequency, resource utilization of each storage area, storage costs, network bandwidth fees, and data access requirements.

[0071] In the specific implementation, the data acquisition module of the control layer can acquire the data to be migrated, as well as at least one target storage area for cross-regional data migration.

[0072] Step 202: Based on the data to be migrated, perform multi-factor analysis on each of the target storage areas to obtain the multi-factor analysis results for each target storage area;

[0073] The execution process of step 202 is mainly based on the multi-factor analysis module. That is, the multi-factor analysis module analyzes and processes the collected data to clarify the degree of influence of each factor on the data migration decision. Figure 3 A schematic diagram of the logical framework for multi-factor analysis is shown.

[0074] Combination Figure 3 Multifactor analysis can mainly include:

[0075] Location matching analysis. For example, if there are no restrictions on the location of the target storage area, the score is 0.5 points; a perfect match scores 1 point; a partial match scores 0.7 points; and a no match scores 0 points.

[0076] Storage resource utilization analysis. Specifically, this involves calculating the ratio of available storage in the target storage area to total storage.

[0077] Storage cost analysis. Specifically, for all target storage areas, the following function (storageCost - minCost) / (maxCost - minCost) is executed to perform storage cost analysis based on the maximum storage cost (maxCost), minimum storage cost (minCost), and storage cost (storageCost).

[0078] Network bandwidth analysis. Network bandwidth analysis combines the amount of data to be migrated and the unit data transfer cost. Specifically, based on the amount of data to be migrated and the traffic cost, the required bandwidth is determined, and the bandwidth cost is used as a decision factor in the decision calculation. For example, assuming a single device needs to migrate 10GB of data, public network traffic costs 0.15 yuan / GB (traffic prices vary by region), and the migration needs to be completed within 24 hours, the required bandwidth is approximately 10 / 24 * 8 ≈ 4Mbps, and the traffic cost is 0.15 * 10 = 1.5 yuan. The higher the traffic cost and bandwidth, the lower their weight in the decision calculation.

[0079] Data access requirements analysis. This analysis primarily calculates penalty values ​​based on the regional storage costs and energy consumption collected by the data acquisition module. For example, if the IP address of data D's access request originates from region A, but region A has higher storage costs and energy consumption than region B, then the weighted penalty value for migrating data D to region A will be higher than that for region B.

[0080] Finally, by weighted summation and integration of the analysis results from each factor, and incorporating a penalty mechanism (e.g., penalizing unbalanced loads or regions that are full by reducing the weight of the analysis results), a multi-factor analysis of user-stored data for each target storage region in the migration area is completed. That is, during the calculation process, the initial weights corresponding to the analysis results of each factor can be dynamically adjusted using a penalty mechanism to obtain the final multi-factor analysis results.

[0081] The weights of each factor (location, storage resources, cost, bandwidth, and access requirements) sum to 1. The weight values ​​vary depending on the region and data center configuration where each application is deployed (i.e., each weight value is a configurable item that can be manually adjusted based on the current regional situation). For example, the weight values ​​for APLLO configuration items differ depending on the service deployment region. Weight settings directly affect the specific scores of each factor analysis, thus influencing the decision-making results.

[0082] Based on the preceding content, in the specific implementation, the execution flow of performing multi-factor analysis on each target storage area based on the data to be migrated, and obtaining the multi-factor analysis results for each target storage area, may include: for each target storage area, based on the data to be migrated, performing location matching degree analysis, storage resource utilization analysis, storage cost analysis, data access demand analysis, and network bandwidth analysis on the target storage area, obtaining the location matching degree analysis results, storage resource utilization analysis results, storage cost analysis results, data access demand analysis results, and network bandwidth analysis results; combining a weighted penalty mechanism, the location matching degree analysis results, storage resource utilization analysis results, storage cost analysis results, data access demand analysis results, and network bandwidth analysis results are weighted and summed to obtain the multi-factor analysis results for the target storage area.

[0083] Step 203: Based on the multi-factor analysis results of each target storage area, perform optimized scheduling of the data to be migrated by combining divide-and-conquer and parallel genetic evolution to obtain the optimal data migration scheme;

[0084] The execution flow of step 203 is mainly implemented based on the optimized scheduling algorithm module. Specifically, it uses a genetic algorithm combined with a divide-and-conquer + parallel evolution approach for optimized scheduling. A simplified diagram of the optimized scheduling process using the divide-and-conquer + parallel evolution approach is shown below. Figure 4 As shown.

[0085] Combination Figure 4 For all users whose existing data needs to be migrated across regions, they are first grouped (e.g., Group 1 includes users 1-1000, Group 2 includes users 1001-2000, and so on). Next, a parallel genetic algorithm is used to optimize each group of users. Finally, the optimization results are merged. The merged optimization results can also be used for global coordination. Figure 5A flowchart of a GA (Genetic Algorithm) optimizer based on a genetic algorithm is shown. Figure 5 As shown, the process includes the following steps:

[0086] First, the `regionId` of the decision variable data storage area (i.e., each target storage area) is encoded to generate an initial population. The size of each user population is based on the target storage area. For example, if there are `m` target storage areas, the size of each user population is `m`. If there are `n` users in the current group, there are `m*n` populations. Second, the fitness function (i.e., objective function Z = fitness value f - multifactor analysis constraint r + load balancing degree S) is used to evaluate the quality of each individual. The multifactor analysis constraint `r` is the result of multifactor analysis for each target storage area. The load balancing degree S of the target storage area is 0 by default. In special cases, such as when the capacity is full, it is 1. When S=1, it indicates that the target storage area can no longer write data, and the area cannot be selected. Then, through genetic operations such as selection, crossover, and mutation, the population is iteratively optimized until the optimal solution for each user (i.e., the optimal migration sub-scheme for that user) is found. Finally, the output results of each GA optimizer are merged as the global optimization scheduling result, i.e., the globally optimal data migration scheme.

[0087] Furthermore, considering the real-time nature of data migration, new user (incremental user) data migration requests may be continuously received during the migration process. At this time, based on the currently calculated global optimization scheduling results, adjustments can be made between regions through the global coordinator with the aim of capacity balancing, to obtain the final data migration allocation scheme.

[0088] Figure 6 A schematic diagram of a logical framework for inter-region adjustments via a global coordinator is shown. Combined with... Figure 6 Inter-region adjustments mainly include periodically performing global optimization based on genetic algorithms and updating the decision cache. For new user data migration requests, a simple decision verification is performed based on the amount of data to be migrated. If the amount of data to be migrated is small, a simple decision is executed using a weighted scoring method. Otherwise, genetic algorithm optimization is required.

[0089] This allows for periodic global optimization, which, through pre-warming caching, can reduce the computational pressure on distributed genetic algorithms, improve the efficiency of computational models, and enhance the real-time migration of stored data across regions, thereby reducing user perception.

[0090] Based on the above, in the specific implementation, the execution flow of optimizing the data migration scheme by combining divide-and-conquer and parallel genetic evolution based on the multi-factor analysis results of each target storage area can include: first, extracting all users who need to migrate data across regions from the data to be migrated; then, grouping all users according to a preset grouping method to obtain multiple groups of users; and finally, performing parallel genetic optimization scheduling on the multiple groups of users based on the multi-factor analysis results of each target storage area to obtain the optimal data migration scheme.

[0091] Furthermore, based on the multi-factor analysis results of each target storage area, parallel genetic optimization scheduling is performed on multiple user groups to obtain the optimal data migration scheme. Specifically, this may include: for each user group, based on the multi-factor analysis results of each target storage area and combined with the load balancing degree of each target storage area, constructing the fitness function of each user in each target storage area; performing genetic optimization scheduling on the users in each group with the goal of maximizing the fitness function to obtain the optimal scheduling result of each user in each group, which serves as the optimal migration sub-scheme for the user; and merging the optimal migration sub-schemes of each user in each user group as the global optimal data migration scheme.

[0092] Step 204: Based on the optimal data migration scheme, migrate the data to be migrated to the at least one target storage area.

[0093] Step 204's execution process is primarily implemented based on the migration execution module. That is, according to the optimal data migration plan output by the optimized scheduling algorithm module, the cross-regional migration operation of business data is executed. The cross-regional migration process of business data is as follows: Figure 7 As shown.

[0094] Combination Figure 7 First, based on the previously saved user retrieval routes, business data is accurately retrieved from the original data storage area, and then copied to the target storage area using a batch-by-batch write mechanism. During the write process, the data transmission status, system energy consumption, and cost fluctuations are monitored in real time to ensure the stable progress of the migration task. Once the data write is complete, the data volume before and after the migration is immediately compared to complete the migration integrity verification.

[0095] Once the verification is successful, the business data migration is considered complete. The user retrieval route update is then executed. After the route update takes effect, user-initiated business data retrieval requests will automatically be directed to the migrated target storage area, ensuring a seamless switch in business access.

[0096] Finally, by creating a scheduled task to add a delayed deletion instruction for the original data, the storage space of the original data storage area is cleaned up after the target data is confirmed to be available, so as to achieve efficient resource recycling.

[0097] During the data migration process, if some data becomes unavailable (migration fails), the migration of that portion of data will be automatically re-executed. If multiple retries (e.g., 3 times) fail, an alert will be sent via email or SMS. Upon receiving the alert, technical personnel can manually verify the data.

[0098] For data migration integrity verification, depending on the amount of business data to be migrated, integrity verification can be performed after all business data has been written, or it can be performed after each batch of data has been written.

[0099] Specifically, when the volume of business data is small, integrity checks can be performed after all data has been written. When the volume of business data is large, integrity checks are performed immediately after each batch of data has been written. During data migration, if a migration anomaly occurs in a certain batch of data, the migration of only that batch of data needs to be automatically retried to improve migration efficiency.

[0100] Based on the foregoing, the optimal data migration scheme can include optimal migration sub-schemes for all users whose data needs to be migrated across regions. In the specific implementation, the execution process of migrating data to at least one target storage area based on the optimal data migration scheme can include: for each user whose data needs to be migrated across regions, retrieving the user's business data from the data to be migrated, and selecting a target storage area from at least one target storage area based on the optimal migration sub-scheme; writing the business data to the target storage area in batches and multiple times, and performing migration integrity verification; after passing the migration integrity verification, immediately performing a retrieval route update operation for the user, and deleting the business data stored in the original data storage area after a preset time.

[0101] Furthermore, the execution process of migration integrity verification can be as follows: the data volume of business data is judged based on a preset data volume threshold, and the data migration process of business data is verified for integrity based on the data volume judgment result to obtain the migration integrity verification result.

[0102] Furthermore, based on a preset data volume threshold, the data volume of the business data is determined, and based on the data volume determination result, the integrity of the business data migration process is verified to obtain the migration integrity verification result, which may specifically include:

[0103] First, determine whether the amount of business data is less than or equal to the preset data volume threshold.

[0104] In the first scenario, if the amount of business data is less than or equal to the preset data volume threshold, an integrity check can be performed again after all the business data has been written to the storage area to be migrated.

[0105] In the second scenario, the volume of business data exceeds a preset data volume threshold. In this case, an integrity check is performed after each batch of business data is written to the storage area to be migrated. Only after the integrity check of the current batch of data passes will the writing operation for the next batch of data be executed. It is understood that if the integrity check of the current batch of data fails, the migration of a portion of the current batch of data will be automatically re-executed, and the integrity check will be performed again. If multiple retries (e.g., 3 retries) still fail, an alert can be sent via email or SMS so that technical personnel can manually verify the data.

[0106] In any of the above scenarios, when confirming the execution of the integrity verification process, the data volume before and after the successful migration of the business data is compared. If the data volume before and after the migration is consistent, the current migration integrity verification is deemed to have passed. If the data volume before and after the migration is inconsistent, the current migration integrity verification is deemed to have failed.

[0107] In this embodiment of the invention, a method for cross-regional migration of stored data based on a genetic algorithm is provided, comprehensively considering multiple factors such as storage resources, storage costs, network bandwidth costs, energy consumption, and data access requirements. By comprehensively considering multiple factors and optimizing scheduling, the total cost of data storage and transmission can be effectively reduced, and the rational allocation and efficient utilization of storage resources can be achieved, avoiding secondary human intervention and improving the continuity of user experience. Furthermore, by solving the mathematical model based on the improved genetic algorithm, the optimal data migration scheme can be obtained, realizing scientific and intelligent decision-making for data migration, significantly improving data migration efficiency, and making it suitable for different storage environments and data access requirements.

[0108] For better explanation, refer to Figure 8 This diagram illustrates the overall flow of a data cross-region migration method provided by an embodiment of the present invention. It should be noted that this embodiment only provides a brief description of the general flow of data cross-region migration. The specific implementation process of each step can be understood by referring to the relevant content in the foregoing embodiments, and will not be elaborated upon here. It is understood that the present invention does not impose any limitations on this.

[0109] Step 801: Obtain the data to be migrated, and at least one target storage area for the cross-regional data migration;

[0110] Step 802: For each target storage area, based on the data to be migrated, perform location matching analysis, storage resource utilization analysis, storage cost analysis, data access demand analysis, and network bandwidth analysis on the target storage area respectively, and obtain the results of location matching analysis, storage resource utilization analysis, storage cost analysis, data access demand analysis, and network bandwidth analysis.

[0111] Step 803: Combining the weighted penalty mechanism, the results of the location matching degree analysis, storage resource utilization analysis, storage cost analysis, data access demand analysis, and network bandwidth analysis are integrated by weighted summation to obtain the multi-factor analysis results of the target storage area;

[0112] Step 804: Extract all users who need to migrate data across regions from the data to be migrated, and group all users according to a preset grouping method to obtain multiple groups of users;

[0113] Step 805: Based on the multi-factor analysis results of each target storage area, perform parallel genetic optimization scheduling for multiple groups of users to obtain the optimal migration sub-scheme for all users who need to migrate data across regions;

[0114] Step 806: For each user who needs to migrate data across regions, retrieve the user's business data from the data to be migrated, and select the storage area to be migrated from at least one target storage area based on the optimal migration sub-scheme;

[0115] Step 807: Write the business data to the storage area to be migrated in batches and perform migration integrity verification. After passing the migration integrity verification, immediately perform the retrieval route update operation for the user, and delete the business data stored in the original data storage area after a preset time.

[0116] Unlike current cross-storage area data migration methods, the technical solution provided by this invention is based on a genetic algorithm. It quantifies the weight of user data migration to different storage areas, provides the optimal data migration strategy for each user, and achieves load balancing on the resource side. Unlike current single-factor strategies, the technical solution provided by this invention is based on multi-factor analysis. Through optimization of the genetic algorithm, a mathematical model is constructed, and a corresponding migration algorithm is designed to calculate the optimal cross-storage data migration scheme.

[0117] In summary, compared with traditional technologies, the main advantages of the technical solution provided by this invention are:

[0118] Multi-factor analysis: This paper proposes an optimized scheduling method for cross-regional migration of storage data, which comprehensively considers multiple factors such as storage resources, storage costs, network bandwidth costs, energy consumption, and data access requirements. By comprehensively considering multiple factors and optimizing scheduling, it can avoid secondary human intervention, effectively reduce the total cost of data storage and transmission, and achieve the rational allocation and efficient utilization of storage resources, thereby achieving the goal of cost reduction and efficiency improvement.

[0119] Improved Genetic Algorithm: Based on the improved genetic algorithm, mathematical models are solved to obtain the optimal data migration scheme, realizing scientific and intelligent decision-making for data migration. It is applicable to different storage environments and data access needs, and improves the continuity of user experience.

[0120] The technical solution provided by this invention is applicable, on the one hand, to scenarios requiring balanced cross-regional migration of user business data and optimal selection of storage regions for data writing based on multi-factor decision-making during the construction of multi-AZ and unitized systems. On the other hand, it is also applicable to scenarios requiring improvement of resource utilization efficiency and economic benefits of distributed cloud computing storage systems based on multiple factors.

[0121] Reference Figure 9 The diagram illustrates a structural block diagram of a data cross-region migration device provided in an embodiment of the present invention, which may specifically include:

[0122] The data acquisition unit 901 is used to acquire the data to be migrated, and at least one target storage area for cross-regional data migration;

[0123] The multi-factor analysis unit 902 is used to perform multi-factor analysis on each of the target storage areas based on the data to be migrated, and obtain the multi-factor analysis results of each target storage area.

[0124] The optimization scheduling unit 903 is used to perform optimization scheduling of the data to be migrated by combining divide-and-conquer and parallel genetic evolution based on the multi-factor analysis results of each target storage area, so as to obtain the optimal data migration scheme.

[0125] The data migration unit 904 is used to migrate the data to be migrated to the at least one target storage area based on the optimal data migration scheme.

[0126] In one optional embodiment, the multi-factor analysis unit 902 is specifically used for:

[0127] For each target storage area, based on the data to be migrated, the target storage area is subjected to location matching degree analysis, storage resource utilization rate analysis, storage cost analysis, data access demand analysis, and network bandwidth analysis to obtain the location matching degree analysis results, storage resource utilization rate analysis results, storage cost analysis results, data access demand analysis results, and network bandwidth analysis results.

[0128] By combining the weighted penalty mechanism, and integrating the results of the location matching degree analysis, the storage resource utilization analysis, the storage cost analysis, the data access demand analysis, and the network bandwidth analysis by weighted summation, the multi-factor analysis results of the target storage area are obtained.

[0129] In one optional embodiment, the optimized scheduling unit 903 includes:

[0130] The user extraction unit is used to extract all users whose data needs to be migrated across regions from the data to be migrated.

[0131] The user grouping unit is used to group all users according to a preset grouping method to obtain multiple grouped users;

[0132] The parallel genetic optimization scheduling subunit is used to perform parallel genetic optimization scheduling on the multiple grouped users based on the multi-factor analysis results of each target storage area to obtain the optimal data migration scheme.

[0133] In one alternative embodiment, the parallel genetic optimization scheduling subunit includes:

[0134] The fitness function construction unit is used to construct the fitness function of each user in the group under each target storage area based on the multi-factor analysis results of each target storage area and the load balancing degree of each target storage area for each group user.

[0135] The genetic optimization scheduling subunit is used to perform genetic optimization scheduling on the group of users with the goal of maximizing the fitness function, and to obtain the optimal scheduling result for each user in the group of users, which is used as the optimal migration sub-scheme for the user.

[0136] The scheme merging unit is used to merge the optimal migration sub-schemes of each user in each of the aforementioned user groups as the global optimal data migration scheme.

[0137] In one optional embodiment, the optimal data migration scheme includes an optimal migration sub-scheme for all users who need to migrate data across regions; the data migration unit 904 includes:

[0138] The user business data extraction unit is used to retrieve the user's business data from the data to be migrated for each user who needs to migrate data across regions, and select the storage area to be migrated from the at least one target storage area based on the optimal migration sub-scheme.

[0139] The data writing and verification unit is used to write the business data to the storage area to be migrated in batches and multiple times, and to perform migration integrity verification.

[0140] The data update unit is used to immediately perform a retrieval route update operation for the user after passing the migration integrity verification, and delete the business data stored in the original data storage area after a preset time.

[0141] In one optional embodiment, the data writing and verification unit includes:

[0142] The migration integrity verification unit is used to determine the data volume of the business data based on a preset data volume threshold, and to perform integrity verification on the data migration process of the business data according to the data volume determination result, so as to obtain the migration integrity verification result.

[0143] In one optional embodiment, the migration integrity verification unit is specifically used for:

[0144] Determine whether the amount of the business data is less than or equal to a preset data amount threshold;

[0145] If so, then after all the business data is written to the storage area to be migrated, an integrity check will be performed again;

[0146] If not, an integrity check will be performed after each batch of business data is written to the storage area to be migrated, and the next batch of data will be written only after the integrity check of the current batch of data passes.

[0147] When the integrity verification process is confirmed to be executed, the data volume before and after the migration of the currently successfully migrated business data is compared.

[0148] If the data volume before and after the migration is consistent, the current migration integrity verification is considered to have passed.

[0149] If the data volume before and after the migration is inconsistent, the current migration integrity check is deemed to have failed.

[0150] As the device embodiment is basically similar to the method embodiment, it is described in a relatively simple way. For relevant details, please refer to the description of the method embodiment above.

[0151] This invention also provides an electronic device, which includes a processor and a memory:

[0152] The memory is used to store program code and transfer the program code to the processor;

[0153] The processor is used to execute the data cross-region migration method of any embodiment of the present invention according to the instructions in the program code.

[0154] This invention also provides a computer-readable storage medium for storing program code for executing the data cross-region migration method of any embodiment of this invention.

[0155] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0156] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units through some interfaces, and may be electrical, mechanical, or other forms.

[0157] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0158] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0159] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0160] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data cross-region migration method, characterized in that, The method comprises the following steps: acquiring to-be-migrated data and at least one target storage area for cross-region data migration; performing multi-factor analysis on each target storage area based on the to-be-migrated data to obtain a multi-factor analysis result of each target storage area; performing combined divide-and-conquer and parallel genetic evolution optimization scheduling on the to-be-migrated data based on the multi-factor analysis result of each target storage area to obtain an optimal data migration scheme; migrating the to-be-migrated data to the at least one target storage area based on the optimal data migration scheme.

2. The data cross-zone migration method according to claim 1, characterized in that, The multi-factor analysis on each target storage area based on the to-be-migrated data to obtain a multi-factor analysis result of each target storage area comprises the following steps: for each target storage area, performing local matching degree analysis, storage resource utilization rate analysis, storage cost analysis, data access demand analysis, and network bandwidth analysis on the target storage area based on the to-be-migrated data to obtain a local matching degree analysis result, a storage resource utilization rate analysis result, a storage cost analysis result, a data access demand analysis result, and a network bandwidth analysis result; integrating the local matching degree analysis result, the storage resource utilization rate analysis result, the storage cost analysis result, the data access demand analysis result, and the network bandwidth analysis result by weighted summation to obtain the multi-factor analysis result of the target storage area by combining a weight penalty mechanism.

3. The data cross-zone migration method of claim 1, wherein, The combined divide-and-conquer and parallel genetic evolution optimization scheduling on the to-be-migrated data based on the multi-factor analysis result of each target storage area to obtain an optimal data migration scheme comprises the following steps: extracting all users of the to-be-migrated data that need to be migrated cross-regionally; grouping the all users according to a preset grouping manner to obtain a plurality of grouped users; performing parallel genetic optimization scheduling on the plurality of grouped users based on the multi-factor analysis result of each target storage area to obtain an optimal data migration scheme.

4. The data cross-zone migration method according to claim 3, characterized in that, The parallel genetic optimization scheduling on the plurality of grouped users based on the multi-factor analysis result of each target storage area to obtain an optimal data migration scheme comprises the following steps: for each grouped user, constructing an fitness function of each user in the grouped user under each target storage area based on the multi-factor analysis result of each target storage area and the load balancing degree of each target storage area; performing genetic optimization scheduling on the grouped user to maximize the fitness function to obtain an optimal scheduling result of each user in the grouped user as an optimal migration sub-scheme of the user; merging the optimal migration sub-scheme of each user in each grouped user as a global optimal data migration scheme.

5. The data cross-zone migration method of claim 1, wherein, The optimal data migration scheme comprises optimal migration sub-schemes of all users of the to-be-migrated data that need to be migrated cross-regionally. The migrating the to-be-migrated data to the at least one target storage area based on the optimal data migration scheme comprises the following steps: For each of the users requiring cross-region migration of data, service data of the user is retrieved from the data to be migrated, and based on the optimal migration sub-scheme, a to-be-migrated storage area is selected from the at least one target storage area; The service data is written to the to-be-migrated storage area in a batched and multiple-time manner, and migration integrity verification is performed; When passing the migration integrity verification, a retrieval routing update operation is immediately performed for the user, and the service data stored in the original data storage area is deleted after a preset time length.

6. The data cross-zone migration method according to claim 5, characterized in that, The execution process of the migration integrity verification includes: Based on a preset data amount threshold, the data amount of the service data is judged, and according to the data amount judgment result, the data migration process of the service data is integrity-verified to obtain a migration integrity verification result.

7. The data cross-zone migration method of claim 6, wherein, The judgment of the data amount of the service data based on the preset data amount threshold and the integrity verification of the data migration process of the service data according to the data amount judgment result to obtain the migration integrity verification result includes: Judging whether the data amount of the service data is less than or equal to the preset data amount threshold; If yes, after all the data of the service data is written to the to-be-migrated storage area, integrity verification is performed again; If no, after each batch of data of the service data is written to the to-be-migrated storage area, integrity verification is performed once, and after the integrity verification of the current batch of data passes, the writing operation of the next batch of data is performed; When it is confirmed that the integrity verification process is performed, the data amount before and after the current successful migration of the service data is compared; If the data amount before and after the migration is consistent, it is determined that the current migration integrity verification passes; If the data amount before and after the migration is inconsistent, it is determined that the current migration integrity verification fails.

8. A data cross-area migration apparatus characterized by comprising: It includes: A data acquisition unit is configured to acquire to-be-migrated data and at least one target storage area for cross-region migration of data; A multi-factor analysis unit is configured to perform multi-factor analysis on each of the target storage areas based on the to-be-migrated data to obtain multi-factor analysis results of each of the target storage areas; An optimization scheduling unit is configured to perform optimization scheduling on the to-be-migrated data based on the multi-factor analysis results of each of the target storage areas by combining divide-and-conquer and parallel genetic evolution to obtain an optimal data migration scheme; A data migration unit is configured to migrate the to-be-migrated data to the at least one target storage area based on the optimal data migration scheme.

9. An electronic device, comprising: The device includes a processor and a memory: The memory is configured to store program code and transmit the program code to the processor; The processor is configured to execute the cross-region migration method of any one of claims 1-7 according to instructions in the program code.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is configured to store program code for executing the cross-region migration method of any one of claims 1-7. The computer-readable storage medium is configured to store program code for executing the cross-region migration method of any one of claims 1-7.