Data migration method and device and storage medium
By selecting cold data with a longer survival time and combining the deletion timestamp prediction model and global scheduling equipment, the problem of low returns on moving cold data in the prior art is solved, and efficient utilization and cost optimization of storage resources are achieved.
Patent Information
- Application Number
- CN202410616506.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-08
- Filing Date
- 2024-05-14
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, when moving cold data with low access popularity from a high-cost storage system to a low-cost storage system, the benefits are lower, especially because cold data with short remaining survival time is quickly deleted after the transfer, resulting in high network resource consumption.
By selecting cold data with low access popularity and the longest remaining survival time or cold data with exceeding the duration threshold, it is moved from a high-cost storage system to a low-cost storage system, and using the deletion timestamp prediction model or data deletion rules to calculate the remaining survival time, combined with the global scheduling equipment configuration quota, the migration process is optimized.
It improves the benefits of moving data, optimizes the utilization rate of storage resources, reduces the consumption of network resources, and reduces the negative impact of cost differences.
Smart Images

Figure CN120447823A_ABST
Abstract
Description
[0001] This application claims priority to Chinese patent application No. 202410178074.8, filed on February 8, 2024, entitled “A method, device and other equipment for data processing”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present application relates to the field of storage, and in particular to a method, device and storage medium for moving data. Background Art
[0003] Cloud storage uses distributed storage, storing data across multiple regional storage systems. Due to the varying locations of each region, some regions have lower data storage costs, while others have higher costs. For example, a region with abundant electricity may have lower data storage costs.
[0004] In related technologies, cloud storage vendors have introduced archival storage services. This service allows for a high-cost first region and a low-cost second region to identify less-popular cold data from the data stored in a first storage system in the first region and move this cold data from the first storage system to a second storage system in the second region. This allows the more popular hot data to remain in the first storage system while the less-popular cold data is stored in the second storage system, improving the cost-effectiveness of data storage.
[0005] For some cold data stored in the first storage system, the benefit of moving all of the cold data from the first storage system to the second storage system is low. Summary of the Invention
[0006] This application provides a method, device, and storage medium for migrating data to increase the benefits of migrating data. The technical solution is as follows:
[0007] In a first aspect, the present application provides a method for migrating data, the method being applied to a management device in a first storage system. In the method, at least one cold data item whose access popularity satisfies a cold data condition is determined from multiple data items stored in the first storage system. Based on the remaining lifetime of each cold data item in the at least one cold data item, at least one target cold data item with the longest remaining lifetime, or whose remaining lifetime exceeds a lifetime threshold, is selected from the at least one cold data item. The at least one target cold data item is then migrated from the first storage system to a second storage system.
[0008] If cold data with a short remaining lifetime is moved from the first storage system to the second storage system, it will be quickly deleted from the second storage system, resulting in a low benefit from moving the cold data with a short remaining lifetime. However, since at least one target cold data item with the longest remaining lifetime, or whose remaining lifetime exceeds a threshold, is selected from the at least one cold data item based on the remaining lifetime of each cold data item, the benefit from moving the at least one target cold data item to the second storage system is increased.
[0009] In one possible implementation, a time-to-live calculation policy is received from a policy management device. The policy management device includes at least one time-to-live calculation policy, which is used to calculate the remaining time-to-live of cold data. The remaining time-to-live of each cold data item in the at least one cold data item is obtained based on the time-to-live calculation policy. This allows the time-to-live calculation policy to be configured on the policy management device and centrally distributed to storage systems requiring data migration, thereby improving the efficiency of distributing the time-to-live calculation policies.
[0010] In another possible implementation, the survival time calculation strategy uses a deletion timestamp prediction model. Attribute information for each cold data item is obtained. Based on the deletion timestamp prediction model and the attribute information for each cold data item, the deletion timestamp for each cold data item is determined. Based on the deletion timestamp for each cold data item, the remaining survival time for each cold data item is determined. In this way, the deletion timestamp prediction model can be used to determine the deletion timestamp for each cold data item, ensuring that the deletion timestamp can be determined for all cold data, regardless of whether the data is preserved or not.
[0011] In another possible implementation, the survival time calculation strategy is a data deletion rule. This data deletion rule is used to calculate the remaining survival time of each cold data item based on its metadata. Based on the data deletion rule, the management device obtains metadata for each cold data item, which includes a deletion timestamp for each cold data item. Based on the deletion timestamp of each cold data item, the management device determines the remaining survival time of each cold data item. This allows the remaining survival time of each cold data item to be calculated based on the deletion timestamp in the metadata, simplifying the process of calculating the remaining survival time.
[0012] In another possible implementation, a first quota is received from a global scheduling device, where the first quota indicates the total amount of data that can be moved from the first storage system during a current migration cycle. When the total amount of the at least one target cold data and the total amount of data already moved from the first storage system during the current migration cycle do not exceed the first quota, the at least one target cold data is moved from the first storage system to the second storage system. This allows the global scheduling device to centrally allocate quotas to storage systems that need to migrate data, thereby improving quota allocation efficiency.
[0013] In another possible implementation, statistical information is sent to the global scheduling device. The statistical information includes one or more of the following information: the total amount of data moved from the first storage system during the previous migration cycle, or the average lifetime of cold data moved from the first storage system during the previous migration cycle. The first quota is obtained by the global scheduling device by adjusting the second quota based on the statistical information. The second quota indicates the total amount of data that can be moved from the first storage system during the previous migration cycle. This allows the quota allocated to each storage system to be dynamically adjusted based on the statistical information within each cycle, ensuring that the total network resources consumed by moving data within each cycle do not fluctuate significantly.
[0014] In another possible implementation, the cost of storing data in the first storage system is higher than the cost of storing data in the second storage system, or the price of the electric energy used to supply power to the first storage system is higher than the price of the electric energy used to supply power to the second storage system, or the storage performance of the storage medium included in the first storage system is higher than the storage performance of the storage medium included in the second storage system.
[0015] In another possible implementation, the first storage system and the second storage system are two cloud storage systems.
[0016] In another possible implementation, the first storage system is located in a first region, and the second storage system is located in a second region, where the first region and the second region are two different administrative regions.
[0017] In another possible implementation, the first storage system and the second storage system are storage systems located in two computer rooms.
[0018] In another possible implementation, the migration condition is to select at least one cold data with the longest remaining survival time, or the migration condition is to select at least one cold data with a remaining survival time exceeding a time threshold.
[0019] In a second aspect, the present application provides a device for moving data, configured to execute the method in the first aspect or any possible implementation of the first aspect. Specifically, the device includes a unit for executing the method in the first aspect or any possible implementation of the first aspect.
[0020] In a third aspect, the present application provides a computing device cluster, the computing device cluster comprising at least one computing device, each computing device comprising a processor and a memory;
[0021] The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method in the first aspect or any possible implementation manner of the first aspect.
[0022] In a fourth aspect, the present application provides a computer program product comprising instructions, which, when executed by a computing device cluster, causes the computing device cluster to execute the method in the first aspect or any possible implementation of the first aspect.
[0023] In a fifth aspect, the present application provides a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method in the first aspect or any possible implementation of the first aspect.
[0024] In a sixth aspect, the present application provides a chip comprising a memory and a processor, wherein the memory is used to store computer instructions, and the processor is used to call and run the computer instructions from the memory to execute the method in the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a schematic diagram of a network architecture provided by an embodiment of the present application;
[0026] Figure 2 This is a schematic diagram of another network architecture provided by an embodiment of the present application;
[0027] Figure 3 This is a schematic diagram of another network architecture provided by an embodiment of the present application;
[0028] Figure 4 This is a schematic diagram of another network architecture provided by an embodiment of the present application;
[0029] Figure 5 This is a flow chart of a method for moving data provided by an embodiment of the present application;
[0030] Figure 6This is a flow chart of a method for training a deletion timestamp prediction model provided by an embodiment of the present application;
[0031] Figure 7 This is a flow chart of another method for moving data provided in an embodiment of the present application;
[0032] Figure 8 This is a schematic diagram of the structure of a device for moving data provided in an embodiment of the present application;
[0033] Figure 9 is a structural diagram of a computing device provided in an embodiment of the present application;
[0034] Figure 10 This is a schematic diagram of a data processing cluster structure provided by an embodiment of the present application;
[0035] Figure 11 This is another schematic diagram of a data processing cluster structure provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] See also Figure 1 An embodiment of the present application provides a network architecture 100, which includes multiple storage systems. For example, the network architecture 100 includes a first storage system 101 and a second storage system 102.
[0037] Optionally, the network architecture 100 may also include more storage systems. For example, the network architecture 100 may also include a third storage system 103 .
[0038] The second storage system 102 can communicate with the first storage system 101, and the second storage system 102 can communicate with the third storage system 103. The cost of storing data in the first storage system 101 is higher than the cost of storing data in the second storage system 102, and the cost of storing data in the third storage system 103 is higher than the cost of storing data in the second storage system 102.
[0039] The cost of storing data in the first storage system 101 is high. For at least one piece of cold data in the first storage system 101 that has low access popularity and a long remaining lifetime, this at least one piece of cold data remains unused in the first storage system 101 for a long time, occupying valuable storage resources in the first storage system 101. This at least one piece of cold data can be moved from the first storage system 101 to the second storage system 102. Since the cost of storing data in the second storage system 102 is lower than that of the first storage system 101, storing this at least one piece of cold data in the second storage system 102 frees up more storage resources in the first storage system 101 to store hot data with high access popularity, thereby improving the effective utilization of storage resources in the first storage system 101.
[0040] Similarly, the cost of storing data in the third storage system 103 is high. Cold data with low access popularity and long remaining survival time in the third storage system 103 can also be moved to the second storage system 102, so that the third storage system 103 can free up more storage resources to store hot data with high access popularity, thereby improving the effective utilization of storage resources in the third storage system 103.
[0041] If cold data with a short remaining lifespan is moved to the second storage system 102, it will be quickly deleted from the second storage system 102. This process consumes network resources, resulting in low returns from moving the cold data. Therefore, cold data with a longer remaining lifespan in the first storage system 101 and cold data with a longer remaining lifespan in the third storage system 103 are moved to the second storage system 102. This allows the cold data to be stored in the second storage system 102 for a longer period of time after being moved, thereby increasing the returns from moving the cold data.
[0042] In some embodiments, the price of the electric energy supplied to the first storage system 101 is higher than the price of the electric energy supplied to the second storage system 102. And / or, the price of the electric energy supplied to the third storage system 103 is higher than the price of the electric energy supplied to the second storage system 102.
[0043] For example, the first storage system 101 is located in the first region, the second storage system 102 is located in the second region, the first region and the second region are two different administrative regions, and the price of electric energy in the first region is higher than that in the second region, so that the price of electric energy supplied to the first storage system 101 is lower than that of electric energy supplied to the second storage system 102. And / or, the third storage system 103 is located in the third region, the third region and the second region are two different administrative regions, and the price of electric energy in the third region is higher than that in the second region, so that the price of electric energy supplied to the third storage system 103 is lower than that of electric energy supplied to the second storage system 102.
[0044] In some embodiments, the storage performance of the storage medium included in the first storage system 101 is higher than the storage performance of the storage medium included in the second storage system 102. And / or, the storage performance of the storage medium included in the third storage system 103 is higher than the storage performance of the storage medium included in the second storage system 102.
[0045] For example, the storage medium included in the first storage system 101 is a solid-state drive, while the storage medium included in the second storage system 102 is a magnetic disk. The price of a solid-state drive is much higher than that of a magnetic disk, and the read and write speeds of a solid-state drive are much higher than those of a magnetic disk. As a result, the storage performance of the storage medium included in the first storage system 101 is higher than the storage performance of the storage medium included in the second storage system 102. And / or, the storage medium included in the third storage system 103 is also a solid-state drive, so the storage performance of the storage medium included in the third storage system 103 is higher than the storage performance of the storage medium included in the second storage system 102.
[0046] In some embodiments, the first storage system and the second storage system are two different cloud storage systems, and / or the third storage system and the second storage system are two different cloud storage systems. The first region and the second region are two different administrative regions, and / or the third region and the second region are two different administrative regions.
[0047] In this case, the network architecture 100 can be a private cloud storage system or a public cloud storage system.
[0048] In some embodiments, the first storage system 101 and the second storage system 102 are storage systems located in two different computer rooms, or the first storage system 101 and the second storage system 102 are two different storage systems located in the same computer room. The third storage system 103 and the second storage system 102 are storage systems located in two different computer rooms, or the third storage system 103 and the second storage system 102 are two different storage systems located in the same computer room.
[0049] In this case, the network architecture 100 can be an offline heterogeneous storage system, the storage performance of the storage medium of the first storage system 101 is higher than the storage performance of the storage medium of the second storage system 102, and the storage performance of the storage medium of the third storage system 103 is higher than the storage performance of the storage medium of the second storage system 102.
[0050] In some embodiments, see Figure 2 The first storage system 101 includes a first management device 1011. The first management device 1011 can obtain at least one cold data with low access popularity and long remaining survival time in the first storage system 101, and move the at least one cold data from the first storage system 101 to the second storage system 102.
[0051] Optionally, the first storage system 101 further includes at least one first storage device 1012. Data stored in the first storage system 101 is stored in the at least one first storage device 1012. The second storage system 102 includes at least one second storage device 1021. Saving the at least one cold data in the second storage system 102 is storing the at least one cold data in the at least one second storage device 1021. In other words, the first management device 1011 moves the at least one cold data from the at least one first storage device 1012 to the at least one second storage device 1021.
[0052] In some embodiments, see Figure 2 The third storage system 103 includes a second management device 1031. The second management device 1031 can obtain at least one piece of cold data with low access popularity and a long remaining lifetime from the third storage system 103, and move the at least one piece of cold data from the third storage system 103 to the second storage system 102. Optionally, the third storage system 103 also includes at least one third storage device 1032. The data stored in the third storage system 103 is: the data stored in the at least one third storage device 1032. In other words, the second management device 1031 moves the at least one piece of cold data from the at least one third storage device 1032 to the at least one second storage device 1021.
[0053] In some embodiments, the first management device 1011 may be a mover device in the first storage system 101 , and the second management device 1031 may be a mover device in the third storage system 103 .
[0054] In some embodiments, see Figure 3The network architecture 100 also includes a policy management device 104, which can communicate with the management device in the storage system that needs to move data included in the network architecture 100. Optionally, the policy management device 104 can send a survival time calculation policy to the management device in the storage system that needs to move data included in the network architecture 100, and the survival time calculation policy defines a method for calculating the remaining survival time of the data to be calculated. For example, the policy management device 104 can send a survival time calculation policy to the first management device 1011 in the first storage system 101 and / or the second management device 1031 in the third storage system 103.
[0055] In this way, the first management device 1011 receives the survival time calculation policy, obtains at least one piece of cold data with low access popularity and a long remaining survival time from the first storage system 101 based on the survival time calculation policy, and moves the at least one piece of cold data from the first storage system 101 to the second storage system 102. And / or, the second management device 1031 receives the survival time calculation policy, obtains at least one piece of cold data with low access popularity and a long remaining survival time from the third storage system 103 based on the survival time calculation policy, and moves the at least one piece of cold data from the third storage system 103 to the second storage system 102.
[0056] In some embodiments, the policy management device 104 includes at least one survival time calculation policy, and the survival time calculation policy sent by the policy management device 104 includes the at least one survival time calculation policy, or part of the survival time calculation policy in the at least one survival time calculation policy.
[0057] In some embodiments, see Figure 3 At least one policy card is plugged into the policy management device 104, and each policy card corresponds to a different survival time calculation policy, so that the policy management device 104 includes the at least one survival time calculation policy. The survival time calculation policy sent by the policy management device 104 includes the survival time calculation policy corresponding to the at least one policy card, or the survival time calculation policy corresponding to some of the at least one policy card. In this way, by plugging different policy cards into the policy management device 104, technicians can flexibly control the policy management device 104 to send different survival time calculation policies to the management device included in the storage system in the network architecture 100 that needs to move data, making it convenient for technicians to flexibly configure the survival time calculation policy.
[0058] Optionally, the policy management device 104 may be a scoring policy maker or other device.
[0059] In some embodiments, see Figure 4The network architecture 100 further includes a global scheduling device 105, which can communicate with a management device in a storage system in the network architecture 100 that needs to move data. Optionally, the global scheduling device 105 can configure a first quota for a storage system in the network architecture 100 that needs to move data, and send the first quota of the storage system to the management device of the storage system. The first quota refers to the total amount of data that can be moved from the storage system within a migration cycle.
[0060] For example, the global scheduling device 105 can configure a quota for the first storage system 101 and send the quota of the first storage system 101 to the first management device 1011 included in the first storage system 101. The global scheduling device 105 can also configure a quota for the third storage system 103 and send the quota of the third storage system 103 to the second management device 1031 included in the third storage system 103. The quota of the first storage system 101 indicates the total amount of data that can be migrated from the first storage system 101 during the current migration cycle. The quota of the third storage system 103 indicates the total amount of data that can be migrated from the third storage system during the current migration cycle.
[0061] The first management device 1011 receives the quota of the first storage system 101 and obtains at least one cold data item with low access popularity and a long remaining lifetime from the first storage system 101. If the total amount of the at least one cold data item and the total amount of data already migrated from the first storage system 101 during the current migration cycle does not exceed the quota, the at least one cold data item is migrated from the first storage system 101 to the second storage system 102.
[0062] The second management device 1031 receives the quota of the third storage system 103 and obtains at least one cold data item with low access popularity and a long remaining lifetime from the third storage system 103. If the total amount of the at least one cold data item and the total amount of data already migrated from the third storage system 103 during the current migration cycle does not exceed the quota, the at least one cold data item is migrated from the third storage system 103 to the second storage system 102.
[0063] In some embodiments, the global scheduling device 105 may be a global scheduler, etc. The policy management device 104 and the global scheduling device 105 may be two different devices, or the policy management device 104 and the global scheduling device 105 may be integrated into one device.
[0064] See also Figure 5 The present invention provides a method 500 for moving data. The method 500 can be applied to Figure 1 、 Figure 2 or Figure 3The method 500 is described by taking the migration of cold data from the first storage system to the second storage system as an example. Similarly, the method 500 can also migrate cold data from other storage systems (such as the third storage system) to the second storage system, which will not be listed one by one. Figure 5 As shown, the method 500 includes the following process.
[0065] Step 501: The first management device obtains m data whose access popularity meets the cold data condition from the data stored in the first storage system as cold data, where m is an integer greater than or equal to 1.
[0066] In some embodiments, the first storage system uses a storage queue to store data. When data in the storage queue is accessed, the first storage system moves the data to the head of the storage queue. Therefore, the closer the data is to the head of the storage queue, the more popular it is, and the closer the data is to the tail of the storage queue, the less popular it is.
[0067] In this case, the cold data condition means that the m data closest to the end of the storage queue are cold data. In other words, the first management device can select the m data closest to the end of the storage queue as the m cold data.
[0068] In some embodiments, each data item stored in the first storage system has metadata. For each data item stored in the first storage system, the metadata includes information such as the timestamp of the data item's last access or the timestamp of each access to the data item. The metadata for the data item may be stored in the first storage system, in the first management device, or in another device.
[0069] In this case, the first management device may obtain m cold data from the first storage system through the following method 1 or method 2.
[0070] Method 1: The first management device obtains a first least recently used (LRU) rule. The first LRU rule defines cold data as data that has not been accessed within a target time period, where the target time period is a time period closest to the current time period and has a specified duration. Based on the first LRU rule and metadata of multiple data stored in the first storage system, m data that have not been accessed within the target time period are obtained as cold data.
[0071] For example, if the specified duration is 30 days and the target time period is the most recent 30 days, the cold data condition defined by the first LRU rule is that data that has not been accessed within the most recent 30 days is considered cold data. In other words, data that has not been accessed for more than 30 days is defined as cold data.
[0072] Optionally, in method one, the first management device determines the target time period based on the first LRU rule, obtains the metadata of each data stored in the first storage system, obtains the last access timestamp of each data from the metadata of each data, and based on the last access timestamp of each data, obtains m data that have not been accessed within the target time period as cold data.
[0073] The first storage system stores a large amount of data, which requires a lot of time and computing resources to obtain the most recent access timestamp of each data. To this end, the first management device randomly samples from the first storage system and obtains cold data from the sampled data. The detailed implementation process is as follows.
[0074] Optionally, in method one, the first management device determines a target time period based on the first LRU rule, randomly samples n data from the data stored in the first storage system, where n is an integer greater than m, obtains metadata for each of the n data, obtains the last access timestamp of each data from the metadata of each data, and based on the last access timestamp of each data, obtains m data that have not been accessed within the target time period from the n data as cold data.
[0075] Method 2: The first management device obtains a second LRU rule. The second LRU rule defines a cold data condition where m pieces of data with the lowest access frequency within a target time period are considered cold data, or data with an access frequency below a frequency threshold is considered cold data. The target time period is a time period closest to the current time period with a specified duration. Based on the second LRU rule and metadata of multiple pieces of data stored in the first storage system, m pieces of data that meet the cold data condition within the target time period are obtained as cold data.
[0076] For example, assuming the specified duration is 30 days and the target time period is the 30 days closest to the current time, the cold data condition defined by the second LRU rule is that the m data with the lowest access frequency within the 30 days closest to the current time are cold data, or the data with an access frequency lower than the threshold within the 30 days closest to the current time are cold data.
[0077] Optionally, in method two, the first management device determines the target time period based on the second LRU rule, obtains the metadata of each data stored in the first storage system, and based on the timestamp of each access included in the metadata of each data, counts the access frequency of each data within the target time period, and selects m data with the lowest access frequency from each data as cold data or selects m data with an access frequency lower than a frequency threshold as cold data.
[0078] The first storage system stores a large amount of data, which results in a large amount of time and computing resources required to obtain the access frequency of each data. Therefore, random sampling is performed from the first storage system to obtain cold data from the sampled data. The detailed implementation process is as follows.
[0079] Optionally, in method two, the first management device determines the target time period based on the second LRU rule, randomly samples n data from the data stored in the first storage system, where n is an integer greater than m, obtains metadata for each of the n data, and based on the timestamp of each access included in the metadata of each data, counts the access frequency of each data within the target time period, and based on the access frequency of each data, selects m data with the lowest access frequency from the n data as cold data or selects m data with an access frequency lower than a frequency threshold as cold data.
[0080] In some embodiments, the first LRU rule or the second LRU rule is a rule stored in the first management device, or the first LRU rule or the second LRU rule is received by the first management device from the policy management device.
[0081] Step 502: The first management device obtains the remaining survival time of each cold data in the m cold data.
[0082] In step 502, the first management device obtains a survival time calculation policy that defines a method for calculating the remaining survival time of data and obtains the remaining survival time of each of the m cold data based on the survival time calculation policy.
[0083] Optionally, the first management device stores a survival time calculation policy, and the first management device obtains the locally stored survival time calculation policy.
[0084] Optionally, the first management device receives the survival time calculation policy sent by the policy management device.
[0085] In some embodiments, the policy management device includes at least one survival time calculation policy, and sends the at least one survival time calculation policy or a portion of the at least one survival time calculation policy to the first management device.
[0086] In some embodiments, the policy management device is plugged into at least one policy card, each policy card corresponding to a different time-to-live calculation policy. The policy management device obtains at least one time-to-live calculation policy corresponding to the at least one policy card, or obtains time-to-live calculation policies corresponding to some of the at least one policy card. The policy management device then sends the obtained time-to-live calculation policy to the first management device.
[0087] Optionally, for each policy card, the policy card stores a survival time calculation policy corresponding to the policy card, and the policy management device may obtain the storage survival time calculation policy corresponding to the policy card from the policy card.
[0088] Optionally, the policy management device stores the survival time calculation policy corresponding to each policy card. For each policy card, the policy card stores identification information of the policy card. The policy management device can obtain the identification information of the policy card from the policy card and, based on the identification information of the policy card, locally obtain the survival time calculation policy corresponding to the policy card. This eliminates the need for the policy card to include a large amount of storage space to store the survival time calculation policy corresponding to the policy card, thereby reducing the cost of the policy card.
[0089] Technicians can configure the survival time calculation policy sent by the policy management device to the management device of the storage system to be migrated in the network architecture by inserting and removing the policy card on the policy management device, thereby improving the flexibility of configuring the survival time calculation policy and facilitating the use of technicians.
[0090] In some embodiments, the survival time calculation policy obtained by the first management device is a deletion timestamp prediction model or a data deletion rule. The deletion timestamp prediction model is used to predict the remaining survival time of data, and the data deletion rule defines how to calculate the remaining survival time of the data to be calculated based on metadata of the data to be calculated, where the metadata of the data to be calculated includes a deletion timestamp of the data to be calculated.
[0091] Optionally, the data deletion rule definition needs to obtain a deletion timestamp of the data to be calculated from metadata of the data to be calculated, and obtain a remaining survival time of the data to be calculated based on the deletion timestamp of the data to be calculated.
[0092] For the sake of convenience, for each of the m cold data, the cold data is referred to as first cold data. The first management device can obtain the remaining survival time of the first cold data through the following method 1 and method 2.
[0093] Method 1: The survival time calculation strategy obtained by the first management device is a deletion timestamp prediction model. The first management device obtains the attribute information of the first cold data, obtains the deletion timestamp of the first cold data based on the deletion timestamp prediction model and the attribute information of the first cold data, and obtains the remaining survival time of the first cold data based on the deletion timestamp of the first cold data.
[0094] In some embodiments, the attribute information of the first cold data may include one or more of the following information: size, type, storage path, user identification information, or creation timestamp of the first cold data. Optionally, the metadata of the first cold data includes the attribute information of the first cold data.
[0095] In method 1, the first management device obtains metadata for the first cold data, obtains attribute information for the first cold data from the metadata, and inputs the attribute information for the first cold data into a deletion timestamp prediction model. The deletion timestamp prediction model receives the attribute information for the first cold data and, based on the attribute information, infers the deletion timestamp for the first cold data. The first management device obtains the deletion timestamp for the first cold data output by the deletion timestamp prediction model and, based on the current timestamp and the deletion timestamp of the first cold data, obtains the remaining lifetime of the first cold data.
[0096] Optionally, the deletion timestamp prediction model can be obtained by training an artificial intelligence (AI) model in advance. The detailed training process will be described in detail in subsequent content and will not be introduced in detail here.
[0097] Method 2: The survival time calculation strategy obtained by the first management device is a data deletion rule. Based on the data deletion rule, the first management device obtains the metadata of the first cold data, which includes the deletion timestamp of the first cold data. Based on the deletion timestamp of the first cold data, the first management device obtains the remaining survival time of the first cold data.
[0098] Optionally, the metadata of the first cold data may include a deletion timestamp of the first cold data. This deletion timestamp may be configured by the user to whom the first cold data belongs when creating the first cold data. Thus, in approach 2, the first management device obtains the metadata of the first cold data based on the data deletion rule. If the metadata includes the deletion timestamp of the first cold data, the first management device obtains the remaining lifetime of the first cold data based on the deletion timestamp of the first cold data and the current timestamp.
[0099] Optionally, if the metadata of the first cold data does not include a deletion timestamp of the first cold data, the first management device obtains the deletion timestamp of the first cold data using method 1. If the metadata of the first cold data includes a deletion timestamp of the first cold data, the first management device obtains the deletion timestamp of the first cold data using method 1 or method 2, which provides greater flexibility.
[0100] Step 503: The first management device selects at least one cold data whose remaining survival time satisfies the migration condition from the m cold data.
[0101] The migration condition is to select at least one cold data item with the longest remaining life span or to select at least one cold data item with a remaining life span exceeding a life span threshold. In other words, the first management device selects at least one cold data item with the longest remaining life span from the m cold data items, or selects at least one cold data item with a remaining life span exceeding a life span threshold.
[0102] In some embodiments, the first management device scores each cold data based on the remaining survival time of each cold data in the m cold data to obtain a score corresponding to each cold data.
[0103] For example, multiple remaining survival time ranges can be defined, and the multiple remaining survival time ranges correspond one-to-one to multiple scores. For each cold data, the remaining survival time range in which the remaining survival time of the cold data lies is determined, and the score corresponding to the remaining survival time range is used as the score of the cold data.
[0104] Optionally, the longer the remaining survival time of cold data, the higher the score of the cold data. The migration condition is to select at least one cold data with the highest score or to select at least one cold data with a score exceeding a score threshold. In other words, the first management device selects at least one cold data with the highest score from the m cold data, or selects at least one cold data with a score exceeding the score threshold.
[0105] Optionally, the longer the remaining survival time of cold data, the lower the score of the cold data. The migration condition is to select at least one cold data with the lowest score or to select at least one cold data whose score does not exceed the score threshold. In other words, the first management device selects at least one cold data with the lowest score from the m cold data, or selects at least one cold data whose score does not exceed the score threshold.
[0106] Step 504: The first management device moves the at least one cold data from the first storage system to the second storage system.
[0107] In step 504, the first management device sends a storage request to the second storage system, the storage request including the at least one cold data item. The second storage system receives the storage request, stores the at least one cold data item, and sends a storage response to the first management device. The first management device receives the storage response and deletes the at least one cold data item from the first storage system, thereby moving the at least one cold data item from the first storage system to the second storage system.
[0108] In some embodiments, the first storage system and the second storage system may be different cloud storage systems located in different regions. The first storage system and the second storage system may be connected via a core network. Therefore, bandwidth resources between the first storage system and the second storage system are limited and very valuable. The first management device selects at least one cold data item whose remaining lifetime meets the migration criteria and migrates the at least one cold data item from the first storage system to the second storage system. This maximizes the value of the bandwidth resources between the first storage system and the second storage system, thereby increasing the benefits of data migration.
[0109] For the above deletion timestamp prediction model, see Figure 6In one embodiment of the present application, a method 600 for training a deletion timestamp prediction model is provided. The method 600 may be executed by a first management device or a policy management device. For example, the first management device trains a deletion timestamp prediction model and stores it locally. When the deletion timestamp prediction model is needed to obtain the remaining survival time of cold data, the deletion timestamp prediction model is obtained locally.
[0110] For another example, after training a deletion timestamp prediction model, the policy management device associates the deletion timestamp prediction model with a policy card and stores the deletion timestamp prediction model locally. Optionally, the policy management device stores the deletion timestamp prediction model locally in correspondence with the identification information of the policy card. When the policy management device needs to send the deletion timestamp prediction model to a management device included in a storage system in the network architecture 100 that needs to move data, the policy management device reads the deletion timestamp prediction model from the local storage based on the identification information of the policy card.
[0111] For another example, after the policy management device trains the deletion timestamp prediction model, it saves the deletion timestamp prediction model in a policy card. When the policy management device needs to send the deletion timestamp prediction model to the storage system in the network architecture 100 that needs to move data, it reads the deletion timestamp prediction model from the policy card.
[0112] like Figure 6 As shown, the method 600 includes the following process from step 601 to step 604.
[0113] Step 601: Acquire at least one training sample. For each training sample in the at least one training sample, the training sample includes attribute information of a cold data and a first deletion timestamp of the cold data.
[0114] In some embodiments, the training sample may be configured by a technician. For example, the technician may obtain the attribute information and deletion timestamp of cold data deleted from the storage system, use the deletion timestamp as the first deletion timestamp of the cold data, and form the training sample with the attribute information and the first deletion timestamp of the cold data.
[0115] Step 602: Based on the deletion timestamp prediction model to be trained and the attribute information of the cold data included in each training sample, obtain the second deletion timestamp of the cold data included in each training sample.
[0116] For each training sample, the second deletion timestamp of the cold data included in the training sample is obtained by reasoning the deletion timestamp prediction model to be trained based on the attribute information of the cold data.
[0117] The deletion timestamp prediction model to be trained includes AI models such as a random forest algorithm, a logistic regression algorithm, or a support vector machine (SVM).
[0118] In step 602, the attribute information of the cold data included in each training sample is input into the deletion timestamp prediction model to be trained, so that the deletion timestamp prediction model to be trained can infer the second deletion timestamp of the cold data included in each training sample based on the description information of the cold data included in each training sample, and obtain the second deletion timestamp of the cold data included in each training sample output by the deletion timestamp prediction model to be trained.
[0119] Step 603: Based on the first deletion timestamp and the second deletion timestamp of the cold data included in each training sample, a loss value is calculated using a loss function, and the deletion timestamp prediction model to be trained is adjusted based on the loss value.
[0120] In step 603, the hyperparameters and / or model structure of the deletion timestamp prediction model to be trained are adjusted based on the loss value.
[0121] Step 604: When it is determined to continue training the deletion timestamp prediction model to be trained, return to step 602 for execution; when it is determined not to continue training the deletion timestamp prediction model to be trained, use the deletion timestamp prediction model to be trained as the deletion timestamp prediction model.
[0122] In some embodiments, when the number of times the deletion timestamp prediction model to be trained is trained reaches a specified number, it is determined not to continue training the deletion timestamp prediction model to be trained.
[0123] In some embodiments, multiple verification samples are used to obtain the accuracy of the deletion timestamp prediction model to be trained in inferring the deletion timestamp. If the accuracy exceeds a specified accuracy threshold, it is determined that the deletion timestamp prediction model to be trained will not be further trained. In implementation:
[0124] Multiple verification samples are obtained. For each of the multiple verification samples, the verification sample includes a cold data and a first deletion timestamp of the cold data. Based on attribute information of the cold data included in the verification sample and a deletion timestamp prediction model to be trained, a second deletion timestamp of the cold data included in the verification sample is obtained. The second deletion timestamp of the cold data included in each verification sample is obtained according to the above process.
[0125] The accuracy is calculated based on the first and second deletion timestamps of the cold data included in each verification sample. If the accuracy does not exceed a specified accuracy threshold, training of the deletion timestamp prediction model to be trained is continued. If the accuracy exceeds the specified accuracy threshold, training of the deletion timestamp prediction model to be trained is discontinued.
[0126] For each verification sample, if the difference between the first deletion timestamp and the second deletion timestamp of the cold data included in the verification sample is less than the difference threshold, it is determined that the deletion timestamp prediction model to be trained correctly inferred the second deletion timestamp of the cold data included in the verification sample. If the difference between the first deletion timestamp and the second deletion timestamp of the cold data included in the verification sample is greater than or equal to the difference threshold, it is determined that the deletion timestamp prediction model to be trained did not correctly infer the second deletion timestamp of the cold data included in the verification sample.
[0127] In some embodiments, the process of obtaining the multiple verification samples is the same as the process of obtaining the training samples described above, and will not be described in detail here.
[0128] In an embodiment of the present application, the first management device obtains m cold data whose access popularity meets the cold data condition from the data stored in the first storage system, obtains the remaining survival time of each cold data based on the survival time calculation strategy, and selects at least one cold data that meets the migration condition from the m cold data based on the remaining survival time of each cold data. The first management device then moves the at least one cold data stored in the first storage system to the second storage system. Since the at least one cold data is the cold data with the longest remaining survival time among the m cold data or the cold data with a remaining survival time exceeding the time threshold among the m cold data, after the first management device moves the at least one cold data from the first storage system to the second storage system, it needs to be stored in the second storage system for a long time before being deleted from the second storage system, thereby avoiding moving cold data with a shorter remaining survival time from the first storage system to the second storage system, thereby increasing the benefits of moving data.
[0129] See also Figure 7 , the embodiment of the present application provides a method 700 for moving data, which can be applied to Figure 4In the network architecture 100 shown. The method 700 is described by taking the migration of cold data in the first storage system to the second storage system as an example. Similarly, the method 700 can also migrate cold data in other storage systems (such as the third storage system) to the second storage system, which will not be listed one by one. In the method 700, the global scheduling device configures a quota for the first storage system for the first management device. The quota of the first storage system is the total amount of data that the first management device can migrate from the first storage system during the migration cycle. The method 700 includes the following process.
[0130] Step 701: The global scheduling device sends a standard quota to a first management device included in a first storage system. The standard quota is the total amount of data that the first management device can migrate from the first storage system in a first migration cycle.
[0131] In step 701, the global scheduling device may send a standard quota to management devices in other storage systems included in the network architecture 100 that need to move data. For example, the global scheduling device may also send a standard quota to the second management device in the third storage system. Alternatively, the standard quota may be a preset quota that can be received by both the first management device in the first storage system and the second management device in the third storage system.
[0132] Step 702: The first management device receives the standard quota and obtains at least one cold data to be migrated in the first migration cycle. The at least one cold data is cold data stored in the first storage system and has a remaining survival time that meets the migration condition.
[0133] For a detailed implementation process of the first management device acquiring the at least one cold data, see Figure 5 The relevant contents of steps 501 to 503 of the method 500 are not described in detail here.
[0134] Step 703: The first management device moves the at least one cold data from the first storage system to the second storage system when the sum of the total data volume of the at least one cold data and the first data volume does not exceed the standard quota, where the first data volume is the total data volume moved from the first storage system in the first migration cycle.
[0135] For a detailed implementation process of the first management device moving the at least one cold data from the first storage system to the second storage system, see Figure 5 The relevant contents of steps 501 to 504 of the method 500 are not described in detail here.
[0136] The first management device further accumulates the first data volume and the total data volume of the at least one cold data to obtain the total data volume that has been moved from the first storage system in the first migration cycle.
[0137] The first management device may further continue to obtain cold data stored in the first storage system whose remaining survival time meets the migration condition, and migrate the cold data from the first storage system to the second storage system.
[0138] Step 704: The first management device sends first statistical information to the global scheduling device at the end of the first migration cycle. The first statistical information includes one or more of the following information: the total amount of data moved from the first storage system during the first migration cycle, or the average survival time of cold data moved from the first storage system during the first migration cycle.
[0139] For other storage systems in the network architecture 100 that need to migrate data, the management devices in the other storage systems also send first statistical information to the global scheduling device at the end of the first migration cycle. For example, the second management device of the third storage system sends first statistical information to the global scheduling device at the end of the first migration cycle. The first statistical information sent by the second management device includes one or more of the following information: the total amount of data migrated from the third storage system during the first migration cycle, or the average survival time of cold data migrated from the third storage system during the first migration cycle.
[0140] Step 705: The global scheduling device receives the first statistical information sent by the first management device, and obtains a second quota based on the first statistical information. The second quota is the total amount of data that the first management device can migrate from the first storage system in the second migration cycle.
[0141] In step 705, the global scheduling device receives first statistical information sent by the management device included in each storage system that needs to move data in the network architecture, obtains the second quota of each storage system based on each received first statistical information, and then sends the second quota of each storage system to each storage system respectively.
[0142] In step 705 , the global scheduling device receives the first statistical information sent by each storage system, and adjusts the standard quota corresponding to each storage system based on each received first statistical information to obtain a second quota corresponding to each storage system.
[0143] Next, an implementation example of adjusting a standard quota is listed. In this implementation example, for first statistical information sent by a first management device in a first storage system, a global scheduling device performs a weighted operation on the total amount of data moved and / or the average survival time included in the first statistical information to obtain a configuration indicator for the first storage system. The configuration indicator of each storage system is obtained in the same manner as described above, an average configuration indicator is calculated based on the configuration indicator of each storage system, and multiple intervals are determined based on the average configuration indicator and the indicator offset. The multiple intervals correspond one-to-one to multiple quota offsets, wherein the farther the interval is from the average configuration indicator, the larger the quota offset corresponding to the interval, and the closer the interval is to the average configuration indicator, the smaller the quota offset corresponding to the interval.
[0144] Among the multiple intervals, the average configuration index is the lower limit value of a certain interval and the upper limit value of another interval. Among the multiple intervals, there are x intervals that are smaller than the average configuration index and y intervals that are larger than the average configuration index. Both x and y are integers greater than or equal to 1.
[0145] For each interval, if the interval is smaller than the average configuration index, if the interval is not the smallest interval, the upper limit of the interval may be the value obtained by subtracting i*S from the average configuration index, where S is the index offset, and the lower limit of the interval may be the value obtained by subtracting (i+1)*S from the average configuration index, where i=0, 1, 2, ..., x-1. If the interval is the smallest interval, the upper limit of the interval may be the value obtained by subtracting (x-1)*S from the average configuration index. Or,
[0146] If the interval is larger than the average configuration index, and if it is not the largest interval, the lower limit of the interval may be the value obtained by adding j*S to the average configuration index, and the upper limit of the interval may be the value obtained by adding (j+1)*S to the average configuration index, where j=0, 1, 2, ..., y-1. If it is the largest interval, the lower limit of the interval may be the value obtained by adding (y-1)*S to the average configuration index.
[0147] For example, assuming the average configuration index is 500G, the index offset S=50G, and assuming the multiple intervals are six intervals, the six intervals are interval 1 less than 400, interval 2 greater than or equal to 400 and less than 450, interval 3 greater than or equal to 450 and less than 500, interval 4 greater than or equal to 500 and less than 550, interval 5 greater than or equal to 550 and less than 600, and interval 6 greater than or equal to 600. Interval 1, interval 2, and interval 3 are intervals below the average configuration index, and interval 4, interval 5, and interval 6 are intervals greater than the average configuration index.
[0148] Assume that the quota offset for interval 1 is 200GB, the quota offset for interval 2 is 50GB, the quota offset for interval 3 is 0GB, the quota offset for interval 4 is 0GB, the quota offset for interval 5 is 50GB, and the quota offset for interval 6 is 200GB. Intervals 3 and 4 are the two intervals closest to the average configuration indicator, and the quota offsets for these two intervals can be 0 or greater.
[0149] Determine the interval in which the configuration indicator of the first storage system lies. If the interval is smaller than the interval of the average configuration indicator, it indicates that the amount of data that needs to be moved by the first storage system in the first migration cycle is small and / or the average remaining survival time of the moved data is short. Based on this situation, it can be predicted that in the second migration cycle, the total amount of cold data in the first storage system that meets the migration conditions may not be large and / or there may not be much cold data with a long remaining survival time that meets the migration conditions. Therefore, the quota offset corresponding to the interval can be reduced from the standard quota to obtain the second quota of the first storage system, so as to allocate more quota to other storage systems.
[0150] If this interval is greater than the average configuration index, it indicates that the first storage system needs to migrate a large amount of data during the first migration cycle and / or the average remaining lifetime of the migrated data is long. Based on this, it can be predicted that during the second migration cycle, the total amount of cold data in the first storage system that meets the migration requirements may be large and / or there may be a large amount of cold data with a long remaining lifetime that meets the migration requirements. Therefore, the quota offset corresponding to this interval can be added to the standard quota to obtain the second quota for the first storage system, allocating more quota to the first storage system. Second quotas for other storage systems can be obtained in this manner.
[0151] In step 705, the global scheduling device adjusts the quota of each storage system based on the statistical information of each storage system to ensure that when moving data in each migration cycle, the total bandwidth resources required for moving data will not fluctuate too much, and at the same time, it can ensure that the storage system with a large amount of cold data that meets the migration conditions or a large amount of cold data with a long remaining storage time has sufficient quota.
[0152] Step 706: The global scheduling device sends the second quota to the first management device in the first storage system.
[0153] In step 706, the global scheduling device may send the second quotas corresponding to other storage systems included in the network architecture 100 that need to migrate data. For example, the global scheduling device may also send the second quota corresponding to the third storage system to the third storage system. Optionally, the first management device in the first storage system may receive the second quota corresponding to the first storage system, and the second management device in the third storage system may receive the second quota corresponding to the third storage system.
[0154] Step 707: The first management device receives the second quota and obtains at least one cold data to be migrated in the second migration cycle. The at least one cold data is cold data stored in the first storage system and has a remaining survival time that meets the migration condition.
[0155] For a detailed implementation process of the first management device acquiring the at least one cold data, see Figure 5 The relevant contents of steps 501 to 503 of the method 500 are not described in detail here.
[0156] Step 708: The first management device moves the at least one cold data from the first storage system to the second storage system when the sum of the total data volume and the second data volume of the at least one cold data does not exceed the second quota, where the second data volume is the total data volume moved from the first storage system in the second migration cycle.
[0157] For a detailed implementation process of the first management device moving the at least one cold data from the first storage system to the second storage system, see Figure 5 The relevant contents of steps 501 to 504 of the method 500 are not described in detail here.
[0158] The first management device further accumulates the second data volume and the total data volume of the at least one cold data to obtain the total data volume that has been moved from the first storage system in the second migration cycle.
[0159] The first management device may further continue to obtain cold data stored in the first storage system whose remaining survival time meets the migration condition, and migrate the cold data from the first storage system to the second storage system.
[0160] Step 709: The first management device sends second statistical information to the global scheduling device at the end of the second migration cycle. The second statistical information includes one or more of the following information: the total amount of data moved from the first storage system during the second migration cycle, or the average survival time of cold data moved from the first storage system during the second migration cycle.
[0161] Step 710: The global scheduling device receives the second statistical information sent by the first management device, and obtains a first quota based on the second statistical information. The first quota is the total amount of data that the first management device can move from the first storage system in the third migration cycle.
[0162] In step 710, the global scheduling device receives second statistical information sent by the management device included in each storage system that needs to move data in the network architecture, obtains the first quota of each storage system based on each received second statistical information, and then sends the first quota of each storage system to the management device included in each storage system.
[0163] In step 710, the global scheduling device receives second statistical information sent by each storage system in the network architecture that needs to move data, and adjusts the second quota corresponding to each storage system based on each received second statistical information to obtain the first quota corresponding to each storage system.
[0164] Next, an implementation example of adjusting the second quota is listed. In this implementation example, for the second statistical information sent by the first management device in the first storage system, the global scheduling device performs a weighted operation on the total amount of data moved and / or the average survival time included in the second statistical information to obtain the configuration index of the first storage system. The configuration index of each storage system is obtained in the same manner as described above, and an average configuration index is calculated based on the configuration index of each storage system. Multiple intervals are determined based on the average configuration index and the index offset. The multiple intervals correspond one-to-one to multiple quota offsets, wherein the intervals farther from the average configuration index have corresponding quota offsets larger, and the intervals closer to the average configuration index have corresponding quota offsets smaller.
[0165] Determine the interval within which the configuration indicator of the first storage system falls. If the interval is smaller than the average configuration indicator, reduce the second quota of the first storage system by the quota offset corresponding to the interval to obtain the first quota of the first storage system. If the interval is larger than the average configuration indicator, increase the second quota of the first storage system by the quota offset corresponding to the interval to obtain the first quota of the first storage system. In this manner, the first quotas of other storage systems can be obtained.
[0166] Step 711: The global scheduling device sends a first quota to a first management device included in the first storage system.
[0167] In step 711, the global scheduling device may send the first quota corresponding to other storage systems to the management devices in other storage systems in the network architecture 100 that need to move data. For example, the global scheduling device may also send the first quota corresponding to the third storage system to the second management device in the third storage system. Optionally, the first management device in the first storage system may receive the first quota corresponding to the first storage system, and the second management device in the third storage system may receive the first quota corresponding to the third storage system.
[0168] The first management device migrates the cold data stored in the first storage system and whose remaining survival time meets the migration condition to the second storage system in the third migration cycle according to the process from step 707 to step 709.
[0169] In an embodiment of the present application, a first management device receives a quota for the first storage system from a global scheduling device and obtains at least one cold data item whose access popularity satisfies both cold data conditions and migration conditions from the data stored in the first storage system. The first management device then obtains the total amount of data that has been migrated from the first storage system during the current migration cycle. When the difference between the quota and the total amount of data exceeds the total amount of the at least one cold data item, the at least one cold data item stored in the first storage system is migrated to a second storage system. Because the at least one cold data item is the cold data item with the longest remaining lifetime among the m cold data items or the cold data item with a remaining lifetime exceeding a lifetime threshold among the m cold data items, after the first management device migrates the at least one cold data item from the first storage system to the second storage system, the at least one cold data item must be stored in the second storage system for a long period of time before being deleted from the second storage system. This avoids migrating cold data with a shorter remaining lifetime from the first storage system to the second storage system, thereby increasing the benefits of migrating data.
[0170] See also Figure 8 , the embodiment of the present application provides a device 800 for moving data, the device 800 is deployed in Figure 1-Figure 4 The storage system of the network architecture 100 shown includes a management device, or is deployed on Figure 5 The method 500 or Figure 7 The management device of the method 700 is shown. The apparatus 800 includes a processing unit 801 and a moving unit 802.
[0171] The processing unit 801 is configured to determine, from among the plurality of data stored in the first storage system, at least one data item whose access popularity satisfies a cold data condition and is cold data;
[0172] The processing unit 801 is further configured to select, from the at least one cold data, at least one target cold data having the longest remaining survival time, or having a remaining survival time exceeding a duration threshold, based on the remaining survival time of each cold data in the at least one cold data;
[0173] The migration unit 802 is configured to migrate the at least one cold data from the first storage system to the second storage system.
[0174] Optionally, the detailed implementation process of the processing unit 801 obtaining cold data whose access popularity meets the cold data condition can be found in Figure 5 The relevant contents in step 501 of the method 500 are not described in detail here.
[0175] Optionally, the detailed implementation process of the processing unit 801 selecting at least one target cold data can be found in Figure 5 The relevant contents in step 503 of the method 500 are not described in detail here.
[0176] Optionally, the migration unit 802 migrates the at least one target cold data from the first storage system to the second storage system. Figure 5 The relevant contents in step 504 of the method 500 are not described in detail here.
[0177] Optionally, the apparatus 800 further includes a first receiving unit 803;
[0178] A first receiving unit 803 is configured to receive a survival time calculation policy sent by a policy management device, where the policy management device includes at least one survival time calculation policy, wherein the survival time calculation policy is used to calculate the remaining survival time of cold data;
[0179] The processing unit 801 is configured to determine the remaining survival time of each cold data in at least one cold data based on the survival time calculation strategy.
[0180] Optionally, the first receiving unit 803 receives the survival time calculation policy sent by the policy management device. Figure 5 The relevant contents in step 502 of the method 500 are not described in detail here.
[0181] Optionally, the processing unit 801 determines the remaining survival time of each cold data based on the survival time calculation strategy. Figure 5 The relevant contents in step 502 of the method 500 are not described in detail here.
[0182] Optionally, the policy management device includes at least one survival time calculation policy, and the survival time calculation policy sent by the policy management device includes the at least one survival time calculation policy, or a portion of the at least one survival time calculation policy. This improves the flexibility of sending the survival time calculation policy.
[0183] Optionally, the survival time calculation strategy is to delete the timestamp prediction model. Processing unit 801 is used to:
[0184] Get the attribute information of each cold data;
[0185] Determine the deletion timestamp of each cold data based on the deletion timestamp prediction model and the attribute information of each cold data;
[0186] Based on the deletion timestamp of each cold data, the remaining survival time of each cold data is determined.
[0187] Optionally, the detailed implementation process of the processing unit 801 determining the attribute information of each cold data can be found in Figure 5 The content related to the method 1 in step 502 of the method 500 will not be described in detail here.
[0188] Optionally, the processing unit 801 determines the deletion timestamp of each cold data based on the deletion timestamp prediction model and the attribute information of each cold data. Figure 5 The content related to the method 1 in step 502 of the method 500 will not be described in detail here.
[0189] Optionally, the processing unit 801 determines the remaining survival time of each cold data based on the deletion timestamp of each cold data. Figure 5 The content related to the method 1 in step 502 of the method 500 will not be described in detail here.
[0190] Optionally, the survival time calculation strategy is a data deletion rule, which is used to calculate the remaining survival time of each cold data based on the metadata of each cold data. The processing unit 801 is used to:
[0191] Based on the data deletion rules, obtain the metadata of each cold data, which includes the deletion timestamp of each cold data;
[0192] Based on the deletion timestamp of each cold data, the remaining survival time of each cold data is determined.
[0193] Optionally, the processing unit 801 determines the metadata of each cold data based on the data deletion rule. Figure 5 The content related to the method 2 in step 502 of the method 500 will not be described in detail here.
[0194] Optionally, the processing unit 801 determines the remaining survival time of each cold data based on the deletion timestamp of each cold data. Figure 5 The content related to the method 2 in step 502 of the method 500 will not be described in detail here.
[0195] Optionally, the apparatus 800 further includes a second receiving unit 804;
[0196] The second receiving unit 804 is configured to receive a first quota sent by the global scheduling device, where the first quota indicates a total amount of data that can be moved from the first storage system in a current migration cycle;
[0197] The migration unit 802 is configured to migrate the at least one target cold data from the first storage system to the second storage system when the sum of the total data volume of the at least one target cold data and the total data volume migrated from the first storage system in the current migration cycle does not exceed the first quota.
[0198] Optionally, the detailed implementation process of the second receiving unit 804 receiving the first quota sent by the global scheduling device can be found in Figure 7 The relevant contents in step 707 of the method 700 are not described in detail here.
[0199] Optionally, the migration unit 802 migrates the at least one target cold data from the first storage system to the second storage system. Figure 7 The relevant contents in step 708 of the method 700 are not described in detail here.
[0200] Optionally, the apparatus 800 further includes a sending unit 805;
[0201] Sending unit 805 is configured to send statistical information to the global scheduling device. The statistical information includes one or more of the following: the total amount of data moved from the first storage system during the previous migration cycle, or the average lifetime of cold data moved from the first storage system during the previous migration cycle. The first quota is obtained by the global scheduling device by adjusting the second quota based on the statistical information. The second quota indicates the total amount of data that can be moved from the first storage system during the previous migration cycle. This allows the quota allocated to each storage system to be dynamically adjusted based on the statistical information within each cycle, ensuring that the total network resources consumed by data migration during each cycle do not fluctuate significantly.
[0202] Optionally, the detailed implementation process of the sending unit 805 sending statistical information to the global scheduling device is shown in Figure 7 The relevant contents in step 709 of the method 700 are not described in detail here.
[0203] Optionally, the cost of storing data in the first storage system is higher than the cost of storing data in the second storage system, or the price of the electric energy used to supply power to the first storage system is higher than the price of the electric energy used to supply power to the second storage system, or the storage performance of the storage medium included in the first storage system is higher than the storage performance of the storage medium included in the second storage system.
[0204] Optionally, the first storage system and the second storage system are two cloud storage systems.
[0205] Optionally, the first storage system is located in a first area, and the second storage system is located in a second area, and the first area and the second area are two different administrative areas.
[0206] Optionally, the first storage system and the second storage system are storage systems located in two computer rooms.
[0207] Optionally, the migration condition is to select at least one cold data with the longest remaining survival time, or the migration condition is to select at least one cold data with a remaining survival time exceeding a time threshold.
[0208] In this embodiment of the present application, the processing unit obtains the remaining lifetime of each cold data item, selects at least one target cold data item from each cold data item with the longest remaining lifetime or a remaining lifetime exceeding a lifetime threshold, and then moves the at least one target cold data item from the first storage system to the second storage system. This ensures that after the migration unit migrates the at least one target cold data item to the second storage system, the at least one target cold data item will be stored in the second storage system for a long period of time, avoiding the need to migrate cold data items with shorter remaining lifetimes from the first storage system to the second storage system, thereby increasing the benefits of migrating the at least one target cold data item.
[0209] See also Figure 9 , the embodiment of the present application provides a computing device 900. For example, the computing device 900 may be Figure 1-Figure 4 The management device in the storage system in the network architecture 100 shown, or the computing device 900 may be Figure 5 The method 500 or Figure 7 The management device in the method 700 is shown, etc.
[0210] like Figure 9 As shown, computing device 900 includes a bus 902, a processor 904, a memory 906, and a communication interface 908. Processor 904, memory 906, and communication interface 908 communicate with each other via bus 902. Computing device 900 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 900.
[0211] The bus 902 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9 The bus 902 may include a path for transmitting information between various components of the computing device 900 (eg, the processor 904, the memory 906, and the communication interface 908).
[0212] The processor 904 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0213] The memory 906 may include volatile memory, such as random access memory (RAM). The memory 906 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0214] See also Figure 9 , the memory 906 stores executable program code, and the processor 904 executes the executable program code to respectively implement Figure 8 The functions of the processing unit 801, the moving unit 802, the first receiving unit 803, the second receiving unit 804 and the sending unit 805 in the device 800 shown in the figure are used to implement the method provided by any of the above embodiments. That is, the memory 906 stores instructions for executing the method provided by any of the above embodiments. Or,
[0215] The communication interface 908 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 900 and other devices or a communication network.
[0216] Embodiments of the present application also provide a data migration cluster. The data migration cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0217] like Figure 10 As shown, the cluster for moving data includes at least one computing device 900. The memory 906 of one or more computing devices 900 in the cluster for moving data may store the same instructions for executing the method provided by any of the above embodiments.
[0218] In some possible implementations, the memory 906 of one or more computing devices 900 in the data migration cluster may also store partial instructions for executing the above-described method for migrating data. In other words, a combination of one or more computing devices 900 can jointly execute instructions for executing the method provided in any of the above-described embodiments.
[0219] In some possible implementations, one or more computing devices in the cluster for moving data may be connected via a network, such as a wide area network or a local area network. Figure 11 A possible implementation is shown. Figure 11 As shown, two computing devices 900A and 900B are connected via a network. Specifically, the connection to the network is achieved through a communication interface in each computing device.
[0220] In some possible implementations, the memory 906 in the computing device 900A stores the execution Figure 8 The processing unit 801 and the moving unit 802 in the embodiment shown in the figure are executed. Figure 8 Instructions for the functions of the first receiving unit 803, the second receiving unit 804 and the sending unit 805 in the illustrated embodiment.
[0221] It should be understood that Figure 11 The functionality of the computing device 900A shown in FIG. 1 may also be implemented by multiple computing devices 900. Similarly, the functionality of the computing device 900B may also be implemented by multiple computing devices 900.
[0222] The embodiment of the present application also provides another cluster for moving data. The connection relationship between the computing devices in the cluster for moving data can be similarly referred to as Figure 11The connection mode of the cluster for moving data is different in that the memory 906 of one or more computing devices 900 in the cluster for moving data may store the same instructions for executing the method provided in any of the above embodiments.
[0223] In some possible implementations, the memory 906 of one or more computing devices 900 in the data migration cluster may also store partial instructions for executing the method provided in any of the above embodiments. In other words, a combination of one or more computing devices 900 can jointly execute instructions for executing the method provided in any of the above embodiments.
[0224] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored on any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the method provided in any of the above embodiments.
[0225] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the method provided in any of the above embodiments.
[0226] The information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0227] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0228] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for moving data, characterized in that: The method is applied to a management device in a first storage system, and includes: Determining, from the plurality of data stored in the first storage system, at least one cold data whose access popularity satisfies a cold data condition; Selecting, according to the remaining survival time of each cold data in the at least one cold data, at least one target cold data having the longest remaining survival time from the at least one cold data, or having a remaining survival time exceeding a time threshold; The at least one target cold data is moved from the first storage system to the second storage system.
2. The method according to claim 1, wherein Before selecting, from the at least one cold data according to the remaining survival time of each cold data, at least one target cold data having the longest remaining survival time, or having a remaining survival time exceeding a time threshold, the method further includes: Receiving a survival time calculation policy sent by a policy management device, where the policy management device includes at least one survival time calculation policy, wherein the survival time calculation policy is used to calculate the remaining survival time of cold data; Based on the survival time calculation strategy, a remaining survival time of each cold data in the at least one cold data is determined.
3. The method according to claim 2, wherein The survival time calculation strategy is a deletion timestamp prediction model, and determining the remaining survival time of each cold data in the at least one cold data based on the survival time calculation strategy includes: Obtaining attribute information of each cold data; Determining a deletion timestamp of each cold data based on the deletion timestamp prediction model and the attribute information of each cold data; Based on the deletion timestamp of each cold data, the remaining survival time of each cold data is determined.
4. The method according to claim 2, wherein The survival time calculation strategy is a data deletion rule, and the data deletion rule is used to calculate the remaining survival time of each cold data based on the metadata of each cold data; The determining, based on the survival time calculation strategy, the remaining survival time of each cold data in the at least one cold data, includes: Based on the data deletion rule, obtaining metadata of each cold data, wherein the metadata of each cold data includes a deletion timestamp of each cold data; Based on the deletion timestamp of each cold data, the remaining survival time of each cold data is determined.
5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: receiving a first quota sent by a global scheduling device, where the first quota is used to indicate a total amount of data that can be migrated from the first storage system in a current migration cycle; The moving the at least one target cold data from the first storage system to the second storage system includes: When the sum of the total data volume of the at least one target cold data and the total data volume that has been moved from the first storage system in the current migration cycle is less than the first quota, the at least one target cold data is moved from the first storage system to the second storage system.
6. The method according to claim 5, wherein Before receiving the first quota sent by the global scheduling device, the method further includes: Send statistical information to the global scheduling device, where the statistical information includes one or more of the following information: the total amount of data moved from the first storage system in the last migration cycle, or the average survival time of cold data moved from the first storage system in the last migration cycle, where the first quota is obtained by the global scheduling device adjusting the second quota based on the statistical information, and the second quota is used to indicate the total amount of data that can be moved from the first storage system in the last migration cycle.
7. The method according to any one of claims 1 to 6, wherein: The cost of storing data in the first storage system is higher than the cost of storing data in the second storage system, or the price of electric energy used to supply power to the first storage system is higher than the price of electric energy used to supply power to the second storage system, or the storage performance of the storage medium included in the first storage system is higher than the storage performance of the storage medium included in the second storage system.
8. A device for moving data, characterized in that: The device is located in a first storage system and includes: a processing unit, configured to determine, from the plurality of data stored in the first storage system, at least one cold data whose access popularity satisfies a cold data condition; The processing unit is further configured to select, from the at least one cold data, at least one target cold data having the longest remaining survival time or having a remaining survival time exceeding a time threshold, based on the remaining survival time of each cold data in the at least one cold data; A migration unit is configured to migrate the at least one target cold data from the first storage system to the second storage system.
9. The device according to claim 8, wherein The device further includes: a first receiving unit; The first receiving unit is configured to receive a survival time calculation policy sent by a policy management device, where the policy management device includes at least one survival time calculation policy, wherein the survival time calculation policy is used to calculate the remaining survival time of cold data; The processing unit is further configured to determine a remaining survival time of each cold data in the at least one cold data based on the survival time calculation strategy.
10. The device according to claim 9, wherein The survival time calculation strategy is to delete the timestamp prediction model, and the processing unit is used to: Obtaining attribute information of each cold data; Determining a deletion timestamp of each cold data based on the deletion timestamp prediction model and the attribute information of each cold data; Based on the deletion timestamp of each cold data, the remaining survival time of each cold data is determined.
11. The device according to claim 9, wherein The survival time calculation strategy is a data deletion rule, and the data deletion rule is used to calculate the remaining survival time of each cold data based on the metadata of each cold data; The processing unit is configured to: Based on the data deletion rule, obtaining metadata of each cold data, wherein the metadata of each cold data includes a deletion timestamp of each cold data; Based on the deletion timestamp of each cold data, the remaining survival time of each cold data is determined.
12. The device according to any one of claims 8 to 11, characterized in that The device further includes: a second receiving unit; The second receiving unit is configured to receive a first quota sent by a global scheduling device, where the first quota is used to indicate a total amount of data that can be moved from the first storage system in a current migration cycle; The moving unit is used for: When the sum of the total data volume of the at least one target cold data and the total data volume that has been moved from the first storage system in the current migration cycle is less than the first quota, the at least one target cold data is moved from the first storage system to the second storage system.
13. The device according to claim 12, wherein The device further includes: a sending unit; The sending unit is used to send statistical information to the global scheduling device, where the statistical information includes one or more of the following information: the total amount of data moved from the first storage system in the previous migration cycle, or the average survival time of cold data moved from the first storage system in the previous migration cycle. The first quota is obtained by the global scheduling device by adjusting the second quota based on the statistical information, and the second quota is used to indicate the total amount of data that can be moved from the first storage system in the previous migration cycle.
14. The device according to any one of claims 8 to 13, characterized in that The cost of storing data in the first storage system is higher than the cost of storing data in the second storage system, or the price of electric energy used to supply power to the first storage system is higher than the price of electric energy used to supply power to the second storage system, or the storage performance of the storage medium included in the first storage system is higher than the storage performance of the storage medium included in the second storage system.
15. A computing device cluster, characterized in that: comprising at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in a memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 7.
17. A computer program product comprising instructions, characterized in that When the instruction is executed by a computing device cluster, the computing device cluster is caused to execute the method according to any one of claims 1 to 7.