Data management method and device based on cold data migration, equipment and storage medium
Patent Information
- Application Number
- CN202411334232.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-09-24
AI Technical Summary
[0007]本发明的主要目的在于提供一种基于冷数据迁移的数据管理方法、装置、设备及存储介质,旨在解决现有技术无法有效识别和管理分布式文件系统中的冷数据,导致冷数据占用计算资源,造成系统资源浪费和性能低下的技术问题
[0051] Beneficial Effects: This invention relates to a data management method based on cold data migration. Metadata is extracted from a distributed file system and imported into a metadata repository of a distributed data processing platform. Through offline analysis tasks, file information in the metadata repository is analyzed to identify cold data, and its file path information is stored in a database. Based on the identified cold data, a data migration plan is generated, including the migration target location, migration priority ranking, and migration schedule. According to this plan, the cold data is migrated from its original storage location to the target remote storage, and the table partition paths in the metadata repository are updated to point to the new storage location. After the migration is complete, the cold data in the original storage location is deleted. This invention achieves effective separation of cold data from computing resources, optimizes the utilization efficiency of storage resources, reduces the overall resource consumption of the system, and improves the scalability and performance of the system.
Smart Images

Figure CN119166592B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data management technology, and in particular to a data management method, apparatus, device and storage medium based on cold data migration. Background Technology
[0002] In the financial sector, with rapid business growth and an explosive increase in data volume, financial institutions face the challenge of efficiently storing and processing massive amounts of data. To address this issue, many financial institutions have adopted the distributed data processing framework Hadoop. Hadoop, as a scalable software framework capable of distributed processing of large amounts of data, is widely used in financial applications such as data analysis, risk management, and customer behavior prediction.
[0003] The Hadoop framework comprises several key components. HDFS (Hadoop Distributed File System) provides efficient distributed storage for massive amounts of data, while YARN (YetAnother Resource Negotiator) handles resource management and scheduling services for various computing frameworks. In traditional Hadoop cluster deployments, to maximize the locality of computing resources, compute and storage nodes are typically deployed on the same machine. This approach reduces data transfer overhead and improves data processing efficiency.
[0004] However, this deployment method faces some practical problems in the financial industry. As business demands increase, financial institutions need to frequently expand the size of their Hadoop clusters to meet ever-growing computing and storage requirements. Typically, horizontally scaling a Hadoop cluster by adding machines can easily improve system performance. However, this scaling method requires a relatively balanced configuration of the host's computing and storage resources; for example, the configuration of CPU, memory, network interface cards (NICs), and storage resources need to be matched. Because these resources must be scaled simultaneously, the overall cost per machine is high, and in some cases, resource waste occurs.
[0005] Specifically, the data stored in HDFS follows the Pareto principle (80 / 20 rule), meaning that 80% of the data is "cold data" with extremely low access frequency. While this cold data occupies a significant amount of storage space, its low access frequency means that the machines storing it have low demands on computing resources such as CPU and memory during actual operation, resulting in wasted computing resources. This problem of resource waste is particularly prominent in the financial industry, as cold data consumes a large amount of expensive computing resources that are not fully utilized, affecting overall resource utilization efficiency.
[0006] Therefore, how to optimize the utilization of computing and storage resources and reduce resource waste while ensuring efficient data processing has become an urgent problem for the financial industry when using the Hadoop framework for big data processing. Summary of the Invention
[0007] The main objective of this invention is to provide a data management method, apparatus, device, and storage medium based on cold data migration, aiming to solve the technical problem that existing technologies cannot effectively identify and manage cold data in distributed file systems, resulting in cold data occupying computing resources, causing system resource waste and low performance.
[0008] To achieve the above objectives, the present invention provides a data management method based on cold data migration, comprising:
[0009] Extract metadata from the distributed file system and import the metadata into the metadata repository of the distributed data processing platform;
[0010] The offline analysis task of the distributed data processing platform is used to analyze the file information of the metadata in the metadata repository, identify cold data, and store the file path information of the cold data in the database.
[0011] Based on the identified cold data, a data migration plan is generated, which includes the migration target location, migration priority ranking, and migration time schedule.
[0012] According to the data migration plan, the cold data is migrated from its original storage location to the target remote storage, and the table partition path in the metadata repository is updated to point to the new storage location after migration.
[0013] After the data migration is complete, delete the migrated cold data from the original storage location.
[0014] In one embodiment, analyzing file information of metadata in the metadata repository and identifying cold data includes:
[0015] Analyze the metadata in the metadata repository to extract file information related to the file, including file path, last access time, last modification time, file size, and creation time;
[0016] Set access frequency thresholds and time point thresholds for identifying cold data;
[0017] Based on file information, the access frequency of the file is counted and compared with an access frequency threshold. If the access frequency of the file is lower than the access frequency threshold, the file is marked as cold data.
[0018] The last access time of the file is compared with a time point threshold. If the last access time of the file is earlier than the time point threshold, the file is marked as cold data.
[0019] Files are classified according to their size, and a comprehensive evaluation is conducted based on the file's access frequency and last access time. If the comprehensive evaluation criteria are met, the file is marked as cold data.
[0020] In one embodiment, extracting metadata from a distributed file system and importing the metadata into a metadata repository of a distributed data processing platform includes:
[0021] Export the current FSImage file from the distributed file system, the FSImage file containing metadata for all files and directories in the file system;
[0022] The exported FSImage file is converted into a file format that conforms to the table structure specification of the distributed data processing platform. The conversion includes converting metadata into a structured table format.
[0023] Select the target table in the distributed data processing platform, and import the converted metadata file into the partition of the target table according to the specified partitioning rules, which are based on date or file type;
[0024] The imported metadata files are stored in the metadata repository of the distributed data processing platform for use by data analysis tasks.
[0025] In one embodiment, a data migration plan is generated based on the identified cold data, including:
[0026] Determine the target location for migrating cold data, including remote storage, cloud storage, or archive storage devices;
[0027] Calculate the migration priority of cold data based on file size, last access time, and access frequency;
[0028] Cold data is divided into multiple migration batches according to migration priority, and a migration schedule is set for each migration batch;
[0029] A data migration plan is generated based on the migration target location, migration priority, and migration time schedule of cold data.
[0030] In one embodiment, the migration priority of cold data is calculated based on file size, last access time, and access frequency, including:
[0031] Based on the amount of storage resources occupied by file size, the reflection of data usage frequency by last access time, and the data usage reflected by access frequency, set the weights of file size, last access time, and access frequency.
[0032] Standardize the data on file size, last access time, and access frequency to convert data of different scales into a unified numerical range;
[0033] Based on the standardized file size, last access time, and access frequency data, a migration priority score is calculated for each cold data file according to the set weights. The priority score is obtained by weighting the file size, last access time, and access frequency according to their respective weights.
[0034] Cold data files are sorted according to priority scores, and a corresponding migration order is assigned to each cold data file.
[0035] In one embodiment, after deleting the migrated cold data from the original storage location, the method further includes:
[0036] Monitor access to migrated cold data and record the access time and frequency of cold data files.
[0037] Determine whether cold data becomes active again based on a preset activity threshold;
[0038] If the access frequency of cold data reaches the set activity threshold, a data migration plan will be generated.
[0039] The data migration plan is executed to migrate the reactivated cold data to its original storage location or high-frequency storage area;
[0040] Update the file path information in the metadata repository and update the new storage location of the migrated files to the corresponding table partition records;
[0041] After the data migration is completed, update the cache content and record the specific information of the migration operation.
[0042] In one embodiment, migrating the cold data from its original storage location to a target remote storage location according to the data migration plan includes:
[0043] Read the data migration plan to obtain the list of cold data files, migration target location, migration priority, and migration schedule;
[0044] Check the availability of the target migration location, verify that the storage space meets the migration requirements, and verify the stability of the network connection;
[0045] Based on the migration priority and migration schedule in the data migration plan, initiate the migration tasks for cold data files sequentially.
[0046] During the migration process, cold data files exceeding a preset size threshold are split and transmitted part by part;
[0047] Real-time monitoring of transmission progress, transmission rate, and network connection stability; recording transmission data and status.
[0048] After the data migration is completed, data verification is performed to check whether the data is intact and undamaged during the transmission process, and detailed log information of the migration process is recorded.
[0049] Furthermore, to achieve the above objectives, the present invention also provides a data management device based on cold data migration, the data management device based on cold data migration including a memory, a processor, and a data management program based on cold data migration stored in the memory and executable on the processor, wherein when the data management program based on cold data migration is executed by the processor, it implements the steps of the data management method based on cold data migration as described above.
[0050] Furthermore, to achieve the above objectives, the present invention also provides a computer storage medium storing a data management program based on cold data migration, wherein the data management program based on cold data migration, when executed by a processor, implements the steps of the data management method based on cold data migration as described above.
[0051] Beneficial Effects: This invention relates to a data management method based on cold data migration. Metadata is extracted from a distributed file system and imported into a metadata repository of a distributed data processing platform. Through offline analysis tasks, file information in the metadata repository is analyzed to identify cold data, and its file path information is stored in a database. Based on the identified cold data, a data migration plan is generated, including the migration target location, migration priority ranking, and migration schedule. According to this plan, the cold data is migrated from its original storage location to the target remote storage, and the table partition paths in the metadata repository are updated to point to the new storage location. After the migration is complete, the cold data in the original storage location is deleted. This invention achieves effective separation of cold data from computing resources, optimizes the utilization efficiency of storage resources, reduces the overall resource consumption of the system, and improves the scalability and performance of the system. Attached Figure Description
[0052] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0053] Figure 1 This is a flowchart illustrating an embodiment of the data management method based on cold data migration according to the present invention.
[0054] Figure 2This is a schematic diagram of the functional modules of a preferred embodiment of the data management device based on cold data migration of the present invention;
[0055] Figure 3 This is a schematic diagram of the hardware operating environment of the data management device based on cold data migration according to an embodiment of the present invention. Detailed Implementation
[0056] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0057] In the financial sector, with rapid business growth and an explosive increase in data volume, financial institutions face the challenge of efficiently storing and processing massive amounts of data. To address this issue, many financial institutions have adopted the distributed data processing framework Hadoop. Hadoop, as a scalable software framework capable of distributed processing of large amounts of data, is widely used in financial applications such as data analysis, risk management, and customer behavior prediction.
[0058] The Hadoop framework comprises several key components. HDFS (Hadoop Distributed File System) provides efficient distributed storage for massive amounts of data, while YARN (YetAnother Resource Negotiator) handles resource management and scheduling services for various computing frameworks. In traditional Hadoop cluster deployments, to maximize the locality of computing resources, compute and storage nodes are typically deployed on the same machine. This approach reduces data transfer overhead and improves data processing efficiency.
[0059] However, this deployment method faces some practical problems in the financial industry. As business demands increase, financial institutions need to frequently expand the size of their Hadoop clusters to meet ever-growing computing and storage requirements. Typically, horizontally scaling a Hadoop cluster by adding machines can easily improve system performance. However, this scaling method requires a relatively balanced configuration of the host's computing and storage resources; for example, the configuration of CPU, memory, network interface cards (NICs), and storage resources need to be matched. Because these resources must be scaled simultaneously, the overall cost per machine is high, and in some cases, resource waste occurs.
[0060] Specifically, the data stored in HDFS follows the Pareto principle (80 / 20 rule), meaning that 80% of the data is "cold data" with extremely low access frequency. While this cold data occupies a significant amount of storage space, its low access frequency means that the machines storing it have low demands on computing resources such as CPU and memory during actual operation, resulting in wasted computing resources. This problem of resource waste is particularly prominent in the financial industry, as cold data consumes a large amount of expensive computing resources that are not fully utilized, affecting overall resource utilization efficiency.
[0061] Therefore, how to optimize the utilization of computing and storage resources and reduce resource waste while ensuring efficient data processing has become an urgent problem for the financial industry when using the Hadoop framework for big data processing.
[0062] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the data management method based on cold data migration provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0063] like Figure 1 As shown, the data management method based on cold data migration proposed in this invention includes the following steps:
[0064] S10, Extract metadata from the distributed file system and import the metadata into the metadata repository of the distributed data processing platform;
[0065] In this embodiment, the FSImage file of the current system is obtained from a distributed file system (such as HDFS). The FSImage file is a metadata snapshot of the file system, containing metadata such as the structure, file paths, permissions, and block information of all files and directories in the system. The FSImage file is parsed, and its content is converted into a structured data format to facilitate subsequent data processing and analysis. Based on the requirements of the distributed data processing platform (such as Hadoop or Spark), the parsed metadata file is converted into a table structure format compatible with the target platform for easy import and use.
[0066] Select the target table (such as an HBase table or a Hive table), and import the transformed metadata file into the metadata repository of the distributed data processing platform according to pre-defined partitioning rules (such as by date, file type, directory structure, etc.). Execute the data import operation to import the structured metadata file into the target table in batches, ensuring data consistency and integrity during the import process.
[0067] In one specific implementation, under certain scenarios (such as when the system data volume is very large), a custom script or program is used to directly extract and parse the FSImage file, and the parsed data is stored directly in JSON format. Subsequently, the JSON data is imported into HBase or other distributed databases using Spark or other distributed computing frameworks to ensure efficient data reading and processing in high-concurrency scenarios.
[0068] By extracting metadata from the distributed file system and importing it into the metadata repository of the distributed data processing platform, data structures and file information in the distributed system can be effectively managed and analyzed centrally. This not only improves data manageability and traceability but also provides fundamental support for subsequent data processing and optimization. Especially in large-scale data processing scenarios, it achieves efficient metadata management, enhances system scalability and flexibility, and significantly reduces the complexity of data management and system maintenance costs.
[0069] S20, through the offline analysis task of the distributed data processing platform, analyze the file information of the metadata in the metadata repository, identify cold data, and store the file path information of the cold data in the database;
[0070] In this embodiment, an offline analysis task is defined on a distributed data processing platform (such as Hadoop or Spark) to periodically scan and analyze file information in a metadata repository using batch processing jobs.
[0071] The analyzed file information includes file path, file size, last access time, last modification time, creation time, and access frequency.
[0072] Use a scheduler (such as Oozie or Airflow) to periodically trigger offline analysis tasks and run them automatically at preset time intervals, thereby ensuring that the analysis process has minimal impact on the real-time system load.
[0073] The extracted file information is categorized and filtered according to a predetermined set of rules. These rules may include: file access frequency below a certain threshold; file last access time earlier than a set time threshold; file size exceeding or falling below a certain preset standard.
[0074] Based on the above rules, cold data is automatically identified, and its file path information is extracted as the basis for subsequent data migration and management. The identified cold data file path information is organized into a structured data format (such as tables) and stored in a dedicated database (such as MySQL, HBase, etc.) to facilitate subsequent querying, statistics, and data migration operations. Data consistency and integrity are ensured during storage, and redundancy or backup mechanisms are used to prevent data loss when necessary.
[0075] In one specific implementation, Hadoop's MapReduce programming model is used. The required file information is extracted from the metadata repository in HDFS during the Map phase, and analysis and cold data identification are performed during the Reduce phase. To improve efficiency, preprocessing can be performed during the Map phase, such as filtering out files that are clearly not cold data (e.g., recently accessed files), reducing the computational burden of the Reduce phase.
[0076] In MapReduce tasks, analysis is performed not only on the overall information of a file but also on its individual data blocks. By identifying cold data characteristics in certain blocks, it may be possible to migrate only some cold data blocks instead of the entire file, thus saving storage space. When storing the path information of the identified cold data files in an HBase table, its storage structure is expanded by adding redundant fields such as access frequency, file size, and last access time. This information can be further used in subsequent decision-making or migration strategies.
[0077] By executing offline analysis tasks on a distributed data processing platform, cold data in the system can be automatically identified and classified, and its file path information can be stored. This not only significantly reduces the system's reliance on high-performance computing resources and minimizes waste, but also provides a precise data foundation for subsequent data migration and storage optimization. By regularly analyzing and identifying cold data, the system can more flexibly adjust storage strategies and optimize resource allocation, thereby improving the efficiency and performance of overall storage management.
[0078] S30, Based on the identified cold data, a data migration plan is generated, which includes the migration target location, migration priority ranking, and migration time schedule;
[0079] In this embodiment, the system automatically generates a data migration plan based on the cold data identified in the previous step. This plan specifies the migration target location, priority, and time schedule for each cold data file to ensure the efficient execution of the data migration operation.
[0080] Based on the characteristics and storage requirements of cold data, determine the target location for migration. The target location can include remote storage devices, cloud storage services, archive storage systems, etc. Selection criteria include factors such as storage cost, data access frequency, storage capacity, and data security requirements.
[0081] Based on the attributes of cold data, such as file size, last access time, and access frequency, the migration priority of each cold data file is calculated. Prioritization is typically based on a comprehensive consideration of resource utilization, storage urgency, and system load, ensuring that high-priority data is migrated first.
[0082] After prioritizing the data, a suitable migration window is scheduled for each cold data file based on the system's operating status and resource load. The scheduling needs to consider off-peak hours, bandwidth utilization, and potential business interruption risks.
[0083] In one specific implementation, within a local data center, cold data can be migrated to different storage media (such as SSDs, HDDs, and tape storage) based on storage needs. The selection of the migration destination location is based on data access requirements and cost-effectiveness; for example, high-value cold data is migrated to low-speed HDDs, and low-value cold data is migrated to tape storage.
[0084] Priority can be dynamically adjusted. For example, during peak data center load periods, the migration priority of non-critical cold data can be automatically reduced, while the migration of cold data for critical business operations can be prioritized to ensure the overall performance and stability of the system.
[0085] By analyzing historical load data from the data center, migration times can be intelligently scheduled. For example, cold data migration tasks can be scheduled for nighttime or off-peak hours to minimize disruption to other system tasks.
[0086] In another specific implementation, in a cloud environment, priority can be assigned based on storage tier. For example, frequently accessed cold data can be migrated to higher-tier cloud storage (such as standard storage), while infrequently accessed cold data can be migrated to lower-tier storage (such as cold storage or archive storage).
[0087] By generating a data migration plan based on identified cold data, data storage management can be effectively optimized. This migration plan clearly defines the target location, priority, and migration timeline for cold data, making its storage and management more systematic and refined. It reduces unnecessary computational resource consumption, improves storage space utilization, and avoids the impact of frequent access to hot data on system performance, ultimately enhancing the overall system efficiency and the effectiveness of storage resource utilization.
[0088] S40, according to the data migration plan, the cold data is migrated from the original storage location to the target remote storage, and the table partition path in the metadata repository is updated to point to the new storage location after migration.
[0089] In this embodiment, based on the previously generated data migration plan, the system initiates a data migration task to migrate the identified cold data from its original storage location to a designated target remote storage. This step ensures that the cold data is removed from the original storage location, which consumes a significant amount of computing resources. During the migration process, an appropriate transmission protocol (such as FTP, SFTP, HTTP, etc.) can be selected to ensure the security and stability of data transmission.
[0090] Cold data can be migrated to different target remote storage locations, selected based on factors including storage cost, storage requirements, data access frequency, data security, and compliance requirements. Target remote storage can be cloud storage, archive storage devices, or off-site storage systems. During the data migration process, the system may divide large files into chunks for transmission to ensure efficiency. After migration, the system will perform data integrity checks to ensure that no data has been damaged or lost during transmission.
[0091] After data migration is complete, the system automatically updates the table partition path information in the metadata repository, pointing the table partitions to the newly migrated storage location. This operation ensures that subsequent access to the data can correctly locate the new storage location, thereby guaranteeing data consistency and accessibility.
[0092] While updating the table partition path, file-related metadata information, such as storage location and file status, can also be updated simultaneously. This ensures the system uses the latest metadata for future data retrieval, analysis, and management. This step further guarantees the accuracy and reliability of the system's data management.
[0093] In one specific implementation, within a local data center, cold data can be transmitted in segments via the internal network. By utilizing the high transmission rate of the internal network bandwidth, large files can be divided into smaller data packets for transmission, reducing network congestion during the transmission process and improving transmission efficiency.
[0094] During data migration, you can optimize the local storage path based on file size, type, and frequency of use, and update the path information in the metadata repository to ensure efficient use of local data storage.
[0095] By migrating cold data from its original storage location to the target remote storage according to a data migration plan, and updating the table partition paths in the metadata repository in a timely manner, efficient and secure data migration can be ensured, while guaranteeing the correctness of data access during the migration process. The migrated cold data can automatically locate the new storage location, improving the utilization of system storage resources, reducing storage costs, enhancing system scalability and flexibility, and reducing the coupling between computing and storage resources, thereby improving the overall system efficiency and performance.
[0096] S50, after completing the data migration, deletes the migrated cold data from the original storage location.
[0097] In this embodiment, after the cold data is successfully migrated to the target remote storage, the system confirms the integrity and consistency of the migrated data through a verification mechanism (such as a checksum or file hash value) to ensure that the data is not damaged or lost during the migration process.
[0098] After the migration is confirmed to be complete, the system deletes the migrated cold data from the original storage location. The deletion operation can use the standard deletion command of the file system or perform specific deletion operations (such as file erasure, marking as overwriteable, etc.) based on the characteristics of the storage medium.
[0099] The deletion process employs secure deletion technology to ensure that migrated cold data cannot be recovered by unauthorized access. This may include measures such as multiple overwrites and file fragmentation cleanup, and is particularly suitable for scenarios containing sensitive or compliant data.
[0100] After the deletion operation is completed, the system will remark the released storage space as available space and update the space utilization of the storage system to ensure that new data can make full use of the released storage resources.
[0101] In another specific implementation, within a cloud environment, the system utilizes automated deletion mechanisms provided by cloud services (such as Amazon S3's lifecycle management) to periodically clean up migrated cold data, automatically freeing up space in the original storage location. This approach reduces manual intervention and ensures that deletion operations are performed according to predetermined strategies. After data deletion, the cloud service can retain versions or snapshots of the deleted data for a period of time as a means of recovery backup. Once the data is verified to be correct, these snapshots are also automatically deleted, completely freeing up storage space.
[0102] By deleting migrated cold data from its original storage location after migration is complete, storage space is effectively freed up and optimized. This not only ensures the security and integrity of cold data during the migration process but also avoids the risk of data leakage through secure deletion technology, meeting data security and compliance requirements. Furthermore, the freed-up storage space can be reused, thereby improving the overall resource utilization of the storage system, reducing system storage costs, and enhancing system operating efficiency.
[0103] This invention relates to a data management method based on cold data migration. It extracts metadata from a distributed file system and imports it into a metadata repository of a distributed data processing platform. Through offline analysis tasks, file information in the metadata repository is analyzed to identify cold data, and its file path information is stored in a database. Based on the identified cold data, a data migration plan is generated, including the migration target location, migration priority ranking, and migration schedule. According to this plan, the cold data is migrated from its original storage location to the target remote storage, and the table partition paths in the metadata repository are updated to point to the new storage location. After the migration is complete, the cold data in the original storage location is deleted. This invention achieves effective separation of cold data from computing resources, optimizes the utilization efficiency of storage resources, reduces the overall resource consumption of the system, and improves the scalability and performance of the system.
[0104] In one embodiment, in S20 above, analyzing the file information of metadata in the metadata repository and identifying cold data includes:
[0105] S201, Analyze the metadata in the metadata repository and extract file information related to the file, including file path, last access time, last modification time, file size and creation time;
[0106] S202, set the access frequency threshold and time point threshold for identifying cold data;
[0107] S203, based on file information, calculate the access frequency of the file, compare the access frequency of the file with the access frequency threshold, and if the access frequency of the file is lower than the access frequency threshold, mark the file as cold data.
[0108] S204, compare the last access time of the file with the time point threshold. If the last access time of the file is earlier than the time point threshold, then mark the file as cold data.
[0109] S205 classifies files according to their size and performs a comprehensive evaluation based on the file's access frequency and last access time. If the comprehensive evaluation criteria are met, the file is marked as cold data.
[0110] In this embodiment, file-related metadata is extracted from a metadata repository. This metadata includes file path, last access time, last modification time, file size, and creation time. The system iterates through the records in the metadata repository to obtain detailed information for each file, providing foundational data for subsequent analysis and identification.
[0111] System administrators or automation systems set predefined rules for identifying cold data based on business needs. These rules include:
[0112] Access frequency rule: Calculate the access frequency based on the file's access records.
[0113] Last access time rule: The judgment is based on the difference between the last access time of the file and the current time.
[0114] File size rule: Based on the file's storage size, determine whether it is classified as cold data.
[0115] The system calculates the access frequency of each file based on the extracted access record data. The calculated access frequency value is compared with a preset frequency threshold; if the frequency is lower than the threshold, the file is marked as cold data. This step typically involves the application of statistical analysis tools or algorithms to ensure the accuracy of the frequency calculation.
[0116] The system compares the file's last access time with a preset time threshold. If the file's last access time is earlier than this threshold, the system considers the file to have not been accessed for a long time and should also mark it as cold data. This step can be achieved through a simple timestamp comparison.
[0117] Files are categorized based on size. The system then combines the file's access frequency and last access time to make a comprehensive judgment and determine whether a file meets the criteria for cold data.
[0118] The system first categorizes files based on their size. File size is typically divided into several ranges, each with a different storage strategy and criterion. For example: small files (less than 1MB); medium files (between 1MB and 1GB); large files (between 1GB and 10GB); and very large files (more than 10GB).
[0119] Files within each size range may have different criteria for being classified as cold data. For example, large files that are not frequently accessed may be more easily marked as cold data, while small files require more criteria to be labeled as cold data. For files that have already been classified, the system further analyzes them by combining access frequency and last access time. For example: Access frequency: The system calculates the number of times each file is accessed within a certain time period and compares it with a preset access frequency threshold. If the access frequency of a file is lower than the threshold, it is marked as potentially cold data. Last access time: The system compares the last access time of a file with the current time. If the last access time of a file is earlier than a set time threshold, the file is further marked as cold data.
[0120] The system comprehensively analyzes the results of file size, access frequency, and last access time. This comprehensive analysis can be performed in the following ways: Weighted method: A weight is assigned to each indicator, and a weighted average score is calculated. For example, for large files, file size may have a larger weight, while for small files, access frequency and last access time may have greater weight. Files with a weighted average score below a certain threshold are marked as cold data. Logical combination method: Logical conditions are set. For example, only files with a file size greater than 1GB, an access frequency below a specific threshold, and a last access time exceeding a certain period are marked as cold data. This method ensures more accurate cold data identification.
[0121] After the above comprehensive assessment, files that meet the cold data criteria will be marked as cold data. If a file meets any rule or meets the conditions for cold data after comprehensive evaluation, it will be marked as cold data.
[0122] All files that meet the predetermined rules or are determined to be cold data are marked. The marking information usually includes the file path and the identification criteria, which facilitates subsequent data migration and management.
[0123] In one specific implementation, the system can combine weighted summation and logical combination methods for comprehensive judgment. First, a weighted summation method is used to calculate the comprehensive score, and then a secondary filtering is performed based on logical conditions to ensure the accuracy of the judgment result. Simultaneously, the system can establish a rule base and dynamically update the judgment rules based on actual operating conditions. For example, under different time periods or varying storage pressures, the system will automatically adjust the judgment criteria to ensure the timeliness and accuracy of cold data identification.
[0124] This embodiment analyzes file information in a metadata repository and identifies cold data based on predetermined rules, enabling efficient and accurate identification and classification of cold data. This not only optimizes system resource utilization and reduces unnecessary storage and computational overhead, but also provides an accurate basis for subsequent data migration and management.
[0125] In one embodiment, S10 includes:
[0126] S101, Export the current FSImage file from the distributed file system, the FSImage file containing metadata of all files and directories in the file system;
[0127] S102, convert the exported FSImage file into a file format that conforms to the table structure specification of the distributed data processing platform. The conversion includes converting metadata into a structured table format.
[0128] S103, Select the target table in the distributed data processing platform, and import the converted metadata file into the partition of the target table according to the specified partitioning rules, wherein the partitioning rules are based on date or file type;
[0129] S104 stores the imported metadata file in the metadata repository of the distributed data processing platform for use by data analysis tasks.
[0130] In this embodiment, the system exports the current FSImage file from a distributed file system (such as HDFS). The FSImage file is a metadata snapshot of the Hadoop Distributed File System, containing metadata information for all files and directories in the entire file system. This information includes file paths, sizes, block information, permissions, modification times, etc.
[0131] An FSImage file is a binary format file that is generated periodically to back up and restore file system metadata. The system obtains the latest FSImage file through command-line tools (such as Hadoop's hdfs dfsadmin-fetchImage) or API interfaces.
[0132] The exported FSImage file needs to be converted to a table structure format suitable for distributed data processing platforms such as Hive or HBase. This conversion process parses the binary data in the FSImage into readable structured data and converts it into a table format (such as CSV, JSON, Parquet, etc.) that conforms to the target platform.
[0133] The transformation process includes mapping metadata entries to columns in a table, such as mapping file paths to a path column and file sizes to a size column. Simultaneously, the system may normalize the data as needed to make it suitable for storage and subsequent analysis.
[0134] The converted metadata file needs to be imported into the target table in the distributed data processing platform. The system selects the appropriate target table based on the characteristics of the data and the analysis requirements.
[0135] During the import process, the system divides the metadata files into multiple partitions according to specified partitioning rules and imports them into the table. Partitioning rules can be based on file dates, file types, or directory structures, thereby optimizing query and analysis performance. For example, file data can be imported into different partitions by date to quickly locate data for specific time periods during subsequent analysis.
[0136] The metadata files imported into the target table are ultimately stored in the metadata repository of the distributed data processing platform for use in subsequent data analysis tasks. The metadata repository can be a dedicated database for storing and managing metadata (such as a Hive Metastore or HBase table), providing efficient metadata access and query support for data processing tasks.
[0137] This embodiment achieves centralized management and efficient analysis of metadata by extracting metadata from the distributed file system and importing it into the metadata repository of the distributed data processing platform. It converts all file and directory information in the file system into structured data, facilitating rapid querying and analysis within the distributed data processing platform. By importing metadata files in partitions according to specific rules, the system improves the organization and access efficiency of data storage, providing a reliable data foundation for subsequent data processing and optimization. Simultaneously, the automated metadata processing and storage mechanism significantly reduces the complexity of manual operations, enhancing the overall management efficiency of the system.
[0138] In one embodiment, in step S30 above, generating a data migration plan based on the identified cold data includes:
[0139] S301, determine the target location for the migration of cold data, wherein the target location includes remote storage, cloud storage, or archive storage devices;
[0140] S302, calculate the migration priority of cold data based on file size, last access time, and access frequency;
[0141] S303: Divide cold data into multiple migration batches according to migration priority, and set a migration schedule for each migration batch;
[0142] S304 generates a data migration plan based on the migration target location, migration priority, and migration time schedule of cold data.
[0143] In this embodiment, the system determines the migration target location for each piece of cold data based on the storage requirements and business strategies. The target location may include remote storage (such as off-site data centers), cloud storage (such as Amazon S3 or Google Cloud Storage), or archival storage devices (such as tape libraries or optical disc storage). The selection criteria may include storage costs, data access requirements, data retention strategies, and data security requirements.
[0144] The system calculates the migration priority for each piece of cold data by analyzing its file size, last access time, and access frequency. Priority calculation is typically based on predefined weights. For example:
[0145] File size: Larger files may be given higher migration priority because they take up more storage space.
[0146] Last access time: Files with an earlier last access time may be given higher priority, indicating that these files are unlikely to be accessed in the near future.
[0147] Access frequency: Files with lower access frequency should be migrated first, as they have less impact on system usage.
[0148] The system divides cold data into multiple migration batches based on the calculated migration priority. Each batch contains a certain number of cold data files, and the system sets a corresponding migration schedule for each batch. The migration schedule typically takes into account system load, network bandwidth, and off-peak business periods to ensure that the migration process does not have a significant impact on normal business operations.
[0149] The system generates a complete data migration plan based on the target location, priority, and timeline of the cold data migration. This plan includes the migration path, priority, target storage location, and specific migration timeline for each cold data file, providing detailed execution guidelines for the migration task.
[0150] In one specific implementation, within a cloud computing environment, the system can choose to migrate cold data to cloud storage in different geographical regions based on data storage costs, access latency, and compliance requirements. For example, for non-critical cold data, the system can choose to store it in regions with lower storage costs.
[0151] This embodiment generates a data migration plan based on identified cold data, enabling effective management of cold data. By determining the migration target location (such as remote storage, cloud storage, or archive storage devices) and calculating the migration priority of cold data based on file size, last access time, and access frequency, the system can rationally arrange the migration order and timing of data. This not only optimizes the use of storage resources but also effectively reduces storage costs and improves the efficiency of data storage and access. Furthermore, the detailed data migration plan ensures the orderliness and security of the data migration process, avoiding potential data loss or system performance degradation during migration.
[0152] In one embodiment, S301 includes:
[0153] S3011, based on the amount of storage resources occupied by file size, the reflection of data usage frequency by last access time, and the data usage reflected by access frequency, set the weights of file size, last access time, and access frequency.
[0154] S3012 standardizes data on file size, last access time, and access frequency, converting data of different scales into a unified numerical range.
[0155] S3013, based on the standardized file size, last access time and access frequency data, calculates the migration priority score for each cold data file according to the set weights. The priority score is obtained by weighting the file size, last access time and access frequency according to their respective weights.
[0156] S3014 Sort cold data files according to priority scores and assign corresponding migration orders to cold data files.
[0157] In this embodiment, the system first analyzes three key metrics: file size, last access time, and access frequency, to determine their importance in migration priority calculation. This step forms the basis for weight setting. The weight setting rules typically follow these principles:
[0158] File size weighting: File size reflects how much storage resources a file occupies. Larger files occupy more storage space, so they are usually given a higher weight when the system needs to release storage resources. This means that in priority calculations, larger, colder data may be migrated first.
[0159] Weighting of Last Access Time: Last access time refers to the time when a file was last accessed. Files that haven't been accessed for a long time are generally considered cold data because they are less relevant to current business operations or user needs. Therefore, the weight of last access time can be higher to ensure that files that haven't been accessed for a long time are processed first during the migration process.
[0160] Access frequency weighting: Access frequency represents the number of times a file is accessed within a certain period. Files with low access frequency may also be considered cold data. The system assigns weights to access frequencies based on historical access records to ensure that files with low access frequency receive appropriate migration priority.
[0161] When setting weights, the system typically seeks a balance between freeing up storage resources and maintaining data activity. For example, if current storage space is very tight, the system may increase the weight of file size; while when focusing on data availability, the system may increase the weight of access frequency and last access time.
[0162] The system can dynamically adjust the weights of various metrics based on actual storage resource usage and business needs. For example, when the system detects that storage space is about to run out, it can temporarily increase the weight of file size to release space as soon as possible; during system idle periods, the system may decrease the weight of file size to prioritize the migration of files that have not been accessed for a long time.
[0163] To ensure that metrics at different scales can be reasonably weighted and summed, the system needs to standardize data on file size, last access time, and access frequency. The purpose of standardization is to transform data with different units and ranges into a unified numerical range (e.g., between 0 and 1), thereby eliminating differences in data magnitude and making them comparable during weighted summation. Standardization methods include, but are not limited to:
[0164] Min-Max Normalization: The value of each indicator is scaled to between 0 and 1 by subtracting the minimum value and then dividing by the difference between the maximum and minimum values.
[0165] Z-score standardization: Subtract the mean from the value of each indicator and then divide by the standard deviation to make the data distribution conform to a standard normal distribution (mean is 0, standard deviation is 1).
[0166] Through standardization, the system ensures that the numerical comparisons of file size, last access time, and access frequency are meaningful, thus making the priority score obtained by weighted summation more accurate.
[0167] After completing the data standardization process, the system calculates a migration priority score for each cold data file based on preset weights. This process includes the following steps:
[0168] Weighted summation: The standardized file size, last access time, and access frequency are each multiplied by their corresponding weights to obtain a weighted value for each metric. These weighted values are then summed to obtain the file's overall priority score. The formula is as follows:
[0169] Priority score = W size × S size + W time × S time + W frequency × S frequency.
[0170] Where W represents the weight and S represents the standardized value.
[0171] Priority scores are typically values between 0 and 1; lower values indicate that the file is more likely to be cold data and should be migrated first. The score directly determines the order of data migration.
[0172] After obtaining priority scores for all cold data files, the system sorts these files from lowest to highest score. The file with the lowest score is considered to best meet the criteria for cold data and is therefore prioritized for migration. After sorting, the system assigns a migration order to each file based on the sorting results, ensuring that the files most in need of migration are processed first during the actual data migration task.
[0173] The system will execute migration tasks for each cold data file in sequence according to the preset migration time window and resource availability, ensuring that the data can be safely and effectively transferred to the designated storage location as planned.
[0174] This embodiment calculates the migration priority of cold data based on file size, last access time, and access frequency. The system can effectively identify the coldest data that should be migrated first among numerous data files. By setting reasonable weights and standardization processes, the system can comprehensively consider factors such as storage resource utilization, data activity, and system performance, ensuring that the data migration plan is executed efficiently and systematically.
[0175] In one embodiment, after deleting the migrated cold data from the original storage location in step S60 above, the method further includes:
[0176] S701 monitors the access status of migrated cold data and records the access time and frequency of cold data files.
[0177] S702 determines whether cold data has become active again based on a preset activity threshold;
[0178] S703: If the access frequency of cold data reaches the set active threshold, a data migration plan will be generated.
[0179] S704, execute the data migration plan to migrate the reactivated cold data to the original storage location or high-frequency storage area;
[0180] S705, update the file path information in the metadata repository, and update the new storage location of the migrated file to the corresponding table partition record;
[0181] S706: After completing the data migration, update the cache content and record the specific information of the migration operation.
[0182] In this embodiment, after the cold data is migrated to the target storage location (such as remote storage or archive storage), the system continues to monitor this data. The system records the access time and access frequency of each cold data file, and this data is used to determine the activity status of the cold data.
[0183] Monitoring systems typically run in real-time or periodically to ensure timely detection of re-access to cold data. The system records each access event using access logs or access counters and updates the access frequency and last access time of the cold data.
[0184] The system determines whether cold data has become active again based on preset activity thresholds. These thresholds can include multiple metrics such as access frequency and last access time. For example, if cold data is accessed frequently within a short period of time, and the access frequency reaches or exceeds the preset threshold, the system will determine that the cold data has become active again.
[0185] The criteria for judging system activity are usually based on business needs and historical data analysis, and can be adjusted according to specific use cases. For example, some businesses may require greater sensitivity to the frequency of file access, while others may focus more on the last access time.
[0186] Once the system determines that cold data has become active again, it will automatically generate a data migration plan. The migration plan includes moving the cold data from its current storage location back to its original storage location or other high-frequency storage areas to more efficiently support subsequent access needs.
[0187] The migration plan will define in detail the target location, priority, and specific time schedule for migration. The system will prioritize and migrate high-priority files based on the importance and activity level of the cold data to meet real-time business needs.
[0188] The system executes the data migration task according to the data migration plan. Reactivated cold data will be migrated to its original storage location or other high-frequency storage areas to ensure data can quickly respond to user access requests.
[0189] During the data migration process, the system may compress or transmit files in chunks to improve migration efficiency and reduce network bandwidth consumption. Simultaneously, the system will monitor the transmission progress and data integrity throughout the migration process.
[0190] After the data migration is complete, the system needs to update the file path information in the metadata repository. The migrated files will be relocated to a new storage location, and the system needs to reflect this change in the metadata repository so that the data processing platform can correctly locate the files. Specifically, the system will update the table partition records in the metadata repository to point to the new storage location after migration. This step ensures the accuracy of data access paths and the consistency of system metadata.
[0191] After the data migration is complete, the system will update the relevant cache content to ensure that the file path information in the cache is consistent with the latest metadata, thus avoiding access problems caused by cache expiration or errors.
[0192] In addition, the system records detailed information about the data migration operation, including migration time, target location, list of files to be migrated, and logs during the migration process. These records provide a basis for subsequent auditing, analysis, and troubleshooting.
[0193] This embodiment monitors data access after cold data migration, enabling the system to promptly determine if the data has become active again and automatically generate and execute a data migration plan when necessary. This ensures that data storage locations are dynamically adjusted based on actual access needs, optimizing storage resource utilization and improving data accessibility and system responsiveness. By updating metadata and cached content, the system guarantees the accuracy of data access paths, avoiding access failures or performance issues caused by incorrect paths.
[0194] In one embodiment, in step S50 above, migrating the cold data from its original storage location to a target remote storage location according to the data migration plan includes:
[0195] S501, Read the data migration plan to obtain the list of cold data files, migration target location, migration priority, and migration schedule;
[0196] S502, check the availability of the migration target location, verify whether the storage space meets the migration requirements, and verify the stability of the network connection;
[0197] S503, according to the migration priority and migration time schedule in the data migration plan, start the migration tasks of cold data files in sequence;
[0198] S504, during the migration process, cold data files exceeding a preset size threshold are split and transmitted part by part;
[0199] S505 monitors transmission progress, transmission rate, and network connection stability in real time, and records transmission data and status.
[0200] S506 After the data migration is completed, data verification is performed to check whether the data is complete and undamaged during the transmission process, and detailed log information of the migration process is recorded.
[0201] In this embodiment, the system first reads a pre-generated data migration plan to obtain a list of cold data files, migration target locations, migration priorities, and migration schedules. This step ensures that the system can accurately identify the files that need to be migrated and their corresponding target storage locations, while also clarifying the priority and time window of the migration task.
[0202] Before initiating the migration task, the system needs to verify the availability of the target storage location. This includes checking whether the target storage location has sufficient storage space to accommodate the cold data files to be migrated, and whether the network connection is stable enough to support large-scale data transfers. The purpose of this step is to ensure that the migration process can proceed smoothly and avoid migration failures due to insufficient storage space or network instability.
[0203] The system initiates the migration of cold data files sequentially according to the migration priority and schedule in the data migration plan. Files with higher priority will be migrated first to ensure that important data is processed in a timely manner. The system starts the migration within the appropriate time window as scheduled to minimize interference with other system operations.
[0204] For cold data files exceeding a preset size threshold, the system will split them and then transmit them part by part. This method helps improve transmission efficiency, especially when network bandwidth is limited. Chunked transmission can reduce the amount of data transmitted in a single transmission, reduce network load, and reduce the risk of transmission failure.
[0205] Specific methods for file splitting may include cutting to a fixed size or logically splitting based on the file's structure. The split file portions will be reassembled at the target storage location to ensure data integrity.
[0206] During the migration process, the system monitors the data transmission progress, transmission rate, and network connection stability in real time. The purpose of monitoring is to ensure smooth transmission, promptly detect and handle any anomalies (such as network jitter or reduced transmission speed), thereby guaranteeing the quality of data transmission.
[0207] The system records the data status and transmission logs during the transmission process to enable rapid location and resolution of problems. After data transmission is complete, the system verifies the migrated data to ensure that no data has been damaged or lost during transmission. Common data verification methods include checksums and hash value comparisons (such as MD5 or SHA-256).
[0208] After data verification is complete, the system will record detailed log information for the entire migration process, including migration time, transmission rate, and data integrity verification results. This log information provides crucial information for subsequent auditing, analysis, and troubleshooting.
[0209] This embodiment achieves efficient and reliable cold data migration by migrating cold data from its original storage location to a target remote storage location according to a data migration plan. By checking the availability of the target storage location and the stability of the network connection, the system ensures the smooth progress of the migration process. By splitting large files and transmitting them part by part, the system optimizes network bandwidth utilization and reduces the risk of transmission failure. Real-time monitoring and data verification further guarantee the integrity and security of the data during transmission. Finally, detailed migration logs provide strong support for system auditing, analysis, and troubleshooting, improving system management efficiency and data security.
[0210] The present invention also provides a data management device based on cold data migration, with reference to Figure 2 , Figure 2 This is a functional module diagram of a preferred embodiment of the data management device based on cold data migration of the present invention. The data management device based on cold data migration includes:
[0211] Metadata extraction and import module 10 is used to extract metadata from the distributed file system and import the metadata into the metadata repository of the distributed data processing platform.
[0212] The cold data identification and storage module 20 is used to analyze the file information of metadata in the metadata repository through the offline analysis task of the distributed data processing platform, identify cold data, and store the file path information of the cold data in the database.
[0213] The data migration plan generation module 30 is used to generate a data migration plan based on the identified cold data. The data migration plan includes the migration target location, migration priority ranking, and migration time schedule.
[0214] The data migration and metadata update module 40 is used to migrate the cold data from the original storage location to the target remote storage according to the data migration plan, and update the table partition path in the metadata repository to point to the new storage location after migration.
[0215] The data cleaning module 50 is used to delete the migrated cold data from the original storage location after the data migration is completed.
[0216] The specific implementation of the data management device based on cold data migration of the present invention is basically the same as the embodiments of the data management method based on cold data migration described above, and will not be repeated here.
[0217] This invention also provides a data management device based on cold data migration, such as... Figure 3 As shown, the data management device based on cold data migration may include: a processor 1001, such as a CPU; a communication bus 1002; a user interface 1003; a network interface 1004; and a memory 1005. The communication bus 1002 is used to establish communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or stable non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0218] Those skilled in the art will understand that Figure 3 The hardware structure of the data management device based on cold data migration shown in the figure does not constitute a limitation on the data management device based on cold data migration. It may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0219] like Figure 3As shown, the memory 1005, as a storage medium, may include an operating system, a network communication module, a user interface module, and a data management program based on cold data migration. The operating system is a program that manages and controls the data management devices and software resources based on cold data migration, supporting the operation of the network communication module, the user interface module, the data management program based on cold data migration, and other programs or software. The network communication module manages and controls the network interface 1004; the user interface module manages and controls the user interface 1003.
[0220] exist Figure 3 In the hardware structure of the data management device based on cold data migration shown, the network interface 1004 is mainly used to connect to the backend server and communicate with the backend server; the user interface 1003 is mainly used to connect to the client and communicate with the client; the processor 1001 can call the data management program based on cold data migration stored in the memory 1005 and perform the same operation as the data management method based on cold data migration.
[0221] The specific implementation of the data management device based on cold data migration of the present invention is basically the same as the various embodiments of the data management method based on cold data migration described above, and will not be repeated here.
[0222] Furthermore, this embodiment of the invention also proposes a computer storage medium storing a data management program based on cold data migration. When the data management program based on cold data migration is executed by a processor, it implements the steps of the data management method based on cold data migration as described above.
[0223] The specific implementation of the computer storage medium of the present invention is basically the same as the various embodiments of the data management method based on cold data migration described above, and will not be repeated here.
[0224] It should be noted that if any software tools or components not belonging to our company appear in the embodiments of this application, they are merely for illustrative purposes and do not represent actual use.
Claims
1. A data management method based on cold data migration, characterized in that, Includes the following steps: Extract metadata from the distributed file system and import the metadata into the metadata repository of the distributed data processing platform; The offline analysis task of the distributed data processing platform is used to analyze the file information of the metadata in the metadata repository, identify cold data, and store the file path information of the cold data in the database. The migration target location for cold data is determined, including remote storage, cloud storage, or archive storage devices. Based on the storage resource consumption of file size, the reflection of data usage frequency by last access time, and the data usage reflected by access frequency, weights are assigned to file size, last access time, and access frequency. The data of file size, last access time, and access frequency are standardized to convert data of different scales into a unified numerical range. Based on the standardized data of file size, last access time, and access frequency, a migration priority score is calculated for each cold data file according to the set weights. The priority score is obtained by weighted summation of file size, last access time, and access frequency according to their respective weights. Cold data files are sorted according to their priority scores, and corresponding migration orders are assigned to them. Cold data is divided into multiple migration batches according to migration priority, and a migration schedule is developed for each batch. Based on the migration target location, migration priority, and migration schedule of the cold data, a data migration plan is generated. The data migration plan includes the migration target location, migration priority ranking, and migration schedule. According to the data migration plan, the cold data is migrated from its original storage location to the target remote storage, and the table partition path in the metadata repository is updated to point to the new storage location after migration. After the data migration is complete, delete the migrated cold data from the original storage location.
2. The data management method based on cold data migration as described in claim 1, characterized in that, Analyze the file information of metadata in the metadata repository and identify cold data, including: Analyze the metadata in the metadata repository to extract file information related to the file, including file path, last access time, last modification time, file size, and creation time; Set access frequency thresholds and time point thresholds for identifying cold data; Based on file information, the access frequency of the file is counted and compared with an access frequency threshold. If the access frequency of the file is lower than the access frequency threshold, the file is marked as cold data. The last access time of the file is compared with a time point threshold. If the last access time of the file is earlier than the time point threshold, the file is marked as cold data. Files are classified according to their size, and a comprehensive evaluation is conducted based on the file's access frequency and last access time. If the comprehensive evaluation criteria are met, the file is marked as cold data.
3. The data management method based on cold data migration as described in claim 1, characterized in that, Extracting metadata from the distributed file system and importing that metadata into the metadata repository of the distributed data processing platform includes: Export the current FSImage file from the distributed file system. The FSImage file is a metadata snapshot of the file system, containing metadata for all files and directories in the file system. The exported FSImage file is converted into a file format that conforms to the table structure specification of the distributed data processing platform. The conversion includes converting metadata into a structured table format. Select the target table in the distributed data processing platform, and import the converted metadata file into the partition of the target table according to the specified partitioning rules, which are based on date or file type; The imported metadata files are stored in the metadata repository of the distributed data processing platform for use by data analysis tasks.
4. The data management method based on cold data migration as described in claim 1, characterized in that, After deleting the migrated cold data from the original storage location, the following is also included: Monitor access to migrated cold data and record the access time and frequency of cold data files. Determine whether cold data becomes active again based on a preset activity threshold; If the access frequency of cold data reaches the set activity threshold, a data migration plan will be generated. The data migration plan is executed to migrate the reactivated cold data to its original storage location or high-frequency storage area; Update the file path information in the metadata repository and update the new storage location of the migrated files to the corresponding table partition records; After the data migration is completed, update the cache content and record the specific information of the migration operation.
5. The data management method based on cold data migration as described in claim 1, characterized in that, According to the data migration plan, the cold data is migrated from its original storage location to the target remote storage, including: Read the data migration plan to obtain the list of cold data files, migration target location, migration priority, and migration schedule; Check the availability of the target migration location, verify that the storage space meets the migration requirements, and verify the stability of the network connection; Based on the migration priority and migration schedule in the data migration plan, initiate the migration tasks for cold data files sequentially. During the migration process, cold data files exceeding a preset size threshold are split and transmitted part by part; Real-time monitoring of transmission progress, transmission rate, and network connection stability; recording of transmission data and status. After the data migration is completed, data verification is performed to check whether the data is intact and undamaged during the transmission process, and detailed log information of the migration process is recorded.
6. A data management device based on cold data migration, characterized in that, The data management device based on cold data migration includes: The metadata extraction and import module is used to extract metadata from the distributed file system and import the metadata into the metadata repository of the distributed data processing platform. The cold data identification and storage module is used to analyze the file information of metadata in the metadata repository through the offline analysis task of the distributed data processing platform, identify cold data, and store the file path information of the cold data in the database. A data migration plan generation module is used to determine the migration target location for cold data. The migration target location includes remote storage, cloud storage, or archive storage devices. Based on the storage resource usage of file size, the reflection of data usage frequency by last access time, and the data usage reflected by access frequency, weights are assigned to file size, last access time, and access frequency. The data of file size, last access time, and access frequency are standardized to convert data of different scales into a unified numerical range. Based on the standardized data of file size, last access time, and access frequency, a migration priority score is calculated for each cold data file according to the set weights. The priority score is obtained by weighted summation of file size, last access time, and access frequency according to their respective weights. The cold data files are sorted according to the priority scores, and a corresponding migration order is assigned to each cold data file. The cold data is divided into multiple migration batches according to the migration priority, and a migration schedule is formulated for each migration batch. Based on the migration target location, migration priority, and migration schedule of the cold data, a data migration plan is generated. The data migration plan includes the migration target location, migration priority ranking, and migration schedule. The data migration and metadata update module is used to migrate the cold data from the original storage location to the target remote storage according to the data migration plan, and update the table partition path in the metadata repository to point to the new storage location after migration. The data cleaning module is used to delete migrated cold data from the original storage location after the data migration is completed.
7. A data management device based on cold data migration, characterized in that, The data management device based on cold data migration includes a memory, a processor, and a data management program based on cold data migration stored in the memory and executable on the processor. When the data management program based on cold data migration is executed by the processor, it implements the steps of the data management method based on cold data migration as described in any one of claims 1-5.
8. A computer storage medium, characterized in that, The storage medium stores a data management program based on cold data migration, which, when executed by a processor, implements the steps of the data management method based on cold data migration as described in any one of claims 1-5.
Citation Information
Patent Citations
Data migration method, equipment, storage medium and device
CN117827748A
Partial migration of an object to another storage location in a computer system
US20050097126A1