Data return method and device, equipment and medium
By analyzing the target file information in the data request and obtaining archived metadata information, the data relocation automatically reads and writes back from the tape library and is solved, and the data relocation operation in the existing technology is inconvenient, which improves the relocation efficiency.
Patent Information
- Application Number
- CN202510103780.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, data relocation operation is inconvenient, and the required files and relocation paths need to be manually specified, which affects the relocation efficiency.
By analyzing the target file information in the data request, we judge whether the file is an archived file, and obtaining the archived metadata information of the target file from the archive record metadata database. Based on this information, file data is read from the tape library and written back to online storage to realize automatic data relocation.
It improves the efficiency of data relocation, reduces manual operation steps, and realizes an automated data relocation process.
Smart Images

Figure CN119961214A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of storage technology, and in particular to a data migration method, device, equipment and medium. Background Art
[0002] Currently, after a file is archived, it is deleted from the online storage and the business system cannot traverse to the file. When you need to access an archived file, you need to manually specify the required file, migration path and other information. After the migration is completed, the user can query and access the file. Since you need to manually specify the required file and migration path and other information, data migration operations are inconvenient and affect migration efficiency.
[0003] Therefore, how to improve the efficiency of data migration is a problem that technical personnel in this field need to solve. Summary of the invention
[0004] The purpose of the embodiments of the present invention is to provide a data migration method, device, equipment and medium, which can improve the efficiency of data migration. The specific scheme is as follows:
[0005] In a first aspect, the present invention discloses a data migration method, comprising:
[0006] Parse the target file information carried in the data request;
[0007] Based on the target file information, the target file is read from the online storage. If the read file metadata information indicates that the target file is an archived file, metadata information of the target file after archiving is obtained from the archive record metadata database to obtain target metadata information, wherein the target metadata information includes the location of the target file in the tape library, the file metadata information in the online storage is metadata information retained in the online storage and added with an archive mark when the file is archived, and the archived file is a file archived from the online storage to the tape library;
[0008] The file data of the target file is read from the tape library based on the target metadata information, and the file data is written back to the online storage.
[0009] Optionally, also include:
[0010] Regularly scan files in online storage based on preset archiving policies and write files that meet archiving conditions into the tape library;
[0011] Truncate the files written to the tape library, retain the file metadata information, and add an archive tag to mark the files written to the tape library as archive files.
[0012] Optionally, before writing files that meet the archiving criteria to the tape library, it also includes:
[0013] Aggregate files with the same characteristics among the files that meet the archiving conditions to obtain aggregated files;
[0014] Accordingly, the files meeting the archiving condition are written into the tape library, including: continuously writing the aggregated files into the same tape in the tape library.
[0015] Optionally, the step of aggregating files having the same characteristics among the files meeting the archiving condition to obtain an aggregated file includes:
[0016] Aggregate files with the same characteristics among files that meet the archiving conditions based on a preset aggregation strategy to obtain an aggregated file;
[0017] The same features configured in the preset aggregation policy include one or more of belonging to the same user, belonging to the same directory, and the same file type.
[0018] Optionally, write the aggregate files consecutively to the same tape in a tape library, including:
[0019] When the preset aggregate data volume threshold is reached, the aggregate files are continuously written to the same tape in the tape library.
[0020] Optionally, the step of reading the target file from the online storage based on the target file information, and if the read file metadata information indicates that the target file is an archived file, acquiring the metadata information of the target file after archiving from the archive record metadata database to obtain the target metadata information includes:
[0021] Using the storage service to read the target file from the online storage based on the target file information, if it is found according to the file metadata information that the target file has been truncated and marked as archived, initiating a migration instruction to the master control service;
[0022] Using the master control service to read the archive record metadata database, and obtain the metadata information of the target file after archiving;
[0023] Accordingly, reading the file data of the target file from the tape library based on the target metadata information and writing the file data back to the online storage includes:
[0024] Using the master control service to generate a migration job based on the metadata information of the target file after archiving;
[0025] Using the archiving service to call the tape library read interface based on the relocation operation, the file data of the target file is retrieved from the tape library and rewritten into the online storage;
[0026] The master control service is used to notify the storage service that the data has been migrated back, so that the storage service reads the file data of the target file from the online storage and returns it to the data requester.
[0027] Optionally, also include:
[0028] When any file is stored in the online storage, the metadata information of the file is synchronized to the search engine, wherein the metadata information of the file includes the file path information;
[0029] Correspondingly, the parsing of the target file information carried in the data request includes:
[0030] The target file path information carried in the data request is parsed, wherein the target file path information is the path information of the file meeting the search condition returned by the search engine to the requesting party.
[0031] In a second aspect, the present invention discloses a data migration device, comprising:
[0032] A file information acquisition module, used to parse the target file information carried in the data request;
[0033] a target file reading module, configured to read the target file from the online storage based on the target file information, and if the read file metadata information indicates that the target file is an archived file, obtain the metadata information of the target file after archiving from the archive record metadata database to obtain the target metadata information, wherein the target metadata information includes the location of the target file in the tape library, the file metadata information in the online storage is the metadata information retained in the online storage and added with an archive mark when the file is archived, and the archived file is a file archived from the online storage to the tape library;
[0034] The target file relocation module is used to read the file data of the target file from the tape library based on the target metadata information, and write the file data back to the online storage.
[0035] In a third aspect, the present invention discloses an electronic device, comprising:
[0036] Memory for storing computer programs;
[0037] The processor is used to execute the computer program to implement the steps of the aforementioned data migration method.
[0038] In a fourth aspect, the present invention discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the aforementioned data migration method are implemented.
[0039] In a fifth aspect, the present invention provides a computer program product, including a computer program / instruction, which implements the steps of the aforementioned disclosed data migration method when executed by a processor.
[0040] It can be seen from the above scheme that the present invention provides a data migration method, including: parsing the target file information carried in the data request; reading the target file from the online storage based on the target file information, if the read file metadata information indicates that the target file is an archived file, then obtaining the metadata information of the target file after archiving from the archive record metadata database to obtain the target metadata information, wherein the target metadata information includes the position of the target file in the tape library, the file metadata information in the online storage is the metadata information retained in the online storage and added with the archiving mark when the file is archived, and the archived file is the file archived from the online storage to the tape library; reading the file data of the target file from the tape library based on the target metadata information, and writing the file data back to the online storage.
[0041] It can be seen that the beneficial effects of the present invention are: based on the target file information carried in the data request, the target file is read from the online storage. Since the archived file retains the metadata information in the online storage and adds an archive mark, the target metadata information is read to determine whether the file is an archived file. If the target file is an archived file, the file data of the target file is read from the tape library based on the metadata information of the target file after archiving obtained from the archive record metadata database, and rewritten to the online storage, thereby realizing automatic data migration from the tape library and improving data migration efficiency.
[0042] Correspondingly, a data migration device, equipment and medium provided by the present invention also have the above technical effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0044] Figure 1 A flow chart of a data migration method provided by an embodiment of the present invention;
[0045] Figure 2 A schematic diagram of a data archiving process provided by an embodiment of the present invention;
[0046] Figure 3 A schematic diagram of data migration provided by an embodiment of the present invention;
[0047] Figure 4 A schematic diagram of data migration based on a search engine provided by an embodiment of the present invention;
[0048] Figure 5 A schematic diagram of the structure of a data migration device provided by an embodiment of the present invention;
[0049] Figure 6 A structural diagram of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0051] Data archiving refers to the classification of long-term idle massive unstructured data into hot, warm, and cold data according to strategies. The software system automatically migrates a large amount of unused data from online storage to quasi-online, near-line, and offline storage devices to achieve long-term and secure archiving storage of massive historical data, and support convenient data retrieval, recovery, and reuse. Global data is growing rapidly and is expected to reach 180ZB soon, of which 90% is unstructured data, and 85%-90% of unstructured data is cold data. Among cold data, according to the rules of different industries, 80% of the data needs to be stored for more than 10 years, 70% of the data needs to be stored for more than 20 years, and 55% of the data needs to be stored for more than 30 years. The change in data access frequency and data volume will change inversely with the migration of time. According to this inverse proportional change, the value of data is evaluated, low-cost data storage methods are used for a large amount of low-value information, and high-value data is stored online safely. Store data in different types of storage devices (solid-state disks, mechanical disks, tapes, and Blu-ray discs), and migrate data between different storage devices through hierarchical storage, thereby reducing data center storage costs and energy consumption. After traditional archiving technology archives files to a tape library, when you need to access the files later, you need to manually specify the file to be migrated, the migration path, and other information. The business system can only access the file after the migration is completed.
[0052] Currently, after a file is archived, it is deleted from the online storage, and the business system cannot traverse to the file. When you need to access an archived file, you need to manually specify the required file, migration path and other information. After the migration is completed, the user can query and access the file. Since you need to manually specify the required file and migration path and other information, the data migration operation is inconvenient, which affects the migration efficiency. To this end, the present invention provides a data migration solution that can improve the data migration efficiency.
[0053] See also Figure 1 As shown, an embodiment of the present invention provides a data migration method, including:
[0054] Step S11: parse out the target file information carried in the data request.
[0055] The data request may be a request sent by a business system, and may be a data request such as reading, downloading, or previewing. A business system may be understood as a system composed of a series of interrelated components used to execute specific business functions and processes. The target file information is information that can identify the target file, such as file path information.
[0056] In an optional implementation, when any file is stored in the online storage, the metadata information of the file is synchronized to the search engine, wherein the metadata information of the file includes file path information; accordingly, parsing the target file information carried in the data request includes: parsing the target file path information carried in the data request, wherein the target file path information is the path information of the file that meets the search criteria and is returned by the search engine to the requester.
[0057] Among them, online storage is working-level storage. The storage device and the stored data are always online and can be read and modified at any time to meet the speed requirements of the front-end application server or database for data access. For example, online storage is a hard disk pool. The search engine can be a component with search function, such as ES (i.e. Elasticsearch, an open source distributed search engine). ES can be used to implement functions such as search, log statistics, analysis, and system monitoring. The metadata information of the file can include information such as the file path and file size.
[0058] That is, in the embodiment of the present invention, after the file is stored online, the original data information can be synchronized to the search engine. When performing data retrieval, the search request first determines the file that meets the search criteria based on the request of the data requester as the target file, and returns the file path information of the target file. The data requester reads the target file from the online storage based on the file path information. In this way, unified retrieval of online storage and tape libraries can be achieved through the search engine.
[0059] Step S12: reading the target file from the online storage based on the target file information; if the read file metadata information indicates that the target file is an archived file, obtaining the metadata information of the target file after archiving from the archive record metadata database to obtain the target metadata information, wherein the target metadata information includes the location of the target file in the tape library; the file metadata information in the online storage is metadata information retained in the online storage and added with an archive mark when the file is archived; the archived file is a file archived from the online storage to the tape library.
[0060] In an optional implementation, the storage service can be used to read the target file from the online storage based on the target file information. If the target file is found to have been truncated and marked as archived according to the file metadata information, a migration instruction is sent to the master control service; the master control service is used to read the archive record metadata database to obtain the metadata information of the target file after archiving. That is, in the embodiment of the present invention, after the file is archived, a truncation operation will be performed to truncate the file body, retain the file metadata information, and add an archived mark.
[0061] In an optional implementation, files in online storage can be scanned periodically based on a preset archiving strategy, and files that meet the archiving criteria can be written to a tape library; files written to the tape library can be truncated, file metadata information can be retained, and an archiving tag can be added to mark the files written to the tape library as archived files.
[0062] The embodiment of the present invention can complete data archiving in a permanent incremental manner. The permanent incremental manner is that after the initial archiving is completed, only the changed data is archived in the subsequent archiving process. The archiving condition is the condition for filtering archived files set in the archiving policy. The archiving policy may include but is not limited to policies based on archiving time, files or directories, file time, etc. After archiving, the file is truncated to release the storage space of the online storage.
[0063] Among them, the archiving strategy based on archiving time is a strategy for archiving based on the time dimension. For example, you can set an archiving cycle, such as archiving specific data or the data of the entire system every day, week, month, or year. The archiving strategy based on archiving time helps to ensure the timeliness and integrity of the data, while facilitating subsequent data recovery and search. The archiving cycle can be set according to the data life cycle and business needs. For data that needs to be stored for a long time but is less frequently accessed, a longer archiving cycle can be set; for data that needs to be accessed frequently, a shorter archiving cycle can be set.
[0064] The archiving strategy based on files or directories is a strategy for archiving based on the specific attributes of files or directories. It can archive based on the specific attributes of files or directories (such as name, type, size, etc.). For example, you can select the data to be archived based on attributes such as file type (such as documents, pictures, videos, etc.), file size, and file creation or modification date. In this way, you can more accurately control the content and scope of the archive, avoiding unnecessary data redundancy and storage waste. At the same time, it also helps to improve the efficiency and accuracy of data management.
[0065] The archiving strategy based on file time is an archiving strategy based on the time attribute of the file itself. It can be archived based on the attributes such as the creation time, modification time or access time of the file. Compared with the archiving time strategy, it focuses more on archiving based on the time attribute of the file itself. For example, you can set it to archive only files within the recent period of time, or only archive files created or modified within a certain period of time. This helps to ensure the timeliness and relevance of the data, and also helps to reduce the storage space occupied. In actual operation, you can set specific time ranges and archiving conditions based on business needs and data characteristics.
[0066] In addition to the above archiving strategies, archiving strategies can also be combined and customized according to actual needs. For example, two strategies, namely, archiving by time and by file or directory, can be used simultaneously to formulate a more comprehensive archiving strategy. The embodiment of the present invention can be comprehensively considered according to data characteristics and business needs. Through a reasonable archiving strategy, the integrity, security and accessibility of the data can be ensured, while also helping to improve the efficiency and accuracy of data management.
[0067] In an optional implementation, before writing the files that meet the archiving conditions into the tape library, the files that meet the archiving conditions and have the same characteristics may be aggregated to obtain an aggregated file; accordingly, writing the files that meet the archiving conditions into the tape library includes: continuously writing the aggregated file into the same tape in the tape library. Aggregating files with the same characteristics and continuously writing them into the same tape can improve data reading efficiency.
[0068] In an optional implementation, files having the same characteristics among files that meet the archiving conditions may be aggregated based on a preset aggregation strategy to obtain an aggregated file; wherein the same characteristics configured in the preset aggregation strategy include one or more of belonging to the same user, belonging to the same directory, and the same file type.
[0069] In the embodiment of the present invention, the preset aggregation policy includes but is not limited to policies corresponding to the same user, the same directory, and the same file type.
[0070] In an optional implementation, the aggregate files may be continuously written to the same tape in the tape library, including: when a preset aggregate data volume threshold is reached, the aggregate files are continuously written to the same tape in the tape library.
[0071] In this way, when the aggregate data volume threshold is reached, the aggregate files are continuously written to the tape library, thereby avoiding frequent writing to the tape library.
[0072] In an optional implementation, the storage service can periodically scan the data in the online storage based on a preset archiving strategy, aggregate the files that meet the archiving conditions, and initiate an archiving instruction to the master control service after reaching the aggregated data volume threshold. After receiving the archiving instruction, the master control service generates an archiving job. The master control service notifies the archiving service to start the archiving job. The archiving service first completes the pre-actions such as tape loading and head positioning, and then calls the storage service interface to read the aggregated archived files. The archiving service calls the tape library write interface. After the writing is completed, the archiving service truncates the file, retains the file's metadata information, and adds an archived mark to indicate that it has been archived to the tape library. At the same time, the master control service records the metadata information of the archived file after archiving in the archive record metadata database, including information such as the location of the file data on the tape, so that the files recorded on the tape can be read during the back-to-source operation.
[0073] Step S13: reading the file data of the target file from the tape library based on the target metadata information, and writing the file data back to the online storage.
[0074] In an optional implementation, the master control service can be used to generate a migration operation based on the metadata information of the target file after archiving; the archiving service can be used to call the tape library read interface based on the migration operation to retrieve the file data of the target file from the tape library and rewrite it into the online storage; the master control service can be used to notify the storage service that the data has been migrated, so that the storage service reads the file data of the target file from the online storage and returns it to the data requester. Among them, the location corresponding to the file path information in the online storage can be written.
[0075] It can be seen that the embodiment of the present invention reads the target file from the online storage based on the target file information carried in the data request. Since the archived file retains the metadata information in the online storage and adds an archive mark, the target metadata information is read to determine whether the file is an archived file. If the target file is an archived file, the file data of the target file is read from the tape library based on the metadata information of the target file after archiving obtained from the archive record metadata database, and rewritten to the online storage, thereby realizing automatic data migration from the tape library and improving data migration efficiency.
[0076] The embodiment of the present invention can realize data hierarchical storage and data migration based on the master control service, archiving service, archiving record database and archiving strategy. In addition, combined with the search engine, the business system can uniformly search for files in online storage and tape library.
[0077] Master control service: Receives archiving and relocation instructions from the business system, and issues the task of migrating data files between different storage devices to the archiving service, which then completes the final data migration. The master control service is responsible for allocating and scheduling tasks between multiple archiving services, so that they can work in parallel and coordinate with each other, thereby ensuring efficient and stable operation of the archiving operation. After the archiving operation is completed, the master control service updates the archiving operation status, including the archiving time, file size, location of the file in the tape library, and tape identifiers of multiple copies, etc., so that the file can be migrated back later based on this information.
[0078] Archiving service: This module receives archiving and relocation tasks issued by the master control service, scans the specified file directory regularly according to the set archiving strategy, saves the files that meet the archiving strategy to the tape library through the tape library interface, and truncates the files to release the storage space occupied by the files. At the same time, it updates the file metadata information and marks it as archived. After archiving is completed, the master control service is notified, and the master control service completes the metadata update of the archived record.
[0079] Archive record database: the metadata database of the archive system, which records the metadata information of the archived data, including but not limited to: file name, file size, storage location, archive time and other information.
[0080] Archiving strategy: This module can configure archiving strategy and archiving cycle. Create automatic archiving strategy to complete data archiving in permanent increment mode; archiving strategies include: archiving time, file or directory, file time, etc.
[0081] For further information, see Figure 2 As shown, an embodiment of the present invention provides a schematic diagram of a data archiving process. Specifically, the following steps may be included:
[0082] 21. Initiate archiving tasks: According to the pre-configured archiving strategy, the storage service regularly scans the data in the associated hard disk pool, aggregates the data that meets the archiving conditions, and initiates an archiving instruction to the archiving master service after reaching the preset number of aggregates. The archiving instruction carries the corresponding archiving task information. Aggregation is for better efficiency in subsequent source return and to avoid long client waits due to long tape addressing time. Therefore, during the archiving process, the relevant data is aggregated, data with the same characteristics are aggregated together, and in subsequent archiving tasks, it is ensured that it is continuously written to a tape to achieve continuous storage on the tape.
[0083] 22. Generate archiving job: After receiving the archiving instruction, the master control service generates the archiving job.
[0084] 23. Read the data to be archived: The master control service notifies the archiving service to start the archiving operation. The archiving service first completes the pre-actions such as tape loading and head positioning, and then calls the storage service interface to read the aggregated archived files.
[0085] 24. Archive data is written into the tape library: The archive service calls the tape library write interface. After writing is completed, the archive service truncates the file and updates the file metadata information to indicate that it has been archived to the tape library.
[0086] At the same time, the master control service records the location of the archived file on the tape and other information in the archive record metadata database so that the files recorded on the tape can be read during the return operation.
[0087] For further information, see Figure 3 As shown, an embodiment of the present invention discloses a schematic diagram of data migration. The data migration process may include the following steps:
[0088] 31. Read the file.
[0089] 32. Initiate a relocation command: When the business system reads a file through a storage service (such as NFS, CIFS protocol), the storage service finds that the file has been truncated and marked as archived based on the file metadata information. The storage service initiates a relocation command to the master control service.
[0090] 33. Query archive metadata: The master control service reads the archive metadata service to obtain metadata information of archived files (such as archive tape location, etc.).
[0091] 34. Generate a migration job: The master control service generates a migration job based on the migration instruction and the metadata information of the archived files.
[0092] 35. Read data from the tape library.
[0093] 36. Write to online storage. The master control service calls the archive service to perform the migration operation. The archive service calls the tape library read interface to retrieve the file from the tape library and rewrite it to the online storage.
[0094] After completing the data retrieval, the master control service notifies the storage service that the data has been migrated back. The storage service continues the previous data reading process, reads the actual data from the hard disk, and returns it to the business system.
[0095] For further information, see Figure 4As shown, the embodiment of the present invention discloses a schematic diagram of data migration based on a search engine. The full-text search engine service is implemented through ES, thereby realizing the function of fast retrieval of file metadata. When the file is stored on the disk, the storage service synchronously reports the metadata to ES, and the metadata retrieval function of the file is realized through ES.
[0096] 401. Retrieve data: The business system sends a retrieval request to retrieve data.
[0097] 402. Query file metadata: The retrieval service queries metadata through ES.
[0098] 403. Return file path information: The retrieval service returns the queried file path information.
[0099] That is, in the embodiment of the present invention, when the business system retrieves data according to the retrieval condition, the retrieval service queries the list of files that meet the condition through the ES and returns the file path information.
[0100] 404. Download data request: When the business system needs to preview or download a file, it sends a download data request to the storage service based on the file path information.
[0101] 405. Read download data: The storage system reads files from the online storage pool and determines whether the files have been archived to the tape library based on the file metadata information. If the files have not been archived, the download data is directly returned to the business system. When reading files based on the file path information, the file metadata information is read first.
[0102] 406. Initiate a relocation command: If the file has been archived to the tape library, the storage service initiates a relocation command to the master control service, and the master control service reads the metadata service to obtain the metadata information of the archived file (such as the location of the archived tape, etc.).
[0103] 407. Generate a relocation job: The master control service calls the archive service to generate a relocation job.
[0104] 408. Reading migrated data from the tape library: The archive service calls the tape library read interface to retrieve the file from the tape library.
[0105] 409. Write the retrieved data to the online storage: After retrieving the file from the tape library, write it back to the online storage.
[0106] 410. Notify archiving job: After completing data migration, the archiving service notifies the master control service that the archiving job has been completed.
[0107] 411: Return the downloaded data. The master control service updates the migration task information. At the same time, the storage service reads the migrated data and returns the downloaded data to the business system. The migration task information may include the task completion time, etc.
[0108] The embodiment of the present invention realizes hierarchical storage and global retrieval of data, and automatically archives cold data that is not frequently accessed into a tape library according to the archiving strategy, thereby reducing storage costs. At the same time, data with the same characteristics are aggregated together during the archiving process, and are ensured to be continuously written to a tape in subsequent archiving tasks, so as to realize continuous storage on the tape and achieve higher reading efficiency during the migration process. After archiving is completed, the file is truncated to release the storage space occupied by the file, and is marked as archived in the file metadata. When the business system needs to access the data, it can be automatically migrated from the tape library to avoid the inconvenience caused by manual migration, and the business system is unaware. At the same time, combined with ES, unified data retrieval across online storage and tape libraries is completed, and file retrieval and preview download operations can be completed unaware of the retrieval conditions.
[0109] In this way, cold data is archived to the tape library according to the archiving strategy, reducing storage costs. During the archiving process, data with the same characteristics are aggregated together, and in subsequent archiving tasks, they are continuously written to a tape to achieve continuous storage on the tape, so as to achieve higher reading efficiency during the migration process. Unified data retrieval across online storage and tape libraries is achieved, and file migration operations are completed imperceptibly based on the retrieval results, improving ease of use.
[0110] See also Figure 5 As shown, an embodiment of the present invention discloses a data migration device, including:
[0111] The file information acquisition module 51 is used to parse the target file information carried in the data request;
[0112] The target file reading module 52 is used to read the target file from the online storage based on the target file information. If the read file metadata information indicates that the target file is an archived file, the metadata information of the target file after archiving is obtained from the archive record metadata database to obtain the target metadata information, wherein the target metadata information includes the location of the target file in the tape library, the file metadata information in the online storage is the metadata information retained in the online storage and added with an archive mark when the file is archived, and the archived file is a file archived from the online storage to the tape library;
[0113] The target file relocation module 53 is used to read the file data of the target file from the tape library based on the target metadata information, and write the file data back to the online storage.
[0114] The device further comprises a data archiving module, which is used to:
[0115] Based on the preset archiving strategy, the files in the online storage are scanned regularly, and the files that meet the archiving conditions are written to the tape library; the files written to the tape library are truncated, the file metadata information is retained, and an archiving tag is added to mark the files written to the tape library as archived files.
[0116] The data archiving module also includes an aggregation submodule, which is used to aggregate files with the same characteristics among the files that meet the archiving conditions before writing them into the tape library to obtain aggregated files; accordingly, the data archiving module is used to continuously write the aggregated files to the same tape in the tape library.
[0117] In an optional embodiment, the aggregation submodule can be used to aggregate files with the same characteristics among files that meet the archiving conditions based on a preset aggregation strategy to obtain aggregated files; wherein the same characteristics configured in the preset aggregation strategy include one or more of belonging to the same user, belonging to the same directory, and the same file type.
[0118] In an optional implementation, the data archiving module may be configured to continuously write aggregate files to the same tape in the tape library when a preset aggregate data volume threshold is reached.
[0119] In an optional embodiment, the target file reading module 52 can be used to use the storage service to read the target file from the online storage based on the target file information. If it is found according to the file metadata information that the target file has been truncated and marked as archived, a migration instruction is initiated to the master control service; the archive record metadata database is read using the master control service to obtain the metadata information of the target file after archiving.
[0120] Correspondingly, the target file migration module 53 can be used to utilize the master control service to generate a migration operation based on the metadata information of the target file after archiving; utilize the archiving service to call the tape library read interface based on the migration operation to retrieve the file data of the target file from the tape library and rewrite it into the online storage; utilize the master control service to notify the storage service that the data has been migrated, so that the storage service reads the file data of the target file from the online storage and returns it to the data requester.
[0121] The device also includes:
[0122] A metadata synchronization module, used to synchronize the metadata information of any file to the search engine when the file is stored in the online storage, wherein the metadata information of the file includes the file path information;
[0123] Correspondingly, the file information acquisition module 51 can be used to parse out the target file path information carried in the data request, wherein the target file path information is the path information of the file meeting the search condition returned by the search engine to the requesting party.
[0124] It can be seen that the embodiment of the present invention reads the target file from the online storage based on the target file information carried in the data request. Since the archived file retains the metadata information in the online storage and adds an archive mark, the target metadata information is read to determine whether the file is an archived file. If the target file is an archived file, the file data of the target file is read from the tape library based on the metadata information of the target file after archiving obtained from the archive record metadata database, and rewritten to the online storage, thereby realizing automatic data migration from the tape library and improving data migration efficiency.
[0125] Figure 6 A structural diagram of an electronic device provided by an embodiment of the present invention, such as Figure 6 As shown, the electronic device includes: a memory 60 for storing a computer program;
[0126] The processor 61 is used to implement the steps of the data migration method in the above embodiment when executing a computer program.
[0127] The electronic device provided in this embodiment may include but is not limited to a smart phone, a tablet computer, a laptop or desktop computer, a server, etc.
[0128] Among them, the processor 61 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 61 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 61 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 61 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 61 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.
[0129] The memory 60 may include one or more computer-readable storage media, which may be non-transitory. The memory 60 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 60 is at least used to store the following computer program 601, wherein, after the computer program is loaded and executed by the processor 61, it can implement the relevant steps of the data migration method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 60 may also include an operating system 602 and data 603, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 602 may include Windows, Unix, Linux, etc. Data 603 may include but is not limited to the configuration for executing the data migration method, etc.
[0130] In some embodiments, the electronic device may further include a display screen 62 , an input / output interface 63 , a communication interface 64 , a power supply 65 , and a communication bus 66 .
[0131] Those skilled in the art will understand that Figure 6 The structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure.
[0132] It is understandable that if the data migration method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the current technology or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and executes all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, magnetic disk or optical disk and other media that can store program codes.
[0133] Based on this, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned data migration method are implemented.
[0134] A computer program product provided by an embodiment of the present invention is introduced below. The computer program product described below can be referenced to other embodiments described in this document.
[0135] A computer program product includes a computer program / instruction, which implements the steps of the aforementioned disclosed data migration method when executed by a processor.
[0136] The above is a detailed introduction to a data migration method, device, equipment and medium provided in an embodiment of the present invention. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0137] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0138] The above is a detailed introduction to a data migration method, device, equipment and medium provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A data migration method, characterized in that: include: Parse the target file information carried in the data request; Based on the target file information, the target file is read from the online storage. If the read file metadata information indicates that the target file is an archived file, metadata information of the target file after archiving is obtained from the archive record metadata database to obtain target metadata information, wherein the target metadata information includes the location of the target file in the tape library, the file metadata information in the online storage is metadata information retained in the online storage and added with an archive mark when the file is archived, and the archived file is a file archived from the online storage to the tape library; The file data of the target file is read from the tape library based on the target metadata information, and the file data is written back to the online storage.
2. The data migration method according to claim 1, characterized in that: Also includes: Regularly scan files in online storage based on preset archiving policies and write files that meet archiving conditions into the tape library; Truncate the files written to the tape library, retain the file metadata information, and add an archive tag to mark the files written to the tape library as archive files.
3. The data migration method according to claim 2, characterized in that: Before writing files that meet the archiving criteria to the tape library, it also includes: Aggregate files with the same characteristics among the files that meet the archiving conditions to obtain aggregated files; Accordingly, the files meeting the archiving condition are written into the tape library, including: continuously writing the aggregated files into the same tape in the tape library.
4. The data migration method according to claim 3, characterized in that: The step of aggregating files having the same characteristics among the files meeting the archiving condition to obtain an aggregated file includes: Aggregate files with the same characteristics among files that meet the archiving conditions based on a preset aggregation strategy to obtain an aggregated file; The same features configured in the preset aggregation policy include one or more of belonging to the same user, belonging to the same directory, and the same file type.
5. The data migration method according to claim 3, characterized in that: Write aggregate files consecutively to the same tape in a tape library, including: When the preset aggregate data volume threshold is reached, the aggregate files are continuously written to the same tape in the tape library.
6. The data migration method according to claim 2, characterized in that: The step of reading the target file from the online storage based on the target file information, and if the read file metadata information indicates that the target file is an archived file, acquiring the metadata information of the target file after archiving from the archive record metadata database to obtain the target metadata information, includes: Using the storage service to read the target file from the online storage based on the target file information, if it is found according to the file metadata information that the target file has been truncated and marked as archived, initiating a migration instruction to the master control service; Using the master control service to read the archive record metadata database, and obtain the metadata information of the target file after archiving; Accordingly, reading the file data of the target file from the tape library based on the target metadata information and writing the file data back to the online storage includes: Using the master control service to generate a migration job based on the metadata information of the target file after archiving; Using the archiving service to call the tape library read interface based on the relocation operation, the file data of the target file is retrieved from the tape library and rewritten into the online storage; The master control service is used to notify the storage service that the data has been migrated back, so that the storage service reads the file data of the target file from the online storage and returns it to the data requester.
7. The data migration method according to any one of claims 1 to 6, characterized in that: Also includes: When any file is stored in the online storage, the metadata information of the file is synchronized to the search engine, wherein the metadata information of the file includes the file path information; Correspondingly, the parsing of the target file information carried in the data request includes: The target file path information carried in the data request is parsed, wherein the target file path information is the path information of the file meeting the search condition returned by the search engine to the requesting party.
8. A data migration device, characterized in that: include: A file information acquisition module, used to parse the target file information carried in the data request; a target file reading module, configured to read the target file from the online storage based on the target file information, and if the read file metadata information indicates that the target file is an archived file, obtain the metadata information of the target file after archiving from the archive record metadata database to obtain the target metadata information, wherein the target metadata information includes the location of the target file in the tape library, the file metadata information in the online storage is the metadata information retained in the online storage and added with an archive mark when the file is archived, and the archived file is a file archived from the online storage to the tape library; The target file relocation module is used to read the file data of the target file from the tape library based on the target metadata information, and write the file data back to the online storage.
9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is used to execute the computer program to implement the steps of the data migration method as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data migration method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Data archiving storage method and device, electronic equipment and storage medium
CN121116911A