File migration method and device, equipment and storage medium

By building a file system graph and working together with the dynamic sharding verification pipeline, the efficiency bottlenecks, metadata integrity and real-time issues in traditional file migration methods are resolved, achieving efficient and real-time file migration and consistency verification. Dynamic sharding reduces I/O operations and manages and repairs file metadata.

CN120849346APending Publication Date: 2025-10-28JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510849453.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Traditional file migration methods suffer from efficiency bottlenecks, lack of metadata integrity, insufficient real-time performance, resource waste, and difficulty in complex permission management when handling large-scale data migration. Furthermore, hard links are easily lost during migration, leading to the destruction of data correlation.

Method used

By obtaining the metadata of the source files, the file system graph is constructed, dynamically sharded and transmitted to the target end, and the file shards are verified and repaired in real time. The metadata topology modeling and real-time verification pipeline work together to achieve consistency verification and dynamic sharding of file migration.

Benefits of technology

Improves file migration efficiency and consistency verification capabilities, reduces the number of I/O operations, effectively manages metadata, detects and repairs differences during the migration process in real time, and ensures the correctness of the file system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849346A_ABST
    Figure CN120849346A_ABST
Patent Text Reader

Abstract

The invention discloses a file migration method, which relates to the technical field of computers, and comprises the following steps: obtaining metadata of a plurality of files in a source file list of a source end, and constructing a first file system diagram according to the metadata; classifying the plurality of files according to the sizes of the plurality of files, and fragmenting the plurality of files according to the classification to obtain a plurality of file fragments; transmitting the plurality of file fragments to a target end in sequence, verifying and repairing the file fragments, and determining that the file fragments are transmitted correctly; in response to completion of file fragment transmission of the plurality of files, constructing a second file system diagram corresponding to the plurality of files after transmission is completed; and comparing the first file system diagram with the second file system diagram, and in response to the fact that the first file system diagram and the second file system diagram are the same, completing file migration, so that the technical problems of resource waste, insufficient real-time performance and weak automatic repair capability are solved, and migration and consistency verification of mass files are efficiently processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a file migration method, apparatus, device, and storage medium. Background Technology

[0002] In the field of information technology, file migration is the core operation for data flow across storage systems, cloud platforms or data centers. With the exponential growth of data scale (such as medical images, financial transaction logs, and IoT device data), traditional file migration methods usually adopt static sharding strategies and perform unified verification after full migration, which face severe challenges. Specific problems include: (1) Efficiency bottleneck: performing hash comparison on a file-by-file basis will cause the verification of millions of small files to take several days, resulting in huge I / O pressure; (2) Lack of metadata integrity: existing tools (such as rsync) ignore metadata such as permissions (ACL), hard links, and extended attributes (xattr), resulting in permission confusion or broken links after migration; (3) Insufficient real-time performance: incremental migration scenarios lack real-time verification mechanisms, and full comparison is still required when switching, resulting in a high risk of business interruption; (4) Waste of resources: static sharding strategies use fixed shard sizes, resulting in uneven shard load in small file scenarios and low utilization of computing resources; (5) Lack of complex permission management: traditional methods cannot handle multi-level permissions; (6) Risk of broken hard links: hard links across directories are easily lost during migration, resulting in the destruction of data correlation.

[0003] Therefore, in view of the shortcomings of existing technical solutions, the present invention provides a file migration method. Summary of the Invention

[0004] This application provides a file migration method, apparatus, device, and storage medium to at least address the problems of resource waste, insufficient real-time performance, and weak automated repair capabilities in related technologies.

[0005] This application provides a file migration method, which includes: obtaining metadata of multiple files in a source file list at the source end, and constructing a first file system graph based on the metadata; classifying the multiple files according to their size, and fragmenting the multiple files according to their classification to obtain multiple file fragments; sequentially transmitting the multiple file fragments to the target end, verifying and repairing the file fragments, and confirming that the file fragment transmission is correct; in response to the completion of the file fragment transmission of the multiple files, constructing a second file system graph corresponding to the transmitted multiple files; comparing the first file system graph and the second file system graph, and in response to the first file system graph and the second file system graph being the same, the file migration is completed.

[0006] This application also provides a file migration apparatus, comprising: a first processing module for acquiring metadata of multiple files in a source file list at the source end, and constructing a first file system graph based on the metadata; a second processing module for classifying the multiple files according to their size, and fragmenting the multiple files according to the classification to obtain multiple file fragments; a third processing module for sequentially transmitting the multiple file fragments to the target end, verifying and repairing the file fragments, and confirming that the file fragment transmission is correct; a fourth processing module for constructing a second file system graph corresponding to the transmitted multiple files in response to the completion of the file fragment transmission; and a fifth processing module for comparing the first file system graph and the second file system graph, and completing the file migration in response to the first file system graph and the second file system graph being identical.

[0007] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program and implementing the following steps: obtaining metadata of multiple files in a source file list at the source end, and constructing a first file system graph based on the metadata; classifying the multiple files according to their size, and fragmenting the multiple files according to their classification to obtain multiple file fragments; sequentially transmitting the multiple file fragments to the target end, verifying and repairing the file fragments, and confirming that the file fragment transmission is correct; in response to the completion of the file fragment transmission of the multiple files, constructing a second file system graph corresponding to the transmitted multiple files; comparing the first file system graph and the second file system graph, and in response to the first file system graph and the second file system graph being the same, the file migration is completed.

[0008] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it performs the following steps: obtaining metadata of multiple files in a source file list at the source end, and constructing a first file system graph based on the metadata; classifying the multiple files according to their size, and fragmenting the multiple files according to their classification to obtain multiple file fragments; sequentially transmitting the multiple file fragments to the target end, verifying and repairing the file fragments, and confirming that the file fragment transmission is correct; in response to the completion of the file fragment transmission of the multiple files, constructing a second file system graph corresponding to the multiple files that have been transmitted; comparing the first file system graph and the second file system graph, and in response to the first file system graph and the second file system graph being the same, the file migration is completed.

[0009] This application also provides a computer program product, including a computer program, which, when executed by a processor, performs the following steps: obtaining metadata of multiple files in a source file list at the source end, and constructing a first file system graph based on the metadata; classifying the multiple files according to their size, and fragmenting the multiple files according to their classification to obtain multiple file fragments; sequentially transmitting the multiple file fragments to the target end, verifying and repairing the file fragments, and confirming that the file fragment transmission is correct; in response to the completion of the file fragment transmission of the multiple files, constructing a second file system graph corresponding to the multiple files that have been transmitted; comparing the first file system graph and the second file system graph, and in response to the first file system graph and the second file system graph being the same, the file migration is completed.

[0010] This application achieves the following: First, it obtains metadata from multiple files in the source file list at the source end and constructs a first file system graph based on the metadata. Second, it categorizes the multiple files according to their size and then fragments them based on the categories, resulting in multiple file fragments. Third, it sequentially transmits these file fragments to the target end, verifies and repairs the fragments, and confirms that the transmission was correct. Fourth, in response to the completion of the file fragment transmission, it constructs a second file system graph corresponding to the transmitted files. Fifth, it compares the first and second file system graphs; if they match, the file migration is complete. Therefore, through the collaborative work of dynamic fragmentation, metadata topology modeling, and a real-time verification pipeline, it can efficiently handle the migration and consistency verification of massive amounts of files. Dynamic fragmentation effectively reduces the number of I / O operations, metadata topology modeling effectively manages and queries file metadata, and the real-time verification pipeline quickly detects and repairs differences during the file migration process. Attached Figure Description

[0011] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a file migration method provided in an embodiment of this application;

[0013] Figure 2 A schematic diagram of a file permission tree for a file migration method provided in an embodiment of this application;

[0014] Figure 3 This application provides a schematic diagram of a file migration method fragmentation process.

[0015] Figure 4This application provides a schematic diagram of a file fragment verification process for a file migration method.

[0016] Figure 5 This application provides a schematic diagram of a file verification process for a file migration method.

[0017] Figure 6 A structural block diagram of a file migration device provided in an embodiment of this application;

[0018] Figure 7 This is an internal structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0020] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0021] It should be noted that the terms "S1," "S2," etc., are used only for descriptive purposes and do not specifically refer to the order or sequence, nor are they intended to limit this application. They are merely for the convenience of describing the method of this application and should not be construed as indicating the sequential order of the steps. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0022] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] The system implementing the file migration method may include a source end, a file migration service, and a target end. The file migration service may be a standalone server or an application deployed on the source end or the target end. Preferably, the file migration service in this embodiment is an application deployed on the source end. It runs on the source end, obtains the files to be migrated in the source end through the file access interface provided by the source end, and then migrates them to the target end.

[0024] The source and destination can communicate via a network, which may include a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, or a combination thereof, and is connected to a website, user equipment (e.g., computing devices), and backend systems.

[0025] The source or target end can be a node of a cloud computing system, or each server can be a separate cloud computing system, including multiple computers interconnected by a network and operating as a distributed processing system, without limitation in the embodiments herein.

[0026] The embodiments of this application provide a file migration method, and the method is described in detail below in conjunction with the execution flow of the file migration method.

[0027] S101: Obtain the metadata of multiple files in the source file list from the source end, and construct the first file system graph based on the metadata.

[0028] Here, the metadata for each file can include one metadata item or multiple metadata items.

[0029] Here, file metadata can include basic information, permission information, location information, content information, and source information. Basic information can include filename, file type, file size, creation time, modification time, and access time; permission information can include users who can read, write, or execute the file; location information can include the file's path or location in the file system; content information can include the file's author and title; and source information can include the file's original source or copyright information.

[0030] Here, a file system graph can be an architectural diagram that includes the relationships between multiple files. One piece of metadata can correspond to one file system graph, and multiple pieces of metadata can correspond to multiple file system graphs.

[0031] Here, the source can be an electronic device or a node in a cloud platform.

[0032] Specifically, the metadata of each file in the source file list is obtained, and a file system graph is constructed for each file based on the type of metadata.

[0033] S102: Classify multiple files according to their size, and then divide the multiple files into fragments according to their classification to obtain multiple file fragments.

[0034] Here, the size of multiple file fragments can be the same or different.

[0035] Here, files can be categorized into small files, large files, etc.

[0036] Here, file fragmentation can include merged fragmentation, independent fragmentation, redundant fragmentation, encrypted fragmentation, and verification fragmentation.

[0037] S103: Transmit multiple file fragments to the target end in sequence, verify and repair the file fragments, and confirm that the file fragment transmission is correct.

[0038] Here, the verification methods may include hash value verification, cyclic redundancy check, erasure coding verification, fragment sequence number and length check, etc.

[0039] Here, repair methods may include retransmitting fragments, retransmitting after a power outage, and resuming transmission after a power failure.

[0040] Here, the target can be an electronic device or a node in a cloud platform.

[0041] The verification and repair process can be a loop until the file fragments are transmitted correctly.

[0042] Specifically, multiple file fragments corresponding to multiple files are transmitted sequentially. The target end receives each file fragment and verifies it. If the verification passes, the file fragment is confirmed to have been transmitted correctly. If the verification fails, the file is repaired and the file fragment is verified again until the verification passes, confirming that the file fragment has been transmitted correctly.

[0043] S104: In response to the completion of file fragment transfer of multiple files, construct a second file system graph corresponding to the multiple files that have been transferred.

[0044] The transmission process ensures that the content of each file slice in each file is consistent with the content of the corresponding file slice in the source file.

[0045] Specifically, after the file fragments of multiple files in the target end are transmitted, the files are reassembled according to the same fragmentation rules, the metadata of each file is obtained, and the corresponding second file system graph is constructed with reference to the first file system graph.

[0046] S105: Compare the first file system map and the second file system map. If the first file system map and the second file system map are the same, the file migration is completed.

[0047] Here, file migration refers to the process of moving files from one storage location to another.

[0048] This involves generating a unique image identifier for each file system graph, and then comparing the image identifiers to determine if the images are identical.

[0049] Among these methods, file system graphs can be compared through metadata comparison, hash value verification, manual comparison, and automatic comparison using tools.

[0050] In one embodiment, the first file system diagram and the second file system diagram are different. The file system diagrams are compared one by one to determine the location of the transmission error and the corresponding files are retransmitted.

[0051] It should be noted that this application, through the collaborative work of dynamic sharding, metadata topology modeling, and real-time verification pipeline, can efficiently handle the migration and consistency verification of massive files. Dynamic sharding can effectively reduce the number of I / O operations, metadata topology modeling can effectively manage and query file metadata, and the real-time verification pipeline can quickly detect and repair differences during the file migration process.

[0052] In some specific implementations, metadata of multiple files in the source file list is obtained, and a first file system graph is constructed based on the metadata, including:

[0053] Based on the metadata of multiple files, determine the permission information, inodes, and hard links of multiple files;

[0054] Based on the permission information, construct a file permission tree corresponding to multiple files;

[0055] Construct a directed graph by using the index nodes of multiple files as nodes and hard links as edges;

[0056] Determine the first file system graph based on the file permission tree and the directed graph.

[0057] Here, permission information includes whether a user can access a file and what operations a user can perform on the file.

[0058] The file permission tree can include multiple levels of permissions, such as owner, group, and user.

[0059] Specifically, obtain the permission information for each file and generate a file permission tree based on the permission information.

[0060] For example, Figure 2 This is a schematic diagram of the file permission tree in an embodiment of this application, such as... Figure 2As shown, the file permission tree in this application may include: owner, user group, and user. Specifically, the owner is admin, the user group includes developers and auditors, the user group developers also includes users john and lily, and the user group auditors also includes user lucy.

[0061] Here, a hard link refers to a link relationship where multiple filenames point to the same inode.

[0062] Here, an inode is a data structure that stores file metadata. All information except the filename is stored in the inode, which may include file size, file owner, permissions, timestamps, and the location of file data blocks. Each file has a unique inode number for identification.

[0063] Here, a directed graph is a graph data structure consisting of a set of vertices and a set of edges connecting the vertices, with each edge having a direction.

[0064] Specifically, directed graphs can be stored using adjacency lists, for example:

[0065] {

[0066] "INODE_1001": [" / data / file1.txt", " / backup / file1_copy.txt"],

[0067] "INODE_1002":[" / data / report.pdf"]

[0068] }

[0069] In this graph, “INODE_1001” and “INODE_1002” are nodes in the directed graph, and “ / data / file1.txt”, “ / backup / file1_copy.txt”, and “ / data / report.pdf” are edges in the directed graph.

[0070] This solves the problems of existing technologies being unable to handle multi-level permissions and the easy loss of hard links across directories during migration, which can lead to the destruction of data consistency and improve the efficiency of file migration consistency verification.

[0071] In some specific implementations, multiple files are classified according to their size, and then fragmented according to their classification to obtain multiple file fragments, including:

[0072] Sort the multiple files according to their size to obtain a set of file sizes;

[0073] Based on the preset quantile, obtain the target file corresponding to the preset quantile in the file size set, and use the size of the target file as the segmentation threshold;

[0074] Compare the file size of multiple files with the fragmentation threshold;

[0075] If the file size is greater than or equal to the fragmentation threshold, the file is determined to be the first file.

[0076] Based on the file size of the first file, determine the file fragment size and perform independent fragmentation on the first file;

[0077] If the file size is less than the fragmentation threshold, the file is determined to be the second file.

[0078] The number of second files in a file segment is determined based on a preset fixed segment size and segment threshold.

[0079] Sort multiple second files according to their metadata;

[0080] Based on the number of second files in the file fragment, multiple second files are merged and fragmented in sequence.

[0081] Here, quantiles refer to the division of a dataset into equal parts. Common quantiles include percentiles, decimals, and quartiles. For example, a quantile could be 75%.

[0082] Here, metadata can be INODEs, sorted in ascending order of INODEs.

[0083] The first file is the large file, and the second file is the small file.

[0084] Specifically, all files in the file list are sorted by size, and the file corresponding to the 75th percentile is obtained. The size of the file is used as the fragmentation threshold. Files greater than or equal to the fragmentation threshold are large files, and large files are fragmented independently. The fragmentation size is determined based on the size of the large files. Files smaller than the fragmentation threshold are small files, and small files are merged into fragments. Based on the preset fragmentation size and the threshold size, the number of small files that can be included in a fragment is determined. Small files are merged in ascending order according to INODE, and multiple small files are divided into a fragment based on the number.

[0085] For example, Figure 3 This is a schematic diagram of the fragmentation process in an embodiment of this application, such as... Figure 3 As shown, the fragmentation process in this application may include: obtaining a list of source files, statistically analyzing the file size distribution, calculating a dynamic threshold, merging and distributing small files, independently fragmenting large files, attaching metadata and calculating fragment hashes, and transmitting the data to the target end.

[0086] In one embodiment, factors such as historical data, network bandwidth, and storage performance can be combined to determine the sharding threshold.

[0087] In one embodiment, partial metadata of multiple small files that are merged into the same file fragment can be merged.

[0088] In one embodiment, after the file is divided into large and small file segments, the hash value of each file segment is calculated and stored in the file segment, and then transmitted to the target end along with the file segment transmission.

[0089] Specifically, for small files, a holistic hash algorithm is used to calculate the file's hash value. For large files, the large file is divided into multiple smaller blocks, and a hash value is calculated for each block. These hash values ​​are then combined according to a certain hierarchical structure (e.g., using a binary tree) to finally obtain the hash value corresponding to the large file. This approach is applicable to different scenarios and can achieve an optimal balance between efficiency and accuracy in various application scenarios.

[0090] In this way, small files can be adaptively merged based on file size and INODE order, reducing I / O pressure. At the same time, for large files, the fragment size can be dynamically adjusted, which can effectively reduce the number of I / O operations during file migration, thereby improving the efficiency of file migration.

[0091] In some specific implementations, based on a preset quantile, the target file corresponding to the preset quantile in the file size set is obtained, and the size of the target file is used as the fragmentation threshold, including:

[0092] Compare the size of the target file with the preset protection threshold;

[0093] If the size of the target file is less than the protection threshold, the protection threshold will be used as the fragmentation threshold.

[0094] If the size of the target file is greater than or equal to the protection threshold, the size of the target file is used as the fragmentation threshold.

[0095] Specifically, assume the protection threshold is 10KB, the fragmentation threshold is 10KB, and the target file size is the larger of the two values.

[0096] In this way, even when files are generally small or have particularly low percentiles, the system can maintain a certain level of performance.

[0097] In some specific implementations, multiple file fragments are transmitted sequentially to the target end, the file fragments are verified and repaired, and the correct transmission of the file fragments is confirmed, including:

[0098] By comparing the hash values ​​of the file fragments at the target end with the hash values ​​of the file fragments at the source end, it can be determined whether the file fragment transmission is correct.

[0099] If the hash value of the file fragment at the destination end is different from the hash value of the file fragment at the source end, the file fragment is retransmitted from the source end.

[0100] Compare the hash values ​​of the retransmitted file fragments at the source end with the hash values ​​of the file fragments at the destination end;

[0101] If the hash value of the retransmitted file fragment at the source is different from the hash value of the file fragment at the destination, the problematic file is located by binary search and then retransmitted from the source.

[0102] Here, a hash value is a fixed-length numerical value obtained by calculating data using a hash function.

[0103] Here, binary search is an efficient algorithm for finding a specific element in an ordered array. It quickly locates the target value by gradually halving the search range.

[0104] For example, Figure 4 This is a schematic diagram of the file fragment verification process in an embodiment of this application, such as... Figure 4 As shown, the file fragment verification process in this application may include: the target end monitors changes to files or data through FANOTIFY. When a change is detected, the relevant information is placed in a message queue such as Kafka. The information retrieved from the Kafka queue is used to determine the specific location or fragment of the data to be processed. The hash value of the located file fragment is calculated, and it is determined whether the hash value of the target end and the hash value of the source end are consistent. If they are consistent, the marking is completed. If they are inconsistent, the repair is triggered. The file fragment is retransmitted first, and the hash value of the retransmitted fragment is verified. If the hash value verification passes, the marking is completed. If the hash value verification still fails, the file is located by binary search and the file is retransmitted.

[0105] In this way, differences during the file migration process can be monitored and repaired in real time.

[0106] In some specific implementations, comparing the first file system graph and the second file system graph includes:

[0107] Traverse and compare the file permission trees in the first file system graph and the second file system graph layer by layer;

[0108] Since the file permission tree in the first file system graph is the same as the file permission tree in the second file system graph, a graph identifier corresponding to the directed graph is generated based on the directed graph in the second file system graph using a graph hashing algorithm.

[0109] Compare the graph identifiers in the first file system graph and the graph identifiers in the second file system graph;

[0110] If the graph identifier in the first file system graph is the same as the graph identifier in the second file system graph, it is determined that the first file system graph and the second file system graph are the same.

[0111] Here, the graph hashing algorithm can be the Weisfeiler-Lehman algorithm, which is an efficient algorithm for graph isomorphism detection. It is mainly used to determine whether two graphs have the same structure, that is, whether there is a one-to-one node mapping relationship so that the edges of the two graphs are completely consistent.

[0112] Specifically, the target file system is traversed to generate a file permission tree, the target INODE is scanned to generate a directed graph of hard links, and then the file permission trees of the source and target are compared layer by layer. At the same time, the WL hash of the directed graph of hard links of the source and the target is calculated and quickly compared, and permission errors and broken links are automatically repaired.

[0113] For example, a directed graph can be constructed, each node in the graph can be assigned an initial label, and the label of each node can be updated according to the WL algorithm. After the last iteration, the final labels of all nodes can be collected and sorted into a string to obtain the graph identifier.

[0114] In this way, based on the confirmation that the transmitted content is correct, file system differences can be detected to ensure the correctness of the file system.

[0115] In some specific implementations, the method further includes:

[0116] In response to file updates in the source file list, retrieve the metadata of the updated file in the source file list and update the first file system graph;

[0117] In response to the completion of file fragment transfer of the updated file, update the second file system graph according to the completed updated file;

[0118] The updated first file system map and the updated second file system map are compared. If the updated first file system map and the updated second file system map are the same, the updated files are migrated.

[0119] Here, an update can include modifications or additions.

[0120] Among these features, file updates can be detected through FANOTIFY monitoring.

[0121] Specifically, the update files are obtained, and the file permission tree and hard link directed graph of the source end are updated according to the update files. The update files are sorted by file size to obtain a set of update files. A new fragmentation threshold is determined according to a preset quantile, and the update files are fragmented. The update files are transmitted in fragment form. After the target end receives the file fragments, a hash value comparison is performed. After all fragments are compared by hash value, the file permission tree and hard link directed graph of the target end are updated according to the file fragments. The file permission tree and hard link directed graph of the source end and the target end are compared.

[0122] This enables real-time verification of incremental migration scenarios.

[0123] In one embodiment, Figure 5 This is a schematic diagram of the file verification process in an embodiment of this application, such as... Figure 5 As shown, the file verification process in this application includes: pre-migration preparation, including dynamic fragmentation rule calculation, construction of source-side permission graph and construction of source-side hard link graph; full migration, for the first migration, dynamic fragmentation processing, fragmented transmission, and target-side fragmented reconstruction, for incremental migration, time monitoring, dynamic fragmentation update, and simultaneous update of permissions and / or hard link graph, and local verification and repair; final verification, construction of target-side permission and / or hard link graph, global comparison, and difference repair.

[0124] Specifically, in the pre-migration preparation stage, this step mainly involves statistically analyzing file size distribution and calculating dynamic thresholds; traversing the source file system to generate a permission tree; scanning INODEs and file paths to construct a hard link graph in the form of an adjacency list; and outputting dynamic thresholds, a snapshot of the source permission graph, and a snapshot of the source hard link graph.

[0125] Specifically, during the full migration phase, the full migration data is merged into small files according to a dynamic threshold to generate logical shards, and shard metadata (file list, INODE order) is recorded; then the shard data and hash values ​​are transmitted to the target end; the target end reassembles the files according to the same sharding rules to generate the target end file structure.

[0126] Specifically, in the incremental synchronization phase after the full migration is completed, file changes (creation / modification / deletion) are captured via FANOTIFY or database logs; newly added / modified files are re-sharded according to the latest threshold; the source permission tree and hard link graph are updated in real time; hashes are recalculated only for affected shards and compared with those on the target side; if inconsistencies are found, a binary search method is triggered to locate and repair the issue.

[0127] Specifically, in the final verification phase, the target file system is traversed to generate a permission graph, and the target INODE is scanned to generate an adjacency hard link graph. Then, the source permission graph is traversed layer by layer to compare with the target permission graph. Simultaneously, the WL hash of the source and target hard link graphs is calculated and quickly compared to automatically repair permission errors and broken links. After a full migration, a full verification of this step is required. Subsequent incremental migrations can be configured to perform full verification periodically, such as every 3 days or once a week.

[0128] It should be understood that, although Figure 1-5 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1-5 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0129] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0130] Embodiments of this application also provide a file migration apparatus, comprising: a first processing module 601, configured to acquire metadata of multiple files in a source file list at the source end, and construct a first file system graph based on the metadata; a second processing module 602, configured to classify the multiple files according to their size, and fragment the multiple files according to the classification to obtain multiple file fragments; a third processing module 603, configured to sequentially transmit the multiple file fragments to the target end, verify and repair the file fragments, and determine that the file fragment transmission is correct; a fourth processing module 604, configured to construct a second file system graph corresponding to the transmitted multiple files in response to the completion of the file fragment transmission; and a fifth processing module 605, configured to compare the first file system graph and the second file system graph, and in response to the first file system graph and the second file system graph being the same, the file migration is completed.

[0131] As a preferred implementation, in this embodiment of the application, the first processing module 601 is specifically used to: determine the permission information, index nodes, and hard links of multiple files based on the metadata of multiple files; construct a file permission tree corresponding to multiple files based on the permission information; construct a directed graph by using the index nodes of multiple files as nodes and the hard links as edges; and determine a first file system graph based on the file permission tree and the directed graph.

[0132] In a preferred embodiment of this application, the second processing module 602 is specifically configured to: sort multiple files according to their sizes to obtain a set of file sizes; obtain a target file corresponding to a preset quantile in the set of file sizes, and use the size of the target file as a fragmentation threshold; compare the file sizes of multiple files with the fragmentation threshold; determine a file as a first file in response to a file size greater than or equal to the fragmentation threshold; determine a file fragment size based on the file size of the first file, and perform independent fragmentation on the first file; determine a file as a second file in response to a file size less than the fragmentation threshold; determine the number of second files in the file fragments based on a preset fixed fragment size and fragmentation threshold; sort the multiple second files according to their metadata; and merge and fragment the multiple second files in order according to the number of second files in the file fragments.

[0133] In a preferred embodiment of this application, the second processing module 602 is further configured to: compare the size of the target file with a preset protection threshold; in response to the size of the target file being less than the protection threshold, use the protection threshold as a fragmentation threshold; in response to the size of the target file being greater than or equal to the protection threshold, use the size of the target file as a fragmentation threshold.

[0134] In a preferred embodiment of this application, the third processing module 603 is specifically used to: determine whether the file fragment transmission is correct by comparing the hash value of the file fragment at the target end with the hash value of the file fragment at the source end; in response to the fact that the hash value of the file fragment at the target end and the hash value of the file fragment at the source end are different, retransmit the file fragment from the source end; compare the hash value of the retransmitted file fragment at the source end with the hash value of the file fragment at the target end; in response to the fact that the hash value of the retransmitted file fragment at the source end and the hash value of the file fragment at the target end are different, locate the problematic file by binary search and retransmit the problematic file from the source end.

[0135] In a preferred embodiment of this application, the fifth processing module 605 is specifically used to: traverse and compare the file permission trees in the first file system graph and the second file system graph layer by layer; in response to the file permission trees in the first file system graph and the second file system graph being the same, generate a graph identifier corresponding to the directed graph based on the directed graph in the second file system graph using a graph hash algorithm; compare the graph identifiers in the first file system graph and the second file system graph; in response to the graph identifiers in the first file system graph and the second file system graph being the same, determine that the first file system graph and the second file system graph are the same.

[0136] In a preferred embodiment of this application, the apparatus further includes an update module, which is specifically configured to: in response to a file update in the source file list, obtain the metadata of the updated file in the source file list and update the first file system graph; in response to the completion of file fragment transmission of the updated file, update the second file system graph according to the completed updated file; compare the updated first file system graph and the updated second file system graph, and in response to the updated first file system graph and the updated second file system graph being the same, complete the migration of the updated file.

[0137] For a description of the features in the embodiment corresponding to the file migration device, please refer to the relevant description in the embodiment corresponding to the file migration method, which will not be repeated here.

[0138] Embodiments of this application also provide an electronic device, which may be a terminal, and its internal structure diagram may be as follows: Figure 7 As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a file migration method. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.

[0139] Those skilled in the art will understand that Figure 7The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0140] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it performs the following steps: S1: Obtaining metadata of multiple files in a source file list at the source end, and constructing a first file system graph based on the metadata; S2: Classifying the multiple files according to their size, and fragmenting the multiple files according to their classification to obtain multiple file fragments; S3: Transmitting the multiple file fragments sequentially to the target end, verifying and repairing the file fragments, and confirming that the file fragment transmission is correct; S4: In response to the completion of the file fragment transmission of the multiple files, constructing a second file system graph corresponding to the transmitted multiple files; S5: Comparing the first file system graph and the second file system graph, and in response to the first file system graph and the second file system graph being the same, the file migration is completed.

[0141] In one embodiment, when the processor executes the computer program, it further performs the following steps: determining the permission information, inodes, and hard links of multiple files based on the metadata of the multiple files; constructing a file permission tree corresponding to the multiple files based on the permission information; constructing a directed graph using the inodes of the multiple files as nodes and the hard links as edges; and determining a first file system graph based on the file permission tree and the directed graph.

[0142] In one embodiment, when the processor executes the computer program, it further performs the following steps: sorting multiple files according to their sizes to obtain a set of file sizes; obtaining a target file corresponding to a preset quantile in the set of file sizes, and using the size of the target file as a fragmentation threshold; comparing the file sizes of the multiple files with the fragmentation threshold; determining a file as a first file in response to a file size greater than or equal to the fragmentation threshold; determining a file fragment size based on the file size of the first file, and performing independent fragmentation on the first file; determining a file as a second file in response to a file size less than the fragmentation threshold; determining the number of second files in the file fragments based on a preset fixed fragment size and the fragmentation threshold; sorting the multiple second files according to their metadata; and merging and fragmenting the multiple second files in order according to the number of second files in the file fragments.

[0143] In one embodiment, when the processor executes the computer program, it further performs the following steps: comparing the size of the target file with a preset protection threshold; in response to the size of the target file being less than the protection threshold, using the protection threshold as a fragmentation threshold; in response to the size of the target file being greater than or equal to the protection threshold, using the size of the target file as a fragmentation threshold.

[0144] In one embodiment, when the processor executes the computer program, it further performs the following steps: determining whether the file fragment transmission is correct by comparing the hash value of the file fragment at the destination end with the hash value of the file fragment at the source end; retransmitting the file fragment from the source end in response to the hash value of the file fragment at the destination end being different from the hash value of the file fragment at the source end; comparing the hash value of the retransmitted file fragment at the source end with the hash value of the file fragment at the destination end; and locating the problematic file by binary search in response to the hash value of the retransmitted file fragment at the source end being different from the hash value of the file fragment at the destination end, and retransmitting the problematic file from the source end.

[0145] In one embodiment, when the processor executes the computer program, it further performs the following steps: traversing and comparing the file permission trees in the first file system graph and the second file system graph layer by layer; in response to the file permission trees in the first file system graph and the second file system graph being the same, generating a graph identifier corresponding to the directed graph based on the directed graph in the second file system graph using a graph hashing algorithm; comparing the graph identifiers in the first file system graph and the second file system graph; in response to the graph identifiers in the first file system graph and the second file system graph being the same, determining that the first file system graph and the second file system graph are the same.

[0146] In one embodiment, when the processor executes the computer program, it further implements the following steps: in response to a file update in the source file list, obtaining the metadata of the updated file in the source file list and updating the first file system graph; in response to the completion of file fragment transmission of the updated file, updating the second file system graph according to the completed updated file; comparing the updated first file system graph and the updated second file system graph, and in response to the updated first file system graph and the updated second file system graph being the same, the updated file migration is completed.

[0147] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it performs the following steps: S1: Obtain metadata of multiple files in a source file list at the source end, and construct a first file system graph based on the metadata; S2: Classify the multiple files according to their size, and fragment the multiple files according to their classification to obtain multiple file fragments; S3: Transmit the multiple file fragments to the target end in sequence, verify and repair the file fragments, and determine that the file fragment transmission is correct; S4: In response to the completion of the file fragment transmission of the multiple files, construct a second file system graph corresponding to the multiple files that have been transmitted; S5: Compare the first file system graph and the second file system graph. In response to the first file system graph and the second file system graph being the same, the file migration is completed.

[0148] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: determining the permission information, inodes, and hard links of multiple files based on the metadata of the multiple files; constructing a file permission tree corresponding to the multiple files based on the permission information; constructing a directed graph using the inodes of the multiple files as nodes and the hard links as edges; and determining a first file system graph based on the file permission tree and the directed graph.

[0149] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: sorting multiple files according to their sizes to obtain a set of file sizes; obtaining a target file corresponding to a preset quantile in the set of file sizes, and using the size of the target file as a fragmentation threshold; comparing the file sizes of the multiple files with the fragmentation threshold; determining a file as a first file in response to a file size greater than or equal to the fragmentation threshold; determining a file fragment size based on the file size of the first file, and independently fragmenting the first file; determining a file as a second file in response to a file size less than the fragmentation threshold; determining the number of second files in the file fragments based on a preset fixed fragment size and the fragmentation threshold; sorting the multiple second files according to their metadata; and merging and fragmenting the multiple second files in order according to the number of second files in the file fragments.

[0150] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: comparing the size of the target file with a preset protection threshold; in response to the size of the target file being less than the protection threshold, using the protection threshold as a fragmentation threshold; in response to the size of the target file being greater than or equal to the protection threshold, using the size of the target file as a fragmentation threshold.

[0151] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: determining whether the file fragment transmission is correct by comparing the hash value of the file fragment at the destination end with the hash value of the file fragment at the source end; retransmitting the file fragment from the source end in response to the fact that the hash value of the file fragment at the destination end is different from the hash value of the file fragment at the source end; comparing the hash value of the retransmitted file fragment at the source end with the hash value of the file fragment at the destination end; and locating the problematic file by binary search in response to the fact that the hash value of the retransmitted file fragment at the source end is different from the hash value of the file fragment at the destination end, and retransmitting the problematic file from the source end.

[0152] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: traversing and comparing the file permission trees in the first file system graph and the second file system graph layer by layer; in response to the file permission trees in the first file system graph and the second file system graph being the same, generating a graph identifier corresponding to the directed graph based on the directed graph in the second file system graph using a graph hashing algorithm; comparing the graph identifiers in the first file system graph and the second file system graph; in response to the graph identifiers in the first file system graph and the second file system graph being the same, determining that the first file system graph and the second file system graph are the same.

[0153] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: in response to a file update in the source file list, obtaining the metadata of the updated file in the source file list and updating the first file system graph; in response to the completion of file fragment transmission of the updated file, updating the second file system graph according to the completed updated file; comparing the updated first file system graph and the updated second file system graph, and in response to the updated first file system graph and the updated second file system graph being the same, the updated file migration is completed.

[0154] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0155] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0156] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A file migration method, characterized in that, The method includes: Obtain metadata from multiple files in the source file list at the source end, and construct a first file system graph based on the metadata; The multiple files are classified according to their size, and then the multiple files are fragmented according to their classification to obtain multiple file fragments; The multiple file fragments are transmitted to the target end in sequence, the file fragments are verified and repaired, and it is confirmed that the file fragments are transmitted correctly. In response to the completion of file fragment transmission of the multiple files, a second file system graph corresponding to the multiple files that have been transmitted is constructed; The file migration is completed when the first file system graph and the second file system graph are compared and the first file system graph and the second file system graph are the same.

2. The file migration method according to claim 1, characterized in that, The step of obtaining metadata from multiple files in the source file list at the source end and constructing a first file system graph based on the metadata includes: Based on the metadata of the multiple files, determine the permission information, inodes, and hard links of the multiple files; Based on the permission information, construct a file permission tree corresponding to the multiple files; Construct a directed graph by using the index nodes of the multiple files as nodes and the hard links as edges; The first file system graph is determined based on the file permission tree and the directed graph.

3. The file migration method according to claim 1, characterized in that, The step of classifying the multiple files according to their size, and then segmenting the multiple files according to the classification to obtain multiple file segments includes: Based on the size of the multiple files, sort the multiple files to obtain a set of file sizes; Based on the preset quantile, obtain the target file corresponding to the preset quantile in the file size set, and use the size of the target file as the fragmentation threshold; Compare the file size of the multiple files with the fragmentation threshold; In response to the file size being greater than or equal to the fragmentation threshold, the file is determined to be the first file; Based on the file size of the first file, determine the file fragment size and perform independent fragmentation on the first file; In response to the file size being less than the fragmentation threshold, the file is determined to be the second file; The number of second files in a file segment is determined based on a preset fixed segment size and the segment threshold. Sort the multiple second files according to the metadata; Based on the number of second files in the file fragments, the multiple second files are merged and fragmented in sequence.

4. The file migration method according to claim 3, characterized in that, The step of obtaining the target file corresponding to the preset quantile in the file size set according to the preset quantile, and using the size of the target file as the fragmentation threshold, includes: The size of the target file is compared with a preset protection threshold; In response to the target file being smaller than the protection threshold, the protection threshold is used as the fragmentation threshold; In response to the target file size being greater than or equal to the protection threshold, the target file size is used as the fragmentation threshold.

5. The file migration method according to claim 1, characterized in that, The step of sequentially transmitting the multiple file fragments to the target end, verifying and repairing the file fragments, and confirming that the file fragment transmission is correct includes: By comparing the hash values ​​of the file fragments at the target end and the hash values ​​of the file fragments at the source end, it is determined whether the file fragment transmission is correct. If the hash value of the file fragment at the target end is different from the hash value of the file fragment at the source end, the file fragment is retransmitted from the source end. Compare the hash values ​​of the retransmitted file fragments at the source end with the hash values ​​of the file fragments at the target end; If the hash value of the retransmitted file fragment at the source end is different from the hash value of the file fragment at the destination end, the problematic file is located by binary search, and the problematic file is retransmitted from the source end.

6. The file migration method according to claim 2, characterized in that, The comparison of the first file system graph and the second file system graph includes: Traverse and compare the file permission trees in the first file system graph and the second file system graph layer by layer; In response to the fact that the file permission trees in the first file system graph and the second file system graph are the same, a graph identifier corresponding to the directed graph is generated based on the directed graph in the second file system graph using a graph hash algorithm; and the graph identifier in the first file system graph and the graph identifier in the second file system graph are compared. In response to the fact that the graph identifier in the first file system graph and the graph identifier in the second file system graph are the same, it is determined that the first file system graph and the second file system graph are the same.

7. The file migration method according to claim 1, characterized in that, The method further includes: In response to a file update in the source file list, obtain the metadata of the updated file in the source file list and update the first file system graph; In response to the completion of file fragment transmission of the updated file, the second file system graph is updated according to the completed updated file; The updated first file system diagram and the updated second file system diagram are compared. If the updated first file system diagram and the updated second file system diagram are the same, the updated file migration is completed.

8. A file migration device, characterized in that, The device includes: The first processing module is used to obtain metadata of multiple files in the source file list from the source end, and construct a first file system graph based on the metadata; The second processing module is used to classify the multiple files according to their size, and to segment the multiple files according to their classification to obtain multiple file segments; The third processing module is used to sequentially transmit the multiple file fragments to the target end, verify and repair the file fragments, and determine that the file fragments are transmitted correctly. The fourth processing module is used to construct a second file system graph corresponding to the multiple files after the file fragment transmission of the multiple files is completed. The fifth processing module is used to compare the first file system graph and the second file system graph. In response to the first file system graph and the second file system graph being the same, the file migration is completed.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing a computer program to implement the steps of the file migration method as claimed in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program, wherein when executed by a processor, the computer program implements the steps of the file migration method as claimed in any one of claims 1 to 7.