Data platform file migration method, computer program product and data platform
Through the comparison of dynamic chunking strategy and Merkel tree of the target reinforcement learning model, the problem of unreasonable chunking strategy in data platform file migration is solved, and efficient data migration and transmission is achieved.
Patent Information
- Application Number
- CN202510759343.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
In the existing data platform file migration methods, unreasonable data chunking strategy leads to low migration efficiency, especially in semi-structured files such as Parquet and HFile, the chunking logic does not match the storage logic, resulting in low differential positioning and transmission efficiency.
The dynamic blocking strategy of the target reinforcement learning model is adopted. By scanning the data file type and using the state set and action set of the target reinforcement learning model, the most suitable blocking strategy is selected, and the Merkel tree is built for differential positioning and migration, and only differential data blocks are transmitted.
It improves data migration efficiency, reduces the number and transmission volume of different data blocks, improves Merkel tree comparison efficiency and data transmission efficiency, and solves the problem of low migration efficiency caused by unreasonable data blocking strategies.
Smart Images

Figure CN120277046A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data migration, and in particular, to a method for migrating files in a data platform, a computer program product, and a data platform. Background Art
[0002] The complex fault assistance disposal technology based on cross-business scenarios needs to integrate fault data from different business fields, including fault types, fault occurrence times, influence ranges, handling measures, etc., to establish a fault solution library. By using association analysis algorithms, the association rules between different faults are mined to identify the potential relationships between faults. Among them, the fault data in different business fields are pre-stored in different data platforms. During the data integration process, it is usually necessary to migrate the fault data between different data platforms so as to integrate the fault data into one data platform. When the file data in the source data platform changes, it needs to be synchronously updated in the target data platform.
[0003] Traditional data migration uses full-volume data transmission. When there are local data changes, it is also necessary to re-migrate the full-volume data, and the migration efficiency is extremely low. In order to improve the data migration efficiency, a file migration method that divides data into blocks and constructs a data tree has gradually emerged. When part of the data changes, the data block where the change occurs can be located in the data tree (the Diff algorithm can be used for difference location), so as to specifically re-migrate the data block with changes. However, the above file migration method usually adopts a fixed block division strategy, that is, all types of files adopt a unified block division strategy, and the characteristics of different types of files are not deeply studied, resulting in at least an unreasonable block division strategy for some types of files, thus affecting subsequent difference location. Exemplarily, semi-structured files such as Parquet and HFile in data warehouses and data lakes are often stored in a combined manner. Fixed block division is likely to cross the logical record boundary and cannot recognize the connection between different files, resulting in the failure of the Diff algorithm and the migration falling back to full-volume transmission. Essentially, the block division logic of a single block division strategy cannot match the storage logic of all data files. When the block division logic is different from the storage logic, it is easy to allocate the local data with changes to different data blocks, resulting in more differential data blocks. Therefore, the location efficiency of differential data blocks and the transmission efficiency of differential data blocks will be reduced. Moreover, when the number of differential data blocks is too large, it will cause the failure of difference location, and the transmission of differential data blocks in data migration will fall back to full-volume transmission, which also greatly reduces the transmission efficiency.
[0004] Aiming at the problem of low data migration efficiency caused by unreasonable data block division strategies in the current data platform file migration method, no effective solution has been proposed yet. Summary of the Invention
[0005] In the present invention, a method for migrating data platform files, a computer program product, and a data platform are provided to solve the problem of low data migration efficiency caused by an unreasonable data chunking strategy in current data platform file migration methods.
[0006] In a first aspect, the present invention provides a method for migrating data platform files for migrating data files to a target end, including: Scanning the source-end data files to be migrated to obtain the types of the source-end data files; If the type of the source-end data file is in the state set of the target reinforcement learning model, taking the type of the source-end data file as the current state of the target reinforcement learning model, and obtaining the current action of the target reinforcement learning model in response to the current state; wherein, the state set of the target reinforcement learning model includes multiple file types, the action set of the target reinforcement learning model includes multiple file chunking strategies; the reward of the target reinforcement learning model includes the same ratio between the data chunks generated by using the current file chunking strategy for the data files of the current file type before and after local changes respectively; Chunking the source-end data files by using the current action of the target reinforcement learning model; Constructing a Merkle tree of the source-end data files based on the chunking results of the source-end data files; Locating the different data chunks by layer-by-layer hashing and comparing the Merkle trees of the source-end data files and the target-end data files; Migrating the different data chunks to the target end.
[0007] In a second aspect, the present invention provides a method for obtaining data platform files for obtaining file data from a source end, including: Scanning the target-end data files to be updated to obtain the types of the target-end data files; If the type of the target-end data file is in the state set of the target reinforcement learning model, taking the type of the target-end data file as the current state of the target reinforcement learning model, and obtaining the current action of the target reinforcement learning model in response to the current state; wherein, the state set of the target reinforcement learning model includes multiple file types, the action set of the target reinforcement learning model includes multiple file chunking strategies; the reward of the target reinforcement learning model includes the same ratio between the data chunks generated by using the current file chunking strategy for the data files of the current file type before and after local data changes respectively; Chunking the target-end data files by using the current action of the target reinforcement learning model; Constructing a Merkle tree of the target-end data files based on the chunking results of the target-end data files; Locate the differential data blocks by comparing the Merkle trees of the target - side data file and the source - side data file layer by layer; Obtain the differential data blocks from the source side.
[0008] In a third aspect, a data platform is provided in the present invention for storing data files. When the data platform is the source side, it uses the data - platform file migration method described in the first aspect to migrate the data files to the target side; when the data platform is the target side, it uses the data - platform file acquisition method described in the second aspect to obtain data files from the target side.
[0009] In a fourth aspect, a computer program product is provided in the present invention. The computer program product includes a computer program, and when the computer program is executed, it implements the data - platform file migration method described in the first aspect or the data - platform file acquisition method described in the second aspect.
[0010] Compared with the related art, the data - platform file migration method provided by the present invention only migrates the differential data blocks, thereby improving the data migration efficiency. It is necessary to block the source - side data file and the target - side data file to be migrated and construct Merkle trees, and then compare the Merkle trees of the source side and the target side to determine the differential data blocks. Compared with the same - type file migration methods in the prior art, when blocking the data files, a dynamic blocking strategy is adopted, that is, a more suitable blocking strategy is targeted for different types of files. When the blocking logic matches the file storage logic, fewer differential data blocks can be generated when there are local data changes in the data file, improving the subsequent comparison efficiency of the Merkle trees, being able to locate the differential data blocks more quickly, and fewer differential data blocks can be transmitted during data migration, improving the data transmission efficiency, and solving the problem of low data migration efficiency caused by unreasonable data - blocking strategies in the current data - platform file migration methods.
[0011] The details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects, and advantages of this application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a flowchart of the data - platform file migration method provided in this embodiment; Figure 2 is a flowchart of the data - platform file acquisition method provided in this embodiment; Figure 3 is a flowchart of the data migration process provided in this embodiment; Figure 4 is an architecture diagram of the data migration system provided in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] To more clearly understand the purpose, technical solution and advantages of this application, the following describes and explains this application in conjunction with the accompanying drawings and embodiments.
[0014] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by those with ordinary skills in the technical field to which this application belongs. In this application, words such as "a", "one", "a kind of", "the", "these" and the like do not indicate a limitation in quantity, and they can be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connection", "connection", "coupling" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether directly connected or indirectly connected. The "multiple" involved in this application means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are an "or" relationship. The terms "first", "second", "third" and the like involved in this application only distinguish similar objects and do not represent a specific order for the objects.
[0015] To meet scenarios such as enterprise cloud migration, data disaster recovery or multi-data center collaboration, it is usually necessary to migrate key file data into or out of the data platform. Generally speaking, the incoming end of the data file is the target end, and the outgoing end of the data file is the source end. As follows, a data platform migration method is provided, which is executed by the source end in data migration.
[0016] In this embodiment, a data platform file migration method is provided. Figure 1 is the flowchart of the data platform file migration method provided in this embodiment, as Figure 1 shown, this process includes step S110, step S120, step S130, step S140, step S150 and step S160.
[0017] Step S110, scan the source-end data files to be migrated to obtain the types of the source-end data files.
[0018] The source end stores a data file, which is defined as the source-end data file. When it is necessary to migrate the source-end data file out, the source-end data file to be migrated is first scanned to obtain the type of the source-end data file. Specifically, it can be first determined whether the source-end data file is a row-based data file. If so, subsequent operations can directly divide it into blocks by rows, that is, every n rows of data are used as a data block, where n is an integer greater than or equal to 1; if not, it continues to be determined whether the source-end data file is a database file, a data lake file, or a data warehouse file. If so, the configuration file of the source-end data file is read and the file header is read to obtain the file format (type) of the source-end data file. The specific types of data files include B-tree index files, heterogeneous database files, Oracle database files, PostgreSQL database files, MySQL database files, data lake files, and data warehouse files, and these specific types are in the state set of the target reinforcement learning model; if not (that is, the type of the source-end data file is not in the state set of the target reinforcement learning model), it continues to be determined whether there is a custom block division strategy for the source-end data file. If so, the custom block division strategy is used to divide the source-end data file subsequently. The custom block division strategy refers to a specific block division strategy set in advance for some special file types; if not, the block division strategy with the block size inversely proportional to the modification frequency is used to divide the source-end data file subsequently. The block division strategy with the block size inversely proportional to the modification frequency is as follows:
[0019] where MB is megabyte, S is the block size of data block division, and N mod is the historical modification times of the source-end data file.
[0020] Step S120: If the type of the source data file is in the state set of the target reinforcement learning model, use the type of the source data file as the current state of the target reinforcement learning model, and obtain the current action of the target reinforcement learning model in response to the current state. Herein, the state set of the target reinforcement learning model includes multiple file types, and the action set of the target reinforcement learning model includes multiple file chunking strategies. The reward of the target reinforcement learning model includes the same ratio (defined as the block cache hit rate) between the data chunks generated by using the current file chunking strategy for the data file of the current file type before and after local changes. By setting this reward, the target reinforcement learning model can match the best action (file chunking strategy) for each state (file type). The principle is as follows: When the chunking of a certain type of data file is reasonable, the number of data chunks with corresponding data changes after local data changes should be minimized. Exemplarily, for a data file using a logical page storage structure, related data is stored in the same logical page, and a single local data change usually occurs within the same logical page. When using the logical page chunking strategy, a single local data change will only cause a change in a single data chunk, and in subsequent data migrations, only this single data chunk needs to be migrated, minimizing the migration quantity. When using other chunking strategies, the same logical page may be divided into two or more data chunks, and a single local data change may cause changes in multiple data chunks. In subsequent data migrations, multiple data chunks need to be migrated, increasing the migration quantity. Therefore, the higher the block cache hit rate, the more similar the chunking logic of the current file chunking strategy is to the storage logic of the current file type, improving the subsequent difference location accuracy and minimizing the subsequent data chunk migration quantity.
[0021] Among the above rewards, it includes the same ratio between the data chunks generated by the current file chunking strategy for the data files of the current file type before and after local data change. The higher the same ratio between the two sets of data chunks, the lower the difference ratio. At this time, the target reinforcement learning model can obtain a higher reward. By continuously traversing different file types as the current file type and continuously traversing different file chunking strategies as the current file chunking strategy, it can drive the target reinforcement learning model to match the most reasonable file chunking strategy for each file type. Exemplarily, the data files of file type A are chunked by file chunking strategy B before and after local data change to obtain 100 data chunks respectively. Assuming that there are 97 exactly the same data chunks and 3 different data chunks, then the same ratio (chunk cache hit rate) between the data chunks is 97%. Also, the data files of file type A are chunked by file chunking strategy C before and after local data change to obtain 50 data chunks. Assuming that there are 49 exactly the same data chunks and 1 different data chunk, then the same ratio (chunk cache hit rate) between the data chunks is 98%. For the data files of file type A, when using file chunking strategy C, the target reinforcement learning model can obtain a higher reward, so the target reinforcement learning model tends to match file chunking strategy C for the data files of file type A.
[0022] By setting the above rewards, the target reinforcement learning model can match the best action for each state.
[0023] Exemplarily, the multiple file types include B-tree index files, heterogeneous database files, Oracle database files, PostgreSQL database files, MySQL database files, data lake files, and data warehouse files. The multiple file chunking strategies include timestamp chunking strategy, logical page boundary chunking strategy, logical page boundary and node hybrid chunking strategy, and fixed chunking strategy. Among them, the target reinforcement learning model is a DQN agent containing a three-layer convolutional neural network.
[0024] Specifically, based on the DQN method in reinforcement learning, an agent with a three-layer convolutional neural network is trained. Using various file types as states, various chunking strategies as action sets, and the data synchronization chunk cache hit rate after data change when adopting this chunking strategy as the reward, to drive the agent to generate decision-making actions. Among them, the agent is implemented by constructing a three-layer shallow neural network based on the convolutional neural network. The states here can be divided into: B-tree index files, heterogeneous database files, Oracle database files, PostgreSQL database files, MySQL database files, data lakes, and data warehouse files. The corresponding chunking strategies can be classified as follows: Timestamp Chunking Strategy: For example, append content every 10 minutes and divide it into a logical block. Set the block boundary at the end of the log line to avoid cross-line cutting.
[0025] Logical Page Boundary Chunking: According to the preset file format parsing algorithm, read the file header, which is a separate block. Then read the content data and perform content chunking according to the parsing algorithms preset for different formats, and record the chunk hash value and related metadata.
[0026] Logical page alignment chunking for data files such as Oracle database files, PostgreSQL database files, and MySQL database files. For some data files using a logical page storage structure, related data is stored in the same logical page, and there are a large number of "data holes" (i.e., empty pages not occupied by valid data) caused by deletion operations. By strictly aligning the logical page boundary for chunking (such as according to Oracle's DB_BLOCK_SIZE or MySQL's 16KB page size), data integrity within a single page is ensured. This approach brings the following advantages: improved cache hit rate, strong data correlation within the same logical page, easier to trigger the locality principle during access after migration, reducing cache misses; optimized compression efficiency: the data structure within the logical page is regular, and the empty areas can be efficiently identified and skipped during transmission. Combining with compression algorithms (such as Zstandard) can significantly reduce the amount of data transmitted; accurate difference detection, single-page modification only affects the corresponding chunk, avoiding the failure of difference detection caused by traditional fixed chunking crossing logical boundaries.
[0027] Optimization of logical page chunking for data lake / warehouse columnar storage. For the frequent multi-file merging scenarios in data lakes (such as Parquet files) and data warehouses, the data in the merged file actually does not change, and its essence is the change of metadata. Traditional chunking algorithms are limited to single-file comparison and cannot identify consistent column chunks across files, resulting in repeated transmission. This solution improves the parsing of the file Schema, extracts common column chunks across files, decouples the chunking strategy from the merging logic, and ensures that when a new file is merged, only the new column chunks need to be compared instead of the entire file.
[0028] Logical Page Boundary and Node Hybrid Chunking: Data files are chunked according to the page size of the tablespace file, and index files use logical page boundary chunking. Combining with the hierarchical characteristics of the B-tree structure, chunking is performed in units of nodes and data blocks.
[0029] Perform B - Tree node chunking on index files such as MySQL indexes and PostgreSQL index files. The index files adopt a hierarchical B - Tree structure, where each node contains a key - value range and pointers to child nodes. The chunking strategy takes B - Tree nodes as the smallest unit. B - Tree updates usually only affect local nodes. After chunking, only the changed nodes need to be transmitted instead of the entire index file.
[0030] Fixed chunking: Split the file into chunks of a preset fixed size.
[0031] First, complete the training of the DQN agent using pre - trained data. Subsequently, select a chunking strategy based on the trained agent for actual operation.
[0032] Exemplarily, the chunking strategy can be expressed as follows:
[0033] Among them, 、 and represent database files, data lake files, and data warehouse files respectively. f represents the data file to be chunked. 、 、 、 and represent the timestamp - based chunking strategy, the logical page boundary - based chunking strategy, the hybrid chunking strategy of logical page boundary and nodes, the fixed chunking strategy, and the chunking strategy where the chunk size is inversely proportional to the modification frequency respectively.
[0034] As follows, continue to explain the above - mentioned chunking strategies.
[0035] Logical page alignment chunking for data files. For some data files using a logical page storage structure, related data in the dataset is stored in the same logical page, and there are a large number of "data holes" (i.e., empty pages not occupied by valid data) generated due to deletion operations. By strictly aligning the chunking at the logical page boundary (such as according to Oracle's DB_BLOCK_SIZE or MySQL's 16KB page size), the data integrity within a single page is ensured. This brings the following advantages: improved cache hit rate. The data within the same logical page has strong relevance, and when accessed after migration, it is more likely to trigger the locality principle, reducing cache misses; optimized compression efficiency: The data structure within the logical page is regular, and the empty areas can be efficiently identified and skipped during transmission. Combined with compression algorithms (such as Zstandard), the amount of data to be transmitted can be significantly reduced; accurate difference detection. Modifications to a single page only affect the corresponding chunk, avoiding the failure of difference detection caused by traditional fixed chunking crossing logical boundaries.
[0036] Index file B-Tree node chunking. The index file adopts a hierarchical B-Tree structure, and each node contains a key range and pointers to child nodes. The chunking strategy takes B-Tree nodes as the smallest unit. B-Tree updates usually only affect local nodes. After chunking, only the changed nodes need to be transmitted instead of the entire index file.
[0037] Chunking optimization for data lake / warehouse columnar storage. For the frequent multi-file merging scenarios in data lakes (such as Parquet files) and data warehouses, the data in the merged files actually does not change. Its essence is the change of metadata. Traditional chunking algorithms are limited to single-file comparison and cannot identify consistent column chunks across files, resulting in repeated transmissions. This solution improves the parsing of the file Schema, extracts common column chunks across files, and decouples the chunking strategy from the merging logic to ensure that when new files are merged, only the newly added column chunks need to be compared instead of the entire files.
[0038] In step S130, the source data file is chunked using the current action of the target reinforcement learning model. The current action of the target reinforcement learning model is the chunking strategy selected by the target reinforcement learning model for the source data file, and the source data file is chunked through this chunking strategy.
[0039] In step S140, a Merkle tree of the source data file is constructed based on the chunking result of the source data file.
[0040] The chunking result of the source data file contains each data chunk generated after the source data file is chunked. The specific construction process of the Merkle tree is as follows: Calculate the hash value of each data chunk in the chunking result of the source data file and store the metadata. The metadata includes the hash value, chunk size, start offset of each data chunk, and the type identifier of the source data file; use the hash value of each data chunk as the bottom-level node in the Merkle tree, and sequentially construct the upper-level nodes until the root node; in the Merkle tree, an upper-level node is the hash value of at least two lower-level nodes. Exemplarily, assuming there are 10 data chunks, there are 10 nodes in the bottom layer of the Merkle tree, which are the hash values of 10 data chunks respectively. Then, the hash values of every two bottom-level nodes can be calculated to obtain 5 nodes in the second-to-last layer, and so on until the root node is generated.
[0041] Among them, the start offset offset of the data chunk is:
[0042] Among them, S is the chunk size of the data chunk, H is the file header size of the source data file, and filesize is the file size of the source data file.
[0043] Step S150: Locate the differential data blocks by comparing the Merkle trees of the source - side data file and the target - side data file layer by layer through hashing.
[0044] Correspondingly, on the target side, the stored data is used to obtain the corresponding Merkle tree by the same Merkle tree construction method, which facilitates the comparison between the two ends. The process of layer - by - layer hashing comparison is as follows: First, compare the root nodes. If the root nodes are the same, it means there are no differential data blocks. If the root nodes are different, determine the different nodes in the lower - level nodes, and in the lower - level nodes with differences, determine the different lower - lower - level nodes until the bottom - most nodes, then the differential data chunks can be determined.
[0045] Merkle tree comparison function is:
[0046] where k is the branching factor of the tree, h is the height of the tree, the number of comparisons is c, and the recursive process satisfies:
[0047] The comparison complexity of the above - mentioned Merkle tree is significantly better than that of the traditional full - volume comparison.
[0048] Step S160: Migrate the differential data blocks to the target side.
[0049] During the data migration process, only the data blocks with differences are migrated. Compared with the full - volume data migration, the data migration efficiency is greatly improved.
[0050] In summary, the data platform file migration method provided in this embodiment only migrates the data blocks with differences, thus improving the data migration efficiency. It is necessary to block the source - side data file and the target - side data file to be migrated and construct Merkle trees, and then compare the Merkle trees of the source side and the target side to determine the differential data blocks. Compared with the same - type file migration methods in the prior art, when blocking the data files, a dynamic blocking strategy is adopted, that is, a more appropriate blocking strategy is targeted for different types of files. When the blocking logic matches the file storage logic, fewer differential data blocks can be generated when there are local data changes in the data file, improving the subsequent comparison efficiency of the Merkle tree, being able to locate the differential data blocks more quickly, and being able to transfer fewer differential data blocks during data migration, improving the data transmission efficiency, and solving the problem of low data migration efficiency caused by unreasonable data blocking strategies in the current data platform file migration methods.
[0051] It should be noted that if there are multiple types of data files in the source data file, multiple Merkle trees can be established for the multiple types of data files respectively (the same construction strategy is adopted for the target data file) to reduce the depth of the Merkle tree and improve the construction speed. The differences are quickly located by comparing the root hashes layer by layer: if the root hashes are different, the child node hashes are compared layer by layer downward to recursively locate the different subtrees. Finally, the list of different blocks corresponding to all leaf nodes in the different subtrees is extracted, and only the hashes, metadata, and data content of these blocks are transmitted, and the Merkle tree at the target end is updated. During the transmission process, each block is encrypted independently, and the key is generated through the key exchange protocol of TLS 1.3. The data blocks are stored in the cache at the target end to avoid transmitting the already cached data blocks. The breakpoint resumption mechanism is adopted to record the hash list of the transmitted blocks to the local log file. After interruption, the log is read to compare the unfinished blocks and resume the transmission.
[0052] In addition, if there are multiple versions of data files in the source end, multiple versions of Merkle trees can be constructed for the multiple versions of data files respectively, so as to facilitate users to selectively update the target end with different versions of data files in the source end. Each version generates a unique identifier (such as v1.0), and stores the corresponding Merkle tree root hash and version metadata (creation time, migration task ID). An example of version metadata is as follows:
[0053] The rollback process loads the root hash of the corresponding Merkle tree according to the version number. The current tree at the target end is compared with the historical tree at the source end to locate the different blocks that need to be restored. The different blocks are retransmitted from the cache or the source end to overwrite the data file at the target end.
[0054] At the same time, expired data cleaning and lifecycle management. Between different versions of Merkle trees, a log file recording which data blocks were deleted compared to the previous version is maintained. When the current version of the Merkle tree expires or is deleted, the corresponding cached data blocks are found and deleted.
[0055] Therefore, in this embodiment, through the dynamic chunking strategy, the chunking mode (such as timestamp chunking, page alignment chunking, logical page chunking) is adaptively selected for row-based files, heterogeneous databases, and data lake warehouse file types, avoiding the problem of differential detection failure caused by cutting across logical boundaries. Combined with the Merkle tree layer-by-layer hash comparison mechanism, only the differential block data needs to be transmitted, significantly reducing the network bandwidth consumption. Deeply binding the Merkle root hash to the migration version realizes the atomicity of migration operations and version management, supports quickly locating differential blocks and resuming transmission after interruption, without full retransmission, and significantly improves the recovery efficiency. By recording the deletion logs of differential blocks between versions and combining with the caching mechanism, automatic cleaning of expired data is realized, reducing the storage resource occupation. At the same time, multi-version coexistence management is supported to meet the requirements of data auditing and historical traceability. In summary, the present invention has achieved improvements in version control, dynamic adaptability, and resource utilization.
[0056] It should be noted that the data migration method in this embodiment needs to construct a Merkle tree in the same way at the source end and the target end. Furthermore, this embodiment also provides a method for obtaining a data platform file, which is executed by the target end in data migration and is used to obtain file data from the source end, including step S210, step S220, step S230, step S240, step S250, and step S260.
[0057] Step S210: Scan the target end data file to be updated to obtain the type of the target end data file.
[0058] Step S220: If the type of the target end data file is in the state set of the target reinforcement learning model, use the type of the target end data file as the current state of the target reinforcement learning model, and obtain the current action of the target reinforcement learning model in response to the current state; wherein, the state set of the target reinforcement learning model includes B-tree index files, heterogeneous database files, Oracle database files, PostgreSQL database files, MySQL database files, data lake files, and data warehouse files, and the action set of the target reinforcement learning model includes timestamp chunking strategy, logical page boundary chunking strategy, logical page boundary and node hybrid chunking strategy, and fixed chunking strategy.
[0059] Step S230: Chunk the target end data file using the current action of the target reinforcement learning model.
[0060] Step S240: Construct a Merkle tree of the target end data file based on the chunking result of the target end data file.
[0061] Step S250: Locate the differential data blocks by layer-by-layer hash comparison of the Merkle trees of the target end data file and the source end data file.
[0062] Step S260: Obtain the differential data blocks from the source end.
[0063] In this data platform file acquisition method, the way to construct the Merkle tree is exactly the same as that in the data platform file migration method, except that the processing objects are different. In the data platform file migration method, the source data files to be migrated are used, while in the data platform file acquisition method, the target data files to be updated are used. Therefore, the way to construct the Merkle tree will not be elaborated here.
[0064] As described above, the data migration process has been explained from the perspective of the source or the target in data migration. Refer to Figure 3 , the overall process of data migration in this embodiment is as follows: 1. Scan metadata. In the source, the metadata is the metadata of the source data files. In the target, the metadata is the metadata of the target data files.
[0065] 2. Determine whether the data file is a row-based data file. If so, execute step 3; if not, execute step 4.
[0066] 3. Chunk the data file according to the row-based chunking strategy.
[0067] 4. Determine whether the data file is a database file. If so, execute step 6; if not, execute step 5.
[0068] 5. Determine whether the data file is a data lake or data warehouse file. If so, execute step 6; if not, execute step 7.
[0069] 6. Read the file header of the configuration file of the data file, identify the file format of the data file and provide it to the DQN agent, and the DQN agent gives a chunking strategy, which includes timestamp chunking strategy, logical page boundary chunking strategy, logical page boundary and node hybrid chunking strategy, and fixed chunking strategy.
[0070] 7. Determine whether the data file has a custom chunking interface. If so, execute step 8; if not, execute step 9.
[0071] 8. Chunk the data file using the custom interface chunking strategy.
[0072] 9. Define the data file as a normal file and chunk the data file according to the update frequency chunking strategy (a chunking strategy where the chunk size is inversely proportional to the modification frequency).
[0073] Refer to Figure 4, in this embodiment, a data system is further provided, which includes a source file system and a target file system. Both of them include a metadata scanning module, a dynamic chunking module, and a comparison module. The metadata scanning module is used to scan the metadata to obtain the type of the data file (perform the above step 1). The dynamic chunking module is used to perform dynamic chunking on the data file to construct a corresponding Merkle tree (perform the above steps 2-9). The comparison module is used to compare the Merkle trees of the source end and the target end to determine the differential data chunks. The cache module is used to manage the cache of the transferred data chunks, avoid repeated transmission, and support the cleaning of expired data.
[0074] In this embodiment, a data platform is further provided for storing data files. When the data platform is the source end, it migrates the data files to the target end by using the data platform file migration method provided in this embodiment; when the data platform is the target end, it obtains the data files from the target end by using the data platform file acquisition method provided in this embodiment.
[0075] In this embodiment, a computer program product is further provided. The computer program product includes a computer program, and when the computer program is executed, it implements the data platform file migration method or the data platform file acquisition method provided in this embodiment.
[0076] It should be understood that the specific embodiments described here are only used to explain this application, rather than to limit it. According to the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of this application.
[0077] Obviously, the accompanying drawings are only some examples or embodiments of this application. For those of ordinary skill in the art, this application can also be applied to other similar situations based on these drawings without creative work. In addition, it can be understood that although the work done during the development process here may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient disclosure of this application.
Claims
1. A data platform file migration method for migrating data files to a target end, characterized in that, It includes: Scanning the source - end data file to be migrated to obtain the type of the source - end data file; If the type of the source - end data file is in the state set of the target reinforcement learning model, using the type of the source - end data file as the current state of the target reinforcement learning model, and obtaining the current action of the target reinforcement learning model in response to the current state; wherein, the state set of the target reinforcement learning model includes multiple file types, and the action set of the target reinforcement learning model includes multiple file chunking strategies; the reward of the target reinforcement learning model includes the same ratio between the data chunks generated by the current file chunking strategy for the data files of the current file type before and after local data changes. Chunking the source - end data file using the current action of the target reinforcement learning model; Constructing a Merkle tree of the source - end data file based on the chunking result of the source - end data file; Locating the differential data chunks by layer - by - layer hash comparison of the Merkle trees of the source - end data file and the target - end data file; Migrating the differential data chunks to the target end.
2. The data platform file migration method according to claim 1, wherein It also includes: Before scanning the source - end data file to be migrated, determining whether the source - end data file is a row - based data file; If so, chunking the source - end data file by rows; if not, scanning the source - end data file; If the type of the source - end data file is not in the state set of the target reinforcement learning model, determining whether there is a custom chunking strategy for the source - end data file; if so, chunking the source - end data file using the custom chunking strategy; if not, chunking the source - end data file using a chunking strategy where the chunk size is inversely proportional to the modification frequency.
3. The data platform file migration method according to claim 1, wherein The target reinforcement learning model is a DQN agent containing a three - layer convolutional neural network.
4. The data platform file migration method according to claim 1, wherein The multiple file types include B - tree index files, heterogeneous database files, Oracle database files, PostgreSQL database files, MySQL database files, data lake files, and data warehouse files, and the multiple file chunking strategies include timestamp chunking strategy, logical page boundary chunking strategy, logical page boundary and node hybrid chunking strategy, and fixed chunking strategy.
5. The data platform file migration method according to claim 1, characterized in that, Constructing a Merkle tree of the source - end data file based on the chunking result of the source - end data file includes: Calculating the hash value of each data chunk in the chunking result of the source - end data file and storing metadata, where the metadata includes the hash value, chunk size, start offset of each data chunk, and the type identifier of the source - end data file; Using the hash value of each data chunk as the bottom - layer node in the Merkle tree and successively constructing the upper - level nodes until the root node; in the Merkle tree, an upper - level node is the hash value of at least two lower - level nodes.
6. The data platform file migration method according to claim 5, wherein The start offset offset of the data chunk is: , where S is the chunk size of the data chunk, H is the file header size of the source - end data file, and filesize is the file size of the source - end data file.
7. The data platform file migration method according to claim 2, characterized in that, The chunking strategy where the chunk size is inversely proportional to the modification frequency is: , Wherein, MB is megabyte, S is the block size of data chunking, and N mod is the historical modification times of the source - side data file.
8. A method for obtaining data platform files, which is used to obtain file data from a source end, and is characterized in that, It includes: Scanning the target - end data file to be updated to obtain the type of the target - end data file; If the type of the target - end data file is in the state set of the target reinforcement learning model, use the type of the target - end data file as the current state of the target reinforcement learning model, and obtain the current action of the target reinforcement learning model in response to the current state; wherein, the state set of the target reinforcement learning model includes multiple file types, and the action set of the target reinforcement learning model includes multiple file chunking strategies; the reward of the target reinforcement learning model includes the same ratio between the data chunks generated by the current file chunking strategy for the data file of the current file type before and after the local data change. Chunk the target - end data file using the current action of the target reinforcement learning model. Construct a Merkle tree of the target - end data file based on the chunking result of the target - end data file. Locate the differential data chunks by hashing and comparing the Merkle trees of the target - end data file and the source - end data file layer by layer. Obtain the differential data chunks from the source - end.
9. A data platform for storing data files, characterized in that: When the data platform is the source - end, it uses the data platform file migration method described in any one of claims 1 - 7 to migrate the data file to the target - end. When the data platform is the target - end, it uses the data platform file acquisition method described in claim 8 to acquire the data file from the target - end.
10. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is executed, it implements the data platform file migration method described in any one of claims 1 - 7 or the data platform file acquisition method described in claim 8.
Citation Information
Patent Citations
Internet of vehicles edge cache decision method and system based on multi-agent federated learning
CN116915803A
Parallel synchronization method and system for unstructured files
CN119513057A
Big data-based file analysis, storage and division mode and distributed storage system
CN119861881A
Cloud backup platform data deduplication method based on distributed storage
CN119938406A
Replicating and migrating files to secondary storage sites
US20190042595A1
Cited By
Power grid power transformation engineering knowledge graph construction and retrieval method and system
CN121561116A
A power grid substation engineering knowledge graph construction and retrieval method and system
CN121561116B
Charging pile firmware online upgrading method and system
CN121567571A
INP file analysis method
CN121807316A