A method, system, device, storage medium and product for processing data blocks
By establishing data block processing methods between the metadata management center and the client in the cloud data center, generating and parsing metadata snapshots, the problem of low supply efficiency of large-scale small file data sets is solved, and the computing performance and efficiency of deep learning training jobs are improved.
Patent Information
- Application Number
- CN202510332559.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-20
AI Technical Summary
The prior art processes large-scale small file data sets in cloud data centers, and the supply efficiency is low, resulting in limited computing performance and efficiency of deep learning training jobs.
By establishing a data block processing method between the metadata management center and the client, obtaining a tree structure storage table, leaf node storage table and data block storage table, establishing a hierarchical relationship, generating a metadata snapshot, and loading and parsing the metadata snapshot from the local cache at the start of the training job to obtain the required small file data set.
It improves the supply efficiency of small file data sets, reduces the access pressure to the metadata management center, accelerates the access efficiency of metadata, and reduces the consumption of storage and computing resources.
Smart Images

Figure CN119848050B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data storage, and particularly to a method, system, device, storage medium and product for processing data blocks. Background Art
[0002] Currently, more and more artificial intelligence training jobs are running in cloud data centers. The elastic GPU computing resource capabilities provided by cloud data centers can meet the diverse computing resource requirements of artificial intelligence training jobs. For training data, especially large-scale small file datasets, their supply efficiency is directly related to the computing performance and efficiency of the distributed deep learning job system in cloud data centers.
[0003] In related technologies, for large-scale small file datasets, usually a high-performance storage is equipped on the computing nodes of deep learning jobs, and the small file datasets in the remote storage system are copied onto the high-performance storage to provide data for deep learning jobs to train. However, due to the limited capacity of high-performance storage, small file datasets exceeding the capacity cannot use this method to accelerate data supply; or, the small file datasets are stored in a parallel file system, and the data in the small file datasets is pulled from the parallel file system and fed to the computing nodes of deep learning jobs. However, due to the large performance bottlenecks in data querying and random reading of the parallel file system, there will be a mismatch between the data supply efficiency and the computing speed of the computing nodes. Therefore, neither of the above two methods can effectively improve the supply efficiency of small file datasets in deep learning training jobs. Summary of the Invention
[0004] This application provides a method, system, device, storage medium and product for processing data blocks to at least solve the problem of low supply efficiency of small file datasets in deep learning training jobs in related technologies.
[0005] This application provides a method for processing data blocks, which is applied to a metadata management center. The method includes: obtaining a tree-structured storage table, a leaf node storage table, and a data block storage table. The tree-structured storage table stores the first node identifier of the first node. The leaf node storage table includes the attribute information of each leaf node and each node identifier in the tree structure corresponding to multiple data blocks. The data block storage table includes the data block information, data block identifier, and node identifier of each leaf node for each data block stored in each of the multiple leaf nodes, where the first node is the root node or any one of the leaf nodes in the tree structure, and the multiple data blocks are obtained by aggregating multiple small files required for a training job.
[0006] Search in the leaf node storage table and the data block storage table respectively to find out whether there is a node identifier that matches the first node identifier; if node identifiers that match are found in both, establish a first-level relationship between the first node, the attribute information of the first node, and the data block information of the first node; according to the first data block identifier of the data block corresponding to the first node, determine the first small file list information included in the data block corresponding to the first node, and add the first small file list information to the first-level relationship to obtain a second-level relationship; generate a metadata snapshot according to the second-level relationship and at least one third-level relationship of other nodes, and send the metadata snapshot to the client so that the client stores it in the local cache and obtains the metadata snapshot when the training job starts. The other nodes are at least one of the other leaf nodes in the tree structure except the first node.
[0007] The present application also provides another data block processing method, which is applied to the client. The method includes: after the training job starts, load the metadata snapshots corresponding to multiple data blocks from the local cache. The metadata snapshots are generated based on the tree structure storage table, the leaf node storage table, and the data block storage table. The tree structure storage table stores the first node identifier of the first node. The leaf node storage table includes the attribute information and each node identifier of each leaf node in the tree structure. The data block storage table includes the data block information, data block identifier, and the node identifier of each leaf node of each leaf node in the multiple leaf nodes. The first node is the root node or any one of the leaf nodes in the tree structure. The multiple data blocks are obtained by aggregating multiple small files required for the training job; detect whether the snapshot timestamp of the metadata snapshot is consistent with the latest snapshot timestamp of the metadata management center; if so, after the metadata snapshot is loaded, read the multiple data blocks from the backend aggregated file system based on the metadata snapshot; perform a de-aggregation operation on the multiple data blocks to obtain multiple small files.
[0008] The present application also provides a data block processing device, which includes: a first transceiver module, configured to obtain a tree structure storage table, a leaf node storage table, and a data block storage table. The tree structure storage table stores the first node identifier of the first node. The leaf node storage table includes the attribute information and each node identifier of each leaf node in the tree structure corresponding to the multiple data blocks. The data block storage table includes the data block information, data block identifier, and the node identifier of each leaf node of each leaf node in the multiple leaf nodes. The first node is the root node or any one of the leaf nodes in the tree structure. The multiple data blocks are obtained by aggregating multiple small files required for the training job.
[0009] The first processing module is used to separately search in the leaf node storage table and the data block storage table to check if there are node identifiers that match the first node identifier; if matching node identifiers are found in both, establish a first-level relationship between the first node, the attribute information of the first node, and the data block information of the first node; determine the first small file list information included in the data block corresponding to the first node according to the first data block identifier of the data block corresponding to the first node, and add the first small file list information to the first-level relationship to obtain a second-level relationship; generate a metadata snapshot based on the second-level relationship and at least one third-level relationship of other nodes, and send the metadata snapshot to the client so that the client stores it in the local cache and obtains the metadata snapshot when the training job is started. The other nodes are at least one other leaf node in the tree structure except the first node.
[0010] This application also provides another data block processing device, which includes: a second transceiver module, used to load metadata snapshots corresponding to multiple data blocks from the local cache after the training job is started. The metadata snapshots are generated based on a tree structure storage table, a leaf node storage table, and a data block storage table. The tree structure storage table stores the first node identifier of the first node. The leaf node storage table includes the attribute information and each node identifier of each leaf node in the tree structure. The data block storage table includes the data block information, data block identifier, and the node identifier of each leaf node of each of the multiple leaf nodes for each data block stored by each leaf node. The first node is the root node or any one leaf node in the tree structure. The multiple data blocks are obtained by aggregating multiple small files required for the training job.
[0011] The second processing module is used to detect whether the snapshot timestamp of the metadata snapshot is consistent with the latest snapshot timestamp of the metadata management center; if so, after the metadata snapshot is loaded, read multiple data blocks from the backend aggregated file system based on the metadata snapshot; perform a de-aggregation operation on the multiple data blocks to obtain multiple small files.
[0012] This application also provides a data block processing system, which includes a metadata management center and a client. The metadata management center is used to obtain a tree structure storage table, a leaf node storage table, and a data block storage table. The tree structure storage table stores the first node identifier of the first node. The leaf node storage table includes the attribute information and each node identifier of each leaf node in the tree structure corresponding to multiple data blocks. The data block storage table includes the data block information, data block identifier, and the node identifier of each leaf node of each of the multiple leaf nodes for each data block stored by each leaf node. The first node is the root node or any one leaf node in the tree structure. The multiple data blocks are obtained by aggregating multiple small files required for the training job.
[0013] The metadata management center is further configured to separately search in the leaf node storage table and the data block storage table to check whether there is a node identifier that matches the first node identifier; if node identifiers that match are found in both, establish a first-level relationship between the first node, the attribute information of the first node, and the data block information of the first node; determine, according to the first data block identifier of the data block corresponding to the first node, the first small file list information included in the data block corresponding to the first node, and add the first small file list information to the first-level relationship to obtain a second-level relationship; generate a metadata snapshot according to the second-level relationship and at least one third-level relationship of other nodes, and send the metadata snapshot to the client, so that the client stores it in the local cache and obtains the metadata snapshot when the training job is started. The other nodes are at least one leaf node other than the first node in the tree structure.
[0014] The client is configured to, after the training job is started, load the metadata snapshots corresponding to multiple data blocks from the local cache. The metadata snapshots are generated based on the tree structure storage table, the leaf node storage table, and the data block storage table. The tree structure storage table stores the first node identifier of the first node. The leaf node storage table includes the attribute information of each leaf node in the tree structure and each node identifier. The data block storage table includes the data block information, the data block identifier, and the node identifier of each leaf node of each data block stored by each leaf node among multiple leaf nodes. The first node is the root node or any one leaf node in the tree structure. The multiple data blocks are obtained by aggregating multiple small files required for the training job; detect whether the snapshot timestamp of the metadata snapshot is consistent with the latest snapshot timestamp of the metadata management center; if so, after the metadata snapshot is loaded, read the multiple data blocks from the backend aggregated file system based on the metadata snapshot; perform a de-aggregation operation on the multiple data blocks to obtain multiple small files.
[0015] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above data block processing methods when executing the computer program.
[0016] This application also provides a computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps of any of the above data block processing methods.
[0017] This application also provides a computer program product including a computer program, where the computer program, when executed by a processor, implements the steps of any of the above data block processing methods.
[0018] With this application, since the metadata management center can send the metadata snapshot of the training dataset (multiple data blocks) to the computing node where the client is located (which is also the computing node where the application is located), the application can directly read the metadata snapshot from the local cache, reducing the access pressure on the metadata management center and accelerating the metadata access efficiency. The client can quickly obtain the metadata snapshot from the local cache and read the required metadata information from it, without frequently accessing the metadata management center, thereby reducing network overhead and improving the metadata access efficiency, and further improving the supply efficiency of the small file dataset.
[0019] In addition, during the execution of the training job, the dataset is usually read-only, and there will be no update of the metadata snapshot due to the update of the training dataset. Therefore, the localized metadata snapshot cache is applicable to deep learning training jobs. Through the metadata snapshot mechanism, during the training of large-scale small file datasets, high-efficient metadata access performance can be maintained, while reducing the consumption of storage and computing resources and accelerating the parsing and supply efficiency of the training dataset based on data blocks. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To more clearly illustrate the embodiments of this application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1 It is a topological structure diagram of a data block processing system provided by an embodiment of this application;
[0022] Figure 2 It is a schematic diagram of a high-performance data reading pipeline provided by an embodiment of this application;
[0023] Figure 3 It is a schematic flow diagram of a data block processing method provided by an embodiment of this application;
[0024] Figure 4 It is a schematic diagram of metadata construction for data blocks provided by an embodiment of this application;
[0025] Figure 5 It is a schematic flow diagram of another data block processing method provided by an embodiment of this application;
[0026] Figure 6 It is a schematic diagram of reading the number of mini-batches corresponding to data blocks in an embodiment of this application;
[0027] Figure 7 It is a structural view of a data block provided by an embodiment of this application;
[0028] Figure 8 Block diagram of a data block processing device provided by an embodiment of the present application;
[0029] Figure 9 Block diagram of another data block processing device provided by an embodiment of the present application;
[0030] Figure 10 Schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0032] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and not to describe a specific order or sequence.
[0033] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0034] In combination with a specific application environment architecture or a specific hardware architecture on which the execution of a data block processing method depends, the specific application environment architecture or the specific hardware architecture will be described herein.
[0035] The embodiments of the present application are applied to a scenario where a storage system needs to provide a small file data set as a training set to a distributed deep learning job system for model training.
[0036] For large-scale small file data sets, the mainstream processing method is to re-aggregate a large number of small files into a new data format to improve data read performance. However, this data format is only supported by individual programming frameworks and does not have universality. Moreover, the original data needs to be converted in format before the training job starts, and after the training starts, the data format needs to be converted back to the original data for calculation. This method limits its application scenarios.
[0037] For deep learning training jobs, it is common to download or collect private data from public libraries and convert it into a dataset format that can be used for deep learning training. Then, the prepared small file dataset is fed in batches to the computing processes of the computing nodes for calculation. For small file datasets, the file access characteristics of the training job include: each training task operates only on one small file dataset, and only read-only access to the small file dataset is performed during training. The small file dataset remains fixed during this process; the small file dataset is read sequentially by epoch, and all small files are read once in each epoch; between different epochs, multiple small files are randomly shuffled to avoid model overfitting.
[0038] Specifically, at the beginning of a typical deep learning training job, all the file names of the small file dataset are loaded into memory. Before each epoch calculation, the training framework (such as PyTorch) randomly shuffles the indices of the file names to generate a random order for reading file data. After that, the mini-batch dataset is read in batches in an iterative manner. Each mini-batch dataset is fed to the computing node (for example, the computing node can be a Graphics Processing Unit (GPU)) for calculating the parameter updates of the deep neural network. After the calculation is completed, the next batch of the dataset is read until all small files have been read once to complete the current epoch calculation. Before the model parameters converge, the training job often requires multiple rounds of epoch calculations. The overall training time depends on the training data supply efficiency and the GPU computing efficiency. In a large computing cluster, many training jobs run concurrently, and each training job may use different datasets. These training jobs will send a large number of small file read requests to the storage system where the dataset is located. At this time, if the storage system is inefficient in obtaining small files and the training data supply efficiency is slow, it will cause the GPU to wait frequently, resulting in a slowdown in the training process.
[0039] To solve the above technical problems, an embodiment of the present application proposes a data block processing method, which includes: obtaining a tree-structured storage table, a leaf node storage table, and a data block storage table, and when node identifiers matching the first node identifier are found in both the leaf node storage table and the data block storage table, establishing a first-level relationship between the first node, the attribute information of the first node, and the data block information; determining, according to the first data block identifier of the data block corresponding to the first node, the first small file list information included in the data block corresponding to the first node, and adding the first small file list information to the first-level relationship to obtain a second-level relationship; generating a metadata snapshot according to the second-level relationship and at least one third-level relationship of other nodes, and sending the metadata snapshot to the client, so that the client stores it in the local cache and obtains the metadata snapshot when the training job is started, and based on the metadata snapshot, quickly obtains multiple data blocks required for the training job from the file system, and then obtains the required small file dataset, improving the supply efficiency of the small file dataset in the training job.
[0040] The following uses Figure 1 the data block processing system shown as an example to describe the method provided by the embodiment of the present application.
[0041] As Figure 1 shown, Figure 1 it is a topological structure diagram of a data block processing system provided by an embodiment of the present application. Figure 1 In it, the data block processing system 100 includes a metadata management center 101, a client 102, and a backend aggregated file system 103. Optionally, the data block processing system 100 further includes a global shared file system 104.
[0042] The metadata management center 101 (also called CacheFS-Server) in the embodiment of the present application is used to manage and store the metadata information of the entire small file dataset. During the aggregation and writing of the small file dataset into the backend aggregated file system 103, the metadata management center 101 receives the metadata information extracted from the client 102 and stores the metadata information in the distributed database. To improve the access speed, the metadata management center 101 also caches a copy of the metadata information in the memory. The metadata management center 101 supports the metadata snapshot function, constructs a metadata snapshot in a preset format, and distributes it to the client 102.
[0043] The metadata management center 101 has backup and fault tolerance functions. In a distributed storage system, the availability and integrity of metadata directly affect the operation of the entire system. Therefore, the metadata management center adopts measures such as regular backup, transaction logging, data replication, and multi-node deployment to ensure the stability and reliability of the metadata management center, minimize the impact of system failures on the business, guarantee the security of metadata, and maintain the continuous service ability of the file system.
[0044] The metadata management center 101 is configured with full backup and incremental backup strategies to meet different recovery requirements. The full backup completely replicates all metadata in the database, usually executed during periods of low system load, with a relatively low execution frequency to reduce the impact on normal business. The incremental backup only records the data that has changed since the last backup, with a relatively high execution frequency, which not only saves storage space but also improves backup efficiency.
[0045] In terms of fault tolerance, the metadata management center relies on the replication function of the database and its master-slave switchover mechanism to achieve synchronization and fault tolerance processing of metadata. The replication function of the database allows metadata to be synchronized between multiple database instances, providing data redundancy and load balancing. When the primary database instance fails, the system can automatically switch to the standby database instance to ensure service continuity. In the master-slave replication mode, the primary database synchronizes the updated metadata to the slave database in real time. When the primary database is unavailable, the slave database can immediately take over to avoid service interruption.
[0046] It can be understood that when the metadata management center encounters data corruption or other unforeseen system failures, the backup and fault tolerance mechanisms can quickly initiate the recovery process. By using the backup data, the metadata can be restored in a short time to ensure that the business is not interrupted.
[0047] The client 102 (also known as CacheFS-Client) in the embodiments of this application is used to perform read and write operations on small file datasets. The client 102 consists of three parts: a write operation module, a read operation module, and a general operation module, each undertaking different functions.
[0048] Among them, the write operation module is used to scan small files stored in the global shared file system, aggregate the small files into larger data blocks according to a specific aggregation algorithm, store the aggregated data blocks in the backend aggregated file system, and report the metadata information of each data block to the metadata management center 101 in real time.
[0049] The read operation module is used to directly read the aggregated data blocks from the backend aggregated file system using the constructed high-performance data reading pipeline after the training task starts, and perform de-aggregation operations on them. In addition, the read operation module also integrates a data randomization function, enabling the data to meet the requirements of data randomization in deep learning training during the de-aggregation process. The data reading, de-aggregation, and randomization can be integrated into a continuous and efficient operation process, improving the access efficiency of small file data sets and ensuring the efficient supply of data during model training.
[0050] Optionally, the high-performance data reading pipeline integrates a data block parser (Data Reader), a metadata snapshot interpreter (Meta Interpreter), a data block address parser (Chunk Path Parser), and a chunk-based shuffle mechanism (Chunk-Based Shuffle) to achieve efficient reading of training data. As Figure 2 shown, Figure 2 is a schematic diagram of the high-performance data reading pipeline provided by an embodiment of the present application. In Figure 2 , the data reading pipeline obtains a metadata snapshot from the CacheFS metadata of the CacheFS metadata management center, and based on the metadata snapshot, it parses out the original multiple small file data from the stored data blocks in real time to form a mini-batch data, and feeds it to the training job through the POSIX interface. The data reading pipeline also integrates a chunk-based shuffle mechanism, which randomly shuffles the order of all small files at the start of each epoch of training.
[0051] The general operation module is used to perform general file system operations, such as data update, cache management, and directory structure reconstruction. The general operation module can periodically evict expired small file data sets and provide a conversion from the data block storage view to a user-friendly tree structure view. After the training job starts, the general operation module reads data from the backend aggregated file system 103 with data blocks as the smallest granularity, that is, pre-reads the small file data set, and performs cache management on the pre-read small file data set.
[0052] The backend aggregated file system 103 in the embodiment of the present application can be a high-performance parallel file system formed by aggregating high-speed storage resources scattered on each computing node. For example, the parallel file system can be a BeeOND system.
[0053] The global shared file system 104 in the embodiment of the present application can be any storage system. The global shared file system 104 is used to store a large number of small file data sets.
[0054] In the embodiments of the present application, the metadata management center 101 can be deployed as a container service on the Master node in the cluster, and the full life cycle management capability of the Master node for services can be fully utilized. The client 102 can be directly deployed on the computing nodes of the cluster and interact with the metadata management center through the Restful API. The backend aggregated file system 103 can also be deployed on the computing nodes of the cluster to provide data block storage functions through the client.
[0055] When the training job is scheduled to the target computing node, before the container corresponding to the training job is started, the job management system of the container cloud platform calls the client 102 to aggregate the small file data sets required for the training job from the remote global shared file system 104 into large data blocks and write them into the backend aggregated file system 103. When the container corresponding to the training job is started, the data directory corresponding to the large data block is mounted into the container, and the training program in the container can access the data blocks stored in the backend aggregated file system 103 just like accessing the local directory, and supply data to the training job efficiently and quickly through the high-performance data reading pipeline constructed by the read operation of the client.
[0056] In the embodiments of the present application, the metadata management center 101, the client 102, and the backend aggregated file system 103 together constitute a distributed cache system (CacheFS).
[0057] Figure 1 The data block processing system shown is only for illustration and is not used to limit the technical solutions of the present application. Those skilled in the art should understand that in the specific implementation process, the data block processing system may further include more devices, which are not limited.
[0058] Next, the method will be described in detail in combination with the execution process of the data block processing method.
[0059] Embodiments of the present application provide a data block processing method, which is applied to Figure 1 the metadata management center shown, as Figure 3 shown, Figure 3 FIG. is a schematic flow chart of a data block processing method provided by an embodiment of the present application. The data block processing method includes the following steps:
[0060] S301: Obtain a tree-structured storage table, a leaf node storage table, and a data block storage table.
[0061] Among them, the tree - structured storage table is also called the Leaf table. The Leaf table is used to store the leaf - node information in the tree structure corresponding to the dataset directory of multiple data blocks in the backend aggregated file system. The leaf nodes include data - block files and directories. The Leaf table stores the first - node identifier of the first node, the node identifiers of other nodes (also called inodes), the leaf names of each leaf node, the node identifier of the parent - directory node, and the node type. The inode in the Leaf table is the unique identifier of each data - block file or directory, and the inode of the parent directory is used to represent the hierarchical relationship in the tree structure. The node type is used to indicate whether each leaf node is a data - block file or a directory. Each data - block file and directory node is associated with the parent - directory node through the node identifier.
[0062] The leaf - node storage table is also called the LeafAttributes table. The LeafAttributes table is an extension of the Leaf table and stores the attribute information of each leaf node and each node identifier in the tree structure corresponding to multiple data blocks. The attribute information includes the inode of the data - block file or directory, the inode of the parent directory, the node type, the modification time, the creation time, the size, and the read - write permission.
[0063] It can be understood that the LeafAttributes table is associated with the Leaf table through the inode to ensure that all relevant metadata information of the file or directory can be queried.
[0064] The data - block storage table is also called the Chunk table. The Chunk table is used for the metadata information of multiple data blocks. The Chunk table stores the data - block information of each data block stored in each leaf node among multiple leaf nodes and the node identifier of each leaf node among multiple leaf nodes. Among them, the data - block information includes the data - block identifier (chunkid), the data - block name, and the data - block size. One data block corresponds to the node identifier of one leaf node. Among them, the data - block identifier is used to quickly locate the specific data block.
[0065] It can be understood that the Chunk table is associated with the Leaf table and the LeafAttributes table through the inode.
[0066] The first node is the root node or any one leaf node in the tree structure.
[0067] The multiple data blocks are obtained by aggregating multiple small files required for the training job.
[0068] It can be understood that the metadata snapshot is generated based on the update timestamp of the metadata management center. Whenever the metadata of the dataset changes, the metadata management center will automatically refresh the update timestamp. When it is detected that the update timestamp has not changed within a time window (such as 1 hour), it indicates that the metadata of multiple data blocks has been constructed, that is, the tree structure storage table, the leaf node storage table, and the data block storage table have been constructed, and the metadata management center can be triggered to obtain the tree structure storage table, the leaf node storage table, and the data block storage table.
[0069] S302: Search in the leaf node storage table and the data block storage table respectively to check if there is a node identifier that matches the first node identifier.
[0070] S303: If matching node identifiers are found in both, establish a first-level relationship between the first node, the attribute information of the first node, and the data block information of the first node.
[0071] In one example, based on the initial hierarchy in a preset format, the metadata management center writes the attribute information of the first node and the data block information of the first node to the positions corresponding to the preset first keyword fields in the initial hierarchy, and establishes a first-level relationship.
[0072] Among them, the preset format can be the JSON format.
[0073] The preset first keyword field (also known as the Key field) can be the attr field.
[0074] Exemplarily, based on the initial hierarchy in JSON format, the metadata management center writes the first node identifier, the leaf name, the inode of the parent directory, the node type, the modification time, the creation time, the file size, and the read-write permission to the corresponding positions in the attr field of the initial hierarchy in JSON format, and establishes a first-level relationship.
[0075] It can be understood that since the Chunk table is associated with the Leaf table and the LeafAttributes table through the inode, that is, based on the first node identifier of the first node, the first node can be associated with the attribute information of the first node and the data block information of the first node, and a first-level relationship is established.
[0076] S304: According to the first data block identifier of the data block corresponding to the first node, determine the first small file list information included in the data block corresponding to the first node, and add the first small file list information to the first-level relationship to obtain a second-level relationship.
[0077] In one example, the metadata management center obtains a data block attribute storage table; searches in the data block attribute storage table to check if there is a data block identifier that matches the first data block identifier; if a matching data block identifier is found, obtains the information of the first small file list included in the data block corresponding to the first data block identifier in the data block attribute storage table.
[0078] Among them, the data block attribute storage table is an extension of the content of the Chunk table. The data block attribute storage table includes the data block identifier of each data block and the information of the small file list included in each data block. Among them, the small file list information includes the name, size, and modification time of each small file in each data block. The small file list information is a series of strings. If the number of small files included in a data block is very large, the metadata management center compresses the strings of the small file list information into large object (Blob) binary format data through the zlib compression algorithm.
[0079] It can be understood that the compressed small file list information can reduce the storage space occupation and is also conducive to improving the metadata query efficiency.
[0080] In one example, the metadata management center writes the first small file list information to the position corresponding to the preset second keyword field in the first hierarchical relationship to obtain the second hierarchical relationship.
[0081] Among them, the preset second keyword field can be the chunks field.
[0082] Exemplarily, the metadata management center writes the inode, data block identifier, data block name, data block size, and small file list information of the data block to the corresponding positions in the chunks field to establish the second hierarchical relationship.
[0083] It can be understood that the ChunkAttributes table and the Chunk table are associated through the chunkid, enabling accurate tracking of the metadata information of each small file while managing the data blocks.
[0084] In one example, as Figure 4 shown, Figure 4 is a schematic diagram of the metadata construction for data blocks provided by the embodiments of the present application. In Figure 4 the metadata information of the data block is stored in the form of a tree structure storage table (Table 1), a leaf node storage table (Table 2), a data block storage table (Table 3), and a data block attribute storage table (Table 4). The tree structure storage table, the leaf node storage table, the data block storage table, and the data block attribute storage table correspond to the tree-like directory structure (also known as the tree structure) in the file system under Linux.
[0085] All the metadata stored in the tree - structured storage table, leaf - node storage table, data - block storage table, and data - block attribute storage table is stored in a database, a relational database, which can be deployed on a high - speed storage medium of a single server node or distributedly deployed on high - speed storage media of multiple server nodes.
[0086] It can be understood that deploying the database on high - speed storage media of multiple server nodes can avoid metadata loss caused by single - point failures and ensure high - concurrency performance and scalability of metadata queries. The metadata management center uses a database with table - based storage. By dispersing different types of metadata into multiple storage tables, it improves the query efficiency of metadata and the scalability of the system. Each storage table focuses on storing specific metadata information, reducing unnecessary data access and accelerating query speed. At the same time, the table - based storage structure allows for flexible expansion when the data volume increases, avoiding the complexity of reconstructing the entire database. In addition, table - based storage also enhances maintainability and usability, helping to reduce the risk of errors.
[0087] S305: Generate a metadata snapshot according to the second - level relationship and at least one third - level relationship of other nodes, and send the metadata snapshot to the client so that the client stores it in the local cache and obtains the metadata snapshot when the training job starts.
[0088] Among them, the other nodes are at least one other leaf node in the tree - structured except the first node.
[0089] In one example, the tree - structured directory structure of the metadata snapshot in JSON format is as follows:
[0090] cachefs├── attr (inode: 2, type: "directory", mode: 493,...)└── entries├── imagenet│ ├── attr (inode: 39051, type: "directory",mode: 493,...)│ └── entries│ ├── imagenet-train_100│ │ ├── attr(inode: 39152, type: "regular", mode: 420,...)│ │ └── chunks│ │ ├──chunkid: 49283│ │ └── slices│ │ ├── size: 4054528│ │ └── files│ │ ├── "imagenet / train / n10148035 / n10148035_11304.JPEG"│ │ ├── "imagenet / train / n10148035 / n10148035_11195.JPEG"│ │ ├── "imagenet / train / n10148035 / n10148035_11268.JPEG"
[0091] The tree - like directory structure is used to represent the hierarchical relationship in the metadata snapshot. In this tree - like directory structure, starting from the root node (cachefs), it unfolds to the directory imagenet in sequence, then to the aggregated data block chunk imagenet - train_100, and finally shows the list information of small files in the data block. For example, "imagenet / train / n10148035 / n10148035_11304.JPEG".
[0092] In some alternative embodiments, the metadata management center obtains the number of accesses to the metadata snapshot within the first time period and the latest update time of the metadata snapshot; when the second time period between the latest update time and the current moment is greater than or equal to the first threshold and the number of accesses is less than or equal to the second threshold, the metadata snapshot is cleared.
[0093] Based on the above Figure 3In the method shown, the metadata management center can obtain a tree-structured storage table, a leaf node storage table, and a data block storage table. When node identifiers matching the first node identifier are found in both the leaf node storage table and the data block storage table, a first-level relationship is established between the first node, the attribute information of the first node, and the data block information of the first node. Based on the first data block identifier of the data block corresponding to the first node, the first small file list information included in the data block corresponding to the first node is determined, and the first small file list information is added to the first-level relationship to obtain a second-level relationship. A metadata snapshot is generated according to the second-level relationship and at least one third-level relationship of other nodes, and the metadata snapshot is sent to the client.
[0094] Since the metadata management center can send the metadata snapshot of the training data set (multiple data blocks) to the computing node where the client is located (which is also the computing node where the application is located), the application can directly read the metadata snapshot from the local cache, reducing the access pressure on the metadata management center and accelerating the metadata access efficiency.
[0095] After the training job is started, the client first downloads the metadata snapshot with the latest timestamp from the metadata management center to the local storage. All metadata information query operations after the model training job is started are queried from the metadata snapshot cached locally. Through the metadata snapshot mechanism, high-efficient metadata access performance can be maintained during the training of large-scale small file data sets, while reducing the consumption of storage and computing resources and accelerating the parsing and supply efficiency of the training data set based on data blocks.
[0096] An embodiment of the present application provides a data block processing method, which is applied to Figure 1 the client shown, as Figure 5 shown, Figure 5 is a schematic flowchart of another data block processing method provided by an embodiment of the present application. The data block processing method includes the following steps:
[0097] S501: After the training job is started, load the metadata snapshots corresponding to multiple data blocks from the local cache.
[0098] In one example, after the training job is started, the client loads the metadata snapshots corresponding to multiple data blocks from the local cache, and performs preprocessing operations during the loading process of the metadata snapshots.
[0099] Among them, the preprocessing operations include file integrity check and format verification. It can be understood that the preprocessing operations can ensure that the metadata snapshot is not damaged and conforms to the predefined JSON format.
[0100] S502: Detect whether the snapshot timestamp of the metadata snapshot is consistent with the latest snapshot timestamp of the metadata management center.
[0101] In one example, the metadata snapshot interpreter of the client detects whether the snapshot timestamp of the metadata snapshot is consistent with the latest snapshot timestamp of the metadata management center.
[0102] It can be understood that to ensure the timeliness of the metadata snapshot, the metadata snapshot interpreter is built with a timestamp check mechanism. The metadata snapshot interpreter can compare the snapshot timestamp of the metadata snapshot with the latest snapshot timestamp of the metadata management center.
[0103] If the snapshot time of the metadata snapshot is inconsistent with the latest snapshot timestamp, the client downloads the latest metadata snapshot from the metadata management center and updates the local cache to ensure that the training job obtains multiple small files based on the latest metadata snapshot.
[0104] If the snapshot time of the metadata snapshot is consistent with the latest snapshot timestamp, the client executes S503.
[0105] S503: If so, after the metadata snapshot is loaded, based on the metadata snapshot, read multiple data blocks from the backend aggregated file system.
[0106] S504: Perform a de-aggregation operation on the multiple data blocks to obtain multiple small files.
[0107] In some alternative embodiments, the client parses the metadata snapshot to obtain the small file list information and data block identifiers that make up each data block; based on the small file list information and data block identifiers, read multiple data blocks from the backend aggregated file system.
[0108] In one example, the client parses from the positions corresponding to the preset first keyword fields in the metadata snapshot to obtain the data block identifiers and data block names of each data block among the multiple data blocks, and stores the multiple data block identifiers corresponding to the multiple data blocks in the global data block list; parses from the positions corresponding to the preset second keyword fields in the metadata snapshot to obtain the small file list information included in each data block among the multiple data blocks, and the small file list information includes the data block identifiers of each data block and the small file list included in each data block.
[0109] It can be understood that during the process of the client parsing the metadata snapshot, the metadata snapshot interpreter also has a perfect error handling mechanism. When encountering abnormal situations such as parsing failures or data inconsistencies, the interpreter will trigger the error handling logic (for example, logging or warning prompts), and if necessary, re-obtain the latest metadata snapshot from the metadata management center and parse it again.
[0110] In one example, the client can establish a first hash table between a data block identifier and multiple small file names; establish a second hash table between a small file name and at least one data block, as well as the data block name of each data block in the at least one data block.
[0111] Among them, the first hash table is also called the hash_files_to_chunk table. The second hash table is also called the hash_files_to_chunkid table.
[0112] In some alternative embodiments, the client obtains multiple file indexes; based on the multiple file indexes, generates a list of sub-files belonging to the first data block; obtains the total amount of the multiple first small files; if the total amount is less than the first threshold, creates a data block object for the first data block, and obtains the small file header information and the length of the small file header information of each first small file; according to the small file size and the length of the small file header information of each first small file, obtains the address offset of each first small file in the first data block; according to the address offset, obtains the multiple first small files, and stores the data block identifier of the first data block, the small file name of each first small file, and the address offset in the global hash table corresponding to the data block object.
[0113] Among them, the multiple file indexes are the indexes of each first small file included in the small batch data set required by the training program for executing the training job in the small file list. The small batch data set is the training data of a mini-batch.
[0114] The list of sub-files corresponds to multiple first small files. The small file header information includes the small file size and the small file name of each first small file.
[0115] Exemplarily, as Figure 6 shown, Figure 6 is a schematic diagram of the number of small batches corresponding to the data block read in the embodiment of the present application. In Figure 6 , the training data parser of the client shuffles the data of the data block (chunk) and parses it into the original small file data, forming mini-batch data and feeding it to the model training program.
[0116] In the initialization stage of small file reading, the metadata snapshot interpreter generates a second hash table that maps small files to the data blocks they are in. The key value is the small file name, and the value is the data block identifier and data block name. When the training program needs a mini-batch of training data, the client takes out the file indexes of a new mini-batch from the randomized small file list, maps the multiple file indexes to the corresponding data blocks, generates a list of sub-files belonging to the same data block, then creates a data block object (chunk object) for each hit data block, and when initializing the data block object, traverses all the small file header information within the data block.
[0117] It can be understood that the client records the addressing information (address offset) in the global hash table, where the key is the file name and the value is the object containing the addressing information. In this way, when accessing the file subsequently, the position of the file data can be quickly located without parsing the header information again. The global addressing operation for each data block is performed only once, and the result is stored in the global hash table. Other processes of model training can directly extract the address offset of the small file from the cached chunk object when accessing the data block, thus avoiding repeated addressing operations. This can improve the addressing efficiency and reduce unnecessary I / O operations.
[0118] Optionally, the training data parser of the client supports batch addressing. When creating a chunk object, the client can pass the sub-file list parameter and determine whether the current file name is in the sub-list when traversing the data block header information. If it hits, the addressing information is obtained and an addressing object is generated. When the addressing of all files in the sub-file list is completed, the chunk object is initialized without triggering an actual data reading operation.
[0119] In some alternative embodiments, when the client needs to read the first target small file, it obtains the first target name of the first target small file, reads the first address offset of the first target small file corresponding to the first target name from the global hash table; based on the first address offset, loads the target data block including the first target small file into the local cache.
[0120] In some alternative embodiments, when the client needs to read the second target small file within the target data block, it obtains the second target name of the second target small file, reads the second address offset of the second target small file corresponding to the second target name from the global hash table; based on the first address offset, reads the second target small file from the target data block stored in the local cache.
[0121] It is understandable that when any small file included in a data block is triggered to be read, the client will load all the data of the entire data block into the local cache. If other small files within the same data block are hit again in the current mini-batch or subsequent mini-batches, the corresponding file data will be directly read from the local cache without having to repeatedly read from the backend aggregated file system. This approach is equivalent to introducing a data prefetching mechanism, which improves the reading efficiency of small files and reduces the I / O burden on the system.
[0122] Throughout the training cycle, as long as a data block is loaded into the local cache, subsequent access to files within that data block no longer requires interaction with the backend aggregated file system, unless the cached data block is cleared or invalidated. In subsequent epochs, if the target file is hit in the local cache, the system will directly load the data from the local cache, improving the efficiency of training data supply.
[0123] Based on the above Figure 5 As shown in the method, after the training job starts, the client loads the metadata snapshots corresponding to multiple data blocks from the local cache, and checks whether the snapshot timestamps of the metadata snapshots are consistent with the latest snapshot timestamps of the metadata management center. If so, after the metadata snapshots are loaded, based on the metadata snapshots, the client reads multiple data blocks from the backend aggregated file system and performs a de-aggregation operation on the multiple data blocks to obtain multiple small files.
[0124] Since the dataset is usually read-only during the execution of the training job and there will be no update of the metadata snapshot due to the update of the training dataset, the localized metadata snapshot cache is suitable for deep learning training jobs. The client can quickly obtain the metadata snapshot from the local cache and read the required metadata information from it without having to frequently access the metadata management center, thus reducing the network overhead and improving the efficiency of metadata access, and further enhancing the supply efficiency of the small file dataset.
[0125] Optionally, the client can also aggregate multiple small files stored in the global shared file system into data blocks and store them in the backend aggregated file system. This specifically includes the following stages:
[0126] Stage 1: The client scans the dataset directory stored in the remote global shared file system and collects the metadata information of each data block in the dataset directory.
[0127] It is understandable that the scanning process is carried out in units of subdirectories of the dataset directory to avoid files in different directories being aggregated into the same data block, thereby reducing the system implementation complexity. To improve the scanning efficiency, a multi-worker concurrent processing strategy can be adopted, and multiple scanning workers can be started to work in parallel according to factors such as the scale of the dataset, the number of files, the directory structure, and the preset data block size. Each scanning worker is responsible for scanning the files in a part of the subdirectories and sending the collected file information to the scanning channel.
[0128] Phase 2: A dynamic channel allocation algorithm is adopted to enable each scanning worker to staggeringly send file information to multiple channels, avoiding channel congestion and data skew.
[0129] Specifically, the client maintains a channel list for each scanning worker and records the usage status of each channel. When a scanning worker needs to send file information, it selects an unused channel from the channel list and sends the file information to that channel. After sending, the scanning worker updates the usage status of the channel and moves it to the end of the channel list. In addition, each channel has a buffer queue to handle instantaneous data traffic peaks.
[0130] It is understandable that this dynamic channel allocation algorithm can improve the scanning efficiency and the utilization rate of the scanning channels.
[0131] Phase 3: The aggregation worker receives the file information from the scanning channel and aggregates small files into a data block list according to the directory and the preset data block size.
[0132] It is understandable that the aggregation process only involves the processing of metadata information and does not perform actual data reading and writing operations. Each aggregation worker is responsible for processing the file information in a part of the directories and sending the aggregated data block list to the packaging channel. Similar to the scanning channel, the packaging channel also adopts a dynamic channel allocation algorithm to ensure the load balancing and efficient utilization of the channel.
[0133] Phase 4: The packaging worker receives the data block list from the packaging channel, reads the data of the corresponding small files from the remote global shared file system according to the file information in the list, and generates the corresponding header information (header) based on the data of each small file. Then, the header information and the small file data are packaged into a data block (chunk).
[0134] As Figure 7 shown, Figure 7 is the structural view of the data block provided by the embodiment of the present application; in Figure 7Among them, the data block size is a pre-configured value (such as 4MB). The header information of each small file includes: file name (name), file size (size), file modification time (mtime), and file read permission (mode). The length of the header file is fixed at 512 bytes, and the insufficient part is automatically filled.
[0135] Phase 5: The packaging worker writes the generated data blocks to the new dataset directory of the backend aggregation file system deployed on the computing node through the client. The client utilizes the FUSE file system interface to synchronize the metadata information of the generated data blocks by intercepting write operations, and reports it to the synchronization channel of the metadata management center in a batch processing manner.
[0136] Phase 6: The synchronization worker receives the metadata information of the data blocks from the synchronization channel and writes them to the database of the metadata management center in batches. The aggregated file list of each data block will be compressed before being written to the database to improve efficiency and reduce storage space occupancy.
[0137] It can be understood that the client adopts multi-worker concurrent processing for tasks in each phase to improve processing efficiency. Secondly, a dynamic channel allocation algorithm is adopted to ensure that each worker can evenly and alternately use multiple channels, avoid channel congestion and data skew, and improve channel utilization. Thirdly, a batch processing mechanism is adopted to reduce the number of interactions with the metadata center database and reduce the database load pressure. Finally, a metadata compression technology is adopted to reduce the storage space occupancy of metadata and the data transmission volume, and improve the metadata reading efficiency.
[0138] Optionally, the client can also convert the original small file dataset stored in a tree directory structure to a flat directory structure based on data blocks through small file aggregation operations and store it in the backend aggregation file system. Correspondingly, the client can also perform directory structure reconstruction. The directory structure reconstruction function is to re-parse the data blocks stored in the backend aggregation file system into the same small file dataset directory structure as the original stored in the global shared file system to provide the application with a consistent dataset directory structure view.
[0139] Specifically, the directory structure reconstruction is dynamically generated based on the data block metadata. The client obtains the metadata information from the Leaf table, LeafAttributes table, Chunk table, and ChunkAttributes table in the metadata management center, and dynamically restores and constructs the tree-like directory structure of small files. For example, first, the hierarchical relationship of the directory tree structure is restored through the Leaf table, and the attribute information of each leaf node is obtained from the LeafAttributes table. For directories containing data blocks, the system performs an associated query to find the data block identifiers of all aggregated data blocks under the directory using the inode in the Leaf table and the Chunk table, and finds the list of aggregated small files from the ChunkAttributes table through the data block identifiers. Through the relative path information retained in the data blocks by the small file list, the directory hierarchy is restored layer by layer, and each level of directory is added to the corresponding leaf node of the tree-linked list to construct a tree-like directory structure, and the tree-linked list structure body is returned in accordance with the standard format of the Linux file system to complete the directory structure reconstruction.
[0140] To effectively handle the problem of data inconsistency between the original small file dataset in the global shared file system and the data blocks stored in the backend aggregated file system, the client can perform a complete small file aggregation operation on the dataset used for training before the start of a new training job. This method of full-scale update is simple and reliable, and can ensure that the latest dataset is used for each training job.
[0141] However, if the new dataset only makes some local changes to the old dataset, full-scale update may result in a large number of repeated small file aggregation operations, thus wasting unnecessary time. Therefore, in this case, the client can first detect whether the dataset used for training has been cached in the backend aggregated file system before the start of a new training job. If the dataset with the same name has been cached, the modification time will be compared with the smallest granularity of the directory. If the modification time of a directory has changed, that is, the data in the directory has been modified, the small file aggregation operation will be re-executed for that directory. On the contrary, if the modification time of the directory has not changed, there is no need to re-aggregate the directory. In this way, by only updating the changed part, the full-scale update of the entire dataset is avoided, thus significantly improving the efficiency.
[0142] Furthermore, the client can also perform cache data management. Cache data management includes operations to clean up obsolete data, garbage data, and invalid data to ensure the data in the metadata management center and the backend data block storage always remains valid. First, mark the target small file or small file dataset as "deleted" in the metadata. If all files within a data block are marked as "deleted", then mark that data block as "deleted" as well. When the entire dataset is marked as "deleted", uniformly mark all data blocks under that dataset as "deleted". This enables effective tracking of data blocks that are no longer referenced, providing a basis for subsequent garbage collection operations.
[0143] The garbage data cleaning mechanism relies on a regularly running garbage collection task. The client can periodically scan the metadata to identify data blocks that are not referenced by any file and have been marked as "deleted". And after confirming that these data blocks are no longer needed, mark them as recyclable and delete the corresponding records from the metadata management center. To ensure the consistency of data deletion operations, relevant metadata and data blocks are locked during the deletion process to prevent other operations from interfering with the deletion process. This locking mechanism ensures the safety and reliability of deletion operations in a concurrent environment, avoiding the risks of data inconsistency or accidental loss.
[0144] From the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0145] The embodiments of the present application also provide a data block processing system, which includes a metadata management center and a client.
[0146] The metadata management center is used to obtain a tree structure storage table, a leaf node storage table, and a data block storage table. The tree structure storage table stores the first node identifier of the first node. The leaf node storage table includes the attribute information and each node identifier of each leaf node in the tree structure corresponding to multiple data blocks. The data block storage table includes the data block information, data block identifier, and the node identifier of each leaf node for each leaf node stored in multiple leaf nodes. The first node is the root node or any one leaf node in the tree structure. The multiple data blocks are obtained by aggregating multiple small files required for a training job.
[0147] The metadata management center is also used to separately search in the leaf node storage table and the data block storage table to check whether there is a node identifier that matches the first node identifier; if node identifiers that match are found in both, a first-level relationship is established between the first node, the attribute information of the first node, and the data block information of the first node; according to the first data block identifier of the data block corresponding to the first node, the first small file list information included in the data block corresponding to the first node is determined, and the first small file list information is added to the first-level relationship to obtain a second-level relationship; according to the second-level relationship and at least one third-level relationship of other nodes, a metadata snapshot is generated and sent to the client, so that the client stores it in the local cache and obtains the metadata snapshot when the training job is started. The other nodes are at least one other leaf node in the tree structure except the first node.
[0148] The client is used to, after the training job is started, load the metadata snapshots corresponding to multiple data blocks from the local cache. The metadata snapshots are generated based on the tree structure storage table, the leaf node storage table, and the data block storage table. The tree structure storage table stores the first node identifier of the first node. The leaf node storage table includes the attribute information of each leaf node in the tree structure and each node identifier. The data block storage table includes the data block information, the data block identifier, and the node identifier of each leaf node of each data block stored by each of the multiple leaf nodes. The first node is the root node or any one leaf node in the tree structure. The multiple data blocks are obtained by aggregating multiple small files required for the training job; detect whether the snapshot timestamp of the metadata snapshot is consistent with the latest snapshot timestamp of the metadata management center; if so, after the metadata snapshot is loaded, based on the metadata snapshot, read multiple data blocks from the backend aggregated file system; perform a de-aggregation operation on the multiple data blocks to obtain multiple small files.
[0149] For the description of the features in the embodiments corresponding to the data block processing system, reference can be made to the relevant descriptions in the embodiments corresponding to the data block processing method, which will not be elaborated here one by one.
[0150] Embodiments of the present application also provide a data block processing device, as Figure 8 shown Figure 8Block diagram of a data block processing device provided by an embodiment of the present application; the device includes: a first transceiver module 801, configured to obtain a tree - structured storage table, a leaf node storage table, and a data block storage table. The tree - structured storage table stores a first node identifier of a first node. The leaf node storage table includes attribute information of each leaf node and each node identifier in the tree structure corresponding to multiple data blocks. The data block storage table includes data block information, data block identifiers, and node identifiers of each leaf node for each data block stored in each of the multiple leaf nodes. The first node is a root node or any one of the leaf nodes in the tree structure, and the multiple data blocks are obtained by aggregating multiple small files required for a training job.
[0151] A first processing module 802, configured to respectively check whether there are node identifiers matching the first node identifier in the leaf node storage table and the data block storage table; if node identifiers that match are found in both, establish a first - level relationship between the first node, the attribute information of the first node, and the data block information of the first node; determine, according to the first data block identifier of the data block corresponding to the first node, a first small file list information included in the data block corresponding to the first node, and add the first small file list information to the first - level relationship to obtain a second - level relationship; generate a metadata snapshot according to the second - level relationship and at least one third - level relationship of other nodes, and send the metadata snapshot to the client, so that the client stores it in the local cache and obtains the metadata snapshot when the training job is started. The other nodes are at least one other leaf node in the tree structure except the first node.
[0152] In some optional implementation manners, the first processing module 802 is specifically configured to obtain a data block attribute storage table, where the data block attribute storage table includes data block identifiers of each data block and small file list information included in each data block; check whether there is a data block identifier matching the first data block identifier in the data block attribute storage table; if a matching data block identifier is found, obtain the first small file list information included in the data block corresponding to the first data block identifier in the data block attribute storage table.
[0153] In some optional implementation manners, the first processing module 802 is further specifically configured to write the attribute information of the first node and the data block information of the first node into positions corresponding to preset first key fields in the initial hierarchy based on the initial hierarchy in a preset format, and establish the first - level relationship.
[0154] In some optional implementation manners, the first processing module 802 is further specifically configured to write the first small file list information into positions corresponding to preset second key fields in the first - level relationship to obtain the second - level relationship.
[0155] In some alternative embodiments, the first transceiver module 801 is further configured to obtain the number of accesses to the metadata snapshot within the first time period and the latest update time of the metadata snapshot; the first processing module 802 is further configured to clear the metadata snapshot when the second time period between the latest update time and the current time is greater than or equal to the first threshold and the number of accesses is less than or equal to the second threshold.
[0156] An embodiment of the present application further provides another data block processing device, as Figure 9 shown Figure 9 is a structural block diagram of another data block processing device provided by an embodiment of the present application; the device includes:
[0157] A second transceiver module 901 is configured to, after the training job is started, load metadata snapshots corresponding to a plurality of data blocks from the local cache. The metadata snapshots are generated based on a tree-structured storage table, a leaf node storage table, and a data block storage table. The tree-structured storage table stores a first node identifier of a first node. The leaf node storage table includes attribute information of each leaf node in the tree structure and each node identifier. The data block storage table includes data block information, data block identifiers, and node identifiers of each leaf node of each of the plurality of leaf nodes for each data block stored in each leaf node. The first node is the root node or any one leaf node in the tree structure. The plurality of data blocks are obtained by aggregating a plurality of small files required for the training job.
[0158] A second processing module 902 is configured to detect whether the snapshot timestamp of the metadata snapshot is consistent with the latest snapshot timestamp of the metadata management center; if so, after the metadata snapshot is loaded, read a plurality of data blocks from the backend aggregated file system; perform a de-aggregation operation on the plurality of data blocks to obtain a plurality of small files.
[0159] In some alternative embodiments, the second processing module 902 is specifically configured to parse the metadata snapshot to obtain a small file list information and a data block identifier for each data block; based on the small file list information and the data block identifier, read a plurality of data blocks from the backend aggregated file system.
[0160] In some alternative embodiments, the second processing module 902 is further specifically configured to parse, from a position corresponding to a preset first keyword field in the metadata snapshot, the data block identifier and the data block name of each data block among the plurality of data blocks, and store the data block identifiers corresponding to the plurality of data blocks in a global data block list; parse, from a position corresponding to a preset second keyword field in the metadata snapshot, the small file list information included in each data block among the plurality of data blocks. The small file list information includes the data block identifier of each data block and the small file list included in each data block.
[0161] In some alternative embodiments, the second transceiver module 901 is further configured to obtain a plurality of file indexes, where the plurality of file indexes are indexes of each first small file included in the mini-batch dataset required by the training program for executing the training job in the small file list; the second processing module 902 is further configured to generate a sub-file list belonging to the first data block based on the plurality of file indexes, and the sub-file list corresponds to the plurality of first small files; the second transceiver module 901 is further configured to obtain the total file amount of the plurality of first small files; the second processing module 902 is further configured to, if the total file amount is less than the first threshold, create a data block object of the first data block, and obtain the small file header information and the length of the small file header information of each first small file, where the small file header information includes the small file size and the small file name of each first small file; obtain the address offset of each first small file in the first data block according to the small file size and the length of the small file header information of each first small file; obtain the plurality of first small files according to the address offset, and store the data block identifier of the first data block, the small file name of each first small file, and the address offset in the global hash table corresponding to the data block object.
[0162] In some alternative embodiments, when it is necessary to read the first target small file, the second transceiver module 901 is further configured to obtain the first target name of the first target small file, and read the first address offset of the first target small file corresponding to the first target name from the global hash table; the second processing module 902 is further configured to load the target data block including the first target small file into the local cache based on the first address offset.
[0163] In some alternative embodiments, when it is necessary to read the second target small file in the target data block, the second transceiver module 901 is further configured to obtain the second target name of the second target small file, and read the second address offset of the second target small file corresponding to the second target name from the global hash table; the second processing module 902 is further configured to read the second target small file from the target data block stored in the local cache based on the first address offset.
[0164] For the description of the features in the corresponding embodiments of the data block processing device, reference may be made to the relevant description in the corresponding embodiments of the data block processing method, which will not be elaborated here one by one.
[0165] An embodiment of the present application further provides an electronic device, as Figure 10 shown, Figure 10 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application; the electronic device includes a processor 10 and a memory 20, a computer program is stored in the memory 20, and the processor 10 is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the data block processing method.
[0166] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any of the above-described embodiments of the data block processing method when running.
[0167] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0168] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the data block processing method are implemented.
[0169] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the data block processing method are implemented.
[0170] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0171] The above has introduced in detail a data block processing method, system, device, storage medium, and product provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A data block processing method, characterized in that: Applied to a metadata management center, the method comprises: Obtain a tree structure storage table, a leaf node storage table, and a data block storage table, wherein the tree structure storage table stores a first node identifier of a first node, the leaf node storage table includes attribute information and node identifiers of each leaf node in a tree structure corresponding to a plurality of data blocks, the data block storage table includes data block information, a data block identifier, and a node identifier of each leaf node of the plurality of leaf nodes, the first node is a root node or any leaf node in the tree structure, and the plurality of data blocks are obtained by aggregating a plurality of small files required for a training job; Searching in the leaf node storage table and the data block storage table respectively whether there is a node identifier matching the first node identifier; If matching node identifiers are found, a first hierarchical relationship between the first node and the attribute information of the first node and the data block information of the first node is established; Determine, according to the first data block identifier of the data block corresponding to the first node, first small file list information included in the data block corresponding to the first node, and add the first small file list information to the first hierarchical relationship to obtain a second hierarchical relationship; A metadata snapshot is generated based on the second hierarchical relationship and at least one third hierarchical relationship between other nodes, and the metadata snapshot is sent to the client so that the client stores the metadata snapshot in a local cache and obtains the metadata snapshot when the training job starts. The other nodes are at least one leaf node other than the first node in the tree structure.
2. The method according to claim 1, characterized in that The determining, according to the first data block identifier of the data block corresponding to the first node, the small file list information included in the data block corresponding to the first node comprises: Acquire a data block attribute storage table, wherein the data block attribute storage table includes a data block identifier of each data block and small file list information included in each data block; Searching in the data block attribute storage table whether there is a data block identifier that matches the first data block identifier; If a matching data block identifier is found, the first small file list information included in the data block corresponding to the first data block identifier is obtained in the data block attribute storage table.
3. The method according to claim 2, characterized in that The establishing of a first hierarchical relationship between the first node and the attribute information of the first node and the data block information of the first node includes: Based on an initial hierarchical structure in a preset format, the attribute information of the first node and the data block information of the first node are written into a position corresponding to a preset first key field in the initial hierarchical structure to establish the first hierarchical relationship.
4. The method according to claim 3, characterized in that The adding the first small file list information to the first hierarchical relationship to obtain a second hierarchical relationship includes: The first small file list information is written into a position corresponding to a preset second key field in the first hierarchical relationship to obtain the second hierarchical relationship.
5. The method according to claim 4, characterized in that The method further comprises: Obtaining the number of times the metadata snapshot is accessed within a first time period and the latest update time of the metadata snapshot; When a second time period between the latest update time and the current time is greater than or equal to a first threshold, and the number of accesses is less than or equal to a second threshold, the metadata snapshot is cleared.
6. A data block processing method, characterized in that: Applied to a client, the method comprises: When the training job is started, metadata snapshots corresponding to multiple data blocks are loaded from the local cache, wherein the metadata snapshots are generated based on a tree structure storage table, a leaf node storage table, and a data block storage table, wherein the tree structure storage table stores a first node identifier of a first node, the leaf node storage table includes attribute information and node identifiers of each leaf node in the tree structure, and the data block storage table includes data block information, a data block identifier, and a node identifier of each of the multiple leaf nodes stored in each leaf node of the multiple leaf nodes, wherein the first node is a root node or any leaf node in the tree structure, and the multiple data blocks are obtained by aggregating multiple small files required for the training job; Detecting whether the snapshot timestamp of the metadata snapshot is consistent with the latest snapshot timestamp of the metadata management center; If so, after the metadata snapshot is loaded, the plurality of data blocks are read from the backend aggregate file system based on the metadata snapshot; A de-aggregation operation is performed on the multiple data blocks to obtain the multiple small files.
7. The method according to claim 6, characterized in that After the metadata snapshot is loaded, reading the multiple data blocks from the backend aggregate file system based on the metadata snapshot includes: Parsing the metadata snapshot to obtain small file list information constituting each data block and the data block identifier; Based on the small file list information and the data block identifiers, the multiple data blocks are read from a backend aggregate file system.
8. The method according to claim 7, characterized in that The step of parsing the metadata snapshot to obtain the small file list information and the data block identifier constituting each data block includes: parse the location corresponding to the first key field preset in the metadata snapshot to obtain a data block identifier and a data block name of each of the multiple data blocks, and store the multiple data block identifiers corresponding to the multiple data blocks in a global data block list; The small file list information included in each of the multiple data blocks is obtained by parsing the position corresponding to the preset second key field in the metadata snapshot, and the small file list information includes the data block identifier of each data block and the small file list included in each data block.
9. The method according to claim 8, characterized in that The method further comprises: Acquire multiple file indexes, where the multiple file indexes are indexes in the small file list of each first small file included in the small batch data set required for executing the training program of the training job; Based on the multiple file indexes, generating a sub-file list belonging to the first data block, the sub-file list corresponding to the multiple first small files; Obtain the total number of files of the plurality of first small files; If the total amount of the file is less than the first threshold, a data block object of the first data block is created, and small file header information and small file header information length of each of the first small files are obtained, wherein the small file header information includes a small file size and a small file name of each of the first small files; Obtaining an address offset of each of the first small files in the first data block according to the small file size and the small file header information length of each of the first small files; According to the address offset, the multiple first small files are obtained, and the data block identifier of the first data block, the small file name of each of the first small files and the address offset are stored in a global hash table corresponding to the data block object.
10. The method according to claim 9, characterized in that The method further comprises: When it is necessary to read the first target small file, obtain the first target name of the first target small file, and read the first address offset of the first target small file corresponding to the first target name from the global hash table; Based on the first address offset, a target data block including the first target small file is loaded into the local cache.
11. The method according to claim 10, characterized in that The method further comprises: When it is necessary to read the second target small file in the target data block, obtain the second target name of the second target small file, and read the second address offset of the second target small file corresponding to the second target name from the global hash table; Based on the first address offset, the second target small file is read from the target data block stored in the local cache.
12. A data block processing system, characterized in that: The data block processing system includes a metadata management center and a client; The metadata management center is used to execute the data block processing method described in any one of claims 1 to 5; The client is used to execute the data block processing method described in any one of claims 6-11.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data block processing method according to any one of claims 1 to 11 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the data block processing method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the data block processing method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Stand-alone storage engine
CN113377292A
Distributed file system having separate data and metadata and providing a consistent snapshot thereof
US8818951B1