A cache method, apparatus, device and readable storage medium
Patent Information
- Application Number
- CN202310451188.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-21
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-04-21
AI Technical Summary
[0004]虽然,这类高速缓存技术能够满足BeeGFS的缓存需求,但是,这类高速缓存技术中存储的数据具有易失性,只能作为临时存储,无法应对掉电等情况
[0046]应用本申请实施例所提供的方法,获取删除了管理服务和元数据服务,并添加指向BeeGFS元数据服务的BeeOND源码;其中,BeeGFS为分布式文件系统,BeeOND为临时并行文件系统实例;基于BeeOND源码,运行BeeOND程序,以将BeeOND中的计算节点SSD添加到BeeGFS集群,并挂载客户端;利用计算节点本地SSD,创建BeeGFS集群的默认池;在BeeGFS进行数据交互过程中,在默认池中缓存BeeGFS的数据。
Smart Images

Figure CN116700608B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage technology, and in particular to a caching method, apparatus, device, and readable storage medium. Background Technology
[0002] BeeGFS (a distributed file system) is a leading parallel file system developed with a focus on performance. It is designed for ease of use, simple installation, and management, and has experienced continuous growth and significant adoption within the community. BeeGFS has evolved into a globally valuable file system, offering maximum performance, scalability, flexibility, and robustness.
[0003] In practical applications, BeeGFS, as a global storage system, generally uses low-cost storage media. High-speed caching technologies typically employ server-side memory or client-side buffered (using a small static buffer pool for write-back and read-ahead) and native (using the Linux kernel page cache) technologies.
[0004] Although this type of caching technology can meet the caching requirements of BeeGFS, the data stored in this type of caching technology is volatile and can only be used as temporary storage, and cannot cope with situations such as power outages.
[0005] In conclusion, how to effectively solve the problem of data volatility in BeeGFS's cache is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] The purpose of this application is to provide a caching method, apparatus, device, and readable storage medium that enables a unified namespace for BeeOND and BeeGFS; it implements high-speed caching of BeeGFS through BeeOND, significantly improving BeeGFS performance with less expensive all-flash storage media (BeeOND). BeeOND cached data has locality, reducing network latency during data access for local client reads and writes; BeeGFS has global accessibility, allowing compute nodes to access data across nodes. BeeOND data caching is non-volatile, meaning data remains accessible even after a node is powered on or off.
[0007] To solve the above-mentioned technical problems, this application provides the following technical solution:
[0008] A caching method, comprising:
[0009] Obtain the BeeOND source code that has deleted the management service and metadata service, and added a pointer to the BeeGFS metadata service; wherein, BeeGFS is a distributed file system, and BeeOND is a temporary parallel file system instance;
[0010] Based on the BeeOND source code, run the BeeOND program to add the SSD of the compute node in BeeOND to the BeeGFS cluster and mount the client.
[0011] The default pool of the BeeGFS cluster is created using the local SSD of the compute node;
[0012] During the data interaction process of BeeGFS, the data of BeeGFS is cached in the default pool.
[0013] Preferably, the default pool of the BeeGFS cluster is created using the local SSD of the compute node, including:
[0014] Remove the original default pool from the BeeGFS cluster;
[0015] Using the local SSD of the computing node, create a pool named BeeGFS;
[0016] The pool named BeeGFS is designated as the default pool for the BeeGFS cluster.
[0017] Preferably, it further includes:
[0018] Implement the data placement strategy and place the corresponding data in the BeeOND storage pool and BeeGFS storage pool based on the data attributes.
[0019] Preferably, a data placement strategy is implemented, which places corresponding data in the BeeOND storage pool and the BeeGFS storage pool based on data attributes, including:
[0020] Data is categorized into hot data and cold data based on user identifier, group identity, file name, and at least one data attribute in the directory.
[0021] The hot data is stored in the BeeOND storage pool, and the cold data is stored in the BeeGFS storage pool.
[0022] Preferably, it further includes:
[0023] Receive and parse the cache close request to obtain the target compute node to be closed;
[0024] After executing the cache shutdown command corresponding to the target computing node, during the data interaction process in BeeGFS, the data of the target computing node is directly written to the BeeGFS storage medium.
[0025] Preferably, during the data interaction process of BeeGFS, the data of BeeGFS is cached in the default pool, including:
[0026] When the application reads the target data from the BeeGFS, the target data is copied to the default pool.
[0027] Preferably, it further includes:
[0028] Access timing is performed on the target data;
[0029] If the target data is not accessed for a preset period of time, the target data is deleted from the default pool.
[0030] Preferably, during the data interaction process of BeeGFS, the data of BeeGFS is cached in the default pool, including:
[0031] When an application reads target data from BeeGFS, the target data is migrated from BeeGFS to the default pool.
[0032] Preferably, it further includes:
[0033] Access timing is performed on the target data;
[0034] If the target data is not accessed for a preset period of time, the target data will be migrated back from the default pool to the BeeGFS.
[0035] Preferably, it further includes:
[0036] After a power outage and restart, cached data is read from the default pool.
[0037] A caching device, comprising:
[0038] The code acquisition module is used to acquire the BeeOND source code, which has deleted the management service and metadata service and added a pointer to the BeeGFS metadata service; wherein, BeeGFS is a distributed file system and BeeOND is a temporary parallel file system instance.
[0039] The system fusion module is used to run the BeeOND program based on the BeeOND source code, so as to add the SSD computing node in BeeOND to the BeeGFS cluster and mount the client.
[0040] The default pool creation module is used to create a default pool for the BeeGFS cluster using the local SSD of the compute node.
[0041] A caching module is used to cache the data of BeeGFS in the default pool during the data interaction process of BeeGFS.
[0042] An electronic device, comprising:
[0043] Memory, used to store computer programs;
[0044] A processor, used to implement the above-described caching method when executing the computer program.
[0045] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described caching method.
[0046] Using the method provided in this application's embodiments, BeeOND source code, which has had its management service and metadata service deleted and points to the BeeGFS metadata service added, is obtained; wherein BeeGFS is a distributed file system, and BeeOND is a temporary parallel file system instance; based on the BeeOND source code, the BeeOND program is run to add the compute node SSD in BeeOND to the BeeGFS cluster and mount the client; using the compute node's local SSD, a default pool for the BeeGFS cluster is created; during data interaction in BeeGFS, BeeGFS data is cached in the default pool.
[0047] This application first breaks down the physical data isolation between BeeOND and BeeGFS by sharing BeeGFS metadata services, thus achieving a unified namespace. Then, BeeOND's compute node SSDs are added to the BeeGFS cluster, and a default pool for the BeeGFS cluster is created based on these SSDs. Thus, during data interaction within BeeGFS, BeeGFS data is cached in the default pool. In other words, BeeOND acts as a cache for BeeGFS, providing high-speed caching to BeeGFS while ensuring that the cached data is not lost even in the event of power failure due to BeeOND's non-volatility.
[0048] Accordingly, embodiments of this application also provide caching devices, equipment, and readable storage media corresponding to the above-described caching methods, which have the aforementioned technical effects, and will not be described in detail here. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1This is a flowchart illustrating the implementation of a caching method in an embodiment of this application.
[0051] Figure 2 This is a schematic diagram of the structure of a cache device according to an embodiment of this application;
[0052] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application;
[0053] Figure 4 This is a schematic diagram of the specific structure of an electronic device in an embodiment of this application. Detailed Implementation
[0054] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0055] Please refer to Figure 1 , Figure 1 This is a flowchart of a caching method according to an embodiment of this application. The method includes the following steps:
[0056] S101. Obtain the BeeOND source code that has deleted the management service and metadata service, and added a pointer to the BeeGFS metadata service.
[0057] BeeGFS is a distributed file system, and BeeOND is a temporary parallel file system instance.
[0058] In this embodiment of the application, when deploying BeeGFS, the BeeGFS server can be deployed in accordance with the BeeGFS standard implementation manual, including mgmtd management service, meta metadata service, and storage data service.
[0059] BeeGFS is a leading parallel file system developed with a focus on performance. It is designed for ease of use, simple installation, and management, and has experienced continuous growth and significant adoption within the community. BeeGFS has evolved into a globally valuable file system, offering maximum performance, scalability, flexibility, and robustness. One of BeeGFS's most fundamental concepts is the strict avoidance of architectural bottlenecks by partitioning file content across multiple storage servers. Another key feature is the distribution of file system metadata (e.g., directory information) across multiple metadata servers. Large systems and metadata-intensive applications can greatly benefit from this latter feature. BeeGFS is built on efficient and scalable multi-threaded core components and supports native RDMA (Remote Direct Memory Access). File system nodes can simultaneously provide RDMA (InfiniBand, all-path, RoCE) and TCP / IP network connectivity, automatically switching to a redundant connection path if either fails.
[0060] BeeGFS clients and servers can run on the same machine to improve the performance of small clusters or networks. BeeGFS does not require a dedicated file system partition on the server; it uses existing partitions formatted with any standard Linux file system, such as XFS (X File System, a next-generation file system), ext4 (a journaling file system under Linux, the successor to ext3), or ZFS (Zettabyte File System, a dynamic file system, the first 128-bit file system). For larger networks, several different BeeGFS file system partitions can also be created with different configurations.
[0061] Most HPC (high-performance computing) cluster systems use a global storage system based on a parallel file system on dedicated servers to achieve high throughput. Compute nodes are typically equipped with (or easily equipped with) internal hard drives or SSDs (Solid State Disks), which can provide additional performance advantages. The problem with internal drives in compute nodes is that they offer neither the advantages of a single namespace across multiple computers nor the flexibility and performance of a shared parallel file system. BeeOND was developed to enable the dynamic and easy creation of one or more BeeGFS instances, creating a shared parallel file system on all compute nodes for a specific computation job on a per-job basis, aggregating the performance and capacity of internal SSDs or hard drives on compute nodes, providing additional performance and a very elegant burst buffering mechanism.
[0062] Because BeeOND is very easy to start, it is easy to integrate with workload managers such as Torque (a voice assistant application) or Slurm (a job scheduling system). Since BeeOND can start and stop a new BeeGFS instance with just a single command, it can be easily added to these scripts to start when a compute job begins and stop when the job completes.
[0063] BeeOND is non-volatile, but it is data-isolated from BeeGFS, thus preventing the provision of caching technology for BeeGFS within a unified namespace. In other words, BeeGFS and BeeOND are physically isolated; BeeGFS cannot read or write BeeOND data, and vice versa, and there is no overlap between the two.
[0064] To break down the isolation between BeeOND and BeeGFS, this application proposes a metadata sharing service-based approach to achieve a unified namespace for BeeOND and BeeGFS.
[0065] Specifically, this involves first obtaining the BeeOND source code, which has removed the management service and metadata service, and then adding a pointer to the BeeGFS metadata service.
[0066] That is, you can first modify the BeeOND source code, delete the original mgmtd management service and meta metadata service parts, and then add parameters pointing to the BeeGFS management service node to the BeeOND source code. In this way, BeeOND can correspond to all metadata services.
[0067] S102. Based on the BeeOND source code, run the BeeOND program to add the SSD of the compute node in BeeOND to the BeeGFS cluster and mount the client.
[0068] After obtaining the BeeOND source code with metadata service sharing configured, running the BeeOND program based on this source code allows you to add the BeeOND compute node SSDs to the BeeGFS cluster and mount the client. In other words, you can integrate the BeeOND compute node SSDs into the BeeGFS cluster, making the BeeGFS cluster include the BeeOND compute node SSDs. After mounting the client, the BeeGFS cluster can use the BeeOND compute node SSDs.
[0069] S103. Use the local SSD of the compute node to create the default pool of the BeeGFS cluster.
[0070] To enable BeeOND to serve as a high-speed cache for BeeGFS, the default pool of the BeeGFS cluster can be created using the local SSDs of BeeOND's compute nodes.
[0071] In other words, after completing the deployment of steps S101 and S102, the physical isolation between BeeOND and BeeGFS can be broken (i.e., BeeGFS can access the data stored by BeeOND, and BeeOND can also access the data stored by BeeGFS). Therefore, in this embodiment, BeeOND's all-flash storage medium can be used as a component of BeeGFS's default pool, so that when BeeGFS uses the default pool, it uses BeeOND's all-flash storage medium for caching.
[0072] In one specific embodiment of this application, a default pool for the BeeGFS cluster is created using the local SSD of the compute node, including:
[0073] Step 1: Remove the original default pool from the BeeGFS cluster;
[0074] Step 2: Create a pool named BeeGFS using the local SSD of the compute node;
[0075] Step 3: Determine the pool named BeeGFS as the default pool for the BeeGFS cluster.
[0076] For ease of description, the three steps mentioned above will be explained together below.
[0077] BeeGFS itself has a default pool. To avoid confusion between two default pools, after BeeOND and BeeGFS share metadata services, they are initially in the same storage pool. By adjusting the storage pool, the original BeeGFS storage target is removed from the default pool, and a new pool named BeeGFS is created. After this adjustment, all data is preferentially written to the default pool composed of local SSDs on compute nodes, but will not automatically fall into BeeGFS storage media, nor will write caching be implemented in BeeGFS storage media. In other words, BeeGFS cached data is directly stored in BeeOND.
[0078] S104. During the data interaction process in BeeGFS, cache the data of BeeGFS in the default pool.
[0079] Since the cache acts as a buffer for data exchange, and in step S103, a default pool for BeeGFS has been established based on the local SSDs of the compute nodes, BeeGFS data can be cached in the default pool during data interaction with BeeGFS, i.e., when caching is required. In other words, BeeOND is implemented as a high-speed cache for BeeGFS.
[0080] In the buffer, BeeOND serves as a high-speed cache for BeeGFS. Data stored in BeeOND is temporary. To ensure sufficient space in BeeOND and its continued caching capability, data management is necessary. Specifically, BeeOND space needs to be reclaimed. The embodiments in this application provide the following reclamation methods; any one of these methods can be used in practical applications, or other space reclamation methods (such as periodic space reclamation) can be selected to reclaim storage space:
[0081] Reclamation Method 1: If BeeGFS data is cached in the default pool during data interaction with BeeGFS, specifically by copying the target data to the default pool when the application reads the target data from BeeGFS, then space reclamation can be performed by executing the following steps:
[0082] Step 1: Time the access to the target data;
[0083] Step 2: If the target data has not been accessed for more than the preset time, delete the target data in the default pool.
[0084] Since caching acts as a buffer for data exchange, and data is inherently temporary, configuring tiered storage primarily focuses on file modification attributes. For example, files that haven't been accessed for more than 30 minutes (the timeframe is adjustable and not listed here) are automatically migrated to BeeGFS storage. When an application reads data, a copy is made to BeeOND storage; if the data remains inaccessible for more than 30 minutes, the temporary data is deleted from BeeOND.
[0085] In addition, when data is modified in BeeOND, the latest data can be written to BeeGFS before the data stored in BeeOND is deleted.
[0086] Reclamation Method 2: If BeeGFS data is cached in the default pool during data interaction with BeeGFS, specifically including migrating the target data from BeeGFS to the default pool when the application reads the target data from BeeGFS, then space reclamation can be performed by executing the following steps:
[0087] Step 1: Time the access to the target data;
[0088] Step 2: If the target data has not been accessed for more than the preset time, the target data will be migrated back from the default pool to BeeGFS.
[0089] For example, when an application reads data, the data is migrated to BeeOND storage media. If there is no access for more than 30 minutes, the temporary data stored in BeeOND is migrated to BeeGFS.
[0090] Furthermore, when data is modified in BeeOND, the latest data stored in BeeOND will be used when writing it back to BeeGFS.
[0091] Because BeeOND is non-volatile, cached data is read from the default pool after a power outage and restart. Using non-volatile BeeOND as the cache for BeeGFS ensures that cached data is not lost during power outages, effectively improving data reliability and consistency.
[0092] In other words, the embodiments in this application can achieve a unified namespace for BeeOND and BeeGFS; BeeOND is used to implement a high-speed cache for BeeGFS. BeeGFS performance is significantly improved by using less expensive all-flash storage media (BeeOND). BeeOND cached data has locality, reducing network latency during data access for local client reads and writes; BeeGFS has global accessibility, allowing compute nodes to access it across nodes. BeeOND data cache is non-volatile, meaning data remains accessible even after a node is powered on or off.
[0093] Using the method provided in this application's embodiments, BeeOND source code, which has had its management service and metadata service deleted and points to the BeeGFS metadata service added, is obtained; wherein BeeGFS is a distributed file system, and BeeOND is a temporary parallel file system instance; based on the BeeOND source code, the BeeOND program is run to add the compute node SSD in BeeOND to the BeeGFS cluster and mount the client; using the compute node's local SSD, a default pool for the BeeGFS cluster is created; during data interaction in BeeGFS, BeeGFS data is cached in the default pool.
[0094] This application first breaks down the physical data isolation between BeeOND and BeeGFS by sharing BeeGFS metadata services, thus achieving a unified namespace. Then, BeeOND's compute node SSDs are added to the BeeGFS cluster, and a default pool for the BeeGFS cluster is created based on these SSDs. Thus, during data interaction within BeeGFS, BeeGFS data is cached in the default pool. In other words, BeeOND acts as a cache for BeeGFS, providing high-speed caching to BeeGFS while ensuring that the cached data is not lost even in the event of power failure due to BeeOND's non-volatility.
[0095] It should be noted that, based on the above embodiments, the embodiments of this application also provide corresponding improvement schemes. In the preferred / improved embodiments, the same or corresponding steps as in the above embodiments can be referred to each other, and the corresponding beneficial effects can also be referred to each other; however, these will not be elaborated upon in the preferred / improved embodiments herein.
[0096] In one specific embodiment of this application, considering that BeeOND's storage medium is an all-flash storage medium with a high response speed, in order to effectively improve the overall IO performance of BeeGFS, a data placement strategy can also be executed, which places corresponding data in the BeeOND storage pool and the BeeGFS storage pool based on data attributes.
[0097] Specifically, a data placement strategy is implemented, and based on data attributes, corresponding data is placed in the BeeOND storage pool and the BeeGFS storage pool, including:
[0098] Step 1: Based on user identifier, group identity, file name, and at least one data attribute in the directory, classify the data into hot data and cold data;
[0099] Step 2: Store hot data in BeeOND's storage pool and cold data in BeeGFS's storage pool.
[0100] For ease of description, the two steps above will be explained together below.
[0101] By configuring data placement strategies through hierarchical storage, corresponding data can be selectively placed into the BeeOND storage pool or the BeeGFS storage pool based on user-defined UID (user identifier), GID (Group Identification, referring to the identity of users in the shared resource system), file name, directory, and other attributes.
[0102] In one specific embodiment of this application, caching can be enabled or disabled on a per-component basis, specifically on a per-component basis. The specific implementation process includes:
[0103] Step 1: Receive and parse the cache close request to obtain the target compute node to be closed;
[0104] Step 2: After executing the cache shutdown command corresponding to the target compute node, during the data interaction process in BeeGFS, the data of the target compute node is directly written to the BeeGFS storage medium.
[0105] For ease of description, the two steps above will be explained together below.
[0106] Customers can perform operations on the client according to their needs, thereby causing the client to issue a cache disabling request.
[0107] Upon receiving a cache shutdown request, the target compute node to be shut down can be identified by parsing the request. Then, the cache shutdown command corresponding to the target compute node is executed. This allows data from the target compute node to be directly written to the BeeGFS storage medium during data interaction with BeeGFS. In other words, based on BeeOND's characteristics—namely, its extremely simple startup and easy integration with workload managers such as Torque or Slurm—and because BeeOND requires only a single command to start and stop a new BeeGFS instance, it can be easily added to these scripts to start when a compute job begins and stop when the job completes.
[0108] To facilitate those skilled in the art to better apply the caching method provided in the embodiments of this application, the specific application of the caching method will be described in detail below with reference to specific application scenarios.
[0109] In this embodiment, based on the BeeGFS file system, a metadata service sharing approach is proposed to achieve a unified namespace between BeeGFS and BeeOND. The specific implementation process includes the following steps:
[0110] Step 1: Deploy the BeeGFS server according to the BeeGFS standard implementation manual, including mgmtd management service, meta metadata service, and storage data service.
[0111] Step 2: Modify the BeeOND source code and delete the original mgmtd management service and meta metadata service sections.
[0112] Step 3: Add parameters pointing to the management service node to the BeeOND source code.
[0113] Step 3: Run the BeeOND program to automatically add the compute node SSD to the original BeeGFS cluster and mount the client.
[0114] Based on the tiered storage function, a method is proposed to automatically migrate data to BeeGFS storage space and release cache space in a timely manner. The specific implementation process includes the following steps:
[0115] Step 1: By default, BeeOND and BeeGFS are in the same storage pool. By adjusting the storage pool, the original BeeGFS storage target is removed from the default pool, and a new pool named BeeGFS is created. After the adjustment, all data is preferentially written to the default pool composed of local SSDs on compute nodes, but will not automatically fall into BeeGFS storage media, nor will write caching be implemented in BeeGFS storage media.
[0116] Step 2: Configure the data placement strategy through the hierarchical storage function.
[0117] Specifically, based on user-defined attributes such as UID, GID, filename, and directory, corresponding data can be selectively placed into the BeeOND storage pool or the BeeGFS storage pool.
[0118] Step 3: Configure data migration strategy through tiered storage functionality.
[0119] Since the cache is a buffer for data exchange and the data is temporary, configuring tiered storage mainly focuses on the modification attributes of files. Files that have not been accessed for more than 30 minutes (the time is adjustable) will be automatically migrated to BeeGFS storage media. When an application reads data, a copy of the data is copied to BeeOND storage media. If the data has not been accessed for more than 30 minutes, the temporary data will be deleted.
[0120] Step 4: If a computing node task does not require BeeOND caching, it can be temporarily disabled on that node (this can be done via command), and the data will be directly written to the BeeGFS storage medium.
[0121] In other words, applying the caching method provided in this application's embodiments, the BeeOND converged deployment technology: directs the original BeeOND mgmtd management service deployment to the existing BeeGFS cluster, eliminating the meta metadata service deployment, thereby achieving a unified namespace for BeeOND and BeeGFS; the BeeOND high-speed caching technology: through hierarchical storage functionality, enables data I / O interoperability between BeeOND storage media and BeeGFS storage media; the separate deployment method: BeeOND high-speed caching can be enabled on different compute nodes, and is not globally mandatory.
[0122] The caching method provided in this application embodiment can realize a unified namespace for BeeOND and BeeGFS; implement high-speed caching of BeeGFS through BeeOND; realize data distribution and migration strategies through hierarchical storage; and enable cached data to have locality, persistence and global access characteristics.
[0123] Corresponding to the above method embodiments, this application also provides a caching device, and the caching device described below can be referred to in correspondence with the caching method described above.
[0124] See Figure 2 As shown, the device includes the following modules:
[0125] The code acquisition module 101 is used to acquire the BeeOND source code, which has deleted the management service and metadata service and added a pointer to the BeeGFS metadata service; where BeeGFS is a distributed file system and BeeOND is a temporary parallel file system instance.
[0126] The system fusion module 102 is used to run the BeeOND program based on the BeeOND source code, so as to add the SSD of the compute node in BeeOND to the BeeGFS cluster and mount the client.
[0127] The default pool creation module 103 is used to create a default pool for the BeeGFS cluster using the local SSD of the compute nodes.
[0128] The caching module 104 is used to cache BeeGFS data in the default pool during data interaction with BeeGFS.
[0129] Using the apparatus provided in this application embodiment, BeeOND source code is obtained with the management service and metadata service deleted, and a pointer to the BeeGFS metadata service added; wherein, BeeGFS is a distributed file system, and BeeOND is a temporary parallel file system instance; based on the BeeOND source code, the BeeOND program is run to add the SSD of the compute node in BeeOND to the BeeGFS cluster and mount the client; using the local SSD of the compute node, a default pool of the BeeGFS cluster is created; during the data interaction process of BeeGFS, the data of BeeGFS is cached in the default pool.
[0130] This application first breaks down the physical data isolation between BeeOND and BeeGFS by sharing BeeGFS metadata services, thus achieving a unified namespace. Then, BeeOND's compute node SSDs are added to the BeeGFS cluster, and a default pool for the BeeGFS cluster is created based on these SSDs. Thus, during data interaction within BeeGFS, BeeGFS data is cached in the default pool. In other words, BeeOND acts as a cache for BeeGFS, providing high-speed caching to BeeGFS while ensuring that the cached data is not lost even in the event of power failure due to BeeOND's non-volatility.
[0131] In one specific embodiment of this application, the default pool creation module 103 is specifically used to remove the original default pool of the BeeGFS cluster;
[0132] Create a pool named BeeGFS using the local SSD of the compute node;
[0133] The pool named BeeGFS will be designated as the default pool for the BeeGFS cluster.
[0134] In one specific embodiment of this application, it further includes:
[0135] The data placement module is used to execute data placement strategies and place corresponding data in the BeeOND storage pool and BeeGFS storage pool based on data attributes.
[0136] In one specific embodiment of this application, the data placement module is specifically used to divide data into hot data and cold data based on user identifier, group identity, file name, and at least one data attribute in the directory;
[0137] Hot data is stored in BeeOND's storage pool, and cold data is stored in BeeGFS's storage pool.
[0138] In one specific embodiment of this application, it further includes: a cache control module, used to receive and parse a cache shutdown request to obtain the target computing node to be shut down;
[0139] After executing the cache shutdown command corresponding to the target compute node, during the data interaction process in BeeGFS, the data of the target compute node is directly written to the BeeGFS storage medium.
[0140] In one specific embodiment of this application, the cache control module is specifically used to copy the target data to the default pool when the application reads the target data from BeeGFS.
[0141] In one specific embodiment of this application, it further includes:
[0142] Space reclamation module 1 is used to time access to target data;
[0143] If the target data is not accessed for a preset period of time, the target data will be deleted from the default pool.
[0144] In one specific embodiment of this application, the cache control module is specifically used to migrate the target data from BeeGFS to the default pool when the application reads the target data from BeeGFS.
[0145] In one specific embodiment of this application, it further includes:
[0146] Space reclamation module 1 is used to time access to target data;
[0147] If the target data is not accessed for a preset period of time, the target data will be migrated back from the default pool to BeeGFS.
[0148] In one specific embodiment of this application, it further includes:
[0149] The power-down recovery module is used to read cached data from the default pool after a power outage and restart.
[0150] Corresponding to the above method embodiments, this application also provides an electronic device. The electronic device described below and the caching method described above can be referred to in correspondence.
[0151] See Figure 3 As shown, the electronic device includes:
[0152] Memory 332 is used to store computer programs;
[0153] The processor 322 is used to implement the caching method steps of the above method embodiments when executing a computer program.
[0154] For details, please refer to Figure 4 , Figure 4This is a schematic diagram illustrating the specific structure of an electronic device provided in this embodiment. The electronic device can vary significantly due to differences in configuration or performance. It may include one or more central processing units (CPUs) 322 (e.g., one or more processors) and a memory 332. The memory 332 stores one or more computer programs 342 or data 344. The memory 332 can be temporary or permanent storage. The program stored in the memory 332 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the data processing device. Furthermore, the central processing unit 322 may be configured to communicate with the memory 332 and execute the series of instruction operations stored in the memory 332 on the electronic device 301.
[0155] Electronic device 301 may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341.
[0156] The steps in the caching method described above can be implemented by the structure of an electronic device.
[0157] Corresponding to the above method embodiments, this application also provides a readable storage medium. The readable storage medium described below can be referred to in conjunction with the caching method described above.
[0158] A readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the caching method described in the above method embodiments.
[0159] Specifically, the readable storage medium can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other readable storage medium capable of storing program code.
[0160] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0161] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0162] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0163] Finally, it should be noted that in this document, relationships such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "include," "contain," or any other variations are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.
[0164] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A caching method, characterized in that, include: Obtain the BeeOND source code that has deleted the management service and metadata service, and added a pointer to the BeeGFS metadata service; wherein, BeeGFS is a distributed file system, and BeeOND is a temporary parallel file system instance; Based on the BeeOND source code, run the BeeOND program to add the SSD of the compute node in BeeOND to the BeeGFS cluster and mount the client. The default pool of the BeeGFS cluster is created using the local SSD of the compute node; During the data interaction process of BeeGFS, the data of BeeGFS is cached in the default pool; Specifically, the creation of the default pool for the BeeGFS cluster using the local SSD of the computing node includes: Remove the original default pool from the BeeGFS cluster; Using the local SSD of the computing node, create a pool named BeeGFS; The pool named BeeGFS is designated as the default pool for the BeeGFS cluster.
2. The caching method according to claim 1, characterized in that, Also includes: Implement the data placement strategy and place the corresponding data in the BeeOND storage pool and BeeGFS storage pool based on the data attributes.
3. The caching method according to claim 2, characterized in that, Execute the data placement strategy, based on data attributes, and place the corresponding data in the BeeOND storage pool and the BeeGFS storage pool, including: Data is categorized into hot data and cold data based on user identifier, group identity, file name, and at least one data attribute in the directory. The hot data is stored in the BeeOND storage pool, and the cold data is stored in the BeeGFS storage pool.
4. The caching method according to claim 1, characterized in that, Also includes: Receive and parse the cache close request to obtain the target compute node to be closed; After executing the cache shutdown command corresponding to the target computing node, during the data interaction process in BeeGFS, the data of the target computing node is directly written to the BeeGFS storage medium.
5. The caching method according to claim 1, characterized in that, During the data interaction process of BeeGFS, the data of BeeGFS is cached in the default pool, including: When the application reads the target data from the BeeGFS, the target data is copied to the default pool.
6. The caching method according to claim 5, characterized in that, Also includes: Access timing is performed on the target data; If the target data is not accessed for a preset period of time, the target data is deleted from the default pool.
7. The caching method according to claim 1, characterized in that, During the data interaction process of BeeGFS, the data of BeeGFS is cached in the default pool, including: When an application reads target data from BeeGFS, the target data is migrated from BeeGFS to the default pool.
8. The caching method according to claim 7, characterized in that, Also includes: Access timing is performed on the target data; If the target data is not accessed for a preset period of time, the target data will be migrated back from the default pool to the BeeGFS.
9. The caching method according to any one of claims 1 to 8, characterized in that, Also includes: After a power outage and restart, cached data is read from the default pool.
10. A buffer device, characterized in that, include: The code acquisition module is used to acquire the BeeOND source code, which has deleted the management service and metadata service and added a pointer to the BeeGFS metadata service; wherein, BeeGFS is a distributed file system and BeeOND is a temporary parallel file system instance. The system fusion module is used to run the BeeOND program based on the BeeOND source code, so as to add the SSD computing node in BeeOND to the BeeGFS cluster and mount the client. The default pool creation module is used to create a default pool for the BeeGFS cluster using the local SSD of the compute node. The caching module is used to cache the data of BeeGFS in the default pool during the data interaction process of BeeGFS; Specifically, the default pool creation module is used to remove the original default pool of the BeeGFS cluster; create a pool named BeeGFS using the local SSD of the compute node; and determine the pool named BeeGFS as the default pool of the BeeGFS cluster.
11. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the caching method as described in any one of claims 1 to 9 when executing the computer program.
12. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the caching method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Multi-client cache method and system supporting distributed storage
CN111984191A
Storage system management via a remote console
US11340837B1