Storage management method and related apparatus
By achieving metadata synchronization and data flow between distributed file systems and other file systems, the data silos problem is solved, data access convenience and system evolution flexibility are improved, and costs are reduced.
Patent Information
- Application Number
- PCT/CN2024/141874
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-12-24
- Publication Date
- 2025-07-03
AI Technical Summary
The enclosed data management and metadata management of distributed file systems lead to difficulties in data sharing and flow, making it difficult for users to build a global view, limiting the evolution of the system.
Through storage management methods, metadata synchronization and data flow between distributed file systems and other file systems are realized, data migration and access is used to use storage interfaces, the namespace of distributed file systems remains unchanged, and non-invasive data sharing and flow solutions are provided.
Improves the convenience and flexibility of data access, allows data sharing and migration between distributed file systems and other file systems, reduces the limitations of system evolution and reduces implementation costs.
Smart Images

Figure CN2024141874_03072025_PF_FP_ABST
Abstract
Description
A storage management method and related device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 26, 2023, with application number 202311814043.9 and application name “A Storage Management Method and Related Devices”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of storage technology, and in particular to a storage management method and related devices. Background Art
[0003] As user businesses grow in size, single storage devices are no longer sufficient to meet business needs. Consequently, user business data may need to be stored on storage devices that support distributed storage. Data stored on storage devices needs to be organized and accessed through a file system. To meet this need, distributed file systems have been developed. Distributed file systems provide distributed storage services for data storage, and on top of this distributed storage, they offer unified metadata and data management services. While distributed file systems meet user needs for distributed storage, they also introduce challenges such as difficulty in data sharing and mobility.
[0004] The data and metadata management services of a distributed file system are closed to the outside world. This makes it difficult to integrate data stored in the distributed file system with data stored elsewhere, such as in another file system, to create a global view. This can easily lead to user-deployed distributed file systems becoming data silos. Furthermore, the closed management of distributed file systems complicates data mobility, making it difficult for users to decouple from a particular distributed file system, limiting future evolution options. Summary of the Invention
[0005] The present application provides a storage management method and related devices, which can realize metadata synchronization and data flow between a distributed file system and other file systems, solve the data island problem of the distributed file system and improve the convenience of data access.
[0006] In a first aspect, the present application provides a storage management method, comprising: obtaining metadata of a first file system from a first file system, providing the metadata of the first file system to a second file system, and transferring data between the first file system and the second file system via a storage interface provided by the first file system. The first file system is a distributed file system.
[0007] Optionally, the data being streamed includes data belonging to the first file system, so that a client of the second file system accesses the data belonging to the first file system through the second file system.
[0008] The metadata of a file system describes the information (or attributes) of files in the file system and the associations between files. This application provides the metadata of a first file system to a second file system. The second file system can obtain the information of files of the first file system and the associations between files through the metadata of the first file system, so that a view of the files of the first file system can be established outside the first file system, thereby realizing the sharing of the metadata of the first file system. This application also realizes the flow of data between the first file system and the second file system. On the one hand, the client of the second file system can access the data belonging to the first file system through the second file system, which improves the convenience of data access. Users no longer need to deploy multiple clients to access different file systems. On the other hand, since the first file system belongs to a distributed system, this application enables data to be migrated between the two as needed through the flow of data, so that the distributed file system can have more evolutionary options and improve the flexibility of data layout.
[0009] Furthermore, in this application, the data flow of the first file system is implemented through the storage interface provided by the first file system, which allows the first file system to maintain its original namespace internally, and this application does not cause any intrusive modifications to the first file system. Therefore, this application can achieve non-invasive data flow and metadata sharing with the distributed file system, and the services implemented based on the first file system can continue to execute without interruption during the data flow process, which is conducive to business stability and can significantly reduce implementation costs.
[0010] Optionally, the first file system and the second file system are heterogeneous file systems. In this case, data flow and metadata flow can be achieved between the distributed file system and the heterogeneous file systems, thereby improving the convenience of data access and the fluidity of data layout.
[0011] Alternatively, the first file system and the second file system can be isomorphic file systems. For example, the second file system can also be a distributed file system, or even the first file system and the second file system can be two distributed file systems of the same type. In other words, this solution is also applicable to two independently deployed distributed file systems, enabling federation of multiple distributed file systems, improving the convenience of data access and the fluidity of data layout.
[0012] In one possible implementation of the first aspect, metadata of a first file system is used to generate a view of the first file system in a second file system. Through the file view, devices and services can display the file system's hierarchical structure and file information, facilitating user access and flow, and facilitating data management.
[0013] In some implementations, the second file system can provide a view of the files in the first file system to the client, allowing the client to access the data in the first file system through this view. In other words, the above implementation enables devices outside the first file system (such as clients of the second file system) to also have the ability to access data in the distributed file system based on the file view.
[0014] In another possible implementation of the first aspect, the method further includes: providing metadata of the second file system to the first file system, where the metadata of the second file system is used by the first file system to obtain a global file view, where the global file view includes a file view belonging to the first file system and a file view belonging to the second file system.
[0015] In another possible implementation of the first aspect, the metadata obtained from the first file system and the metadata provided to the second file system are in different formats. That is, the above implementation supports converting the metadata format to achieve metadata flow between the first file system and the other file system.
[0016] Optionally, the metadata content may be the same or different. For example, metadata obtained from the first file system may be format-converted to a metadata format recognizable by the second file system. Optionally, the converted metadata format may add new fields (i.e., file attributes).
[0017] Furthermore, after acquiring the metadata, the second file system may also perform format conversion on the metadata to convert the metadata into a metadata format that can be recognized by the second file system.
[0018] In another possible implementation of the first aspect, the first file system includes a first storage device and a metadata management device. The first storage device is configured to provide data storage services for the first file system, and the metadata management device is configured to manage metadata of the first file system. In this case, the metadata of the first file system is obtained from the metadata management device.
[0019] In another possible implementation of the first aspect, the distributed file system includes at least one of the following file systems: a Lustre file system, a general parallel file system (GPFS), a Hadoop distributed file system (HDFS), a Ceph file system, or a Swift file system.
[0020] Exemplarily, when the first file system is a Lustre file system, the first storage device is an object storage service (OSS) of the Lustre file system, optionally including an object storage target (OST), and the metadata management device is a metadata service (MDS), optionally including a metadata storage target (MDT).
[0021] Optionally, the storage interface may include a metadata acquisition interface, an input / output (IO, or I / O) interface, or a migration interface, etc. The migration interface is, for example, a hierarchical storage management (HSM) interface.
[0022] In another possible implementation of the first aspect, the second file system can be a distributed file system, a network file system (NFS), a file system based on the server message block (SMB) protocol, a common internet file system (CIFS), a simple storage service (S3), or an object storage service (OBS), etc.
[0023] Optionally, if the second file system is also a distributed file system, the second file system and the first file system can be the same distributed file system, or different types of distributed file systems. For example, if both the first file system and the second file system are Lustre file systems, then the first file system and the second file system are isomorphic file systems. For another example, if the first file system is a Lustre file system and the second file system is GPFS.
[0024] In another possible implementation of the first aspect, the second file system is a file system in a global file system (GFS). For example, one or more storage devices (these storage devices may also be referred to as GFS sites) are connected to the GFS, and these storage devices have locally deployed file systems (or source file systems). In this case, the second file system may be a local file system of one or more of the storage devices.
[0025] In another possible implementation of the first aspect, obtaining metadata of the first file system from the first file system includes obtaining the metadata of the first file system through a storage interface provided by the first file system. In this case, the first file system provides a storage interface, and the above implementation obtains the metadata of the first file system through the storage interface. This enables non-invasive acquisition of metadata of a distributed file system.
[0026] In some solutions, the first file system provides a metadata acquisition interface to the client, based on which the client can obtain the metadata of the first file system. In this case, the above embodiment obtains the metadata of the first file system as the client through the storage interface provided by the first file system.
[0027] In another possible implementation of the first aspect, obtaining metadata of the first file system through a storage interface provided by the first file system includes: scanning baseline metadata of the first file system through the storage interface provided by the first file system to obtain metadata of the first file system. Scanning refers to continuously obtaining metadata of multiple files in the first file system according to a file system structure hierarchy or according to certain data.
[0028] In another possible implementation of the first aspect, obtaining metadata of the first file system from the first file system includes obtaining change metadata of the first file system, where the change metadata is metadata added to the first file system within a time period, and the metadata of the first file system includes the change metadata. In this implementation, change metadata of the distributed file system can be obtained to maintain metadata consistency.
[0029] In some solutions, the production and consumption of change metadata follows certain rules, such as through a message system or bus system. In this case, the above implementation supports obtaining change metadata by subscribing (or registering) with the message system or bus system.
[0030] In another possible implementation of the first aspect, data of files in the first file system is flowed between the first file system and the second file system via a storage interface provided by the first file system, including: receiving a first data flow request from the second file system, and accessing the data of files in the first file system via the storage interface in response to the first data flow request. Furthermore, the method further includes: feeding back the result of the access to the second file system, so that the second file system feeds back the result of the access to the client of the second file system. The first data flow request is triggered when the client of the second file system accesses a file belonging to the first file system via the second file system.
[0031] The above implementation enables the second file system client to access data of the first file system, achieves the purpose of data sharing with the distributed file system, and improves user convenience.
[0032] In another possible implementation of the first aspect, the first data flow request is triggered when a client of the second file system reads data belonging to a file in the first file system via the second file system. For ease of distinction, taking the reading of a first file as an example, the data of the first file is stored in the first file system. Accessing the data of a file in the first file system via a storage interface includes: reading the data of the first file from the first file system via the storage interface. Feedback of the access result to the second file system includes: providing the data of the first file to the second file system.
[0033] In the above embodiment, the client of the second file system can read data belonging to the first file system. When the data is stored in the first file system, the data of the file is read from the first file system through the storage interface and then fed back to the second file system.
[0034] In another possible implementation of the first aspect, the first data flow request is triggered when a client of the second file system writes a file to the first file system via the second file system. For ease of distinction below, writing a fifth file is used as an example. Accessing data of a file in the first file system through the storage interface includes: writing data of the fifth file to the first file system through the storage interface.
[0035] In the above implementation, the client of the second file system can write the data of the fifth file to the first file system. The data of the fifth file can be written into the first file system for storage through the solution of the above implementation, thereby realizing the writing of the first file system by an external device.
[0036] Furthermore, feeding back the access result to the second file system includes: receiving a write result fed back by the first file system and feeding back the write result to the second file system. Optionally, the write result includes information indicating whether the write was successful, or the write result includes write information (e.g., a write address, operation time, etc.).
[0037] Alternatively, the writing result includes the updated metadata of the fifth file. In this way, the metadata of the fifth file can be updated after the data of the fifth file is written to achieve metadata consistency.
[0038] In another possible implementation of the first aspect, the method further includes: receiving a first metadata update request from a second file system; and, in response to the first metadata update request, submitting a second metadata update request to the first file system via a storage interface. The first metadata update request is triggered when a client of the second file system writes a file via the second file system. For ease of understanding, the second file is described as an example; the second file belongs to the first file system. The first metadata change request includes updated metadata for the second file, and the second metadata update request includes updated metadata for the second file.
[0039] Furthermore, the method further includes: feeding back a result of the metadata update to the second file system.
[0040] In the above implementation, when writing a second file belonging to the first file system, the file is written directly to the second file system. However, the second file still belongs to the first file system, so the metadata of the second file needs to be updated on the first file system side to maintain metadata consistency. In this way, data belonging to the first file system is written to the second file system, allowing the data of the first file system to be gradually transferred to the second file system, providing more options for the evolution of distributed file systems.
[0041] Optionally, the data of the second file is still stored in the second file system after the metadata of the second file is changed. In some scenarios, the first file system needs to archive newly written data belonging to the first file system locally. The above implementation allows the newly written data belonging to the distributed file to be retained in the second file system, allowing the data of the first file system to be gradually transferred to the second file system, providing more options for the evolution of distributed file systems.
[0042] In another possible implementation of the first aspect, the method further includes: intercepting a migration execution instruction from the first file system, so that the data of the second file is still stored in the second file system after the metadata of the second file is updated. The migration execution instruction is used to instruct to migrate the second file to the first file system.
[0043] In this embodiment, some distributed file systems need to migrate the newly written second file to the local area, for example, through a third migration instruction. The above embodiment processes the third migration instruction so that the data of the second file is still stored in the second file system after the metadata of the second file is updated.
[0044] In another possible implementation of the first aspect, the method is applied to a storage management device. The storage management device is further configured to store metadata of the first file system. The method further includes recording changes to the metadata of the second file in the locally stored metadata of the first file system.
[0045] In the above implementation, the storage management device can maintain the metadata of the first file system, which is conducive to ensuring metadata consistency among the second file system, the first file system and the storage management device.
[0046] In another possible implementation of the first aspect, data of a file in the first file system is flowed between the first file system and a second file system via a storage interface provided by the first file system, including: receiving a second data flow request from the first file system, obtaining data of the second file from the second file system in response to the second data flow request, and providing the data of the second file to the first file system via the storage interface, so that a client of the first file system obtains the data of the second file. The second data flow request is triggered when the client of the first file system accesses the second file.
[0047] In this implementation, the first file system can still retain its own file access policy, that is, the implementation achieves non-invasive compatibility with the first file system. When a client of the first file system needs to read data from a second file stored in the second file system, the implementation supports archiving the second file data to the first file system, thereby achieving data consistency for all parties.
[0048] In another possible implementation of the first aspect, transferring data of a file in the first file system between the first file system and the second file system via a storage interface provided by the first file system includes: sending a first migration instruction to the first file system via the storage interface, receiving a first migration execution instruction from the first file system, and, in response to the first migration execution instruction, obtaining data of a third file from the first file system and providing the data of the third file to the second file system. The first migration instruction is used to instruct the migration of data of the third file stored in and belonging to the first file system. Optionally, the first migration instruction may include an identifier of the second file system to instruct the migration of the data of the third file to the second file system.
[0049] In this implementation, data belonging to the first file system can be migrated to the second file system, enabling data migration between the first file system and other file systems, providing more options for the evolution of the first file system. From a user's perspective, this implementation meets the need to manage other file systems as well as the first file system, improving the user experience.
[0050] Optionally, when the first file system is a Lustre file system, the storage interface includes an HSM interface. This allows, from the perspective of the Lustre file system, data to be archived to external devices through tiered storage. The Lustre file system still considers the connected external storage devices to be part of its own namespace system, thus preventing intrusive modifications from occurring on the Lustre file system side. However, from a global perspective, the HSM interface provided by the Lustre file system enables data sharing and coexistence between two previously independently deployed file systems.
[0051] In another possible implementation of the first aspect, transferring data of a file in the first file system between the first file system and a second file system via a storage interface provided by the first file system includes: sending a second migration instruction to the first file system via the storage interface, receiving a second migration execution instruction from the first file system, and obtaining data of a fourth file from the second file system in response to the second migration execution instruction and providing the fourth file data to the first file system. The second migration instruction is used to instruct the migration of the data of the fourth file stored in the second file system to the first file system.
[0052] In the above embodiment, data stored in the second file system can be migrated to the first file system, realizing data migration between the distributed file system and other file systems, providing more data layout possibilities for the storage system containing multiple file systems, and improving the user experience.
[0053] In another possible implementation of the first aspect, the method further includes receiving a migration policy from a data management device, and generating a migration instruction based on the migration policy to trigger data flow between the first file system and the target file. In this implementation, data migration between the first file system and the other file system is performed based on the migration policy issued by the data management device, enabling policy-controlled data migration, facilitating intelligent data management, and improving the user experience.
[0054] Optionally, the data management device is connected to the second file system, or the data management device is connected to the storage management device, and the above method is implemented by the storage management device.
[0055] Exemplarily, the data management device is a data management engine (DME).
[0056] In a second aspect, the present application further provides a data storage method, applied to a second file system, comprising: receiving metadata of a first file system provided by a storage management device, and flowing data between the first file system and the second file system via the storage management device, wherein the first file system is a distributed file system.
[0057] Optionally, the data being streamed includes data belonging to the first file system, so that the client can access the data belonging to the first file system.
[0058] Optionally, the first file system and the second file system are heterogeneous file systems. Alternatively, the first file system and the second file system are homogeneous file systems.
[0059] In a possible implementation of the second aspect, the method further includes: generating a file view based on metadata of the first file system and providing the file view to the client, wherein the file view includes a view of the file of the first file system. The aforementioned access is initiated through the file view.
[0060] In another possible implementation of the second aspect, the method further includes: generating a global view, the global view including a view of the first file system and a file view of the second file system. Optionally, the global view is generated based on metadata of the first file system and metadata of the second file system.
[0061] In another possible implementation of the second aspect, transferring data belonging to a first file system between the first file system and a second file system by the storage management device includes: receiving an access request from a client of the second file system for a file belonging to the first file system, generating a first data transfer request, and sending the first data transfer request to the storage management device. Furthermore, the second file system receives an access result from the storage management device and feeds back the access result to the client of the second file system.
[0062] In another possible implementation of the second aspect, the access request for a file belonging to the first file system includes a read request for data of a first file belonging to the first file system, the data of the first file is stored in the first file system, and the access result includes the data of the first file.
[0063] In another possible implementation of the second aspect, the access request for the file in the first file system includes a write request for a fifth file in the first file system, and the first data flow request includes data of the fifth file. When the storage management device feeds back an access result, the access result includes a write result for the fifth file.
[0064] In another possible implementation of the second aspect, the method further includes: receiving a write request from a client for a second file belonging to the first file system, writing the data of the second file into the second file system and sending a first metadata update request to the storage management device, the first metadata update request including the updated metadata of the second file.
[0065] The write request includes data of the second file. Furthermore, when the storage management device feeds back a result of metadata update, the method further includes: receiving the result of metadata update.
[0066] In another possible implementation of the second aspect, transferring data belonging to the first file system between the first file system and the second file system by the storage management device includes: receiving data of a third file from the storage management device and storing the data of the third file. Optionally, the third file is migrated from the first file system and becomes part of the first file system.
[0067] In another possible implementation of the second aspect, the data belonging to the first file system flows between the first file system and the second file system through the storage management device, including: providing the data of the fourth file to the storage management device to migrate the data of the fourth file to the first file system.
[0068] In a third aspect, an embodiment of the present application provides a storage system comprising a first file system, a storage management device, and a second file system, wherein the storage management device is configured to implement any of the methods described in the first aspect, and the second file system is configured to implement any of the methods described in the second aspect. The first file system is a distributed file system.
[0069] Optionally, the first file system and the second file system are heterogeneous file systems. Alternatively, the first file system and the second file system are homogeneous file systems.
[0070] In one possible implementation of the third aspect, a first file system includes a first storage device and a metadata management device, wherein the first storage device is configured to provide storage space for files of the first file system, and the metadata management device is configured to manage metadata of the first file system. The second file system includes a second storage device, wherein the second storage device is configured to provide storage space for files of the first file system.
[0071] Optionally, the first storage device and the second storage device are heterogeneous storage devices, for example, they organize data in different ways.
[0072] In another possible implementation of the third aspect, the second file system is configured to: receive an access request from a client of the second file system for a file in the first file system, generate a first data flow request, and send the first data flow request to the storage management device. Furthermore, the second file system is configured to: receive an access result from the storage management device and feed back the access result to the client of the second file system.
[0073] In another possible implementation of the third aspect, the access request for a file belonging to the first file system includes a read request for data of a first file belonging to the first file system, the data of the first file is stored in the first file system, and the access result includes the data of the first file.
[0074] In another possible implementation of the third aspect, the second file system is further configured to: receive a write request from a client of the second file system for a second file belonging to the first file system, write data of the second file into the second file system, and send a first metadata update request to the storage management device, the first metadata update request including updated metadata of the second file. The write request includes the data of the second file.
[0075] In a fourth aspect, the present application provides a storage management device, comprising a metadata management module and a data flow module. The metadata management module is configured to obtain and provide metadata, and the processing module is configured to flow data between a first file system and a second file system. The storage management device is configured to implement the method described in the first aspect or any possible implementation of the first aspect.
[0076] The first file system belongs to a distributed file system. Optionally, the flowed data may include data belonging to the first file system.
[0077] Exemplarily, the metadata management module is further used to maintain metadata of the first file system, such as performing metadata format conversion, establishing file views, updating metadata, submitting metadata changes, etc.
[0078] Exemplarily, the data flow module may include a copy module and / or an LFS adapter module. The copy module is used to migrate data, such as responding to migration instructions. The LFS adapter module is used to initiate IO requests or migration requests during data access to implement data access and / or trigger data migration.
[0079] In one possible implementation of the fourth aspect, the storage management device further includes a GFS access module, configured to connect the storage management device and the first file system to the GFS as a source file system. When the second file system is a file system within the GFS, both metadata flow and data flow between the first file system and the second file system may pass through the GFS access module.
[0080] Optionally, the storage management device is also connected to the data management device to receive a migration (including tiering) policy from the data management device.
[0081] In a fifth aspect, the present application provides a storage device comprising a computing device and a storage disk connected to the computing device, wherein a second file system is deployed on the storage device. The connection may be via a wired line or a wireless line. For example, the two may be connected via a bus. In another example, the two may be connected via a switch.
[0082] The storage device is used to implement the data storage method described in the second aspect or any possible implementation manner of the second aspect.
[0083] In a sixth aspect, the present application provides a computing device comprising a processor and a memory, wherein the memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory, so that the computing device implements the method described in the first aspect or any possible implementation of the first aspect, or implements the method described in the second aspect or any possible implementation of the second aspect.
[0084] In a seventh aspect, the present application provides a computer-readable storage medium, which is used to store instructions or computer programs. When the instructions or computer programs are executed, the method described in the first aspect or any possible implementation method of the first aspect is implemented.
[0085] In an eighth aspect, the present application provides a computer-readable storage medium, which is used to store instructions or computer programs. When the instructions or computer programs are executed, the method described in the second aspect or any possible implementation method of the second aspect is implemented.
[0086] In a ninth aspect, the present application provides a computer program product, which, when the instructions or computer program are executed, implements the method described in the first aspect or any possible implementation method of the first aspect.
[0087] In a tenth aspect, the present application provides a computer program product, which, when the instructions or computer program are executed, implements the method described in the second aspect or any possible implementation method of the second aspect.
[0088] The beneficial effects of the solutions in aspects 2 to 10 of this application can be found in the beneficial effects of the first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] FIG1 is a schematic diagram of the architecture of a storage system provided in an embodiment of the present application;
[0090] FIG2 is a schematic diagram of the architecture of another storage system provided in an embodiment of the present application;
[0091] FIG3 is a schematic diagram of the architecture of another storage system provided in an embodiment of the present application;
[0092] FIG4 is a schematic diagram of the architecture of a GFS provided in an embodiment of the present application;
[0093] FIG5 is a schematic diagram of an overall view provided by an embodiment of the present application;
[0094] FIG6 is a flow chart of a storage management method provided in an embodiment of the present application;
[0095] FIG7 is a flow chart of a method for a data flow process provided by an embodiment of the present application;
[0096] FIG8 is a schematic diagram of an operation scenario of a data flow process provided by an embodiment of the present application;
[0097] FIG9 is a flowchart of another method for data flow process provided in an embodiment of the present application;
[0098] FIG10 is a schematic diagram of an operation scenario of writing data provided by an embodiment of the present application;
[0099] FIG11 is a flowchart of another method for data flow process provided in an embodiment of the present application;
[0100] FIG12 is a schematic diagram of an operation scenario of reading data provided by an embodiment of the present application;
[0101] FIG13 is a flowchart of another method for data flow process provided by an embodiment of the present application;
[0102] FIG14 is a flowchart of another method for data flow process provided by an embodiment of the present application;
[0103] FIG15 is a schematic structural diagram of a storage management device provided in an embodiment of the present application;
[0104] FIG16 is a schematic structural diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0105] The following is an introduction to some of the terms involved in the embodiments of this application.
[0106] 1. File System: A file system is a method (i.e., a method for organizing files on a storage device) and its data structures. The primary function of a file system is to allow users to easily read and write files. For example, a user provides the file system with an identifier (such as a file name or path) for a specific file, and the file system then accesses the data in that file.
[0107] File systems include the following types: distributed file systems, NFS, SMB, CIFS, S3, or OBS. Distributed file systems include Lustre, GPFS, HDFS, Ceph, and Swift. File reading and writing are accomplished through the file system's access protocol, which typically differs depending on the file system type.
[0108] The file system involved in this application is a system with a tree-like hierarchical structure that provides storage and access services for multiple files. In some scenarios, systems with similar features may not necessarily be called file systems, but for the convenience of description, they are uniformly referred to as file systems in this article. For example, when some object systems store objects, multiple objects also have a tree-like hierarchical structure, which also falls within the scope of the "file system" in the embodiments of this application. Optionally, the data content of the files stored in the file system is generally unstructured data, such as documents, pictures, videos, or audio, which have no fixed structure.
[0109] In some embodiments of the present application, a file system is also used to represent a node (or site) on which a file system is deployed, or a node cluster of a file system. For example, a distributed file system is used to represent a node on which the distributed file system is deployed, including a computing node and / or a storage node.
[0110] 2. Files, Data, and Metadata: A file, or computer file, is a collection of information. It consists of both data and metadata. Data refers to the file's content; metadata refers to information describing the file, such as its name, size, and type. For example, a file system might contain a file named "001.png." The file's metadata might describe the file's name (i.e., "001.png"), type, size, location, creator, creation time, or permissions.
[0111] In some implementations, metadata can also describe the relationship between files. For example, the metadata of a file includes the identifier of the directory file to which the file belongs. The file system metadata can be used to obtain the hierarchical structure of files in the file system.
[0112] 3. Heterogeneous file systems: Heterogeneous file systems are characterized by different access (and / or control) methods or different metadata formats. Generally speaking, different file systems are considered heterogeneous, as are file systems provided by different vendors. The opposite of heterogeneous is homogeneous file systems, where homogeneous file systems can achieve a global data access system through unified metadata management and data access control.
[0113] 4. Distributed file system: A file system that allows computers to access multiple hosts and share files over a network. Distributed file systems enable multiple users on multiple computers to share files and storage resources. The physical storage resources managed by a distributed file system are not necessarily directly connected to local compute nodes, but are connected to the compute nodes through a network.
[0114] 5. GFS: Also known as the federated file system, it can unite multiple file systems and provide a joint view of multiple file systems, namely a global file view. Users can load the global file view on any device and access the data of files on any storage device of GFS based on the global file view. GFS provides a unified file system namespace and a global file view for applications across data processing systems (for example, across data centers). Business applications do not need to be aware of the location of the data and can see a unified global file view from any device connected to GFS. When an application reads and writes non-local files, GFS will load the data from the remote end to the local device. It does not distinguish between local and non-local data at the application level, and the response speed depends on the network bandwidth and latency between devices.
[0115] For example, in some scenarios, users deploy NFS at Site A and S3 at Site B. Both sites can be connected to GFS. Devices and services connected to GFS can then view the global file view that combines NFS and S3 at any site, such as Site C, and read data from local and non-local files. From the user's perspective, multiple file systems are a unified file system, making it possible to manage data on multiple storage devices in a unified manner. This greatly simplifies file management on multiple storage devices and enables file sharing, mobility, and fast access across multiple storage devices.
[0116] 6. Migration and tiering: Migration refers to the process of migrating data from one device (source device) to another device (destination device). Tiering, or hierarchical storage management, refers to the process of migrating data from one device (source device) to another device (destination device) and deleting the data on the source device. During the tiering process, the source device and the destination device are usually of different levels (or have different storage capabilities) and / or in different regions. For example, the source device and the destination device have different costs and / or data access speeds. For another example, the source device is located in a data center in location A, and the destination device is located in a data center in location B, that is, the two belong to different data centers. For another example, the source device and the destination device are located in different regions and have different levels. For example, the source device is in a local data center and uses a solid-state drive (SSD) as a storage disk, while the target is in the cloud and uses a hard disk drive (HDD) as a storage disk.
[0117] Generally speaking, migration does not focus on how the source device subsequently processes the data stored thereon, whereas staging requires deleting the data stored on the source device to free up storage space.
[0118] In some scenarios, migration and tiering can be used interchangeably. Since this application primarily describes the flow of data, which occurs during both migration and tiering, for ease of understanding, this application uses the term migration to describe it. In other words, the migration described in this application encompasses tiering and / or migration.
[0119] 7. Ownership: File metadata may include ownership information that specifies the file's ownership device, such as the storage device to which the file belongs. The file's ownership device manages the file's data, including but not limited to maintaining the file's complete and up-to-date data content, publishing data changes when the file's data changes, publishing metadata changes when the file's metadata changes, or distributing data (e.g., returning data to a requesting application).
[0120] In some implementations, ownership can be broadly defined, meaning to which file system a file belongs. As a possible solution, file ownership can be determined using predefined rules. For example, a file written to the root directory of a first file system belongs to the first file system. The above terminology can be applied to the following embodiments.
[0121] Distributed file systems' data and metadata management is closed to other file systems. This makes it difficult for files stored in a distributed file system to form a global view with files in other file systems. This results in user-deployed distributed file systems becoming data silos. With the development of GFS, the trend has become to consolidate data from multiple file systems into a unified global view and to enable data flow between multiple file systems. In this situation, when some user data is stored in a distributed file system, it is difficult for this data to flow between other file systems. This makes it difficult for users to decouple from a particular distributed file system, limiting their options for future evolution.
[0122] In view of this, the embodiments of the present application provide a storage management method and related devices, which can realize metadata flow and data flow between a distributed file system and other file systems, solve the data island problem of the distributed file system and improve the convenience of data access.
[0123] The following describes a possible system architecture applicable to this application. It should be understood that the system architecture provided below is an exemplary system listed to facilitate understanding of the application scenarios of this solution. Those skilled in the art should understand that as new business scenarios and new system architectures emerge, this application is equally applicable to solving similar technical problems.
[0124] Please refer to FIG1 , which is a schematic diagram of the architecture of a storage system provided by an embodiment of the present application. The storage system includes a first file system 10, a storage management device 20, and a second file system 30. In particular:
[0125] The first file system 10 is a distributed file system. In some possible designs, the first file system 10 includes a first storage device 103 and a metadata management device 101, and optionally also includes a management device 102. The first storage device 103 is used to provide storage space for storing data of files in the first file system 10, the metadata management device 101 is used to manage the metadata of the first file system 10, and the management device 102 is used to manage the data of the first file system 10. The first file system 10 also provides an external interface (collectively referred to as a storage interface herein) to enable clients and external devices to interact with the first file system. The number of clients of the first file system 10, referred to as first clients for convenience, can be one or more. FIG1 shows two possible first clients, namely, first client 105a and first client 105b. The external device here refers to a device other than the metadata management device 101, the management device 102, the first storage device 103, and the clients of the first file system 10.
[0126] Exemplarily, the first file system 10 can be a Lustre file system, GPFS, HDFS, Ceph file system, or Swift file system. The Lustre file system is introduced below in conjunction with Figure 2. The Lustre file system shown in Figure 2 (regarded as the first file system 10) includes host applications, services, and persistence. Among them, the services include one or more of the object storage service (OSS) 1032, the metadata service (MDS) 1012, and the management service (MGS) 1022. OSS1032 is used to manage the storage of data and provide data storage services for the first client. OSS1032 is used to provide the first client with services related to the metadata of the file system, including one or more of namespace management, maintaining file metadata, or maintaining data distribution views. MGS1022 provides file system-level configuration management, such as managing server information. Persistence includes one or more of an object storage target (OST) 1031, a metadata target (MDT) 1011, and a management target (MGT) 1021. OST 1031 is a storage device (e.g., a block storage device) corresponding to OSS 1032, MDT 1011 is a storage device corresponding to MDS 1012, and MGT 1021 is a storage device corresponding to MGS 1022. In some cases, in conjunction with FIG1 , OSS 1032 and OST 1031 shown in FIG2 can be considered as first storage device 103, MDT 1011 and MDS 1012 can be considered as metadata management device 101, and MGS 1022 and MGT 1021 can be considered as management device 102.
[0127] The storage management device 20 is a device with storage capacity, computing capacity and communication capacity. The storage management device 20 is connected to the first file system 10 and the second file system 30 respectively, and the connection here refers to a channel with data transmission. For example, the storage management device 20 can obtain data and / or metadata from the first file system 10. For another example, the storage management device 20 can provide data and / or metadata to the second file system 30. There are many possible designs for the deployment of the storage management device 20. In some solutions, the storage management device 20 is independently deployed on a physical device (such as the computing device mentioned below) or a physical device cluster. In other solutions, the storage management device 20 and the second file system 30 are deployed on the same physical device (such as the computing device mentioned below) or a physical device cluster. In some solutions, the storage management device 20 and the first file system 10 are deployed on the same physical device (such as the computing device mentioned below) or a physical device cluster.
[0128] The second file system 30 is another file system different from the first file system 10 .
[0129] As a possible implementation, the second file system 30 is a heterogeneous file system from the first file system 10. In this case, the metadata format of the second file system 30 is different from that of the first file system 10, or the access control method of the second file system 30 is different from that of the first file system 10. For example, the first file system 10 is a Lustre file system, and the second file system 30 is NFS.
[0130] As another possible implementation, the second file system 30 is isomorphic with the first file system 10. For example, the first file system 10 is a Lustre file system, while the second file system 30 is another Lustre file system. In this case, this solution can connect two independently deployed, isomorphic distributed file systems.
[0131] Optionally, the second file system 30 is deployed on a site of GFS. For example, a storage device is connected to GFS, and the local file system of the storage device is an NFS file system. In this case, the NFS file system can be regarded as the second file system.
[0132] As mentioned above, the first file system 10 can provide a storage interface. Through this storage interface, devices or services connected to the first file system can archive data in the first file system to external devices. For example, if the first file system is a Lustre file system, the Lustre file system provides an HSM interface to enable data archiving.
[0133] In an embodiment of the present application, in conjunction with Figure 1, the first file system is a distributed file system and provides a storage interface. The storage management device 20 is connected to the first file system 10 via the storage interface, and the storage management device 20 is also connected to other storage devices, such as the second file system 30. The second file system 30 and the first file system 10 are two independently deployed file systems. Therefore, the storage management device 20 is connected to two different file systems. The storage management device 20 can obtain metadata from the first file system 10 and provide it to the second file system 30. At the same time, the second file system 30 also provides storage space. The storage management device 20 can flow data between the first file system 10 and the second file system 30 through the storage interface of the first file system. For example, data stored on the first storage device 103 of the first file system 10 is migrated to the second file system 30 for storage. In some embodiments, the second file system 30 includes a second storage device, which can be used to provide storage space to store data migrated from the first file system 10. The second storage device includes a storage disk on which the second file system is deployed. For another example, data stored on the second file system 30 is migrated to the first file system 10 for storage.
[0134] As can be seen, the embodiment of the present application achieves the coexistence and data flow between the distributed file system and the second file system 30, making the distributed file system no longer a data island. Since data can flow between the first file system 10 and the second file system 30, on the one hand, devices or services can access data belonging to the first file system through the second file system, improving the convenience of data access. Users no longer need to deploy multiple clients to access different file systems. On the other hand, through data flow, data can be migrated between the two as needed, giving the distributed file system more evolution options and improving the flexibility of data layout.
[0135] In addition, in the present application, the flow of data of the first file system is realized through the storage interface provided by the first file system, which allows the first file system to still maintain the original namespace internally, and the present application does not require any modification or adaptation of the first file system. From the perspective of the first file system 10, the storage management device 20 is considered to be a client of the first file system 10 that supports data tiering, and from a global perspective, the second file system 30 is another file system independent of the first file system 10, and the second file system 30 coexists with the first file system 10 through the conversion of the storage management device 20. Therefore, the present application can realize data flow and metadata sharing with the distributed file system non-invasively, and the business implemented based on the distributed file system can still continue to be executed without interruption during the data flow process, which is conducive to business stability and can significantly reduce the implementation cost of the storage system containing heterogeneous file systems.
[0136] In some possible implementations, in the above-mentioned storage system, clients of the first file system 10 continue to access the first file system 10 using the access protocol of the first file system 10, and the storage system supports clients of the first file system 10 to access all data, including data migrated and stored in the second file system 30.
[0137] The following describes some possible implementations of the above-mentioned storage system in conjunction with Figures 3 and 4. In the following description, the second file system 30 is used as a source file system connected to GFS. However, this application is also applicable to the case where the second file system is another file system deployed independently from the first file system 10.
[0138] In some possible implementations, referring to FIG3 , the storage management device 20 includes a metadata management module 201, which is used to store and maintain metadata of the first file system 10, such as metadata of files belonging to the first file system. Optionally, since the format of the metadata stored in the first file system 10 may be different from the format of the metadata of the second file system, the metadata management module 201 also supports conversion of the metadata format. As a possible solution, the metadata management module 201 may convert the format of the metadata into a specified format (e.g., a first format), which may be pre-configured in the storage management device 20 or set by other devices (e.g., GFS, data management device, etc.). Exemplarily, the metadata obtained by the metadata management module 201 from the first file system 10 may be in a log format or a record format. In this case, the metadata management module may convert the obtained metadata into the first format. Further, the metadata management module 201 stores and maintains the metadata in the first format.
[0139] In one possible implementation, the storage management device 20 can also perform multiple format conversions. For example, each site in the GFS can support metadata transfer in a particular metadata format (e.g., the second format). In this case, the storage management device 20 can convert metadata in the first format into metadata in the second format and provide the metadata to the GFS. Optionally, in this implementation, the metadata format conversion can also be performed by the GFS source 203.
[0140] Optionally, the storage management device 20 stores metadata of the first file system 10. For example, the storage management device 20 may have storage space, such as a local storage disk, or the storage management device 20 may be connected to a device having storage space. The storage space is used by the storage management device 20 to store the metadata of the first file system 10. Furthermore, the storage management device 20 supports updating metadata when the metadata changes, such as adding a new metadata record to the changed metadata.
[0141] In some possible implementations, as shown in FIG2 , the storage interface 104 provided by the first file system 10 may also be referred to as a client interface, and may be an application programming interface (API). Optionally, the storage interface 104 includes one or more interfaces, such as a metadata acquisition interface, an input / output (IO, or I / O) interface, or a migration interface. The following examples illustrate several storage interfaces:
[0142] (1) Metadata query interface: This interface is used to query file metadata. The storage management device 20 uses the metadata query interface to perform a baseline scan on the metadata of the first file system 10, thereby obtaining the baseline metadata of the first file system 10. Scanning refers to continuously obtaining metadata for multiple files in the first file system 10 according to the file system structure hierarchy or specific data.
[0143] (2) Interface for consuming file metadata changes: The storage management device 20 consumes metadata changes of the first file system 10 through the interface for consuming file metadata changes.
[0144] (3) Metadata Change Generation Interface: The storage management device 20 submits metadata changes to the first file system 10 through the metadata change generation interface.
[0145] (4) Input / output (IO, or I / O) interface, which is an interface for accessing the data of files in the first file system 10. The storage management device 20 accesses the data of files in the first file system 10 through the IO interface, and the access here includes reading and / or writing data.
[0146] (5) Data migration interface is used for data migration, such as the HSM interface provided by the Lustre file system.
[0147] Of course, the aforementioned interface functions and names are merely examples. In some implementations, the first file system 10 may provide more interfaces or fewer interfaces (e.g., no interface for metadata changes). In other implementations, the names of the interfaces used to implement corresponding functions may be designed differently. The aforementioned names are merely examples.
[0148] In some possible implementations, in conjunction with FIG. 3 , the storage management device 20 includes a copy module 2021 , which is configured to implement migration of file data between the first file system 10 and the second file system 30 .
[0149] In some possible implementations, the storage management device 20 includes a local file system (LFS) adaptation module 2022. The LFS adaptation module 2022 is used to implement interface conversion, and here is used to adapt the storage interface of the first file system.
[0150] In some possible implementations, when the second file system 30 is GFS, the storage management device 20 further includes a GFS access module, such as the GFS source 203 shown in Figure 3. The GFS access module is used to connect the storage management device 20 and the first file system 10 as a source file system to the GFS.
[0151] In some possible implementations, the second file system 30 can be connected to one or more clients. To facilitate distinction, the clients of the second file system 30 are referred to as second clients. FIG3 shows a possible second client 301. The second client 301 can initiate access to data in the GFS.
[0152] In some possible implementations, metadata from a first file system is used to generate a view of the first file system in a second file system. Through this file view, devices and services can display the file system's hierarchical structure and file information, facilitating user access and mobility, and facilitating file sharing, mobility, and rapid access across multiple devices.
[0153] In some possible implementations, the second file system 30 is further configured to generate a global view, for example, based on metadata of the first file system 10 and metadata of the second file system 30 .
[0154] In some possible implementations, the storage management device is further used to provide metadata of the second file system 30 to the first file system 10. The metadata of the second file system 30 is used by the first file system 10 to obtain a global file view, which includes a file view belonging to the first file system and a file view belonging to the second file system.
[0155] The following is an introduction to an exemplary GFS architecture and an introduction to an exemplary global file view.
[0156] Please refer to Figure 4, which is a schematic diagram of the architecture of a GFS provided in an embodiment of the present application. GFS can be regarded as a storage system that combines a first file system F1 (whose root directory is represented by S1) and a second file system F2 (whose root directory is represented by S2). Among them, the first file system F1 is a Lustre file system (only for example), and the files and their hierarchical structure belonging to the first file system F1 are shown in area 401. The second file system F2 is NFS (only for example), and the files and their hierarchical structure belonging to the second file system F2 are shown in area 402. Among them, the first file system F1 (regarded as the first file system 10) is connected to the GFS through the storage management device 20, and the second file system F2 (regarded as the second file system 30) is deployed on the storage device 40 (that is, the second file system F2 is the local file system of the storage device 40) and is connected to the GFS through the storage device 40. GFS can combine the first file system F1 and the second file system F2, and the global file view obtained by the combination is shown in Figure 5. Accessing the GFS device or service, such as the storage management apparatus 20, storage device 40, and device (or service) 403 shown in Figure 4, enables viewing the global file view shown in Figure 5. In conjunction with Figure 5, the global file view includes a view of files of the first file system F1.
[0157] Furthermore, GFS supports access to the data of the first file system through this view. In other words, this application enables devices outside the first file system (such as devices or services connected to GFS) to also have the ability to access data of the first file system based on the file view. Continuing to refer to Figures 4 and 5, after the device or service loads the global file view, the user can access the data of files on any device in GFS. For example, when the user loads the global file view on device (or service) 403, the data of files on the first file system F1 and / or the second file system F2 can be read, or data can be written to the first file system F1 and / or the second file system F2. Optionally, when a user accesses a file in GFS through a device (or service), if the data read is non-local data, the data needs to be loaded from the remote end to the local end. For example, if the user of storage device 40 needs to read 001.data in directory S1, GFS will request that the data be loaded to storage device 40 (or second file system F2) for the user to read.
[0158] It should be noted here that since the second file system F2 is connected to GFS, all devices or services connected to GFS can be regarded as clients of the second file system F2. For example, device (or service) 403 can also be regarded as a client of the second file system F2.
[0159] In some possible implementations, GFS may also be connected to a data management device, such as the DME shown in FIG3 , for issuing data migration (or grading) strategies to support data flow in GFS according to strategies or needs.
[0160] In some possible implementations, GFS provides a message bus that is used to transmit messages between devices. For example, the message bus can be used to update file system metadata between devices. Alternatively, the message bus can be implemented across a network between devices.
[0161] In some possible implementations, GFS also provides a data flow bus for data to flow between devices according to policy or on demand. Optionally, the data flow bus implements its function through a network across devices.
[0162] The aforementioned message bus and / or data flow bus is a virtualized bus, or called a virtual bus, as shown in FIG4 .
[0163] The introduction to the storage management device 20 here can also be combined with the introduction to the storage management device 20 shown in Figure 15 below.
[0164] The method provided in the embodiments of the present application is described below.
[0165] Please refer to Figure 6, which is a storage management method provided by an embodiment of the present application. Optionally, the method can be applied to a storage system, such as one or more storage systems in Figures 1 to 4 above. The storage management method shown in Figure 6 may include one or more steps from step S601 to step S603. It should be understood that for the convenience of description, the description is given in the order of steps S601 to step S603, and it is not intended to limit the execution to the above order. The embodiment of the present application does not limit the order of execution, execution time, number of executions, etc. of the above one or more steps. Steps S601 to step S603 are as follows:
[0166] Step S601: The storage management device obtains metadata of the first file system from the first file system.
[0167] The storage management device is a device with storage, computing, and communication capabilities, such as the aforementioned storage management device 20. The storage management device is connected to the first file system and the second file system, respectively. The first file system is a distributed file system, such as a Lustre file system, GPFS, HDFS, Ceph file system, or Swift file system. The target file can be a distributed file system, NFS, an SMB protocol-based file system, CIFS, S3, or OBS. The first file system and the second file system can be homogeneous or heterogeneous, as described above.
[0168] Exemplarily, the metadata of the first file system includes metadata for files in the first file system, where files include regular files and directories. Metadata is used to describe the identification and attributes of files in the first file system. For example, metadata includes the file's inode number, size, last modification time, creation time, etc. Of course, the file identification can also be considered a file attribute. In some solutions, metadata is also used to describe the relationship between files. For example, the metadata of a file may include the identification of the directory where the file is located, or what is called the parent node identification.
[0169] In some possible implementations, the storage management device can save the metadata of the first file system. For example, the storage management device has storage capabilities and can provide storage space to save the metadata of the first file system. The storage capabilities here can mean that the storage management device includes a storage disk, or that the storage management device is connected to the storage disk. Furthermore, the storage management device can manage the metadata of the first file system, and the management here includes saving, recording its changes, updating file views, etc. In this case, the metadata of the first file system is maintained in the first file system, and the storage management device also maintains a copy of the metadata of the first file system locally.
[0170] In some possible implementations, the storage management device may obtain metadata through a storage interface provided by the first file system. For example, the first file system may provide a metadata acquisition interface to a client, and the storage management device may serve as a client of the first file system and obtain metadata for the first file system based on the metadata acquisition interface.
[0171] The metadata of the first file system can be stored in a variety of ways, and the storage management device can also obtain the metadata of the first file system in a variety of ways. The following lists two ways of obtaining metadata:
[0172] Method 1: The storage management device scans the baseline metadata of the first file system through the storage interface provided by the first file system to obtain the metadata of the first file system. Scanning refers to a method of continuously obtaining the metadata of several files of the first file system according to the file system structure hierarchy or according to certain data. In the first file system having two metadata storage methods, baseline metadata and change metadata (or incremental metadata), baseline metadata refers to the metadata of the first file system saved by the first file system at a certain moment. In addition to the baseline metadata, the distributed file also maintains metadata updated on the basis of the baseline metadata, namely, change metadata. In a subsequent period of time, the first file system can merge the change metadata into the baseline metadata to form a new version of the baseline metadata. In the first file system that does not set the metadata storage method of incremental metadata, the baseline metadata is all the metadata of the files of the first file system.
[0173] In a second approach, the storage management device obtains change metadata for the first file system. Change metadata refers to metadata added to the first file system within a time period. The metadata of the first file system includes change metadata. In some implementations, the production and consumption of change metadata follows certain rules, such as through a messaging system or bus system. In this case, the above implementation supports obtaining change metadata by subscribing to (or registering with) the messaging system or bus system.
[0174] For example, in a Lustre file system, change metadata can be called a ChangeLog, which is recorded in the MDT. The storage management device can register as a client with permission to consume metadata changes and read the ChangeLog using a cursor, a command line, or a bus system.
[0175] It should be understood that the above two methods can be implemented simultaneously. Furthermore, the second method can be executed multiple times periodically or aperiodically to ensure the consistency of the metadata of the first file system.
[0176] In some possible implementations, the storage management device can convert the format of the metadata. For example, after the storage management device obtains the metadata of the first file system, it converts the metadata format into the metadata format supported by GFS. The metadata after format conversion is stored locally in the storage management device or provided to GFS. Furthermore, new fields can be added in the process again. For example, the metadata after format conversion includes one or more of the following fields: node number (inode), name (name), type (type), mode (mode), snapshot identifier (snapid), user identifier (uid), user group identifier (gid), size (size), change operation (action), transaction identifier (tid), soft link (linkto), creation time (ctime), modification time (mtime), last access time (access time, atime), sequence number (sn), data information (datainfo), standard extended attributes, additional extended attributes, access control list (acl), etc. Exemplarily, fields such as transaction identifier and data information are newly added fields of GFS.
[0177] Optionally, the format conversion of the metadata is performed by a metadata management module in the storage management device, such as the metadata management module 201 shown in FIG. 3 .
[0178] In one possible solution, the second file system is a file system connected to the GFS, and the storage management device is used to use the first file system as a source in the GFS to connect the first file system to the GFS. In other words, the storage system including the first file system, the second file system, and the storage management device can be used as a GFS.
[0179] Step S602: The storage management device provides metadata of the first file system to the second file system.
[0180] The second file system is a file system different from the first file system. For example, the first file system is a Lustre file system, and the second file system is a file system such as GFS, NFS, or S3. For another example, the first file system is a Lustre file system, and the second file system is another Lustre file system.
[0181] It should be understood that the storage management device provides the metadata of the first file system to the second file system, and accordingly, the second file system can receive the metadata of the first file system. Based on the metadata of the first file system, the second file system (including the client accessing the second file system) can generate a view containing the files of the first file system.
[0182] In some solutions, the second file system is a file system connected to GFS. In this case, the device or service connected to the global file system can obtain a global file view, which includes a view of the files belonging to the first file system, as shown in Figure 5.
[0183] Optionally, the storage management device can submit metadata changes to the first file system and / or the second file system when metadata changes. For example, when a client of the second file system updates metadata for a file belonging to the first file system, the storage management device obtains the updated metadata for the file and submits the updated metadata to the first file system to maintain consistency between the metadata and the file view.
[0184] Furthermore, during the metadata change submission process, the storage management device may also convert the metadata changes into a specified format and provide it to the corresponding device. For example, when submitting metadata changes to a first file system, the storage management device may convert the metadata changes into metadata supported by the first file system and provide it to the first file system.
[0185] Step S603: The storage management device flows data between the first file system and the second file system.
[0186] Optionally, the data here includes data belonging to the first file system. Exemplarily, data flow includes one or more of data migration, data reading, or data access.
[0187] Optionally, the flow of data is achieved through a storage interface provided by the first file system. The storage interface here may include one or more of an IO interface, a migration interface, etc., and the relevant description can be found in the above.
[0188] Here are some possible scenarios for data flow:
[0189] Case 1: Data flow caused by data access.
[0190] As a possible implementation, the second file system can initiate a first data flow request to the storage management device to request access to the data stored in the first file system, and the storage management device accesses the data of the files of the first file system through the storage interface. Data flow is generated in this process. As an example of data access, the client of the second file system accesses the files belonging to the first file system via the second file system, and the second file system generates a first data flow request and provides it to the storage management device. Accordingly, the storage management device receives the first data flow request from the second file system, and in response to the first data flow request, accesses the data of the files of the first file system through the storage interface. Furthermore, the storage management device also feeds back the access result to the second file system, so that the second file system feeds back the access result to the client of the second file system.
[0191] As a possible implementation, for write requests, data belonging to the first file system can be stored in the second file system, where this storage is either temporary or persistent. This allows data belonging to the first file system to be written to the second file system, allowing it to be gradually transferred to the second file system, providing more options for the evolution of distributed file systems. Furthermore, for clients of the second file system, writing the data of the second file locally improves write request processing efficiency, achieves lower write latency, and provides a better user experience.
[0192] Furthermore, the storage management device submits (or initiates) metadata changes to the first file system to make the metadata consistent with the file system. For example, when the client of the second file system writes the second file, the data of the second file can be stored in the second file system, but the metadata of the second file, including attributes such as file name, size, write time, and directory, can be provided to the first file system, so that the first file system can obtain the latest metadata and maintain a consistent file view. In this way, the client of the first file system can promptly understand the attributes of the newly written file, improving the user experience.
[0193] In some possible implementations, data written to a first file system may be stored in a second file system, and the data of the file may be provided to the first file system synchronously or asynchronously so that the first file system can obtain the complete data content of the file belonging to it. For example, in some scenarios, the device to which the file belongs needs to maintain the complete data content of the file in order to maintain the global consistency of the file data. Therefore, after the data of the written file is stored in the second file system, the first file system can obtain the data of the file to obtain a complete version of the data content of the file, but this data flow process can be synchronous or asynchronous to balance performance and consistency strength and improve the stability of the storage system.
[0194] Exemplarily, a client of a second file system requests to write a second file, and the second file belongs to a first file system. The second file system stores the data of the second file and sends a first metadata update request to the storage management device to update the metadata of the second file. The first metadata change request includes the updated metadata of the second file. Accordingly, the storage management device receives the first metadata update request from the second file system and, in response to the first metadata update request, submits a second metadata update request to the first file system through a storage interface. The second metadata update request includes the updated metadata of the second file. Furthermore, the storage management device can also provide feedback on the results of the metadata update to the second file system.
[0195] Optionally, in the case where the storage management device locally maintains the metadata of the first file system, the storage management device may further record the change of the metadata of the second file in the locally stored metadata of the first file system.
[0196] In some scenarios, the first file system has a need to archive newly written data belonging to the first file system locally. For example, in a Lustre system, when the data of a second file is stored in an external archiving device, if the metadata of the second file is updated, the Lustre system will migrate the file back to the OST for storage. As a possible implementation, the data of the second file is still stored in the second file system after the metadata of the second file is changed. This solution can be implemented in a variety of ways, such as by modifying the archiving logic of the first file system, or the storage management device can intercept the migration execution instruction from the first file system so that the data of the second file is still stored in the second file system after the metadata of the second file is updated. The migration execution instruction is used to instruct the migration of the second file to the first file system.
[0197] As a possible implementation, the first file system can initiate a second data flow request to the storage management device to request access to the data of a file stored in the second file system (still taking the second file as an example), and the storage management device can obtain the data of the second file from the second file system and provide it to the first file system. Data flow is generated in this process. Accordingly, the storage management device receives the second data flow request from the first file system, and in response to the second data flow request, obtains the data of the second file from the second file system, and provides the data of the second file to the first file system through the storage interface, so that the client of the first file system obtains the data of the second file. The second data flow request is triggered when the client of the first file system accesses the second file. Optionally, the access request is a read request.
[0198] Case 2: Data flow due to data migration. Optionally, data migration can be implemented using a migration instruction, which can be executed by a copy module in the storage management device. In some solutions, data migration is implemented using a migration interface provided by the first file system, such as the HSM interface provided by the Lustre file system.
[0199] As a possible implementation, data can be migrated from the first file system to the second file system. Exemplarily, the storage management device can receive a first migration execution instruction from the first file system, and in response to the first migration execution instruction, obtain data of a third file from the first file system and provide the data of the third file to the second file system.
[0200] Optionally, the migration may be triggered by the storage management device or by the first file system. For example, in the former case, the storage management device may send a first migration instruction to the first file system, the first migration instruction being used to instruct the migration of data of the third file stored in and belonging to the first file system, for example, to the second file system.
[0201] As another possible implementation, data may be migrated from the second file system to the first file system. For example, the storage management device may receive a second migration execution instruction from the first file system, and in response to the second migration execution instruction, obtain data of a fourth file from the second file system and provide the data of the fourth file to the first file system.
[0202] Optionally, the migration is triggered by the storage management device, or by the first file system, or by the second file system. For example, in the former case, the storage management device can send a second migration instruction to the first file system, and the second migration instruction is used to instruct the migration of the data of the fourth file stored in the second file system to the first file system. For another example, the client of the second file system can obtain the global file view as shown in Figure 5. If the authority allows, the client of the second file system can update the ownership device of the fourth file from the second file system to the first file system.
[0203] Alternatively, the fourth file may be a file belonging to the distributed system but temporarily stored in the second file system and therefore needs to be migrated to the first file system for storage. Alternatively, the fourth file may be a file belonging to the second file system but needs to be migrated to the first file system for storage and then belongs to the first file system after the migration. In short, the fourth file ultimately belongs to the first file system.
[0204] In one possible implementation, a storage management device can migrate data between a first file system and a target file system based on a migration policy issued by a data management device. This allows data migration to be controlled by the policy, facilitates intelligent data management, and improves the user experience. For example, the storage management device receives the migration policy from the data management device and generates a migration instruction based on the migration policy to trigger the flow of data between the first file system and the target file system.
[0205] In the embodiment shown in Figure 6, the storage management device provides the metadata of the first file system (distributed file system) to the second file system. The second file system can obtain the information of the files of the first file system and the association between the files through the metadata of the first file system, so that the view of the files of the first file system can be established outside the first file system, thereby realizing the sharing of the metadata of the first file system. The storage management device also realizes the flow of data belonging to the first file system between the first file system and the second file system, realizing the coexistence of the distributed file system and other file systems. On the one hand, the client of the second file system can access the data belonging to the first file system through the second file system, which improves the convenience of data access, and the user no longer needs to deploy multiple clients to access different file systems. On the other hand, through the flow of data, data can be migrated between the two as needed, so that the distributed file system can have more evolution options and improve the flexibility of data layout.
[0206] The embodiment shown in FIG6 above introduces a variety of possible implementations of data flows. Some of the implementations are described in detail below in conjunction with FIG7 to FIG14.
[0207] Please refer to Figure 7, which is a flow chart of a method for a data flow process provided by an embodiment of the present application, including steps S71 to S75. It should be understood that for the convenience of description, the order of steps S601 to S603 is described here, and it is not intended to limit the execution to the above order. The embodiment of the present application does not limit the order of execution of one or more steps, the time of execution, the number of executions, etc. Steps S71 to S75 are as follows:
[0208] Step S71: The second file system receives a request to access a target file.
[0209] The target file is a file belonging to the first file system. The request may come from a client of the second file system, i.e., a second client. The request may include a read request and / or a write request. As shown in FIG4 , the second file system may be a file system F2 deployed on a storage device 40, and the client may be a client accessing the storage device 40.
[0210] Optionally, the request may also include an identifier for the target file, such as an inode number, path, or address. For example, the second client may obtain the global file view shown in Figure 5 and request to read 001.data from directory S2, where directory S2 is the root directory of the first file system and 001.data is a file belonging to the first file system. For another example, the second client may also write a new file, such as the aforementioned fifth file, to directory S2. This newly written fifth file belongs to the first file system.
[0211] Please refer to Figure 8, which is a schematic diagram of a process for accessing data provided by an embodiment of the present application. The second client 301 initiates an access request to the first file system 10 (i.e., the solid line with an arrow ① shown in Figure 8), and the access request is used to access files belonging to the first file system. Taking the read request for reading the first file as an example, the first file is stored in the first file system 10, for example, stored in the first storage device 103. Optionally, the second file system 30 belongs to GFS, and in this case, the client's read request first reaches the protocol layer. For a read request, the protocol layer first accesses the local file system (i.e., the second file system 30) to determine whether the data of the first file exists on the local file system (i.e., the solid line with an arrow ② shown in Figure 8). If not, data flow is required.
[0212] Of course, the access request may also be a write request. For example, the client requests to write the data of the fifth file to the first file system 10. The write request may include the data content of the fifth file to be written.
[0213] Step S72: The second file system sends a first data flow request to the storage management device.
[0214] Accordingly, the storage management device receives a first data flow request from the second file system. Optionally, the first data flow request may include a file identifier, such as an inode number, path, or address. The first data flow request is triggered when a client of the second file system accesses a file in the first file system via the second file system. See the access request process described in step S71.
[0215] Step S73: The storage management device accesses the data of the target file in the first file system through the storage interface.
[0216] Taking the target file as the first file and the access request as a request to read the first file as an example, in conjunction with Figure 8, the storage management device reads the data of the first file through the IO interface (③ as shown in Figure 8), for example, reading data from the first storage device 103.
[0217] Alternatively, taking the target file as the fifth file and the access request as a request to write the fifth file as an example, the storage management device writes the data of the fifth file to the first file system through the IO interface, for example, writes it to the first storage device 103.
[0218] Optionally, the storage management device may initiate an access request and submit it to the first file system via the IO interface. Optionally, the first file system may provide feedback on the access result. For a read request, the access result may include an indication of whether the read was successful and / or the data of the file read, such as the data of the first file described above. For a write request, the access result may indicate whether the write was successful and information about the file written.
[0219] Step S74: The storage management device feeds back the access result to the second file system.
[0220] Optionally, for a read request, the storage management device provides the read target file data to the second file system. Furthermore, the second file system may store the read target file data and feed the target file data back to the second client.
[0221] Optionally, for a write request, such as writing the fifth file, the first file system may provide feedback of a write result to the storage management device. Accordingly, the storage management device may receive the write result provided by the first file system and provide feedback of the write result to the second file system. Optionally, the write result includes information indicating whether the write was successful or write information (e.g., the address written, the operation time, etc.).
[0222] Alternatively, the writing result includes the updated metadata of the fifth file. In this way, the metadata of the fifth file can be updated after the data of the fifth file is written to achieve metadata consistency.
[0223] Step S75: The second file system feeds back the access result to the second client.
[0224] The second client includes a client of the second file system.
[0225] For a read request, the second file system may feed back the read data to the client of the second file system, as shown in step ④ in FIG8 .
[0226] For a write request, the fed-back access result includes one or more of information indicating whether the write is successful, write information, metadata of the updated file, and the like.
[0227] As a possible solution, the content of the access result fed back by the second file system to the client of the second file system may be different from the content of the access result received by the second file system from the storage management device.
[0228] For example, after receiving a user's request to write a file, the second file system first writes the file data locally and then returns the first write result to the client. Later, when the file is provided to the first file system, the second file system can receive a second access result from the storage management device. The first and second write results may contain different information, but the present application also applies if the two contain the same information.
[0229] Optionally, in the case of a write request, step S75 may be performed after step S71 and before step S74.
[0230] Optionally, the steps shown in FIG7 can be performed after metadata change acquisition. The metadata change acquisition process includes: the storage management device reads the latest metadata of the first file system through change consumption and provides the updated distributed file system metadata to the second file system, as shown in FIG8 ⑤.
[0231] The data flow process shown in FIG7 and FIG8 enables the second file system client to access the data of the first file system, thereby achieving the purpose of data sharing and improving the user's convenience.
[0232] 7 above illustrates a schematic process of an access request. In some possible implementations, for a write request, the written data belonging to the second file system is stored in the second file system, and the storage management device submits the metadata changes of the second file to the first file system.
[0233] Please refer to FIG9 , which is a flowchart of another method for data flow provided by an embodiment of the present application, including steps S91 to S96 . Specific steps are as follows:
[0234] Step S91: The second file system receives a write request for a second file.
[0235] Optionally, the write request is initiated by a client of the second file system, such as the second client. Optionally, the write request includes data of the second file.
[0236] Please refer to Figure 10, which is a schematic diagram of an operational scenario for writing data provided by an embodiment of the present application. In this scenario, the second client 301 can initiate a write request (i.e., the solid arrow line ① shown in Figure 10) to the first file system 10. This write request is used to write the second file. The write request carries the data content of the second file.
[0237] Step S92: The second file system writes the data of the second file into the second file system.
[0238] Optionally, the second file system 30 belongs to GFS. In this case, the client's write request first reaches the protocol layer and then reaches the second file system (ie, ② shown in FIG10 ).
[0239] Step S93: The second file system sends a first metadata update request to the storage management device. Correspondingly, the storage management device receives the first metadata update request.
[0240] 10 , in the case where the second file system is GFS or a site in GFS, the first metadata update request may pass through the GFS access module and the metadata management module 201 , as shown in step ③ of FIG10 .
[0241] After a file is written, its metadata is often updated. For example, updating attributes such as the last access time, creation time, and data size will generate new values. The first metadata update request can include the updated metadata of the second file. This metadata can only include the updated values, and fields that have not been updated can be omitted to reduce storage and transmission resource consumption.
[0242] Step S94: The storage management device sends a second metadata update request to the first file system.
[0243] Optionally, the storage management device may submit a second metadata update request through the storage interface 104 .
[0244] The second metadata update request includes updated metadata of the second file.
[0245] Optionally, the storage management device further receives an update result fed back by the first file system. The update result is used to indicate, for example, whether the update is successful or whether the update is successfully received.
[0246] Optionally, the method shown in FIG9 further includes step S95 and / or step S96, which are specifically as follows:
[0247] Step S95: The storage management device feeds back the update result to the second file system. Correspondingly, the second file system receives the update result.
[0248] The update result is used to indicate, for example, whether the update is successful or whether the update is successfully received.
[0249] Step S96: The second file system feeds back an update result to the second client. The update result is used to indicate, for example, whether the update is successful or whether the update is successfully received.
[0250] In the data flow process shown in Figures 9 and 10, newly written files on the second file system can be retained in the second file system, and the file view on the first file system is maintained consistent only by updating the metadata. Furthermore, for the client of the second file system, writing the second file data locally can achieve lower write latency and a better user experience.
[0251] Please refer to FIG11 , which is a flowchart of another method for data flow provided by an embodiment of the present application, including steps S111 to S113. The details are as follows:
[0252] Step S111: The first file system sends a second data flow request to the storage management device.
[0253] 11 and 12, the second data flow request is triggered when the client of the first file system (ie, the first client) accesses the second file. The first client can initiate an access request to the second file, as shown in step ① in FIG12.
[0254] Optionally, the request may be a read request, and the data of the second file is stored in the second file system 30 .
[0255] Step S112: The storage management device obtains the data of the second file from the second file system.
[0256] As a possible implementation, in combination with Figure 12, the first file system can send a migration execution instruction to the copy module 2021 through the metadata management device 101, thereby migrating the data of the second file to the local (ie, the first storage device 103 storage), as shown in ② in Figure 12.
[0257] Optionally, in combination with FIG12 , the storage management device may pull the data of the second file from the second file system through the copy module and the GFS source.
[0258] Step S113: The storage management device provides the data of the second file to the first file system through the storage interface.
[0259] Optionally, the storage management device may provide the data of the second file to the first file system through a storage interface (eg, an IO interface), as shown in step ② of FIG12 .
[0260] Alternatively, the data of the second file may also be provided by the copy module 2021 to the first file system through the storage interface.
[0261] Furthermore, the first file system feeds back the data of the second file to the first client, as shown in step ③ of FIG12 .
[0262] During the data flow process illustrated in Figures 11 and 12 , clients of the first file system can still read data stored on the second file system according to the original read logic. From the perspective of the first file system, the storage management module is a hierarchical client, and the second file system is viewed as an "archiving device" for the first file system, belonging to the first file system. However, the two are actually two different file systems. This embodiment enables the coexistence of the first file system and other file systems without requiring modifications to the first file system, ensuring that the services of the first file system can continue to operate normally, and clients of the first file system can continue to access files in the first file system normally.
[0263] Please refer to Figure 13, which is a schematic diagram of a data flow process provided by an embodiment of the present application. The data flow process includes one or more steps from step S131 to step S133. The details are as follows:
[0264] Step S131: The storage management device sends a first migration instruction to the first file system.
[0265] The first migration instruction is used to instruct to migrate the data of the third file stored in and belonging to the first file system. Optionally, the first migration instruction may include an identifier of the second file system to instruct to migrate the data of the third file to the second file system.
[0266] Optionally, the first migration instruction includes an identifier of the third file, such as an inode number, a path, or an address.
[0267] Optionally, the first migration instruction is generated according to a migration policy, which is, for example, issued by the data management device.
[0268] It should be understood that step S131 is an optional step, that is, in some cases, data migration may be triggered without executing step S131.
[0269] Step S132: The storage management device receives a first migration execution instruction from the first file system.
[0270] Step S133: The storage management device obtains the data of the third file from the first file system and provides the data of the third file to the second file system.
[0271] Exemplarily, step S133 may be executed by a copy module in the storage management device.
[0272] In the embodiment shown in Figure 13, data belonging to the first file system can be migrated to the second file system, enabling data migration between the first file system and other file systems, providing more options for the evolution of the first file system. From the user's perspective, this implementation meets the need for joint management of the first and second file systems, improving the user experience.
[0273] Please refer to Figure 14, which is a schematic diagram of a data flow process provided in an embodiment of the present application, including steps S141 to S144.
[0274] The details are as follows:
[0275] Step S141: The storage management device sends a second migration instruction to the first file system.
[0276] The second migration instruction is used to instruct to migrate data of the fourth file stored in the second file system to the first file system.
[0277] Optionally, the second migration instruction includes an identifier of the fourth file, such as an inode number, a path, or an address.
[0278] Optionally, the second migration instruction is generated according to a migration policy, which is, for example, issued by the data management device.
[0279] It should be understood that step S141 is an optional step.
[0280] Step S142: The storage management device receives a second migration execution instruction from the first file system.
[0281] Step S143: The storage management device obtains the data of the fourth file from the second file system and provides the data of the fourth file to the first file system.
[0282] Exemplarily, step S143 may be executed by a copy module in the storage management device.
[0283] In the embodiment shown in FIG14 , data stored in and belonging to the first file system can be migrated back to the first file system, thereby realizing data migration between the first file system and other file systems, providing more data layout possibilities for a storage system containing multiple file systems, and improving the user experience.
[0284] In order to implement the method of the embodiment of the present application, the device of the embodiment of the present application is provided below.
[0285] It can be understood that the multiple devices provided in the embodiments of the present application, such as data processing devices, are for implementing the functions in the above-mentioned method embodiments, and include hardware structures, software modules, or a combination of hardware structures and software structures corresponding to executing each function.
[0286] Those skilled in the art should easily appreciate that the units and algorithm steps of each example described in the embodiments disclosed herein can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different device implementations to implement the aforementioned method embodiments in different usage scenarios, and different implementations of the device should not be considered to exceed the scope of the embodiments of the present application.
[0287] The embodiments of the present application may divide the device into functional modules. For example, each functional module may be divided according to each function, or two or more functions may be integrated into one functional module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules. It should be noted that the division of modules in the embodiments of the present application is schematic and is only a logical functional division. In actual implementation, other division methods may be used.
[0288] Please refer to Figure 15, which is a schematic diagram of the structure of a storage management device provided in an embodiment of the present application. The storage management device 20 includes a metadata management module 201 and a data flow module 202. The metadata management module 201 is used to implement metadata-related operations, including but not limited to one or more of obtaining metadata, storing metadata, maintaining metadata, providing metadata, updating metadata, submitting metadata changes, metadata format conversion, or establishing data views. The data flow module 202 is used to flow data belonging to the first file system between the first file system and the second file system, including but not limited to one or more of initiating data IO requests, receiving data migration instructions, initiating data migration instructions, and executing data copies.
[0289] The storage management device is used to implement the aforementioned data management method, including the method shown in one or more embodiments in Figures 6 to 14 and any optional implementation method thereof.
[0290] Illustratively, the data flow module 202 may include a copy module 2021 and / or an LFS adapter module 2022. The copy module 2021 is configured to migrate data, such as responding to a migration instruction. The LFS adapter module 2022 is configured to initiate an IO request or migration request during data access to implement data access and / or trigger data migration.
[0291] In one possible implementation, the metadata management module 201 is configured to obtain metadata of the first file system from a first file system and provide the metadata of the first file system to a second file system. The data flow module 202 is configured to flow data belonging to the first file system between the first file system and the second file system via a storage interface provided by the first file system. The first file system is a distributed file system.
[0292] Optionally, the first file system and the second file system are heterogeneous file systems. Alternatively, the first file system and the second file system are homogeneous file systems.
[0293] In one possible implementation, metadata of the first file system is used to generate a view of the first file system in the second file system.
[0294] In one possible implementation, the metadata management module 201 is further configured to convert the metadata of the first file system into a format that is recognizable by the second file system, or to be the same format as the metadata of the second file system.
[0295] In a possible implementation, the metadata management module 201 is further configured to obtain metadata of the first file system through a storage interface provided by the first file system.
[0296] In a possible implementation, the metadata management module 201 is further configured to obtain change metadata of the first file system.
[0297] In one possible implementation, the data flow module 202 is further configured to receive a first data flow request from the second file system and, in response to the first data flow request, access data of a file in the first file system through the storage interface. Furthermore, the data flow module 201 is further configured to feedback the access result to the second file system.
[0298] In a possible implementation, the data flow module 202 is further configured to read the data of the first file from the first file system through the storage interface, and provide the data of the first file to the second file system.
[0299] In one possible implementation, the data flow module 202 is further configured to write the data of the fifth file to the first file system via the storage interface. Further, the data flow module 202 is further configured to receive a write result fed back by the first file system and feed back the write result to the second file system.
[0300] In a possible implementation, the metadata management module 201 is further configured to receive a first metadata update request from the second file system, and submit a second metadata update request to the first file system through the storage interface in response to the first metadata update request.
[0301] Furthermore, the metadata management module 201 is further configured to feed back metadata update results to the second file system.
[0302] In a possible implementation, the data flow module 202 is further configured to intercept migration from the first file system.
[0303] In a possible implementation, the metadata management module 201 is further configured to record changes in the metadata of the second file in the metadata of the first file system stored locally.
[0304] In one possible implementation, the data streaming module 202 is further configured to receive a second data streaming request from the first file system, obtain data of the second file from the second file system in response to the second data streaming request, and provide the data of the second file to the first file system via the storage interface, so that the client of the first file system obtains the data of the second file. The second data streaming request is triggered when the client of the first file system accesses the second file.
[0305] In one possible embodiment, the data flow module 202 is also used to send a first migration indication to the first file system through the storage interface, receive a first migration execution instruction from the first file system, and in response to the first migration execution instruction, obtain data of the third file from the first file system and provide the data of the third file to the second file system.
[0306] In one possible embodiment, the data flow module 202 is further used to send a second migration indication to the first file system through the storage interface, receive a second migration execution instruction from the first file system, obtain data of the fourth file from the second file system in response to the second migration execution instruction, and provide the data of the fourth file to the first file system.
[0307] In a possible implementation, the storage management device further includes a GFS access module 204 (eg, the GFS source 203 shown in FIG. 4 ), which is configured to connect the storage management device and the first file system to the GFS as a source file system.
[0308] When the second file system is GFS or the second file system is a file system in GFS, the aforementioned metadata flow and data flow process with the second file system can pass through the GFS access module 204.
[0309] In one possible implementation, the storage management device 20 is further connected to the data management device to receive a migration (including tiering) policy from the data management device.
[0310] In a possible implementation, the data flow module 202 is further configured to receive a migration strategy from a data management device and generate a migration instruction according to the migration strategy.
[0311] Figure 16 is a structural diagram of a computing device provided in an embodiment of the present application. The computing device 50 is a device with computing capabilities. The device here can be a physical device, such as a controller, a processor, a server (such as a rack server), a host, etc., or it can be a virtual device, such as a virtual machine, a container, etc.
[0312] As shown in Figure 16, computing device 50 includes: processor 502 and memory 501, and optionally includes bus 504 and communication interface 503. Processor 502 and memory 501 communicate with each other via bus 504. It should be understood that this application does not limit the number of processors and memories in computing device 50.
[0313] The memory 501 is used to provide storage space, which can optionally store application data, user data, operating systems, and computer programs. The memory 501 may include volatile memory, such as random access memory (RAM). The memory 501 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0314] The processor 502 is a module for performing calculations and may include any one or more of a controller (e.g., a storage controller), a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), a digital signal processor (DSP), a coprocessor (assisting the central processor in completing corresponding processing and applications), an application specific integrated circuit (ASIC), a microcontroller unit (MCU), a virtual machine, a container, and the like.
[0315] The communication interface 503 is used to provide information input or output for the at least one processor. And / or, the communication interface 503 can be used to receive data sent externally and / or send data to the outside. The communication interface 503 can be a wired link interface such as an Ethernet cable, or a wireless link interface (Wi-Fi, Bluetooth, general wireless transmission and other wireless communication technologies, etc.). Optionally, the communication interface 503 can also include a transmitter (such as a radio frequency transmitter, antenna, etc.) coupled to the interface, or a receiver, etc.
[0316] Bus 504 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, for example. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG16 shows only one line, but this does not imply a single bus or type of bus. Bus 504 may include a path for transmitting information between various components of computing device 50 (e.g., memory 501, processor 502, and communication interface 503).
[0317] In the embodiments of the present application, the memory 501 stores executable instructions, and the processor 502 executes the executable instructions to implement the aforementioned storage management method, such as the method shown in one or more embodiments and any optional implementation manners in Figures 6 to 14 , for example, executing the method executed by the storage management device and / or executing the method executed by the second file system. That is, the memory 501 stores instructions for executing the storage management method.
[0318] An embodiment of the present application further provides a computing device cluster, which includes at least one computing device 50, each computing device 50 including a processor 502 and a memory 501. The processor 502 of the at least one computing device 50 is configured to execute instructions stored in the memory 501 of the at least one computing device 50, so that the computing device cluster implements the aforementioned storage management method, such as the method shown in one or more embodiments and any optional implementation manner thereof in FIG. 6 to FIG. 14 , for example, implementing the method performed by the storage management apparatus and / or executing the method performed by the second file system.
[0319] The present application also provides a storage device, comprising a storage disk and a storage management device as shown in FIG15 or a computing device as shown in FIG16. The storage disk is used to provide storage space, and the storage management device or the computing device is used to implement the aforementioned storage management method, such as the method shown in one or more embodiments of FIG6 to FIG14 and any optional implementation thereof, such as implementing the method performed by the storage management device and / or executing the method performed by the second file system.
[0320] The present application also provides a storage device comprising a storage disk and a computing device. The storage disk is used to provide storage space, and the computing device is used to access GFS. Optionally, a second file system is deployed in the storage device. For example, the storage device may be the storage device 40 shown in FIG4 , and the deployed file system may be, for example, file system F1. Of course, file system F1 may be replaced with another file system, such as S3, OBS, or a distributed file system.
[0321] In some scenarios, the aforementioned storage device may be a storage product provided by a storage manufacturer. For example, the storage device may include Dorado or Pacific storage products provided by Huawei.
[0322] An embodiment of the present application also provides a computer-readable storage medium, which stores instructions. When the instructions are executed on at least one processor (or computing device), the aforementioned storage management method is implemented, such as the method shown in one or more embodiments in Figures 6 to 14 and any optional implementation method thereof, for example, the method executed by the storage management device and / or the method executed by the second file system.
[0323] An embodiment of the present application also provides a computer program product, which includes computer instructions, and the computer instructions are used to implement the aforementioned storage management method, such as the method shown in one or more embodiments in Figures 6 to 14 and any optional implementation method thereof, such as implementing the method executed by the storage management device, and / or executing the method executed by the second file system.
[0324] In the embodiments of this application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0325] The “at least one” mentioned in the embodiments of this application refers to one or more, and “plurality” refers to two or more. “At least one of the following items” or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, (a and b), (a and c), (b and c), or (a and b and c), where a, b, c can be single or multiple. “And / or” describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character “ / ” generally indicates that the previous and next associated objects are in an “or” relationship.
[0326] Furthermore, unless otherwise indicated, ordinal numbers such as "first" and "second" in the embodiments of this application are used to distinguish multiple objects and are not used to limit the order, timing, priority, or importance of multiple objects. For example, the terms "first file" and "second file" are merely used to describe files in a certain implementation and do not indicate differences in the importance of the files, data content, etc. In some cases, the first file and the second file can be the same file.
[0327] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
Claims
1. A storage management method, characterized in that, The method includes: Obtaining metadata of the first file system from the first file system, where the first file system is a distributed file system; Providing the metadata of the first file system to the second file system, and enabling data belonging to the first file system to flow between the first file system and the second file system through a storage interface provided by the first file system, so that a client of the second file system can access data belonging to the first file system through the second file system.
2. The method according to claim 1, wherein The second file system and the first file system are heterogeneous file systems.
3. The method according to claim 1 or 2, characterized in that, The metadata of the first file system is used to generate a view of the first file system in the second file system.
4. The method according to any one of claims 1-3, characterized in that, The first file system includes at least one of the following file systems: Lustre file system, General Parallel File System (GPFS), or Hadoop Distributed File System (HDFS).
5. The method according to any one of claims 1-4, characterized in that, The obtaining the metadata of the first file system from the first file system includes: Obtaining the metadata of the first file system through a storage interface provided by the first file system.
6. The method according to claim 5, wherein The obtaining the metadata of the first file system through a storage interface provided by the first file system includes: Scanning baseline metadata of the first file system through a storage interface provided by the first file system to obtain the metadata of the first file system.
7. The method according to any one of claims 1-5, characterized in that, The obtaining the metadata of the first file system from the first file system includes: Obtaining change metadata of the first file system, where the change metadata is metadata added by the first file system within a period of time, and the metadata of the first file system includes the change metadata.
8. The method according to any one of claims 1 to 7, characterized in that, The enabling data of a file in the first file system to flow between the first file system and the second file system through a storage interface provided by the first file system includes: Receiving a first data flow request from the second file system, where the first data flow request is triggered when a client of the second file system accesses a file belonging to the first file system through the second file system; In response to the first data flow request, accessing data of the file in the first file system through the storage interface and feeding back the result of the access to the second file system, so that the second file system feeds back the result of the access to the client of the second file system.
9. The method according to claim 8, characterized in that, The first data flow request is triggered when a client of the second file system reads data of a first file belonging to the first file system through the second file system, and the data of the first file is stored in the first file system; The accessing data of the file in the first file system through the storage interface and feeding back the result of the access to the second file system includes: Reading data of the first file from the first file system through the storage interface and providing the data of the first file to the second file system.
10. The method according to claim 9, wherein The method further includes: Receive a first metadata update request from the second file system, where the first metadata update request is triggered when a client of the second file system writes to a second file via the second file system, and the second file belongs to the first file system: In response to the first metadata update request, submit a second metadata update request to the first file system via the storage interface, where the second metadata update request includes the metadata of the updated second file; Wherein, the data of the second file is still stored in the second file system after the metadata of the second file is updated.
11. The method according to claim 10, characterized in that, The method is applied to a storage management device, and the storage management device is further used to store the metadata of the first file system. The method further includes: Record the change of the metadata of the second file in the metadata of the first file system stored locally.
12. The method according to claim 10 or 11, characterized in that, The flowing of the data of the files in the first file system between the first file system and the second file system through the storage interface provided by the first file system further includes: Receive a second data flow request from the first file system, where the second data flow request is triggered when a client of the first file system accesses the second file; In response to the second data flow request, obtain the data of the second file from the second file system and provide the data of the second file to the first file system via the storage interface, so that the client of the first file system can obtain the data of the second file.
13. The method according to any one of claims 1 to 12, characterized in that, The flowing of the data of the files in the first file system between the first file system and the second file system through the storage interface provided by the first file system includes: Send a first migration instruction to the first file system via the storage interface, where the first migration instruction is used to indicate migrating the data of a third file stored in and belonging to the first file system; Receive a first migration execution instruction from the first file system; In response to the first migration execution instruction, obtain the data of the third file from the first file system and provide the data of the third file to the second file system.
14. The method according to any one of claims 1-13, characterized in that, The flowing of the data of the files in the first file system between the first file system and the second file system through the storage interface provided by the first file system includes: Send a second migration instruction to the first file system via the storage interface, where the second migration instruction is used to indicate migrating the data of a fourth file stored in the second file system to the first file system; Receive a second migration execution instruction from the first file system; In response to the second migration execution instruction, obtain the data of the fourth file from the second file system and provide the data of the fourth file to the first file system.
15. The method according to any one of claims 1-14, characterized in that, The method further includes: Receive a migration policy from a data management device, where the data management device is connected to the second file system; Generate a migration instruction according to the migration policy to trigger the data flow between the first file system and the target file.
16. A storage system, characterized in that, The storage system includes a second file system, a storage management device, and a first file system. The first file system is a distributed file system. The first file system includes a first storage device and a metadata management device. The first storage device is used to provide storage space for the files of the first file system, and the metadata management device is used to manage the metadata of the first file system. The second file system includes a second storage device, and the second storage device is used to provide storage space for the files of the first file system. The storage management device is used to implement the storage management method according to any one of claims 1-15.
17. The storage system according to claim 16, wherein The second file system is used for: Receiving an access request from a client of the second file system for a file belonging to the first file system. Generating a first data flow request and sending the first data flow request to the storage management device. Receiving an access result from the storage management device and feeding back the access result to the client of the second file system.
18. The storage system according to claim 17, wherein The access request for a file belonging to the first file system includes a read request for the data of a first file belonging to the first file system, and the data of the first file is stored in the first file system. The access result includes the data of the first file.
19. The storage system according to claim 17, wherein, The second file system is further used for: Receiving a write request from a client of the second file system for a second file belonging to the first file system, and the write request includes the data of the second file. Writing the data of the second file into the second file system. Sending a first metadata update request to the storage management device, and the first metadata update request includes the updated metadata of the second file.
20. A storage management device, the storage management device includes a metadata management module and a data flow module. The metadata management module is used to obtain and provide metadata. The data flow module is used to flow data between the first file system and the second file system. The storage management device is used to implement the method according to any one of claims 1-15.
21. A computing device, characterized in that, The computing device includes a processor and a memory. Computer instructions are stored in the memory, and the processor is used to call the computer instructions stored in the memory to implement the method according to any one of claims 1-15.
22. A computer-readable storage medium, characterized in that, Computer program instructions are stored in the computer-readable storage medium, and when the computer program instructions are executed by a processor, the method according to any one of claims 1-15 is implemented.
23. A computer program product, characterized in that, The computer program product includes computer program instructions, and when the computer program instructions are executed by a processor, the method according to any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Storage management method and related device
CN120215806A
Method and apparatus for migration of a virtualized file system, data storage system for migration of a virtualized file system, and file server for use in a data storage system
CN105579950A
Metadata management method and related device
CN110825698A
Storage management method and device and storage system
CN113448921A
Data migration method and related device
CN117234412A