File storage method and device, file reading method and device, equipment and medium
By creating logical header files in a distributed storage system, the problem of multi-version files occupying two sets of metadata when written through file protocol is solved, and the effect of reducing metadata storage needs and improving the ease of use of the storage system is achieved.
Patent Information
- Application Number
- CN202510111891.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-16
AI Technical Summary
In the fusion and interoperability scenario, when multiple versions of files are written through the file protocol, they are limited by the file protocol semantics and can only upload a single version of files, resulting in a large number of single version files, occupying two sets of metadata storage space, resulting in a large proportion of supported files in the cluster.
By creating logical header files in a distributed storage system, as the overall external representation of multiple versions of storage files, the metadata storage needs are reduced directly through logical header files to select and switch multi-version states according to the usage scenario.
This has achieved the reduction of metadata storage requirements, improved the supportable file magnitude of the entire cluster, and improved the ease of use and efficiency of the storage system.
Smart Images

Figure CN120011326A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a file storage method, a reading method, a device, a equipment and a medium. Background Art
[0002] The current distributed storage system includes the Simple Storage Service (S3) protocol, the Hadoop Distributed File System (HDFS), and the Network Attached Storage (NAS) protocol. The above protocols need to communicate with each other, that is, data written by any protocol can be accessed or operated by other protocols later. In other words, the three protocols of S3, HDFS, and NAS are integrated and communicated, supporting multiple protocols (NAS / S3 / HDFS) to share a piece of data and communicate with each other at the same time.
[0003] In the converged interoperability scenario, when multiple versions are enabled, files will inevitably be written through the file protocol. Due to the semantic limitations of the file protocol, only a single version of the file can be uploaded, and repeated uploads will overwrite the file. Therefore, when uploading files through the file protocol, there will be a large number of single-version files. In the existing solution, a single-version file will also correspond to a logical header file. Therefore, multi-version files written using the file protocol will occupy two sets of metadata storage space, resulting in a significant decrease in the number of files that can be supported by the entire cluster. Summary of the invention
[0004] In view of this, the present invention provides a file storage method, reading method, device, equipment and medium to solve the problem that multi-version storage files in a fusion and interoperability scenario occupy two sets of metadata storage space, resulting in a large proportion of the supportable file level of the entire cluster is reduced.
[0005] In a first aspect, the present invention provides a file storage method, the method comprising:
[0006] Acquire a first file to be stored and first identification information corresponding to a file to which the first file to be stored belongs;
[0007] Generate first metadata information corresponding to the first file to be stored;
[0008] Determine a transmission protocol for transmitting the first file to be stored;
[0009] When the transmission protocol is determined to be the target protocol, determining whether there is a file corresponding to the first identification information in the distributed storage system;
[0010] When it is determined that no file corresponding to the first identification information exists in the distributed storage system, creating a logical header file in the distributed storage system;
[0011] Identify whether the multi-version function is enabled, and when it is determined that the multi-version function is enabled, generate version information corresponding to the first file to be stored, and configure the version information as the first version information;
[0012] Adding the first metadata information, the first identification information, and the first version information to the logical header file;
[0013] And the storage operation of the first file to be stored is completed in the distributed storage system.
[0014] A file storage method provided by the present invention has the following advantages:
[0015] The method determines whether there is a file corresponding to the first identification information in the distributed storage system by obtaining the first file to be stored and the first identification information corresponding to the file to which the first file to be stored belongs. When it is determined that there is no file corresponding to the first identification information in the distributed storage system, a logical header file is created in the distributed storage system as the overall external representation of multiple versions of the storage file, and the multi-version state is directly selected and switched according to the usage scenario through the logical header file, thereby improving the usability of the storage system. When it is determined that the multi-version function is in the on state, the version information is configured as the first version information, and the first metadata information, the first identification information, and the first version information are added to the logical header file. By adding the metadata information of the first file to be stored to the logical header file, the logical header file and the metadata information file are merged into one file, reducing one metadata information file and increasing the supportable file level of the entire cluster.
[0016] In an optional embodiment, the method further includes:
[0017] When the second file to be stored and the second identification information corresponding to the second file to be stored are acquired, generating second metadata information corresponding to the second file to be stored;
[0018] Searching for identification information matching the second identification information from a logical header file created in the distributed storage system;
[0019] When it is determined that identification information matching the second identification information is stored in the first logical header file, new version information corresponding to the second file to be stored is generated according to the version information stored in the first logical header file, wherein the first logical header file is any one of the logical header files created in the distributed storage system;
[0020] storing the new version information and the second metadata information in the first logical header file;
[0021] The second file to be stored is stored in the distributed storage system.
[0022] Specifically, when the second file to be stored and the second identification information corresponding to the second file to be stored are obtained, new version information corresponding to the second file to be stored is generated according to the version information stored in the first logical header file, which can effectively manage and track the historical versions of the file, and ensure the integrity and consistency of the data for the file system that needs to be frequently updated and backtracked. The new version information and the second metadata information are stored in the first logical header file. By storing the version information and metadata information of the file to be stored in the first logical header file, the time and resource consumption required to find the metadata information corresponding to the file in the distributed storage system are reduced, and the stored file can be quickly located and retrieved.
[0023] In an optional implementation, after storing the new version information and the second metadata information in the logical header file, the method further includes:
[0024] Create an index file and declaration information corresponding to the index file;
[0025] storing the historical metadata information stored in the first logical header file before storing the second metadata information, and the version information of the storage file corresponding to the historical metadata information, in the index file;
[0026] The declaration information corresponding to the index file is stored in the first logical header file, where the declaration information is used to indicate the storage location corresponding to the index file.
[0027] Specifically, create an index file to break through the size limit of the XATTR extended attribute of the logical header file to store the metadata information of historical versions and version information. The index file can adapt to the growing data storage needs, greatly increasing the amount of metadata information corresponding to historical versions that a single logical header file can support, and improving the scalability of the entire system. Create declaration information corresponding to the index file. The declaration information clearly indicates the storage location of the index file. The index file may be distributed in different storage devices or locations. When it is necessary to access or query the index file, the system can quickly locate the storage path of the index file directly from the logical header file based on the declaration information, without the need for complex search or traversal operations, thereby significantly improving the speed and efficiency of data retrieval.
[0028] In an optional embodiment, the method further includes:
[0029] Acquire a third file to be stored and third identification information corresponding to the third file to be stored;
[0030] Generating third metadata information corresponding to the third file to be stored;
[0031] searching for identification information matching the third identification information from a logical header file created in the distributed storage system;
[0032] When it is determined that identification information matching the third identification information is stored in the second logical header file, and it is determined that the acquisition time of the third file to be stored is within the preset time range, identifying whether the version information stored in the second logical header file is the target version information, wherein the second logical header file is any one of the logical header files created in the distributed storage system, and the preset time range is used to indicate that the multi-version function is in a stopped state;
[0033] When it is determined that the version information stored in the second logical header file is the target version information, the third file to be stored overwrites the file to be stored corresponding to the target version information in the distributed storage system, and the version information of the third file to be stored is updated to the target version information;
[0034] The third metadata information and the target version information are updated into the second logical header file.
[0035] Specifically, different states of the multi-version function correspond to different storage methods for storage files, which can meet the needs of different scenarios. In the on state, multiple versions are allowed to coexist, which is helpful for version control and historical tracing, and convenient for users to access and compare different versions of storage files; in the off state, the latest version is retained by overwriting the old version to reduce redundant data; in the stopped state, the third file to be stored in the distributed storage system overwrites the file to be stored corresponding to the target version, and the version information of the third file to be stored is updated to the target version information, only overwriting the storage files uploaded within the preset time range when the multi-version function is in the stopped state, and flexibly overwriting unnecessary storage files. By reasonably configuring the multi-version function state, the performance of the storage system can be optimized, such as improving data access speed when multiple versions coexist, and reducing storage burden when overwriting old versions.
[0036] In an optional embodiment, the method further includes:
[0037] When it is determined that the version information stored in the second logical header file is not the target version, after the third file to be stored is distributedly stored in the distributed storage system, the version information of the third file to be stored is determined as the target version information;
[0038] The third metadata information and the target version information are updated into the second logical header file.
[0039] Specifically, when the multi-version function is in a stopped state, only the storage files with the target version information within the preset time range are overwritten. When there is no storage file of the target version within the preset time range, the version information of the third file to be stored is determined as the target version information. The preset time range is used to indicate that the multi-version function is in a stopped state. By limiting the conditions for version overwriting by putting the multi-version function in a stopped state, the system can accurately control which file versions will be overwritten, improve the utilization efficiency of storage resources, reduce storage costs, and reduce the number of versions that need to be maintained.
[0040] In an optional embodiment, the method further includes:
[0041] When it is determined that a file corresponding to the first identification information exists in the distributed storage system, information indicating that the file already exists is fed back to the transmission end of the first file to be stored, so as to prompt the transmission end to determine whether the first file to be stored is a duplicate upload.
[0042] Specifically, by performing duplication check before uploading files, it is possible to effectively avoid storing the same file multiple times in the system, which not only saves storage space, but also reduces the load on the storage system, and can improve the overall efficiency of the system.
[0043] A file reading method, the method comprising:
[0044] Obtaining a file reading request, where the file reading request includes identification information of a file to be read;
[0045] According to the identification information, searching for a target logical header file storing the identification information from a logical header file created by the distributed storage system;
[0046] Extracting address indication information corresponding to the file to be read from the target logical header file;
[0047] When the address indication information is metadata information corresponding to the file to be read, acquiring the file to be read according to the metadata information;
[0048] The file to be read is fed back to the requester who sent the file reading request.
[0049] Specifically, the identification information included in the file read request is obtained, the target logical header file storing the same identification information is searched from the logical header file created by the distributed storage system, and the address indication information corresponding to the file to be read is extracted from the target logical header file. When the file version information is not specified in the file read request, the address indication information is the metadata information corresponding to the file to be read stored in the logical header file. The storage location of the file to be read is quickly located directly according to the metadata information stored in the logical header file, which reduces additional search steps, thereby improving the efficiency of file retrieval and access, and simplifying the file retrieval process, because the metadata information directly contains the key information required to locate the file, without the need for additional parsing or conversion steps.
[0050] After obtaining the file to be read from the storage location, the file to be read is fed back to the requester who sent the file reading request, maintaining the consistency and accuracy of the data and reducing the errors that may occur during the file retrieval and provision process.
[0051] In an optional implementation, the file read request also includes target version information, and the method further includes:
[0052] When the target version information is not the version information stored in the target logic header file, extracting declaration information from the target logic header file, wherein the declaration information is the address indication information;
[0053] According to the declaration information, obtain the index file corresponding to the declaration information;
[0054] Extracting metadata information of the file to be read corresponding to the target version information from the index file;
[0055] Obtain the file to be read according to metadata information;
[0056] The file to be read is fed back to the requester who sent the file reading request.
[0057] Specifically, when the file reading request also includes target version information, the method determines the storage location of the stored file by using the declaration information as address indication information, that is, through the declaration information, obtaining the index file corresponding to the declaration information, and extracting the metadata information of the file to be read corresponding to the target version information from the index file, thereby determining the storage address of the file to be read, ensuring that in an environment where multiple versions coexist, the correct storage file version can always be accessed, thereby enhancing data consistency, and accurate address indication information helps to improve system reliability.
[0058] In a second aspect, the present invention provides a file storage device, the device comprising:
[0059] An acquisition module, used to acquire a first file to be stored and first identification information corresponding to a file to which the first file to be stored belongs;
[0060] A generating module, used to generate first metadata information corresponding to the first file to be stored;
[0061] A processing module, used to determine a transmission protocol for transmitting the first file to be stored;
[0062] When it is determined that the transmission protocol is the target protocol, determining whether there is a file corresponding to the first identification information in the distributed storage system; when it is determined that there is no file corresponding to the first identification information in the distributed storage system, creating a logical header file in the distributed storage system;
[0063] Identification module, used to identify whether the multi-version function is enabled;
[0064] The generating module is further configured to generate version information corresponding to the first file to be stored and configure the version information as the first version information when it is determined that the multi-version function is in an enabled state;
[0065] The processing module is further used to add the first metadata information, the first identification information, and the first version information into the logical header file;
[0066] And the storage operation of the first file to be stored is completed in the distributed storage system.
[0067] A file storage device provided by the present invention has the following advantages:
[0068] The device determines whether there is a file corresponding to the first identification information in the distributed storage system by obtaining the first file to be stored and the first identification information corresponding to the file to which the first file to be stored belongs. When it is determined that there is no file corresponding to the first identification information in the distributed storage system, a logical header file is created in the distributed storage system as the overall external representation of multiple versions of the storage file, and the multi-version state is directly selected and switched according to the usage scenario through the logical header file, thereby improving the usability of the storage system. When it is determined that the multi-version function is in the on state, the version information is configured as the first version information, and the first metadata information, the first identification information, and the first version information are added to the logical header file. By adding the metadata information of the first file to be stored to the logical header file, the logical header file and the metadata information file are merged into one file, reducing one metadata information file and increasing the supportable file level of the entire cluster.
[0069] The present invention provides a file reading device, the device comprising:
[0070] An acquisition module, used for acquiring a file reading request, wherein the file reading request includes identification information of a file to be read;
[0071] A search module, used for searching a target logical header file storing identification information from a logical header file created by a distributed storage system according to the identification information;
[0072] An extraction module, used for extracting address indication information corresponding to the file to be read from the target logical header file;
[0073] A processing module, configured to obtain the file to be read according to the metadata information when the address indication information is metadata information corresponding to the file to be read;
[0074] The sending module is used to feed back the file to be read to the requester who sends the file reading request.
[0075] The file reading device provided by the present invention has the following advantages:
[0076] The device searches for a target logical header file storing identification information from a logical header file created by a distributed storage system according to the identification information, extracts address indication information corresponding to the file to be read from the target logical header file, and when the file version information is not specified in the file read request, the address indication information is metadata information corresponding to the file to be read stored in the logical header file. The storage location of the file to be read is quickly located directly according to the metadata information stored in the logical header file, thereby reducing additional search steps, thereby improving the efficiency of file retrieval and access, and simplifying the file retrieval process.
[0077] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the file storage method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0078] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the file storage method of the first aspect or any corresponding embodiment thereof.
[0079] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the file storage method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0081] Figure 1 It is a flowchart of a file storage method provided by an embodiment of the present invention;
[0082] Figure 2 It is a structural schematic diagram of a file storage method provided by an embodiment of the present invention;
[0083] Figure 3 is a flowchart of another file storage method provided by an embodiment of the present invention;
[0084] Figure 4 is a flowchart of another file storage method provided by an embodiment of the present invention;
[0085] Figure 5 It is a structural schematic diagram of a file storage method provided by an embodiment of the present invention;
[0086] Figure 6 is a structural diagram of another file storage method provided by an embodiment of the present invention;
[0087] Figure 7 It is a flowchart of another file storage method provided by an embodiment of the present invention;
[0088] Figure 8 It is a flowchart of another file storage method provided by an embodiment of the present invention;
[0089] Fig. 9 It is a flowchart of a file reading method provided by an embodiment of the present invention;
[0090] Fig.10 is a flowchart of another file reading method provided by an embodiment of the present invention;
[0091] Fig.11 is a flowchart of another file reading method provided by an embodiment of the present invention;
[0092] Fig.12 This is a schematic diagram of a structure of switching three states of a multi-version function provided by an embodiment of the present invention;
[0093] Fig.13 It is a structural schematic diagram of a file reading method provided by an embodiment of the present invention;
[0094] Fig.14 is a structural block diagram of a file storage device provided by an embodiment of the present invention;
[0095] Fig.15 It is a structural block diagram of a file reading device provided by an embodiment of the present invention;
[0096] Fig.16 It is a schematic diagram of the hardware structure of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0097] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0098] The embodiments of the present invention relate to the field of distributed storage systems, and more specifically, to a multi-version implementation scheme based on a distributed storage system in an unstructured fusion and interoperability scenario, and propose a multi-version implementation scheme based on a file base, thereby improving the availability, usability, compatibility and semantic integrity of a distributed big data storage system.
[0099] The current market has the Simple Storage Service (S3) protocol, the Hadoop Distributed File System (HDFS), and the Network Attached Storage (NAS) protocol. There is a demand for interoperability through the above protocols, that is, data written through any protocol can be accessed or operated through other protocols later.
[0100] The current distributed storage system already supports the integration and interoperability of the S3 protocol, HDFS protocol, and NAS protocol based on the file base. The S3 protocol has a multi-version function, which means that in the object storage service, multiple different versions of data are saved for the same object. These versions can be different states of the object at different points in time, or different results of the object after different operations (such as uploading, deleting, etc.). With multiple versions of objects, users can easily retrieve and restore each version of object data, so as to quickly restore data in the event of unexpected operations or data loss. The S3 protocol has the following problems with the multi-version function:
[0101] In the converged interoperability scenario, when multiple versions are enabled, files will inevitably be written through the file protocol. Due to the semantic limitations of the file protocol, only a single version of the file can be uploaded, and repeated uploads will overwrite the file. Therefore, when uploading files through the file protocol, there will be a large number of single-version files. In the existing solution, a single-version file will also correspond to a logical header file. Therefore, multi-version files written using the file protocol will occupy two metadata, resulting in a significant decrease in the number of files that can be supported by the entire cluster.
[0102] Specifically, when a file is uploaded to the file system, the file system stores a series of metadata for the file itself. These metadata describe the basic properties of the file, such as file name, size, creation time, modification time, owner, permissions, etc. This information is crucial for the correct operation and management of the file system. In a multi-version control environment, each version of the file may need to be managed and tracked independently. To achieve this, the distributed storage system can create a logical header file for each file version. This header file contains additional information about the version file, such as version number: identifying a specific version of the file; version creation time: recording the time when the version was created; version content summary (such as MD5 or SHA value): providing file integrity verification; version history: may include references to previous and subsequent versions; version access permissions: may have different access controls for different versions, etc. These logical header files themselves also need to store metadata, because the file system needs to manage the information of these header files, including their creation time, last modification time, owner, permissions, etc. Therefore, it is necessary to occupy the storage space of two metadata, one of which is the metadata corresponding to the actual stored file content, such as file name, size, permissions, etc. Another metadata is an additional file, which is metadata corresponding to additional information of the current version of the file, such as creation time, modification time, owner, etc.
[0103] In summary, since each version requires such a logical header file, the file system actually needs to store two sets of metadata for each version: one set is the metadata of the file content, and the other set is the metadata of the logical header file. This increases the storage overhead, resulting in a large decrease in the number of files that can be supported by the entire cluster, and may lead to a decrease in storage capacity and efficiency.
[0104] To solve the above problems, an embodiment of the present invention provides a file storage embodiment. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system (computer device) including a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0105] In this embodiment, a file storage method is provided, which is applied to a distributed storage system and can be used in the above-mentioned terminal devices, such as mobile phones, tablet computers, etc. Figure 1 is a flowchart of a file storage method provided by an embodiment of the present invention, such as Figure 1 As shown, the process includes the following steps:
[0106] Step S101: Acquire a first file to be stored and first identification information corresponding to a file to which the first file to be stored belongs.
[0107] Specifically, the distributed storage system can obtain at least one file to be stored and the corresponding identification information at the same time. In the present application document, only the example of obtaining the first file to be stored and the first identification information is used to explain the operations that need to be performed after obtaining the file to be stored.
[0108] Step S102: Generate first metadata information corresponding to the first file to be stored.
[0109] Specifically, after the distributed storage system obtains the first file to be stored, it generates first metadata information corresponding to the first file to be stored.
[0110] The metadata information corresponding to the stored file describes the basic attributes of the file, such as the name of the stored file, the size of the stored file, the creation time, the modification time, the owner, the permissions, and the specific storage path of the file. The first identification information corresponding to the file to which the first file to be stored belongs is information that can uniquely identify the stored file, such as the name of the stored file. The first identification information corresponding to the file to which the first file to be stored belongs is obtained to perform subsequent storage operations on the stored file according to the identification information.
[0111] Step S103: determining a transmission protocol for transmitting the first file to be stored.
[0112] There are many specific file transfer protocols, such as file protocols and object protocols.
[0113] Step S104: when it is determined that the transmission protocol is the target protocol, it is determined whether there is a file corresponding to the first identification information in the distributed storage system.
[0114] Specifically, after the distributed storage system receives the storage file, it first needs to determine the transmission protocol for transmitting the first file to be stored. When the transmission protocol is determined to be the target protocol, in this application document, for example, when the target protocol is a file protocol, it is determined whether there is a file consistent with the first identification information in the distributed storage system based on the first identification information corresponding to the file to which the first file to be stored belongs obtained in the previous text.
[0115] Step S105: when it is determined that no file corresponding to the first identification information exists in the distributed storage system, a logical header file is created in the distributed storage system.
[0116] Step S106, identifying whether the multi-version function is enabled, and when it is determined that the multi-version function is enabled, after generating version information corresponding to the first file to be stored, configuring the version information as the first version information.
[0117] Specifically, when it is determined that there is no file corresponding to the first identification information in the distributed storage system, it means that the first file to be stored is stored for the first time, so a logical header file needs to be created to store relevant information corresponding to the first file to be stored.
[0118] A logical header file is created in the distributed storage system, and the distributed storage system identifies whether the multi-version function in the system extended attributes (Extended Attributes, referred to as XATTR) is enabled. If the multi-version function is enabled, it means that multiple versions corresponding to the first file to be stored can coexist. Therefore, after generating the version information corresponding to the first file to be stored, Figure 2 As shown, the version information is configured as the first version information, such as FIRST_VERSION_FL.
[0119] Step S107: adding the first metadata information, the first identification information, and the first version information to the logical header file.
[0120] Specifically, the version information configured as the first version information, together with the corresponding first identification information and the first metadata information, are added to the logical header file created above.
[0121] Step S108, completing the storage operation of the first file to be stored in the distributed storage system.
[0122] The file storage method provided in this embodiment determines whether there is a file corresponding to the first identification information in the distributed storage system by obtaining the first file to be stored and the first identification information corresponding to the file to which the first file to be stored belongs. When it is determined that there is no file corresponding to the first identification information in the distributed storage system, a logical header file is created in the distributed storage system as the overall external representation of multiple versions of the storage file, and the multi-version state is directly selected and switched according to the usage scenario through the logical header file, thereby improving the usability of the storage system. When it is determined that the multi-version function is in the on state, the version information is configured as the first version information, and the first metadata information, the first identification information, and the first version information are added to the logical header file. By adding the metadata information of the first file to be stored to the logical header file, the logical header file and the metadata information file are merged into one file, reducing one metadata information file and increasing the supportable file level of the entire cluster.
[0123] As described above, when the first file to be stored and the first identification information corresponding to the file to which the first file to be stored belongs are obtained, and the first metadata information corresponding to the first file to be stored is generated, when the transmission protocol is determined to be the target protocol, it is determined whether there is a file corresponding to the first identification information in the distributed storage system.
[0124] The above embodiment provides a file storage method when there is no file corresponding to the identification information in the distributed storage system. In an optional embodiment, when there is a file corresponding to the identification information in the distributed storage system, the storage operation of obtaining the second file to be stored is as follows: Figure 3 As shown, the method can be used in the above-mentioned mobile terminals, such as mobile phones, tablet computers, etc. Figure 3 : is a flowchart of a file storage method provided by an embodiment of the present invention. The process includes the following steps:
[0125] Step S301: When a second file to be stored and second identification information corresponding to the second file to be stored are acquired, second metadata information corresponding to the second file to be stored is generated.
[0126] Step S302: searching for identification information matching the second identification information from a logical header file created in the distributed storage system.
[0127] Specifically, after obtaining the first file to be stored, files to be stored may be received one after another. As mentioned above, after the distributed system obtains any storage file, it will generate metadata information corresponding to the storage file. Then, according to the identification information, it searches in the distributed storage system whether there is a logical header file storing the corresponding identification information.
[0128] Step S303: when it is determined that identification information matching the second identification information is stored in the first logical header file, new version information corresponding to the second file to be stored is generated according to the version information stored in the first logical header file.
[0129] The first logical header file is any logical header file among the logical header files created in the distributed storage system.
[0130] Specifically, when it is determined that there is a first logical header file storing second identification information in the distributed storage system, it proves that different versions of storage files consistent with the second identification information have been stored in the distributed storage system. At this time, new version information corresponding to the second file to be stored is generated based on the version information stored in the first logical header file.
[0131] For example, when the version information stored in the logic header file is the first version information, the second version information corresponding to the second file to be stored is generated according to a preset version information generation specification.
[0132] Step S304: store the new version information and the second metadata information into the first logical header file.
[0133] Specifically, the latest generated version information and metadata information are stored in a logical header file, and the extended logical header file serves as the overall external representation of the multi-version storage file.
[0134] Step S305: completing the storage operation of the second file to be stored in the distributed storage system.
[0135] Specifically, based on the version information stored in the first logical header file, new version information corresponding to the second file to be stored is generated, which can effectively manage and track the historical versions of the file, and ensure the integrity and consistency of the data for the file system that needs to be frequently updated and backtracked. The new version information and the second metadata information are stored in the first logical header file. By storing the identification information and metadata information in the logical header file, the stored file can be quickly located and retrieved, reducing the time and resource consumption required to find the metadata information corresponding to the file in the distributed storage system.
[0136] This method allows the flexible creation and management of logical header files in a distributed storage system, supporting different types of file storage requirements. Moreover, as the system expands, more logical header files and storage nodes can be easily added to meet the growing storage needs.
[0137] In an optional embodiment, the version information is optimized. In the related art, a completely random 32-bit string is used as the version number, which does not carry any other information. It is not only impossible to sort naturally like version numbers such as numbers or dates, but also difficult to intuitively determine the order of versions, which brings great inconvenience to version management and tracking. Therefore, in this method, the generation specification of multi-version file version information is redefined. The specific generation specification can be the following formula:
[0138] i+current hexadecimal timestamp+v+hexadecimal identification number of the current version file
[0139] Specifically, adding a timestamp element to the version information can record the creation time of the version file to ensure the order between the version information; adding an identification number of the corresponding version file to the version information can be used to directly access the specified version based on the identification number, which can reduce the process of searching for multiple version files and reduce the operation of converting the path to the identification number, thereby improving access performance.
[0140] Based on any of the foregoing embodiments, the new version information and metadata information are stored in the logical header file corresponding to the identification information. In the related art, the XATTR extended attribute of the logical header file is used to store its historical version information and the metadata information of the historical storage file. The XATTR extended attribute has a size limit, which leads to a limited number of supported versions. Calculated based on the common 64K size, the maximum number of historical versions is no more than 1,000. In addition, other related extended attributes are also recorded in the XATTR attribute of the file. Therefore, when the number of historical versions is large, the scalability will be poor due to the limitation of the XATTR attribute, and it will also affect other existing attributes stored in the XATTR.
[0141] Therefore, in an optional embodiment of the present application, an index file is created in the storage pool to store metadata information of historical storage files and corresponding version information to solve the above problem. Figure 4 As shown, Figure 4 It is a flowchart of another file storage method provided by an embodiment of the present invention.
[0142] Specifically, after storing the new version information and the second metadata information in the logical header file, the method may further include the following method steps:
[0143] Step S401: create an index file and declaration information corresponding to the index file.
[0144] The declaration information is used to indicate the storage location corresponding to the index file.
[0145] Specifically, an index file is created in a pre-built storage pool, and declaration information corresponding to the index file is created. For example, an object OMAP may be introduced to point to the index file.
[0146] Step S402: store the historical metadata information stored in the first logical header file before storing the second metadata information, and the version information of the storage file corresponding to the historical metadata information, into the index file.
[0147] Specifically, the index file stores historical metadata information and version information of the storage file corresponding to the historical metadata information. In the index file, the historical metadata information and the version information of the storage file corresponding to the historical metadata information are stored in a key-value format, where the key is the version information of the storage file and the value is the metadata information of the storage file.
[0148] In a specific example, the declaration information of the index file can be called to add, delete, modify, and search the metadata information and version information in the index file.
[0149] Step S403: store declaration information corresponding to the index file into the first logical header file.
[0150] Specifically, the declaration information corresponding to the index file is stored in the first logical header file. At this time, the logical header file includes identification information of the storage file, metadata information and version information of the latest storage file, and the declaration information corresponding to the index file.
[0151] Next, in a specific example, the above embodiments are combined to illustrate a specific embodiment to explain the process of storing the historical metadata information stored in the logical header file 1 before storing the metadata information 2 and the version information of the storage file corresponding to the historical metadata information in the index file after obtaining the file 2 to be stored, such as Figure 5 shown.
[0152] On the basis of the foregoing embodiment, it is determined that the multi-version function is in an enabled state, a file to be stored and identification information one corresponding to the file to be stored are obtained, and the transmission protocol of the file to be stored is determined to be a file protocol. When it is determined that there is no file corresponding to the identification information one in the distributed storage system, a logical header file one is created in the distributed storage system. After the version information corresponding to the file to be stored is generated, the version information is configured as the first version information, and the metadata information one, the identification information one, and the first version information are added to the logical header file one.
[0153] When the second file to be stored and the second identification information corresponding to the second file to be stored are obtained, after the metadata information second is generated, it is determined whether the second identification information is consistent with the first identification information. When it is determined that the first identification information stored in the first logical header file matches the second identification information, the second version information corresponding to the second file to be stored is generated according to the first version information stored in the first logical header file. After the second version information and the second metadata information are stored in the first logical header file, the metadata information in the first logical header file and the first version information are stored in the pre-created index file. Figure 5 The specific operation process is demonstrated, including:
[0154] First, modify the version information stored in the logical header file 1, that is, modify the first version information to version information 1, create an empty file, copy metadata information 1 and version information 1 from the logical header file 1 to the empty file as the metadata information file corresponding to the storage file 1. Finally, update the metadata information and version information 1 corresponding to the storage file 1 to the storage pool specified by OMPA, where OMPA indicates the storage location declaration information corresponding to the index file.
[0155] Based on the above embodiment, the multi-version function status is as follows: Figure 6 As shown, the multi-version function includes not only the on state and the off state, but also the stop state within a preset time range.
[0156] Specifically, when the multi-version function is enabled, the storage method of the storage file is that multiple versions of the storage file coexist. Alternatively, when the multi-version function is disabled, the storage method of the storage file is to overwrite the previously stored version of the storage file, retain the latest version of the storage file, or return an error message indicating that the file already exists.
[0157] When the multi-version function state is in a stopped state within a preset time range, indicating that the multi-version files with the same identification information in the stopped state are to be overwritten and uploaded, in an optional embodiment, the method further includes:
[0158] Step a1: Acquire a third file to be stored and third identification information corresponding to the third file to be stored.
[0159] The third file to be stored is any storage file.
[0160] Step a2: Generate third metadata information corresponding to the third file to be stored.
[0161] Step a3: searching for identification information matching the third identification information from the logical header file created in the distributed storage system.
[0162] Step a4: when it is determined that the second logical header file stores identification information matching the third identification information and it is determined that the acquisition time of the third file to be stored is within the preset time range, identify whether the version information stored in the second logical header file is the target version information.
[0163] The second logical header file is any one of the logical header files created in the distributed storage system, the preset time range is used to indicate that the multi-version function is in a stopped state, and the target version information is the first version information.
[0164] Specifically, when it is determined that there is a logical header file storing third identification information in the distributed storage system, it is identified whether the multi-version function is in a stopped state. When the multi-version function is in a stopped state, it indicates that the storage files with the same identification information uploaded within the preset time range should be overwritten and uploaded, and only the latest storage files should be retained. Therefore, it is necessary to determine whether a storage file with the same identification information has been uploaded before the third file to be stored is uploaded within the preset time range. The determination method is to confirm whether the latest version information stored in the logical header file is the first version information.
[0165] Step a5: when it is determined that the version information stored in the second logical header file is the target version information, the third file to be stored overwrites the file to be stored corresponding to the target version information in the distributed storage system, and the version information of the third file to be stored is updated to the target version information.
[0166] Specifically, as described above, when it is determined that the version information stored in the logical header file is the first version information, the third file to be stored in the distributed storage system will overwrite the file to be stored corresponding to the first version information, and the version information of the third file to be stored will be updated to the first version information, so that the third file to be stored can be uploaded to overwrite the file to be stored with the same identification information after the file to be stored is subsequently obtained, and the cycle is repeated until the multi-version function state changes to other states.
[0167] Step a6: Update the third metadata information and the target version information into the second logical header file.
[0168] Specifically, different states of the multi-version function correspond to different storage methods for storage files, which can meet the needs of different scenarios. In the on state, multiple versions are allowed to coexist, which is helpful for version control and historical tracing, and convenient for users to access and compare different versions of storage files; in the off state, the latest version is retained by overwriting the old version to reduce redundant data; in the stopped state, the third file to be stored in the distributed storage system overwrites the file to be stored corresponding to the target version information, and the version information of the third file to be stored is updated to the target version information, only overwriting the storage files uploaded within the preset time range when the multi-version function is in the stopped state, and flexibly overwriting unnecessary storage files. By reasonably configuring the multi-version function state, the performance of the storage system can be optimized, such as improving data access speed when multiple versions coexist, and reducing storage burden when overwriting old versions.
[0169] Based on the foregoing embodiment, in an optional embodiment, when it is determined that the version information stored in the second logical header file is not the target version information, the method specifically includes:
[0170] Step b1, after the third file to be stored is distributedly stored in the distributed storage system, version information of the third file to be stored is determined as target version information;
[0171] Step b2: Update the third metadata information and the target version information into the second logical header file.
[0172] Specifically, as described above, when it is determined that the version information stored in the second logical header file is not the target version information, it means that no storage file with the same identification information has been uploaded within the preset time range before the third file to be stored is uploaded. Therefore, it is necessary to distribute the third file to be stored in the distributed storage system, and then determine the version information of the third file to be stored as the first version information, so as to subsequently overwrite and upload the third file to be stored based on the first version information in the logical header file, flexibly overwrite unnecessary storage files, and reduce the storage burden.
[0173] In a specific example, when it is determined that the version information stored in the second logical header file is not the target version information, the specific process of updating the version information of the third file to be stored to determine it as the target version information is as follows: Figure 7 As shown, specifically including:
[0174] When the target protocol is a file protocol, the file to be stored is uploaded using the file protocol. According to the second logical header file, it is determined whether the storage file corresponding to the first version information exists. When it is determined that the storage file corresponding to the first version information does not exist, the third file to be stored is distributedly stored in the distributed storage system, and the version information of the third file to be stored is determined as the first version information. The third metadata information and the first version information are updated in the second logical header file, and the process ends.
[0175] When it is determined that the storage file corresponding to the first version information exists, the indication information that the file already exists is fed back to the transmission end of the first file to be stored to prompt the transmission end to determine whether the first file to be stored is a duplicate upload, and the process ends.
[0176] In an optional example, the above method is not only applicable to the storage method corresponding to the file transfer protocol, but also includes the storage method corresponding to the object protocol. Next, a specific example is used to illustrate the specific process of storing files in the object protocol, wherein: Figure 8 A flowchart of a file storage method provided by the present invention specifically includes:
[0177] First, the client uploads the file to be stored using the object protocol, and the server obtains the identification information of the file to be stored. After generating the corresponding metadata information, it searches for the existence of a storage file with the same identification information based on the identification information. When the storage file exists, it searches for the logical header file storing the same identification information in the distributed storage system, and at the same time confirms whether the multi-version function status of the distributed storage system is turned on. When the multi-version function status is turned on, it determines whether the version information stored in the logical header file is the first version information. When the version information stored in the logical header file is the first version information, it is necessary to separate the file metadata information of the first version information stored in the logical header file, and the version information and the logical header file are stored in the pre-created index file. The specific separation process is as described above and will not be repeated here. When the version information stored in the logical header file is not the first version information, new version information corresponding to the file to be stored is generated based on the version information stored in the logical header file, and the new version information and metadata information are stored in the logical header file, and the storage operation of the file to be stored in the distributed storage system is completed.
[0178] When the multi-version function is not in the enabled state, determine whether the multi-version function is in the stopped state. When the multi-version function is in the stopped state, verify whether the first version information and its corresponding metadata information are stored in the logical header file. When the first version information and its corresponding metadata information are stored in the logical header file, configure the version information of the file to be stored as the first version information, overwrite and upload the metadata information and version information in the logical header file, and complete the storage operation of the file to be stored in the distributed storage system. When the first version information and its corresponding metadata information are not stored in the logical header file, determine the version information of the file to be stored as the first version information, and update the metadata information and its corresponding version information to the logical header file. Complete the storage operation of the file to be stored in the distributed storage system.
[0179] When the storage file does not exist, create a logical header file, configure the version information corresponding to the first file to be stored as the first version information, add the metadata information, identification information, and the first version information to the logical header file, and complete the storage operation of the file to be stored in the distributed storage system.
[0180] The storage file is written through the aforementioned file storage method, and then you want to implement multi-version storage file reading in the fusion scenario, that is, the data written through any protocol can be accessed or operated through other protocols later.
[0181] In an embodiment of the present application, a file reading method is also provided, which can be used in the above-mentioned mobile terminal, such as a mobile phone, a tablet computer, etc. Fig. 9 is a flowchart of a file reading method provided by an embodiment of the present invention, such as Fig. 9 As shown, the process includes the following steps:
[0182] Step S901: Obtain a file read request.
[0183] Specifically, the client sends a file reading request, wherein the file reading request includes identification information of the file to be read.
[0184] Step S902: searching, according to the identification information, for a target logical header file storing the identification information from the logical header files created by the distributed storage system.
[0185] Step S903: extracting address indication information corresponding to the file to be read from the target logical header file.
[0186] Specifically, the address indication information corresponding to the file to be read is extracted from the logical header file according to the information included in the file read request. The client may also include version information in the file read request. When the file read request only includes the identification information of the storage file to be read or the version information is the version information of the latest storage file, the address indication information is the metadata information of the latest version of the storage file stored in the logical header file.
[0187] Step S904: when the address indication information is metadata information corresponding to the file to be read, the file to be read is acquired according to the metadata information.
[0188] Specifically, according to the address indication information, the specific storage path of the file is searched from the metadata information corresponding to the storage file to be read, and the file to be read is obtained from the storage path.
[0189] Step S905: Feedback the file to be read to the requesting party that sends the file reading request.
[0190] The file reading method provided in this embodiment obtains identification information included in a file reading request, searches for a target logical header file storing the same identification information from a logical header file created by a distributed storage system, and extracts address indication information corresponding to the file to be read from the target logical header file. When the file version information is not specified in the file reading request, the address indication information is metadata information corresponding to the file to be read stored in the logical header file. The storage location of the file to be read is quickly located directly according to the metadata information stored in the logical header file, thereby reducing additional search steps, thereby improving the efficiency of file retrieval and access, and simplifying the file retrieval process, because the metadata information directly contains the key information required to locate the file, and no additional parsing or conversion steps are required.
[0191] After obtaining the file to be read from the storage location, the file to be read is fed back to the requester who sent the file reading request, maintaining the consistency and accuracy of the data and reducing the errors that may occur during the file retrieval and provision process.
[0192] Based on the foregoing embodiment, in an optional embodiment, when the file read request includes the target version information, it specifically includes:
[0193] Step c1: when the target version information is not the version information stored in the target logic header file, extract declaration information from the target logic header file.
[0194] The declaration information is the address indication information.
[0195] Step c2: according to the declaration information, obtain the index file corresponding to the declaration information.
[0196] Step c3: extracting metadata information of the file to be read corresponding to the target version information from the index file.
[0197] Step c4, obtaining the file to be read according to the metadata information.
[0198] Step c5: Feedback the file to be read to the requesting party that sends the file reading request.
[0199] Specifically, as mentioned above, the metadata information and version information of the historical storage file will be persistently stored in an index file pre-created in a pre-built storage pool. When the file read request includes not only the identification information but also the target version information and this target version information is the target version information corresponding to the historical storage file, it is necessary to call the declaration information according to the version information, find the metadata information of the corresponding version storage file, determine the specific storage path of the corresponding version storage file, and finally feed back the file to be read to the requester who sent the file read request.
[0200] Specifically, when the file reading request also includes target version information, the method determines the storage location of the stored file by using the declaration information as address indication information, that is, through the declaration information, obtaining the index file corresponding to the declaration information, and extracting the metadata information of the file to be read corresponding to the target version information from the index file, thereby determining the storage address of the file to be read, ensuring that in an environment where multiple versions coexist, the correct storage file version can always be accessed, thereby enhancing data consistency, and accurate address indication information helps to improve system reliability.
[0201] The following is a specific example of how to access a multi-version file through a file protocol. Fig.10 As shown:
[0202] First, obtain the file protocol access request, and determine whether the storage file exists according to the identification information of the storage file to be read included in the request. When the storage file does not exist, directly feedback the abnormal alarm information that the storage file does not exist. When the storage file exists, determine whether the storage file is the storage file corresponding to the first version information according to the version information in the logical header file corresponding to the identification information. When the storage file is the storage file corresponding to the first version information, directly read the storage file according to the metadata information in the logical header file. When the storage file is not the storage file of the first version information, determine the multi-version function status of the distributed storage system. When the multi-version function status is turned on, find the metadata information of the corresponding version information according to the previous file reading method, and determine whether the target version is marked as deletemarker. If not, read the target version file and the process ends. Otherwise, feedback that the file does not exist is abnormal, and the process ends. When the multi-version function status is turned off, the version information is empty, and the storage file is read directly, and the process ends.
[0203] In an optional example, the above file reading method is not only applicable to file protocol access to storage files, but also includes object protocol access to storage files. The following will illustrate the process of accessing multi-version files through object protocol in the form of a specific example. The specific process is as follows: Fig.11 As shown, including:
[0204] First, an object protocol access request is obtained, and whether the object exists is determined based on the identification information of the storage file to be read included in the request. If the object does not exist, an object non-existence exception is directly returned, and the process ends.
[0205] When the object exists, determine whether the storage file specified by the object is the storage file corresponding to the first version. When the storage file is the storage file corresponding to the first version, directly read the current file according to the metadata information in the logical header file. When the storage file is not the storage file corresponding to the first version, confirm the multi-version function status. When the multi-version function status is closed, directly read the current file according to the metadata information in the logical header file. When the multi-version function status is on, confirm whether the request includes the target version information. When the target version information is not included, directly read the latest version storage file according to the metadata information in the logical header file. When the file reading request includes the target version information, determine whether the target version storage file exists. When the target version storage file exists, according to the previous file reading method, read the storage file corresponding to the target version according to the index file declaration information. When the storage file corresponding to the target version is marked as a deletemarker version or does not exist, return an object does not exist exception and the process ends.
[0206] Based on the foregoing embodiment, the multi-version function includes not only an on state and a off state, but also a stop state within a preset time range, and the distributed storage system monitors the state changes of the multi-version function in real time.
[0207] See Fig.12 As shown, Fig.12 The diagram in Figure 1 shows the three states of switching between the multi-version function. The multi-version function is in the off state by default and can be switched from the off state to the on state. After the multi-version function is turned on, it cannot be switched back to the off state. However, it can be switched from the on state to the stopped state within the preset time range, but it cannot be switched from the stopped state to the off state within the preset time range.
[0208] The specific reason is that when the multi-version function is in a stopped state, that is, a paused state, within a preset time range, it means that the multi-version storage file coexistence state will be switched to a single-version file storage state. The subsequently received files to be stored with the same identification information will overwrite the previously uploaded storage files with the same identification information. Then, if you switch from the paused state to the closed state, a conflict will occur, because before the multi-version function state is switched from the on state to the stopped state, multiple multi-version storage files may have been stored in the system. The closed state means that the multi-version function is turned off, and there can only be one storage file with the same identification information, and there is no state of multi-version storage files. Therefore, it is not possible to switch from a stopped state within a preset time range to a closed state. However, it is possible to switch from a stopped state within a preset time range to a started state.
[0209] In an optional example, when it is detected that the state of the multi-version function switches from a stopped state within a preset time range to a started state, the method further includes:
[0210] Step d1, obtaining a fourth file to be stored and identification information corresponding to the file to which the fourth file to be stored belongs.
[0211] The fourth file to be stored is the first stored file obtained after it is detected that the multi-version state switches from a stopped state within a preset time range to a started state.
[0212] Step d2, generating metadata information corresponding to the fourth file to be stored.
[0213] Step d3, obtaining the version information of the storage file corresponding to the latest stored historical metadata from the index file.
[0214] Step d4, generating version information corresponding to the fourth file to be stored according to the version information of the storage file corresponding to the latest stored historical metadata.
[0215] Step d5, storing the new version information and metadata information in the logical header file.
[0216] Step d6, overwriting the storage file corresponding to the first version information with the fourth file to be stored in the distributed storage system to complete the storage operation.
[0217] Specifically, as mentioned above, the storage files that are in the stopped state within the preset time range are overwritten and uploaded, and only the first version information and the metadata information corresponding to the first version information are stored in the logical header file. When the multi-version state is switched from the stopped state within the preset time range to the started state, the acquired files to be stored are stored in parallel with the storage files before switching to the stopped state. In this case, it is necessary to overwrite the storage files uploaded within the preset time range and continue the multi-version file storage operation.
[0218] Among them, the switching instruction is executed in the upload interval of each storage file to avoid overwriting necessary storage files. For example, when the client wants to store storage file four in parallel with the storage file before the multi-version function is in the stopped state, it is necessary to switch the multi-version state from the stopped state to the started state within the preset time range before obtaining storage file four and after storing storage file three.
[0219] On the basis of the aforementioned embodiments, an index file is created in a pre-built storage pool to persistently store the metadata information and corresponding version information of historical storage files, and its access mode is not static. With the development of business and the growth of data volume, the access frequency and access mode of index files will change. For example, in operations such as data migration, backup and recovery, or large-scale data query, the access demand for index files may increase significantly. At this time, it is necessary to dynamically adjust the storage strategy and temporarily migrate the index files to high-performance storage devices to meet the needs of high-concurrency access and ensure the stability of system performance.
[0220] In an optional embodiment, the specific dynamic adjustment of the storage strategy includes the following steps:
[0221] Step e1, recording the basic attributes of access to each storage file and index file in real time, and generating an access log.
[0222] Specifically, the access log includes basic information such as access time, access times, visitor identity, etc. By recording the basic information of each storage file, the access frequency of each storage file can be obtained, and the storage location of the storage file can be appropriately adjusted according to the access frequency of each storage file.
[0223] Step e2: presetting the index file access threshold and storage policy adjustment rules.
[0224] For example, when the access frequency of the index file increases significantly in a short period of time and exceeds the set high access frequency threshold, the storage policy adjustment mechanism is triggered.
[0225] Step e3: Adjust the rules according to the storage policy and automatically execute the data migration task.
[0226] Specifically, the storage policy adjustment mechanism is to migrate the index file from the current storage device to the high-performance storage device according to the storage policy adjustment rule to meet the demand of high concurrent access.
[0227] Specifically, Fig.13 The overall operation process of reading storage files with multiple protocols in the fusion scenario is shown. The specific process is as follows Fig.13 As shown:
[0228] As in the file storage method mentioned above, the distributed storage system confirms that the multi-version function status of the storage space is turned on, and the client writes multiple versions of the identification information 1 of the file to the distributed storage system. The distributed storage system generates version information corresponding to the storage file based on the version information stored in the logical header file, and at the same time generates metadata information corresponding to the storage file.
[0229] Assume that the storage file of identification information 1 stores three versions, and generates version 1 metadata information, version 2 metadata information, and version 3 metadata information.
[0230] The access request can be an object protocol access, a file protocol access, or a big data access, etc. The corresponding logical header file is determined according to the identification information in the access request, for example, the logical header file storing the identification information 1. When the access request only contains the identification information but does not contain the target version information or the target version information is the latest version information, the storage file corresponding to version 3 is read directly according to the version 3 metadata information in the logical header file. When the access request includes the target version information, the object OMAP is called to search for the metadata information of the target version and determine the storage location of the storage file corresponding to the target version. The historical metadata information in the index file pointed to by the object OMAP is stored in the form of key-value, where the key is the version information corresponding to the storage file, such as version 1, version 2, version 3, etc., and the value is the metadata information corresponding to the storage file.
[0231] It is also possible to create a logical header file storing identification information 2 by writing multiple storage files of identification information 2 into a distributed storage system. The storage method of storing files with identification information 2 is consistent with the above method and will not be repeated here.
[0232] In this embodiment, a file storage device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0233] This embodiment provides a file storage device, such as Fig.14 As shown, it includes: an acquisition module 1401, a generation module 1402, a processing module 1403, and an identification module 1404.
[0234] An acquisition module 1401 is used to acquire a first file to be stored and first identification information corresponding to a file to which the first file to be stored belongs;
[0235] A generating module 1402, configured to generate first metadata information corresponding to a first file to be stored;
[0236] Processing module 1403, used to determine a transmission protocol for transmitting the first file to be stored;
[0237] When it is determined that the transmission protocol is the target protocol, determining whether there is a file corresponding to the first identification information in the distributed storage system; when it is determined that there is no file corresponding to the first identification information in the distributed storage system, creating a logical header file in the distributed storage system;
[0238] Identification module 1404, used to identify whether the multi-version function is enabled;
[0239] The generating module 1402 is further configured to generate version information corresponding to the first file to be stored and configure the version information as the first version information when it is determined that the multi-version function is enabled;
[0240] The processing module 1403 is further used to add the first metadata information, the first identification information, and the first version information into the logical header file;
[0241] And the storage operation of the first file to be stored is completed in the distributed storage system.
[0242] In an optional embodiment, the acquisition module 1401 is specifically configured to generate second metadata information corresponding to the second file to be stored when the second file to be stored and the second identification information corresponding to the second file to be stored are acquired;
[0243] The processing module 1403 is further used to search for identification information matching the second identification information from the logical header file created in the distributed storage system;
[0244] The generating module 1402 is further configured to generate new version information corresponding to the second file to be stored according to the version information stored in the first logical header file when it is determined that the identification information matching the second identification information is stored in the first logical header file, wherein the first logical header file is any one of the logical header files created in the distributed storage system;
[0245] The processing module 1403 is further used to store the new version information and the second metadata information into the first logical header file;
[0246] The second file to be stored is stored in the distributed storage system.
[0247] In an optional embodiment, the creation module 1405 is specifically used to create an index file and declaration information corresponding to the index file;
[0248] The processing module 1403 is further used to store the historical metadata information stored in the first logical header file before storing the second metadata information, and the version information of the storage file corresponding to the historical metadata information, into the index file;
[0249] In an optional embodiment, the processing module 1403 is further used to store declaration information corresponding to the index file in the first logical header file, where the declaration information is used to indicate a storage location corresponding to the index file.
[0250] The acquisition module 1401 is further used to acquire a third file to be stored and third identification information corresponding to the third file to be stored;
[0251] The generating module 1402 is further used to generate third metadata information corresponding to the third file to be stored;
[0252] The search module 1406 is further configured to search for identification information matching the third identification information from a logical header file created in the distributed storage system;
[0253] The identification module 1404 is further used to identify whether the version information stored in the second logical header file is the target version information when it is determined that the identification information matching the third identification information is stored in the second logical header file and it is determined that the acquisition time of the third file to be stored is within a preset time range, wherein the second logical header file is any one of the logical header files created in the distributed storage system, and the preset time range is used to indicate that the multi-version function is in a stopped state;
[0254] The processing module 1403 is further configured to, when it is determined that the version information stored in the second logical header file is the target version information, overwrite the file to be stored corresponding to the target version information with the third file to be stored in the distributed storage system, and update the version information of the third file to be stored to the target version information;
[0255] The third metadata information and the target version information are updated into the second logical header file.
[0256] In an optional embodiment, the processing module 1403 is further configured to, when it is determined that the version information stored in the second logical header file is not the target version, determine the version information of the third file to be stored as the target version information after the third file to be stored is distributedly stored in the distributed storage system;
[0257] The third metadata information and the target version information are updated into the second logical header file.
[0258] In an optional embodiment, the feedback module 1407 is specifically used to feedback indication information of the existence of the file to the transmission end of the first file to be stored when it is determined that there is a file corresponding to the first identification information in the distributed storage system, so as to prompt the transmission end to determine whether the first file to be stored is a duplicate upload.
[0259] The file storage device in this embodiment is presented in the form of a functional module, where the module refers to an application specific integrated circuit (ASIC), a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0260] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0261] An embodiment of the present invention provides a file storage device, which determines whether there is a file corresponding to the first identification information in a distributed storage system by obtaining a first file to be stored and first identification information corresponding to a file to which the first file to be stored belongs. When it is determined that there is no file corresponding to the first identification information in the distributed storage system, a logical header file is created in the distributed storage system as the overall external representation of multiple versions of storage files, and the multi-version state is directly selected and switched according to the usage scenario through the logical header file, thereby improving the usability of the storage system. When it is determined that the multi-version function is in the on state, after configuring the version information as the first version information, the first metadata information, the first identification information, and the first version information are added to the logical header file. The device adds the metadata information of the first file to be stored to the logical header file, so that the logical header file and the metadata information file are merged into one file, thereby reducing one metadata information file and increasing the supportable file level of the entire cluster.
[0262] In the present embodiment, a file reading device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0263] This embodiment provides a file reading device, such as Fig.15 As shown, it includes: an acquisition module 1501, a search module 1502, an extraction module 1503, a processing module 1504, and a sending module 1505.
[0264] An acquisition module 1501 is used to acquire a file reading request, where the file reading request includes identification information of a file to be read;
[0265] A search module 1502 is used to search for a target logical header file storing the identification information from the logical header files created by the distributed storage system according to the identification information;
[0266] An extraction module 1503 is used to extract address indication information corresponding to the file to be read from the target logical header file;
[0267] The processing module 1504 is used for acquiring the file to be read according to the metadata information when the address indication information is metadata information corresponding to the file to be read;
[0268] The sending module 1505 is used to feed back the file to be read to the requester who sends the file reading request.
[0269] In an optional embodiment, the extraction module 1503 is further used to extract declaration information from the target logic header file when the target version information is not the version information stored in the target logic header file, wherein the declaration information is the address indication information;
[0270] The acquisition module 1501 is further used to acquire the index file corresponding to the declaration information according to the declaration information;
[0271] The extraction module 1503 is further used to extract metadata information of the file to be read corresponding to the target version information from the index file;
[0272] The acquisition module 1501 is also used to acquire the file to be read according to the metadata information;
[0273] The sending module 1505 is specifically configured to feed back the file to be read to the requesting party that sends the file reading request.
[0274] The file reading device in this embodiment is presented in the form of a functional module, where the module refers to an application specific integrated circuit (ASIC), a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0275] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0276] A file reading device provided by an embodiment of the present invention searches for a target logical header file storing identification information from a logical header file created by a distributed storage system according to the identification information, extracts address indication information corresponding to a file to be read from the target logical header file, and when the file version information is not specified in the file reading request, the address indication information is metadata information corresponding to the file to be read stored in the logical header file, and the storage location of the file to be read is quickly located directly according to the metadata information stored in the logical header file, thereby reducing additional search steps, thereby improving the efficiency of file retrieval and access, and simplifying the file retrieval process.
[0277] An embodiment of the present invention further provides a computer device, Fig.16 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Fig.16As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Fig.16 A processor 10 is taken as an example.
[0278] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include an integrated circuit. The integrated circuit may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0279] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.
[0280] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of a computer device based on the presentation of a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0281] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0282] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Fig.16 The example of connecting through bus is taken in the following.
[0283] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0284] The embodiment of the present invention also provides a computer-readable storage medium. The method provided in the above embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or is implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0285] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.
[0286] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A file storage method, characterized in that: The method is applied to a distributed storage system, and the method comprises: Acquire a first file to be stored and first identification information corresponding to a file to which the first file to be stored belongs; Generating first metadata information corresponding to the first file to be stored; Determining a transmission protocol for transmitting the first file to be stored; When it is determined that the transmission protocol is the target protocol, determining whether there is a file corresponding to the first identification information in the distributed storage system; When it is determined that no file corresponding to the first identification information exists in the distributed storage system, creating a logical header file in the distributed storage system; Identify whether the multi-version function is enabled, and when it is determined that the multi-version function is enabled, generate version information corresponding to the first file to be stored, and configure the version information as the first version information; Adding the first metadata information, the first identification information, and the first version information to the logical header file; And completing the storage operation of the first file to be stored in the distributed storage system.
2. The method according to claim 1, characterized in that The method further comprises: When the second file to be stored and the second identification information corresponding to the second file to be stored are acquired, generating second metadata information corresponding to the second file to be stored; Searching for identification information matching the second identification information from a logical header file created in the distributed storage system; When it is determined that identification information matching the second identification information is stored in the first logical header file, new version information corresponding to the second file to be stored is generated according to the version information stored in the first logical header file, wherein the first logical header file is any one of the logical header files created in the distributed storage system; storing the new version information and the second metadata information in the first logical header file; The second file to be stored is stored in the distributed storage system.
3. The method according to claim 2, characterized in that After storing the new version information and the second metadata information in the first logical header file, the method further includes: Creating an index file and declaration information corresponding to the index file; storing the historical metadata information stored in the first logical header file before storing the second metadata information, and the version information of the storage file corresponding to the historical metadata information, in the index file; The declaration information corresponding to the index file is stored in the first logical header file, where the declaration information is used to indicate a storage location corresponding to the index file.
4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Acquire a third file to be stored and third identification information corresponding to the third file to be stored; Generating third metadata information corresponding to the third file to be stored; Searching for identification information matching the third identification information from a logical header file created in the distributed storage system; When it is determined that identification information matching the third identification information is stored in the second logical header file, and it is determined that the acquisition time of the third file to be stored is within a preset time range, identifying whether the version information stored in the second logical header file is the target version information, wherein the second logical header file is any one of the logical header files created in the distributed storage system, and the preset time range is used to indicate that the multi-version function is in a stopped state; When it is determined that the version information stored in the second logical header file is the target version information, in the distributed storage system, the third file to be stored overwrites the file to be stored corresponding to the target version information, and updates the version information of the third file to be stored to the target version information; The third metadata information and the target version information are updated into the second logical header file.
5. The method according to claim 4, characterized in that The method further comprises: When it is determined that the version information stored in the second logical header file is not the target version information, after the third file to be stored is distributedly stored in the distributed storage system, the version information of the third file to be stored is determined as the target version information; The third metadata information and the target version information are updated into the second logical header file.
6. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: When it is determined that a file corresponding to the first identification information exists in the distributed storage system, information indicating that the file already exists is fed back to the transmission end of the first file to be stored, so as to prompt the transmission end to determine whether the first file to be stored is a duplicate upload.
7. A file reading method, characterized in that: The method comprises: Obtaining a file reading request, wherein the file reading request includes identification information of a file to be read; According to the identification information, searching for a target logical header file storing the identification information from a logical header file created by the distributed storage system according to any one of claims 1 to 6; Extracting address indication information corresponding to the file to be read from the target logical header file; When the address indication information is metadata information corresponding to the file to be read, acquiring the file to be read according to the metadata information; The file to be read is fed back to the requesting party that sends the file reading request.
8. The method according to claim 7, characterized in that The file read request also includes target version information, and the method further includes: When the target version information is not the version information stored in the target logic header file, extracting declaration information from the target logic header file, wherein the declaration information is the address indication information; According to the declaration information, obtaining an index file corresponding to the declaration information; Extracting metadata information of a to-be-read file corresponding to the target version information from the index file; Acquire the file to be read according to the metadata information; The file to be read is fed back to the requesting party that sends the file reading request.
9. A file storage device, characterized in that: The device comprises: An acquisition module, configured to acquire a first file to be stored and first identification information corresponding to a file to which the first file to be stored belongs; A generating module, used to generate first metadata information corresponding to the first file to be stored; A processing module, used to determine a transmission protocol for transmitting the first file to be stored; When it is determined that the transmission protocol is the target protocol, determining whether there is a file corresponding to the first identification information in the distributed storage system; when it is determined that there is no file corresponding to the first identification information in the distributed storage system, creating a logical header file in the distributed storage system; Identification module, used to identify whether the multi-version function is enabled; The generating module is further configured to generate version information corresponding to the first file to be stored and configure the version information as the first version information when it is determined that the multi-version function is enabled; The processing module is further used to add the first metadata information, the first identification information, and the first version information into the logical header file; And completing the storage operation of the first file to be stored in the distributed storage system.
10. A file reading device, characterized in that: The device comprises: An acquisition module, used for acquiring a file reading request, wherein the file reading request includes identification information of a file to be read; A search module, configured to search, according to the identification information, a target logical header file storing the identification information from the logical header files created by the distributed storage system according to any one of claims 1 to 6; An extraction module, used for extracting address indication information corresponding to the file to be read from the target logical header file; a processing module, configured to obtain the file to be read according to the metadata information when the address indication information is metadata information corresponding to the file to be read; The sending module is used to feed back the file to be read to the requesting party that sends the file reading request.
11. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the file storage method according to any one of claims 1 to 6, or executes the file reading method according to claim 7 or 8 by executing the computer instructions.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the file storage method described in any one of claims 1 to 6, or to execute the file reading method described in claim 7 or 8.