Data Management Method, Device, Medium and Product for Distributed Storage System
By presetting storage pools of different fault-tolerant types in a distributed storage system, dynamically analyzing the target fault-tolerant types of files, and performing data transfer operations, the problem of degradation of distributed storage systems in the existing technology is solved, and more efficient storage management and performance improvement is achieved.
Patent Information
- Application Number
- CN202510378616.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The existing distributed storage system with a single fault-tolerant mode has been declining performance when facing large-scale storage and data attribute changes, making it difficult to effectively manage and optimize storage resources.
By presetting storage pools of different fault-tolerant types in the distributed storage system, and dynamically analyze the target fault-tolerant type of the file when the data attributes change, perform data transfer operations, and transfer files to the matching target storage pool.
It realizes dynamic adjustment of the storage pool according to the current data attributes and access records of the file, reducing storage costs and improving the performance and stability of the storage system.
Smart Images

Figure CN119883142B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage technology, and particularly to a data management method, device, medium and product for a distributed storage system. Background Art
[0002] With the rapid development of big data and artificial intelligence technologies, the total global data volume has grown exponentially. To achieve the storage management of big data, distributed storage systems have begun to be widely used.
[0003] Based on network technology, a distributed storage system effectively shares the storage load by dispersing data storage across multiple independent physical storage devices. The storage scale of large distributed storage systems is getting larger and the data attributes are constantly changing. Therefore, the performance of existing distributed storage systems with a single fault tolerance mode is continuously declining. Summary of the Invention
[0004] This application provides a data management method, device, medium and product for a distributed storage system to at least solve related technical problems.
[0005] This application provides a data management method for a distributed storage system. The distributed storage system includes a management layer and an application layer, and multiple storage pools with different fault tolerance types are set in the distributed storage system. The method includes:
[0006] In response to the triggering of a preset condition, the management layer obtains the first metadata of the first specified file to determine the current data attributes of the first specified file;
[0007] In response to detecting one or more attribute changes in the current data attributes of the first specified file, the management layer determines the target fault tolerance type of the first specified file according to the access records in the current data attributes and preset rules;
[0008] In response to detecting a change in the target fault tolerance type, the management layer performs a data transfer operation on the first specified file to transfer the first specified file to a matching target storage pool.
[0009] This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above data management methods for a distributed storage system are implemented.
[0010] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any of the above data management methods for a distributed storage system are implemented.
[0011] Based on presetting storage pools with different fault tolerance types in a distributed storage system, this application analyzes the stored first specified file, timely obtains the best fault tolerance type of the specified file after the data attributes change, and further transports the file according to this fault tolerance type to reasonably store the file in the storage pool with the matching fault tolerance type, reducing the storage cost of the storage system while improving the performance of the storage system. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0013] Figure 1 Schematic diagram of a distributed storage system architecture provided by an embodiment of this application;
[0014] Figure 2 Schematic diagram of a data management method provided by an embodiment of this application;
[0015] Figure 3 Schematic diagram of another data management method provided by an embodiment of this application;
[0016] Figure 4 Schematic diagram of a file concept provided by an embodiment of this application;
[0017] Figure 5 Schematic diagram of a data fingerprint array provided by an embodiment of this application;
[0018] Figure 6 Schematic diagram of a data management device provided by an embodiment of this application;
[0019] Figure 7 Schematic diagram of an electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. According to the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of this application.
[0021] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and not to describe a specific order or sequence.
[0022] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be made in conjunction with the accompanying drawings and specific embodiments.
[0023] Regarding the data management method of the distributed storage system disclosed in this application, Figure 1 A typical distributed storage system architecture is shown. A distributed storage system generally includes an application layer, a protocol layer, a management layer, and a hardware layer. Among them, the application layer is usually a program facing users, which can specifically be a client, an application program, or a virtual machine, and the specific form of the application layer is not limited in this application. The protocol layer mainly meets the access requirements of different user protocols, and a protocol library designed by those skilled in the art is deployed therein. This protocol library provides one or more sets of API interfaces that can support different interface protocols and allows the application layer to interact with the distributed storage system through this protocol library. The management layer is the core of the distributed storage system, mainly including a monitor service module, a management service module, an object storage service module, a metadata service module, etc. Among them, the monitor service module is mainly used to maintain the state of the storage cluster, such as how many storage pools are in the cluster, how many data groups are in each storage pool, and the mapping relationship between the storage pool and the data group, etc. The management service module is responsible for managing the information of the cluster and tracking the runtime metrics and current status. There are multiple object storage service nodes in the object storage service module, and each object storage service node manages a disk, responsible for tasks such as handling cluster data replication, recovery, and rebalancing; and provides information to the monitor service module and the management service module by checking the daemon state of other object storage service nodes. The metadata service is mainly used to manage the metadata of the cluster. The storage pool is a logical storage unit of the distributed storage system, used to organize and manage data, and is mapped to specific data groups through the map layer (mapping layer). The data group is a physical data distribution unit, used to disperse data to different object storage service nodes. By splitting the data into multiple parts and distributing them on different object storage service nodes, load balancing is achieved, and the reliability and access efficiency of the data are improved. The hardware layer is composed of physical disks that actually provide storage services.
[0024] It can be understood that in some specific application scenarios, the data management method disclosed in this application can also be adaptively adjusted and applied to a storage cluster, and only need to deploy a corresponding management module in the storage system to execute the data management method disclosed in this application.
[0025] An embodiment of this application provides a data management method for a distributed storage system. It should be noted that different fault-tolerant type storage pools are pre-set in the distributed storage system disclosed in this application. In a specific implementation scenario, it is preferably possible to set two types of fault-tolerant storage pools, such as a multi-copy storage pool and an erasure code storage pool. Of course, in other implementation scenarios, other types of fault-tolerant storage pools can also be set, such as a data sharding fault-tolerant type storage pool; this application does not limit the number and type of specific fault-tolerant types.
[0026] Such as Figure 2 As shown, the above data management method specifically includes the following steps:
[0027] S1. In response to the triggering of a preset condition, the management layer obtains the first metadata of the first specified file to determine the current data attributes of the first specified file.
[0028] In a specific implementation scenario, the above-mentioned preset conditions for response include detecting a file reading request issued by the application layer and reaching a certain period; specifically, the management layer obtains the first metadata of the first specified file to determine the current data attributes of the first specified file in response to detecting a file reading request issued by the application layer; and / or, the management layer obtains the first metadata of the first specified file to determine the current data attributes of the first specified file in response to the time difference between the currently collected time node and the previously collected historical time node being greater than or equal to a preset time difference. Among them, determining the current data attributes of the first specified file according to the first metadata of the first specified file means that the object information, access records, file identifiers (including file paths, file names), and any other information used to describe file attributes included in the metadata of the file are all regarded as current data attributes.
[0029] Optionally, the above-mentioned preset time difference can be set to 30 minutes, 2 hours, etc. The specific preset time difference is set by those skilled in the art according to the scale and performance of the storage system, and this application does not limit it. It should be noted that the above two preset conditions can be selected and used in the storage system, or both can be used. Based on the different triggering mechanisms set above, it is possible to achieve a comprehensive analysis of the files in the storage system, so as to achieve reasonable storage of all files and improve system stability; it is also possible to specifically analyze specific files and save the computing resources of the storage system. Therefore, it can be understood that the above monitoring of the current data attributes can occur at any moment when the distributed storage system is working.
[0030] Explanation of the definition of the first specified file: When responding to a file reading request, the above-mentioned first specified file is the file that matches the file identifier such as the file name or file path included in the file reading request sent by the application layer; when the response time difference is greater than or equal to the preset time difference, the above-mentioned first specified file can be set to all files in the storage system.
[0031] Preferably, the first metadata information includes object information, fault tolerance type, and the most recent ten access records that make up the first specified file; the object information includes a combination of the first object identifier, segmentation information, and the hash value of the object data. It can be understood that in this application, "first" and "second" are only used for distinction. For example, the first specified file and the second specified file both belong to files; as well as the first metadata and the second metadata.
[0032] S2. In response to detecting that one or more attribute changes exist in the current data attributes of the first specified file, the management layer determines the target fault tolerance type of the first specified file according to the access records in the current data attributes and the preset rules.
[0033] It can be understood that as long as any change occurs to the file, that is, any one of the attributes in the data attributes changes, it will trigger the next step of determining the target fault tolerance type. In the method disclosed in the embodiments of this application, the management layer determines whether the current data attributes are consistent with the historical data attributes obtained when responding to the preset conditions according to the current data attributes of the first specified file and the historical data attributes. If they are consistent, it is determined that the current data attributes have not changed. If there is a change in one data attribute, it triggers the next step of determining the target fault tolerance type of the first specified file.
[0034] In the specific implementation scenarios of setting the erasure code fault tolerance storage pool and multi-copy fault tolerance storage, the above-mentioned determination of the target fault tolerance type of the first specified file according to the access records and the preset rules can be achieved by calculating the file heat value. Specifically, the management layer calculates the file heat value according to the access records included in the first metadata corresponding to the first specified file and the preset formula; the management layer determines the target fault tolerance type of the first specified file according to the file heat value.
[0035] Among them, the above-mentioned preset formula is expressed as:
[0036] ;
[0037] Among them, Hot is the file heat value, i is the serial number of the access record, n represents the maximum number of access records matched by the first specified file, T cur represents the current time, T i is the most recenti access time α i indicating i the weight coefficient of the i -th access record. The weight coefficient can be generated by training a weight coefficient determination model based on historical access data through the trained model; or it can be determined by those skilled in the art according to empirical values, and this application does not make any limitations in this regard.
[0038] Among them, the above management layer determines the target fault tolerance type of the first specified file according to the file heat value, including: when the management layer detects that the file heat value is less than or equal to the preset threshold, that is, the heat attribute of the file is hot data, it determines that the target fault tolerance type of the first specified file is the first fault tolerance type (i.e., multi-copy fault tolerance); when the management layer detects that the file heat value is greater than the preset threshold, that is, the heat attribute of the file is cold data, it determines that the target fault tolerance type of the first specified file is the second fault tolerance type (i.e., erasure code fault tolerance). That is, it is defined that the activity of the file stored in the storage pool matching the first fault tolerance type is higher than that of the file stored in the storage pool matching the second fault tolerance type. The preset threshold is set by those skilled in the art according to the actual situation, and this application does not make any limitations in this regard. Of course, if more than two fault tolerance types are set, it can be adaptively adjusted by adding a preset threshold. At this time, it is divided into the first threshold and the second threshold. The target fault tolerance type of the first specified file whose file heat value is greater than the first threshold is the first fault tolerance type, the target fault tolerance type of the first specified file whose file heat value is less than the first threshold and greater than the second threshold is the second fault tolerance type, and the target fault tolerance type of the first specified file whose file heat value is less than the second threshold is the third fault tolerance type.
[0039] This application discloses a specific method for determining the target fault tolerance type of the first specified file, that is, by setting the calculation formula of the heat value, determining the file heat value of different files according to the access record, that is, using the heat value to distinguish the access busy degree of different files, and further improving the accuracy of confirming the best target fault tolerance type of the file.
[0040] Of course, in the implementation scenario of setting storage pools with other fault tolerance types, different rules can also be designed to determine the best target fault tolerance type of each first specified file. For example, by combining the file heat value and the storage cost, calculating the index scores of different files by setting different weights, further setting corresponding thresholds, and determining that the target fault tolerance type of the file whose index score exceeds a specific threshold is the first fault tolerance type, and the target fault tolerance type of the file whose index score does not exceed the specific threshold is the second fault tolerance type; of course, it is also possible to calculate the index scores of different files by combining factors such as data sensitivity and network latency. This application does not impose any constraints on the specific preset rules.
[0041] S3. In response to detecting a change in the target fault tolerance type, the management layer performs a data transfer operation on the first specified file to transfer the first specified file to a matching target storage pool.
[0042] Whether the target fault tolerance type changes is achieved by comparing the target fault tolerance type with the latest fault tolerance type recorded in the first metadata matching the first specified file. If the target fault tolerance type is the same as the latest fault tolerance type, it is determined that the target fault tolerance type has not changed and no data transfer is required; if the target fault tolerance type is different from the latest fault tolerance type, the management layer determines that the target fault tolerance type has changed. The above-mentioned management layer, in response to detecting a change in the target fault tolerance type, performs a data transfer operation on the first specified file to transfer the first specified file to a matching target storage pool, including:
[0043] In response to detecting that the target fault tolerance type of the first specified file changes from the second fault tolerance type to the first fault tolerance type, the management layer performs a data transfer operation on the first specified file to transfer the first specified file to a storage pool matching the first fault tolerance type; in response to detecting that the fault tolerance type of the first specified file changes from the first fault tolerance type to the second fault tolerance type, the management layer obtains the object distribution of the first object included in the first specified file; the management layer determines whether to perform a data transfer operation on the first specified file according to the object distribution, and the object distribution includes single-file distribution and multi-file distribution. That is, when the fault tolerance type changes, it is necessary to consider the fault tolerance types of other files in the object distribution. Only when the fault tolerance types of the first specified file and all files in all object distributions included in the first specified file are consistent with the target fault tolerance type of the first specified file, will the first specified file be transferred to a storage pool matching the target fault tolerance type (i.e., the target storage pool). As long as there is a file with inconsistent fault tolerance type in the object distribution, no data transfer will be performed.
[0044] The above-mentioned management layer determines whether to perform a data transfer operation on the first specified file according to the object distribution, including: in response to detecting that the object distribution is single-file distribution, the management layer transfers the first specified file to a storage pool matching the second fault tolerance type; in response to detecting that the object distribution is multi-file distribution, the management layer obtains the target fault tolerance type of the remaining distributed files other than the first specified file in the first object distribution; if the target fault tolerance type of the remaining distributed files is the second fault tolerance type, the management layer performs a data transfer operation on the first specified file to transfer the first specified file to a storage pool matching the second fault tolerance type; if there is a target fault tolerance type of the remaining distributed files that is the first fault tolerance type, the management layer does not perform a data transfer operation on the first specified file.
[0045] In the specific scenario of setting up a multi-copy fault-tolerant storage pool and an erasure code fault-tolerant storage pool, the above solution specifically includes: when the management layer detects that the target fault-tolerance type of the first specified file changes from erasure code fault-tolerance to multi-copy fault-tolerance, it performs a data transfer operation on the first specified file to transfer the first specified file to the multi-copy storage pool; when the management layer detects that the fault-tolerance type of the first specified file changes from multi-copy fault-tolerance to erasure code fault-tolerance, it obtains the object distribution of the first object included in the first specified file; the management layer determines whether to perform a data transfer operation on the first specified file according to the object distribution, and the object distribution includes single-file distribution and multi-file distribution. When the management layer detects that the object distribution is single-file distribution, it transfers the first specified file to the erasure code storage pool; when the management layer detects that the object distribution is multi-file distribution, it obtains the target fault-tolerance type of the remaining distributed files other than the first specified file in the first object distribution; if the target fault-tolerance type of the remaining distributed files is erasure code fault-tolerance, the management layer performs a data transfer operation on the first specified file to transfer the first specified file to the erasure code storage pool; if there is a target fault-tolerance type of the remaining distributed files that is multi-copy fault-tolerance, the management layer does not perform a data transfer operation on the first specified file.
[0046] In the embodiments of the present application, based on presetting storage pools with different fault-tolerance types in a distributed storage system, the fault-tolerance type of the stored first specified file is analyzed in a timely manner, and the file is reasonably stored in a storage pool with a matching fault-tolerance type, which reduces the storage cost of the storage system and improves the performance of the storage system at the same time.
[0047] Further, as Figure 3 shown, a data management method disclosed in the embodiments of the present application further includes the following steps:
[0048] X1. When the management layer detects a file write request sent by the application layer, it feeds back the second metadata of the second specified file included in the file write request to the application layer.
[0049] For easy understanding, first, some concepts involved in the present application are briefly described. Such as Figure 4As shown, the file is the file that the user needs to store or access; the object is the unit managed inside the distributed storage system, and the maximum disk space occupied by each object is set by the system; when the file exceeds the preset disk space occupancy (usually set as the maximum disk space occupancy), the storage system will split the file into multiple objects for storage. The objects are mapped to different object storage services through data groups. The file is usually split according to the preset disk space occupancy, and the first mapping relationship between the file and the split objects is saved in the metadata management service module; each object is independently mapped to a unique data group, and there is also a second mapping relationship between the data group and the object storage service module. Through a data group, it will be mapped to multiple object storage service modules, and this second mapping relationship will also be stored in the metadata service module. Preferably, the above-mentioned preset disk space occupancy can be the maximum disk space occupancy, and of course, it can also be set by those skilled in the art according to the actual situation.
[0050] Specifically, the composition of the above-mentioned second metadata is the same as that of the first metadata, both including object information, fault tolerance type, and the recent ten access records; for the convenience of description, the object included in the second metadata is defined as the second object. It can be determined that a second specified file includes at least one second object, so the second object information in the second metadata includes at least one; the second object information includes segmentation information (that is, the start position and end position of each second object in the file), the second object identifier, and the combined target hash value corresponding to the second object data.
[0051] The generation of the matching second metadata according to the second specified file included in the file write request specifically includes:
[0052] The management layer splits the second specified file according to the preset disk space occupancy to generate segmentation information and multiple second objects and assigns a unique second object identifier to the second objects; consistent with the embodiments disclosed above, in this embodiment, the above-mentioned preset disk space occupancy is preferably set as the maximum disk space occupancy, and it can also be set by those skilled in the art according to the actual situation. This application does not make any limitations in this regard. The management layer determines the fault tolerance type according to the access type of the file write request, and the access type includes the file modification type and the file creation type.
[0053] Among them, the determination of the fault tolerance type according to the access type of the file write request includes that in response to detecting that the access type is the file creation type, it is determined that the fault tolerance type is the first fault tolerance type; if the file write request is to create a new file and the file activity is high, the created second specified file can be assigned to the storage pool matching the first fault tolerance type. For example, in a scenario where the first fault tolerance type is multi-copy fault tolerance and the second fault tolerance type is erasure code fault tolerance, it is determined that the fault tolerance type of the second specified file is multi-copy fault tolerance and it is assigned to the multi-copy storage pool.
[0054] In response to detecting that the access type is to modify the file type, obtain the historical storage pool address of the file to be modified that matches the second specified file; if the historical storage pool address is the first preset value, determine that the fault tolerance type is the first fault tolerance type; if the historical storage pool address is the second preset value, determine that the fault tolerance type is the second fault tolerance type. The mapping relationship between the preset value and the fault tolerance type is preset and stored. In this embodiment, the first preset value is bound to the first fault tolerance type, and the second preset value is bound to the second fault tolerance type. In other embodiments, replacements can also be made. In embodiments with more than two fault tolerance types, preset values can also be added. This application will not elaborate one by one here.
[0055] That is, if the file write request is to modify an existing file (the file to be modified is the second specified file), then it is necessary to verify the historical storage pool address of the file to be modified; keep the second storage pool address corresponding to the second specified file consistent with the historical storage pool address of the file to be modified, that is, keep the fault tolerance type of the second specified file consistent with the fault tolerance type of the file to be modified. This application determines different fault tolerance types corresponding to different files according to the access type of different file write requests, and ensures the consistency of the file fault tolerance type when responding to requests to modify files, reducing the possibility of file anomalies.
[0056] X2. The application layer performs data repeatability verification on the second specified file based on the second metadata and the preset data fingerprint array, and generates a data write request and sends it to the management layer after the data repeatability verification passes.
[0057] To avoid writing duplicate data in the storage system, this embodiment also proposes to perform data repeatability verification on the second specified file, including:
[0058] The application layer obtains and determines multiple second object data included in the second specified file according to the segmentation information of the second specified file; the application layer calculates the target hash value combination of the second object data; the application layer searches in the preset data fingerprint array to find whether there is an existing hash value combination that matches the target hash value combination; if not, the application layer determines that the data repeatability verification passes; if so, the application layer performs secondary data repeatability verification according to the second object information and the existing object information that matches the existing hash value combination.
[0059] The above application layer performs secondary data duplication verification based on the second object information and the existing object information matched by the existing hash value combination, including: when the application layer detects that the segmentation information in the second object information is inconsistent with the segmentation information in the existing object information matched by the existing hash value combination, it determines that the data duplication verification passes; when the application layer detects that the segmentation information in the second object information is consistent with the segmentation information in the existing object information matched by the existing hash value combination, it determines that the data duplication verification fails, and updates the second object identifier matched by the second object according to the target object identifier included in the object information matched by the existing hash value combination. The segmentation information is obtained by the management layer from the second metadata matched with the second specified file in the metadata service module and sent to the application layer.
[0060] That is, if there is no existing hash value combination in the preset data fingerprint array that matches the target hash value combination, it is determined that the object data of the second specified file is not stored in the storage system. At this time, the application layer generates a data write request for the second specified object storage service node according to the second metadata and sends it to the management layer. If there is an existing hash value combination in the preset data fingerprint array that matches the target hash value combination, it is further determined whether the object data is stored in the storage system according to the segmentation information included in the second metadata matched with the second specified file. If they are inconsistent, it is determined that the object data is not stored in the storage system. At this time, it is still necessary to determine the second specified object storage service node matched with the second specified file, generate a data write request and send it to the second specified object storage service node. If they are consistent, the application layer sends the object information corresponding to the existing hash value that matches the target hash value combination to the metadata service module of the management layer through the protocol, updates the second object identifier in the second metadata of the second specified file to the target object identifier, and determines that the data of the object has been written into the storage system. At this time, there is no need for duplicate storage, and no write storage request is sent to the object storage service.
[0061] Based on the preset data fingerprint array, it is verified whether the object data in the file to be written is repeated. For the repeated object data, only the object identifier in the object information in the file metadata needs to be updated, so that the address where the repeated object data has been stored can be obtained according to the object data, reducing the storage amount of duplicate data in the distributed storage system and reducing the file storage interaction operation, greatly improving the performance of the distributed storage system and reducing the storage pressure.
[0062] The application layer calculates the target hash value combination of the second object data; specifically, the preset first hash algorithm and the preset second hash algorithm are used to calculate the hash values of the same second object data respectively, so as to obtain the first hash value and the second hash value with different numerical values, and then the first hash value and the second hash value are combined to generate the target hash value combination, which is then saved in the metadata information. It can be determined that one object data corresponds to one hash value combination. The present application does not limit the specific types of the first hash algorithm and the second hash algorithm, which can be MD5 (Message Digest Algorithm 5), MurmurHash (non-cryptographic hash function), etc.
[0063] The application layer searches in the preset data fingerprint array to find whether there is an existing hash value combination that matches the target hash value combination. The above data fingerprint array is a two-dimensional data fingerprint array set by the distributed storage system for each storage pool of the fault tolerance type; as Figure 5 shown, one data unit corresponds to one object, and its row index and column index are respectively the first hash value and the second hash value generated according to the object data. The first hash value and the second hash value corresponding to a data unit in the data fingerprint array are called the existing hash value combination; at least object information is stored in this data unit. Of course, for the convenience of subsequent other operations, other information can also be stored in the data unit of the data fingerprint array. It can be understood that since the first hash algorithm and the second hash algorithm are different, the possibility that different object data generate the same first hash value and the same second hash value is very low. In the present application, it is determined that only the same object data can generate the same first hash value and the same second hash value.
[0064] In addition, the application layer determines the target object storage service node based on the second metadata of the received second specified file, and then generates a data write request and sends it to the target object storage service node in the management layer. That is, the present application further determines the corresponding object storage service by searching the metadata of the file to be written, and then the application layer directly sends a data write request to the object storage service, reducing the data access time and realizing the fast writing of the stored data. Specifically, the determination of the target storage service node includes:
[0065] Determine at least one second data address that matches the second specified file according to the second object identifier and fault tolerance type included in the second metadata. According to the data address and the preset data distribution algorithm, determine the second specified object storage service node that matches the second data address; generate a data write request and send it to the second specified object storage service node. That is, after obtaining the second data address, the application layer obtains the second specified object storage service node that matches the second data address according to the preset data distribution algorithm, then sends a data write request to the second object storage service node, and splits the data of the second specified file into data blocks, and then transfers them to the corresponding second specified object storage service node one by one.
[0066] Specifically, the above determination of at least one second data address that matches the second specified file according to the object identifier and fault tolerance type included in the second metadata includes: calling a preset hash function to calculate the transition hash value corresponding to the second object identifier, so as to convert the object identifier into a fixed-length value. The hash function includes but is not limited to non-cryptographic hash functions, cryptographic hash functions, and distributed storage specific hash algorithms. Specifically, hash functions that can be used include MurmurHash (non-cryptographic hash function) and Jenkins hash (Jenkins hash function), etc.; the present application does not limit the selection of specific hash functions. Perform a bitwise AND operation on the transition hash value and the bit mask to generate the second data group address corresponding to the second object identifier, that is, the second data group address PGid = Hash(Oid)&mask, that is, use a bitwise AND operation on the basis of the hash value in combination with a bit mask (mask) to limit the range of the hash value. The application layer can also use methods such as heterogeneous acceleration and pipeline processing to reduce the time of hash calculation and improve the system response performance.
[0067] Then determine the storage pool address according to the fault tolerance type. In the scenarios of the storage pool with multi-copy fault tolerance type and the erasure code fault tolerance type disclosed above, in response to detecting that the fault tolerance type is multi-copy fault tolerance, determine that the first storage pool address is the first preset value; in response to detecting that the fault tolerance type is erasure code fault tolerance, determine that the first storage pool address is the second preset value. Among them, preferably, the first preset value can be set to 1, and the second preset value can be set to 2. Of course, those skilled in the art can also set other values. It should be noted that the first preset value is bound to the multi-copy fault tolerance type, and the second preset value is bound to the erasure code fault tolerance type. Of course, in the case of integrating storage pools with more than two fault tolerance types, a third preset value can be further set and bound to the third fault tolerance type. In this way, the present application does not expand on this.
[0068] Finally, according to the second data group address and the second storage pool address, a second data address is generated. That is, a complete second data address can be expressed as "the first storage pool address: PGid". Through the above steps, the storage address corresponding to the object after the file is partitioned in this application is limited within a specific range, improving the stability of the system storage; on this basis, the final data address of a specific file object can be directly generated in one step according to the object identifier and the fault tolerance type, so that the corresponding data addresses can be generated for multiple fault tolerance types, which is convenient for the subsequent application layer to issue a data read request according to the data address to achieve fast data reading.
[0069] The above data distribution algorithm is used to establish the mapping relationship between the data address and the object storage service node to ensure data reliability and efficient access. Among them, in this embodiment, the data distribution algorithm includes but is not limited to the consistent hashing algorithm, the CRUSH algorithm, the random allocation method, etc. The above consistent hashing algorithm forms a virtual ring with all possible hash values and maps each object storage service node to a certain position on this ring; when a new data group needs to be allocated, first calculate a hash value based on the data group address, and then find the corresponding point on the hash ring; the first object storage service node found in the clockwise direction from this point is the object storage service node corresponding to this data group. The above CRUSH (Controlled Replication Under Scalable Hashing) algorithm is a data distribution algorithm specially designed for distributed storage systems such as Ceph; using the physical and logical hierarchical structure of the storage cluster, it allows administrators to define complex data distribution strategies to optimize performance and reliability; in this algorithm, it is specifically defined how to select object storage services to store the data within a specific data group. The CRUSH Map contains information about all storage devices in the storage cluster and a rule set describing the relationships between these storage devices. It can be determined that the mapping relationship between the above data address and the object storage service is usually stored in the data group for easy calling; of course, in some special scenarios, it can also be stored separately in other modules in the storage system; this application does not make any limitations on this.
[0070] X3. In response to the received data write request, the management layer writes the second specified file to the storage device on the second specified object storage service node that matches the data write request.
[0071] Specifically, after the second designated object storage service node in the management layer and the object service storage node in the management layer receive the data of the second designated file, they will write the data block into the underlying storage device and return a successful write response to the application layer. After the application layer receives the response, the file writing process is completed. This application implements duplicate verification of the data to be written based on the object data fingerprint array set for each storage pool, preventing duplicate data storage and further reducing storage costs.
[0072] It can be understood that while the distributed storage system executes the data management method disclosed in the above embodiments, in order to improve the performance and stability of the storage system and avoid data reading and writing anomalies, this application also proposes to monitor the status of each object storage service node and manage and repair abnormal nodes in a timely manner. Specifically, each object storage service node belonging to the same data group monitors the remaining object storage service nodes corresponding to the data group in real time; each object storage service node reports the health information of the to-be-determined abnormal node collected to the monitor service; the monitor service summarizes the health information of multiple to-be-determined abnormal nodes received; for a to-be-determined abnormal node, when the number of object storage service nodes reporting the to-be-determined abnormal node exceeds the preset number, it is determined that the to-be-determined abnormal node is an abnormal node. Among them, the preset number is set according to the total number of object storage service nodes included in a data group. For example, 80% of the total number is set as the preset number, and specifically, it can be set by those skilled in the art. Then, the monitor service analyzes the failure risk probability of the abnormal node based on the health information of the abnormal node, and further determines the corresponding repair operation according to the failure risk probability. Specifically, if the risk probability of the abnormal node is high but it has not failed, any idle object storage service node is selected to replace the abnormal object storage service node to store the subsequent written data; if the abnormal node has failed, an idle object storage service node with a different fault tolerance type from the current one of the abnormal node is selected for data recovery.
[0073] In addition, in the scenario where the application layer initiates a file reading request to the management layer as described above, the embodiments of this application also disclose that: in response to detecting a file reading request sent by the application layer, the management layer searches for the first designated file according to the file identifier in the file reading request and feeds back the first metadata corresponding to the first designated file to the application layer. In response to receiving the data reading request generated by the application layer for the first designated data storage service node based on the first metadata, it reads the target data in the storage device matching the first designated data storage server and returns it to the application layer.
[0074] The above file reading request includes at least a file identifier for finding a first specified file, such as a file name and a file path. Specifically, the metadata service module of the management layer, after receiving the file reading request initiated by the application layer, finds the first specified file according to the file identifier and returns the first metadata corresponding to the first specified file to the application layer. In addition, the metadata service module also updates the access record in the first metadata of the first specified file to save the access request (i.e., the file reading request) issued by the application layer this time. After the management layer feeds back the first metadata to the application layer, the application layer calculates one or more first specified object storage service nodes that match the first specified file based on the first metadata, that is, the object storage service nodes storing the first specified file can be calculated through the first metadata. At this time, the application layer generates a data reading request and directly initiates it to the determined first specified object storage service node; the first specified object storage service node in the management layer immediately reads the corresponding data block from the corresponding underlying storage device after receiving the data reading request and returns the target data stored in the data block to the application layer; thus, the data reading of the first specified file is completed.
[0075] When the management layer responds to the file reading request, it involves interaction with the application layer. After the management layer feeds back the first metadata to the application layer, for the data reading request generated by the first specified object storage service node according to the first metadata, specifically, at least one first data address that matches the first specified file is determined according to the first object identifier and the fault tolerance type included in the first metadata. According to the data address and the preset data distribution algorithm, the first specified object storage service node that matches the first data address is determined; a data reading request is generated and sent to the first specified object storage service node. The method for determining the second specified object storage service node is the same as above, and will not be elaborated in this application. This application determines the corresponding object storage service by finding the metadata of the file to be read, and then the application layer directly initiates a data reading request to the object storage service, reducing the data access time and realizing the fast reading of the stored data.
[0076] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0077] An embodiment of this application also provides a data management device, such as Figure 6 shown, including a file analysis module 610 and a file management module 620;
[0078] The above file analysis module 610 is used to trigger the management layer to respond to a preset condition trigger, obtain the first metadata of the first specified file to determine the current data attributes of the first specified file;
[0079] The above file analysis module 610 is used to trigger the management layer to determine the target fault tolerance type of the first specified file according to the access record in the current data attribute and the preset rules in response to detecting one or more attribute changes in the current data attribute of the first specified file;
[0080] The above file analysis module 610 is used to trigger the management layer to determine the target fault tolerance type of the first specified file according to the preset rules in response to detecting that the current data attribute is inconsistent with the historical data attribute;
[0081] The above file management module 620 is used to trigger the management layer to perform a data transfer operation on the first specified file to transfer the first specified file to a matching target storage pool in response to detecting a change in the target fault tolerance type.
[0082] The above data management device further includes a file interaction module 630 and a duplicate check module 640;
[0083] The above file interaction module 630 is used to trigger the management layer to feedback the second metadata of the second specified file included in the file write request to the application layer in response to detecting a file write request issued by the application layer;
[0084] The above duplicate check module 640 is used to trigger the application layer to perform data repeatability check on the second specified file based on the second metadata and the preset data fingerprint array, and generate a data write request and send it to the management layer after the data repeatability check passes;
[0085] The above file interaction module 630 is further used for the management layer to write the second specified file to the storage device on the second specified object storage service node matching the data write request in response to receiving the data write request.
[0086] For the description of the features in the corresponding embodiments of the above data management device, reference can be made to the relevant descriptions in the corresponding embodiments of the data management method of the distributed storage system, which will not be elaborated here one by one.
[0087] An embodiment of the present application further provides an electronic device, including: one or more processors; and a memory associated with the one or more processors, the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the following operations are performed:
[0088] Trigger the management layer to obtain the first metadata of the first specified file to determine the current data attribute of the first specified file in response to the trigger of the preset condition;
[0089] Trigger the management layer to determine the target fault tolerance type of the first specified file according to the access record in the current data attribute and the preset rules in response to detecting one or more attribute changes in the current data attribute of the first specified file;
[0090] Trigger the management layer to perform a data transfer operation on the first specified file to transfer the first specified file to a matching target storage pool in response to detecting a change in the target fault tolerance type.
[0091] In some implementation scenarios, the memory is used to store program instructions, and when the program instructions are read and executed by one or more processors, the following operations are also performed:
[0092] Trigger the management layer to feedback the second metadata of the second specified file included in the file write request to the application layer in response to detecting a file write request sent by the application layer;
[0093] Trigger the application layer to perform data duplication verification on the second specified file based on the second metadata and the preset data fingerprint array, and generate a data write request and send it to the management layer after the data duplication verification passes;
[0094] Trigger the management layer to write the second specified file to the storage device on the second specified object storage service node that matches the data write request in response to the received data write request.
[0095] Among them, Figure 7 An exemplary architecture of the electronic device is shown, which may specifically include a processor 710, a video display adapter 711, a disk drive 712, an input / output interface 713, a network interface 714, and a memory 720. The above-mentioned processor 710, video display adapter 711, disk drive 712, input / output interface 713, network interface 714, and the memory 720 can be communicatively connected through a bus 730.
[0096] Among them, the processor 710 can be implemented in a general-purpose CPU (Central Processing Unit, central processor), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solution provided by this application.
[0097] The memory 720 may be implemented in the form of ROM (ReadOnly Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 720 may store the operating system 721 for controlling the execution of the electronic device 700, and the basic input / output system (BIOS) 722 for controlling the low-level operations of the electronic device 700. Additionally, a web browser 723, a data storage management system 724, an icon font processing system 725, etc. may also be stored. The above-mentioned icon font processing system 725 may be the application program that specifically implements the operations of the foregoing steps in the embodiments of the present application. In summary, when implementing the technical solution provided by the present application through software or firmware, the relevant program codes are stored in the memory 720 and are called and executed by the processor 710.
[0098] The input / output interface 713 is used to connect to the input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input devices may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output devices may include a display, a speaker, a vibrator, an indicator light, etc.
[0099] The network interface 714 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module may implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).
[0100] The bus 730 includes a path for transmitting information between various components of the device (such as the processor 710, the video display adapter 711, the disk drive 712, the input / output interface 713, the network interface 714, and the memory 720).
[0101] In addition, the electronic device 700 may also obtain information on specific collection conditions from the virtual resource object collection condition information database for conditional judgment, etc.
[0102] It should be noted that although the above device only shows the processor 710, the video display adapter 711, the disk drive 712, the input / output interface 713, the network interface 714, the memory 720, the bus 730, etc., in the specific implementation process, the device may also include other components necessary for normal execution. In addition, those skilled in the art can understand that the above device may also only include the components necessary for implementing the solution of the present application and does not necessarily include all the components shown in the figure.
[0103] Embodiments of the present application also provide a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in any of the above-described embodiments of the data management method for a distributed storage system when running.
[0104] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to, various media capable of storing a computer program, such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc.
[0105] Embodiments of the present application also provide a computer program product, where the computer program product includes a computer program, and the computer program implements the steps in any of the above-described embodiments of the data management method for a distributed storage system when executed by a processor.
[0106] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and the computer program implements the steps in any of the above-described embodiments of the data management method for a distributed storage system when executed by a processor.
[0107] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0108] The above has introduced in detail a data management method for a distributed storage system provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A data management method for a distributed storage system, wherein the distributed storage system comprises a management layer and an application layer, characterized in that: The distributed storage system is provided with a plurality of storage pools of different fault-tolerant types, and the method comprises: The management layer acquires first metadata of the first designated file in response to a preset condition trigger to determine the current data attribute of the first designated file, wherein the preset condition includes that the management layer detects a file read request issued by the application layer and / or a time difference between a current time node collected and a previous historical time node collected is greater than or equal to a preset time difference; In response to detecting that one or more attribute changes exist in the current data attributes of the first designated file, the management layer determines the target fault tolerance type of the first designated file according to the access records in the current data attributes and preset rules, wherein the preset rules include the management layer calculating the file heat value according to the access records corresponding to the first designated file and a preset formula and determining the target fault tolerance type of the first designated file according to the file heat value; In response to detecting a change in the target fault tolerance type, the management layer performs a data transfer operation on the first designated file to transfer the first designated file to a matching target storage pool.
2. The data management method according to claim 1, characterized in that: The preset formula is: Wherein, Hot is the file heat value, i is the sequence number of the access record, n represents the maximum number of access records matched by the first specified file, T cur Indicates the current time, T i is the most recent i-th access time, α i Represents the weight coefficient of the i-th access record.
3. The data management method according to claim 1, characterized in that: The fault tolerance type includes a first fault tolerance type and a second fault tolerance type, and the management layer determines the target fault tolerance type of the first designated file according to the file heat value, including: In response to detecting that the file heat value is less than or equal to a preset threshold, the management layer determines that the target fault tolerance type of the first designated file is a first fault tolerance type; In response to detecting that the file heat value is greater than the preset threshold, the management layer determines that the target fault tolerance type of the first designated file is the second fault tolerance type.
4. The data management method according to claim 1, characterized in that: In response to detecting a change in the target fault tolerance type, the management layer performs a data transfer operation on the first designated file to transfer the first designated file to a matching target storage pool, including: In response to detecting that the target fault tolerance type of the first designated file is changed from the second fault tolerance type to the first fault tolerance type, the management layer performs a data transfer operation on the first designated file to transfer the first designated file to a storage pool matching the first fault tolerance type; In response to detecting that the fault tolerance type of the first designated file is changed from a first fault tolerance type to a second fault tolerance type, the management layer obtains the object distribution of the first object included in the first designated file; The management layer determines whether to perform a data transfer operation on the first designated file according to the object distribution situation, where the object distribution situation includes single-file distribution and multi-file distribution.
5. The data management method according to claim 4, characterized in that: The management layer determines whether to perform a data transfer operation on the first designated file according to the object distribution, including: In response to detecting that the object distribution is a single-file distribution, the management layer moves the first designated file to a storage pool matching a second fault tolerance type; In response to detecting that the object distribution is a multi-file distribution, the management layer obtains a target fault tolerance type of remaining distribution files of the first object distribution except the first designated file; If the target fault tolerance type of the remaining distributed files is the second fault tolerance type, the management layer performs a data transfer operation on the first designated file to transfer the first designated file to a storage pool matching the second fault tolerance type; If the target fault tolerance type of the remaining distributed files is the first fault tolerance type, the management layer does not perform a data transfer operation on the first designated file.
6. The data management method according to claim 1, characterized in that: The method further comprises: In response to detecting the file write request sent by the application layer, the management layer feeds back second metadata of the second designated file included in the file write request to the application layer; The application layer performs a data duplication check on the second designated file based on the second metadata and the preset data fingerprint array, generates a data write request after the data duplication check passes, and sends the request to the management layer; In response to the received data write request, the management layer writes the second designated file to a storage device on a second designated object storage service node that matches the data write request.
7. The data management method according to claim 6, characterized in that: The second metadata includes second object information and a fault tolerance type, the second object information includes segmentation information and a second object identifier, and the generation of the second metadata includes: The management layer divides the second designated file according to the preset occupied disk space to generate division information and a plurality of second objects, and assigns a unique second object identifier to the second object; The management layer determines the fault tolerance type according to the access type of the file write request, where the access type includes a file modification type and a file creation type.
8. The data management method according to claim 7, characterized in that: The application layer performs data duplication verification on the second designated file based on the second metadata and a preset data fingerprint array, including: the application layer obtains and determines a plurality of second object data included in the second designated file according to the segmentation information of the second designated file; The application layer calculates a target hash value combination of the second object data; The application layer searches a preset data fingerprint array for an existing hash value combination that matches the target hash value combination; If not, the application layer determines that the data repeatability check passes; If so, the application layer performs a secondary data duplication check based on the second object information and the existing object information that matches the existing hash value combination.
9. The data management method according to claim 8, characterized in that: The application layer performs a secondary data duplication check based on the second object information and the existing object information matched by the existing hash value combination, including: The application layer determines that the data duplication check passes in response to detecting that the segmentation information in the second object information and the segmentation information in the existing object information that matches the existing hash value combination are inconsistent; In response to detecting that the segmentation information in the second object information is consistent with the segmentation information in the existing object information that matches the existing hash value combination, the application layer determines that the data duplication check has failed, and updates the second object identifier that matches the second object according to the target object identifier contained in the object information that matches the existing hash value combination.
10. The data management method according to claim 7, characterized in that: The management layer determines the fault tolerance type according to the access type of the file write request, including: In response to detecting that the access type is a create file type, the management layer determines that the fault tolerance type is a first fault tolerance type; In response to detecting that the access type is a file modification type, the management layer obtains a historical storage pool address of a file to be modified that matches the second designated file; If the historical storage pool address is a first preset value, the application layer determines that the fault tolerance type is a first fault tolerance type; If the historical storage pool address is a second preset value, the application layer determines that the fault tolerance type is a second fault tolerance type.
11. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data management method according to any one of claims 1 to 10 when executing the computer program.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the data management method according to any one of claims 1 to 10.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the data management method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
File distributed storage method and device and storage medium
CN117850698A
Data writing method and device of storage system, storage medium and electronic equipment
CN119356628A