A method, device and electronic equipment for storing data based on CEPH
Patent Information
- Application Number
- CN202211722988.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-12-30
AI Technical Summary
又因为脏数据回写阈值为80%,则该缓存池中实际可用于数据缓存的空间小于720GB,当IOPS较高时,缓存池将很快被写满,而在缓存池还未将其中缓存的IO数据写入后端存储盘时,容易发生后续IO直接写入后端存储盘导致性能急剧下降的问题
Smart Images

Figure CN115857830B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud storage technology, and in particular to a method, apparatus and electronic device for storing data based on CEPH. Background Technology
[0002] CEPH (Distributed File System) is a unified distributed storage system that provides three storage functions: block device storage, file system storage, and object storage. Because CEPH only returns confirmation after all replicas have been written (please refer to...), it provides a unique feature to storage. Figure 1 Therefore, its system performance increases linearly with the number of nodes. However, when used in scenarios like cloud computing where high latency and IOPS (IO per sec) performance are required, the underlying CEPH storage cannot provide adequate performance support. For example, in a 3-node CEPH cluster with 36 HDD bays per node, assuming a single-disk write IOPS of 200, the maximum IOPS this CEPH cluster can provide is 21,600. After deducting the performance overhead from replica redundancy strategies and OSD metadata partitioning, the actual IOPS provided by this CEPH is approximately 6,000. Clearly, a performance metric of 6,000 IOPS cannot meet the low-latency requirements of use cases (such as cloud platforms) for the underlying CEPH storage.
[0003] Currently, persistent caching is often used to overcome the above problems: an SSD storage pool is set up as the underlying storage pool for persistent caching. That is, when CEPH performs a write operation, the IO is first mapped to the SSD storage pool, and then the IO in the SSD storage pool is written to the OSD (i.e., dirty data write-back); in actual cache management mechanisms, the dirty data write-back threshold is approximately 80%. This method still has the problem that the cache pool is easily filled due to its small capacity, leading to subsequent IO not being written to the cache pool and being written directly to the HDD storage pool, resulting in a significant decrease in CEPH performance. Generally, the cache pool consists of 3 SSDs; for example, each SSD is 960GB. Please refer to [link / reference]. Figure 2 Since fault recovery and metadata partitioning occupy approximately 5%-10% of the space, the actual usable space per SSD is approximately 900GB or less. Furthermore, because the dirty data write-back threshold is 80%, the actual usable space for data caching in this cache pool is less than 720GB. When IOPS are high, the cache pool will quickly become full. Before the cache pool has written the cached IO data to the backend storage disk, subsequent IO operations may directly write to the backend storage disk, leading to a sharp performance drop. Therefore, the existing CEPH underlying data storage method suffers from high latency. Summary of the Invention
[0004] This invention provides a method for storing data based on CEPH, which reduces the latency when storing data and achieves the goal of efficiently responding to object requests, avoiding the problem that CEPH cannot meet the low latency requirements in existing technologies.
[0005] Firstly, this application provides a method for storing data based on CEPH, the method being applied to a main OSD, comprising:
[0006] Receive data to be stored;
[0007] The system stores the data to be stored and sends a first message to the client; wherein the first message is used to indicate that the data to be stored is written to the main OSD.
[0008] The data to be stored is sent to the slave OSD, so that the slave OSD stores the data to be stored.
[0009] In this embodiment, after the data to be stored is written to the main OSD, the first information is sent directly to the client, and then the data to be stored is written to the slave OSD. This reduces the latency by half, thereby expanding the applicability of CEPH for storing data. It can effectively avoid the high latency caused by the data entering the cache pool first and not being able to be directly stored in the main OSD. It can also avoid the problem of the latency being multiplied by the cache pool in the existing technology, which quickly fills up under continuous write pressure, causing the data to be stored to be directly written to the main OSD.
[0010] In one possible implementation, the master OSD is an SSD and the slave OSD is an HDD.
[0011] One possible implementation includes storing the data to be stored, comprising:
[0012] Determine the length of the data to be stored;
[0013] In response to the fact that the length of the data to be stored is less than a first preset threshold, the data to be stored is stored and added to the first aggregation queue; or,
[0014] In response to the fact that the length of the data to be stored is not less than the first preset threshold, the data to be stored is stored.
[0015] One possible implementation, wherein sending the data to be stored to the OSD, causing the OSD to store the data to be stored, includes:
[0016] Determine the length of the first aggregation queue and the creation time of the first aggregation queue;
[0017] In response to the first aggregation queue being longer than a first preset length threshold, and / or the creation time being longer than a first preset time threshold, the aggregation queue is sent to the slave OSD, causing the slave OSD to write all the data to be stored in the aggregation queue;
[0018] Receive second information; wherein the second information is used to instruct the aggregation queue to write to the OSD.
[0019] One possible implementation includes storing the data to be stored, comprising:
[0020] Determine the length of the data to be stored;
[0021] In response to the fact that the length of the data to be stored is less than a second preset threshold, the data to be stored is stored and added to the first sub-queue; or,
[0022] In response to the fact that the length of the data to be stored is not less than the second preset threshold, the data to be stored is stored and added to the second sub-queue.
[0023] One possible implementation, wherein sending the data to be stored to the OSD, causing the OSD to store the data to be stored, includes:
[0024] Determine the length of the first sub-queue, the creation time of the first sub-queue, and the length of the second sub-queue;
[0025] In response to the first sub-queue's length being greater than a second preset length threshold, and / or the first sub-queue's creation time being greater than a second preset time threshold, the first sub-queue is written to the OSD;
[0026] In response to the second sub-queue's length being greater than the third preset length threshold, the second sub-queue is written to the OSD; the third preset length threshold is less than the second preset length threshold.
[0027] In one possible implementation, the third preset length threshold is equal to the second preset threshold.
[0028] Secondly, embodiments of this application provide a method for storing data based on CEPH, the method being applied to a monitor, including:
[0029] Based on the input parameters in the CRUSH algorithm, the disk type of the master OSD is determined to be SSD; wherein, the input parameters include the disk types of the master OSD and the slave OSD;
[0030] Based on the CRUSH algorithm, the first SSD in the node containing the SSD is selected as the storage disk of the main OSD;
[0031] Based on the input parameters in the CRUSH algorithm, the disk type from the OSD is determined to be HDD;
[0032] Based on the CRUSH algorithm, in a node containing an HDD, the first HDD and the second HDD are selected as the storage disks of the first slave OSD and the second slave OSD, respectively.
[0033] Send an OSD set to the client; wherein the OSD set includes the first SSD, the first HDD and the second HDD, so that the client sends the data to be stored to the main OSD.
[0034] In one possible implementation, if the number of nodes containing HDDs is not less than 2, then the first HDD and the second HDD are selected from different nodes containing HDDs.
[0035] Thirdly, embodiments of this application provide a device for storing data based on CEPH, the device being applied to a main OSD, comprising:
[0036] Receiving unit: Used to receive data to be stored;
[0037] Response unit: used to store the data to be stored and send first information to the client; wherein, the first information is used to indicate that the data to be stored is written to the main OSD;
[0038] Sending unit: Sends the data to be stored to the slave OSD, so that the slave OSD stores the data to be stored.
[0039] In one possible implementation, the master OSD is an SSD and the slave OSD is an HDD.
[0040] In one possible implementation, the response unit is specifically used to determine the length of the data to be stored; in response to the length of the data to be stored being less than a first preset threshold, to store and add the data to be stored to a first aggregation queue; or, in response to the length of the data to be stored being not less than the first preset threshold, to store the data to be stored.
[0041] In one possible implementation, the sending unit is specifically configured to determine the length of the first aggregation queue and the creation time of the first aggregation queue; in response to the length of the first aggregation queue being greater than a first preset length threshold, and / or the creation time of the first aggregation queue being greater than a first preset time threshold, send the aggregation queue to the slave OSD, causing the slave OSD to write all the data to be stored in the aggregation queue; and receive second information; wherein the second information is used to instruct the aggregation queue to be written to the slave OSD.
[0042] In one possible implementation, the response unit is further configured to store and add the data to be stored to a first sub-queue in response to the length of the data to be stored being less than a second preset threshold; or, in response to the length of the data to be stored being not less than the second preset threshold, store and add the data to be stored to a second sub-queue.
[0043] In one possible implementation, the sending unit is further configured to determine the length of the first sub-queue, the creation time of the first sub-queue, and the length of the second sub-queue; in response to the length of the first sub-queue being greater than a second preset length threshold, and / or the creation time of the first sub-queue being greater than a second preset time threshold, the first sub-queue is written to the slave OSD; in response to the length of the second sub-queue being greater than the third preset length threshold, the second sub-queue is written to the slave OSD; the third preset length threshold is less than the second preset length threshold.
[0044] In one possible implementation, the third preset length threshold is equal to the second preset threshold.
[0045] Fourthly, embodiments of this application provide a device for storing data based on CEPH, the device being applied to a monitor, comprising:
[0046] First type of unit: Based on the input parameters in the CRUSH algorithm, determine that the disk type of the master OSD is SSD; wherein, the input parameters include the disk types of the master OSD and the slave OSD;
[0047] First disk unit: Based on the CRUSH algorithm, the first SSD is selected as the storage disk of the main OSD in the node containing the SSD;
[0048] Second type unit: Based on the input parameters in the CRUSH algorithm, determine that the disk type from the OSD is HDD;
[0049] Second disk unit: Based on the CRUSH algorithm, in the node containing HDD, the first HDD and the second HDD are selected as the storage disks of the first slave OSD and the second slave OSD, respectively.
[0050] Collection Unit: Sends an OSD collection to the client; wherein the OSD collection includes the first SSD, the first HDD, and the second HDD, so that the client sends the data to be stored to the main OSD.
[0051] In one possible implementation, if the number of nodes containing HDDs is not less than 2, then the first HDD and the second HDD are selected from different nodes containing HDDs.
[0052] Fifthly, embodiments of this application provide a readable storage medium, including,
[0053] memory,
[0054] The memory is used to store instructions that, when executed by a processor, cause an apparatus including the readable storage medium to perform the method as described in the first aspect to the second aspect and any possible implementation.
[0055] Sixthly, embodiments of this application provide an electronic device, including:
[0056] Memory, used to store computer programs;
[0057] When a processor executes a computer program stored in the memory, it implements the method as described in the first aspect to the second aspect and any possible implementation. Attached Figure Description
[0058] Figure 1 This is a schematic diagram of an existing method for storing data based on CEPH.
[0059] Figure 2 This is a schematic diagram of an existing method for storing data based on CEPH.
[0060] Figure 3 A flowchart illustrating a method for storing data based on CEPH, provided as an embodiment of this application;
[0061] Figure 4 A schematic diagram illustrating another method for storing data based on CEPH provided in this application embodiment;
[0062] Figure 5 A schematic diagram illustrating a method for storing data based on CEPH, provided in an embodiment of this application;
[0063] Figure 6 A schematic diagram of a device for storing data based on CEPH provided in an embodiment of this application;
[0064] Figure 7 A schematic diagram of a device for storing data based on CEPH provided in an embodiment of this application;
[0065] Figure 8 This is a schematic diagram of the structure of an electronic device for storing data based on CEPH, provided as an embodiment of this application. Detailed Implementation
[0066] To address the high latency issue in CEPH data storage in existing technologies, this application provides a method for storing data based on CEPH: After the master OSD receives the data to be stored, it sends a confirmation message (i.e., first message) to the client upon completing the storage on the master OSD, and then initiates a write request to the slave OSD, enabling the slave OSD to store the data to be stored. This avoids the high latency issue caused by sequentially writing to the master OSD and the slave OSD, and then the master OSD only initiating a confirmation message to the client after receiving the confirmation message from the slave OSD.
[0067] It should be noted that in the prior art, the reason why the master OSD only sends an acknowledgment message to the client after receiving the acknowledgment message from the slave OSD is to ensure strong data consistency, that is, to ensure that the data has been written to both the master OSD and the slave OSD. In this application embodiment, while efficiently responding to the client, the data to be stored is still sent to the slave OSD after sending the acknowledgment message to the client, ultimately achieving the goal of storing the same data on both the master OSD and the slave OSD. Therefore, the data storage method provided in this application embodiment does not affect the consistency of the data stored in the master OSD and the slave OSD.
[0068] For ease of understanding, the technical solutions of this application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments and specific features in the embodiments are detailed descriptions of the technical solutions of this application, rather than limitations on the technical solutions of this application. Unless otherwise specified, the embodiments and technical features in the embodiments can be combined with each other.
[0069] Please refer to Figure 3 This application proposes a method for storing data based on CEPH to reduce the latency of CEPH data and achieve high IOPS performance. This method is applied to the main OSD and specifically includes the following implementation steps:
[0070] Step 301: Receive the data to be stored.
[0071] The data to be stored is the data sent by the client.
[0072] OSD (Object Storage Device): One of the core components of Ceph, referring to the process / node responsible for responding to client requests and returning specific data. Each OSD manages one disk.
[0073] Specifically, this application embodiment still adopts a master-slave model. The master OSD and slave OSD can be independently selected from one or more nodes. The number of nodes is preferably three, ensuring that if one node loses power / fails, the other nodes can provide master and slave OSDs for data storage. Correspondingly, when there are multiple master OSDs and / or multiple slave OSDs, selection can be made from different nodes. For example, when setting up three nodes and determining three replicas (OSDs), the first replica (e.g., the master OSD) is selected from any of the three nodes, the second replica is selected from any of the other two nodes, and the third replica is selected from the remaining unselected nodes.
[0074] Furthermore, the disk type for both the primary and secondary OSDs can be either SSD or HDD. Since SSDs and HDDs have different physical structures, and SSDs offer superior storage performance and response speeds compared to HDDs, it is preferable that the primary OSD be an SSD.
[0075] Meanwhile, in order to achieve high economic efficiency, HDD can be selected from OSD; therefore, when the number of nodes is 3, the underlying storage disk of CEPH can be (SSD, HDD, HDD).
[0076] Step 302: Store the data to be stored and send the first information to the client.
[0077] The first piece of information is confirmation information, which is used to indicate that the data to be stored is written to the underlying disk.
[0078] Specifically, in this embodiment, the data to be stored is actually IO (INPUT / OUTPUT), corresponding to request / response. Since the length of the response, i.e., the confirmation information, is fixed, the data to be stored in this embodiment focuses on the request in IO, i.e., the length of the input information.
[0079] The high IOPS performance requirement corresponds to small IO (i.e., the length of data to be stored is less than a first preset threshold), and the corresponding business scenarios also have high latency performance requirements. For example, database applications on the computing cloud. For the aforementioned small IO, if the method of the master OSD receiving the confirmation information from the slave OSD before sending the confirmation information to the client is still used, it will result in high latency and cannot meet the requirements of the scenario.
[0080] Similarly, when the primary OSD writes (i.e. stores) a small IO, it immediately writes to the secondary OSD. Under the pressure of continuous large storage writes, the primary OSD will face an excessive storage write load. To address this, the primary OSD can write the data to be stored to the local object storage engine (BlueStore), aggregate the IO in the OSD to form an aggregate queue, and periodically or at a fixed length (aggregate queue) flush the data to the secondary OSD, thereby further reducing the latency performance of CEPH.
[0081] Therefore, in one embodiment of this application, the length of the data to be stored can be determined first, and whether to add it to the aggregation queue for caching can be determined based on the length of the data to be stored. That is, in response to the length of the data to be stored being less than a first preset threshold (e.g., 64KB, or 128KB), the data to be stored is stored and added to the first aggregation queue. Alternatively, in response to the length of the data to be stored being greater than or equal to the first preset threshold, it is determined that the data to be stored is large, and therefore, only a write operation is performed, that is, the data to be stored is stored in the main OSD.
[0082] In another embodiment of this application, the data to be stored can be cached uniformly, but large and small I / Os can be distinguished according to the length of the data to be stored. Large I / Os and small I / Os are then cached in aggregate queues with different flushing mechanisms or flushing thresholds. Specifically, the length of the data to be stored is determined. If the length of the data to be stored is less than a second preset threshold, the data to be stored is stored and added to a first sub-queue. Alternatively, if the length of the data to be stored is not less than the second preset threshold, the data to be stored is stored and added to a second sub-queue.
[0083] The second preset threshold can be the same as the aforementioned first preset threshold.
[0084] Step 303: Send the data to be stored to the slave OSD, so that the slave OSD stores the data to be stored.
[0085] Specifically, when sending to the slave OSDs, the data to be stored or the aggregation queue is sent simultaneously according to the number of slave OSDs. That is, if aggregation is not performed in step 302, then in this step 303, after sending the first information in step 302, a write request is immediately initiated to the slave OSD, and after the slave OSD writes the data to be stored, confirmation information is received from the slave OSD.
[0086] If, in step 302, only small I / Os with a length less than the first preset threshold are added to the aggregation queue, then large I / Os will be implemented as described above, directly written to the slave OSD. For small I / Os, based on the creation time of the aggregation queue or the aggregation queue itself, the aggregation thread in the master OSD can initiate a write request from the aggregation queue to the slave OSD, and after the I / O is written to the slave OSD, an acknowledgment message is received from the slave OSD. Please refer to [link / reference]. Figure 4 Therefore, in one embodiment of this application, the length of the first aggregation queue and its creation time are determined. In response to the length of the first aggregation queue being greater than a preset length threshold, and / or the creation time being greater than a preset time threshold, the aggregation queue is sent to the slave OSD, causing the slave OSD to write all small IOs, including the aforementioned small IOs, into the aggregation queue. Then, confirmation information from the slave OSD can be received: second information. That is, the second information is used to instruct all IOs in the aggregation queue to be written to the slave OSD, completing the storage on the slave OSD. Please continue to refer to... Figure 4 .
[0087] If, in step 302, the data to be stored is cached in aggregate queues with different flushing mechanisms or flushing thresholds based on the length of the data to be stored, different triggering conditions can be set for different aggregates to promote the writing of large IOs to the slave OSD as soon as possible, avoiding large IOs with low latency requirements from occupying too much cache space and affecting the caching of small IOs. Therefore, in one embodiment of this application, the length of the first sub-queue where the small IO is located and the creation time of the first sub-queue are first determined. At the same time, the length of the second sub-queue for caching large IOs is also determined. Then, in response to the length of the first sub-queue being greater than a first length threshold and / or the creation time of the first sub-queue being greater than a second preset time threshold, all data to be stored in the first sub-queue is written to the slave OSD. Similarly, in response to the length of the second sub-queue being greater than a third length threshold, all data to be stored in the second sub-queue is written to the slave OSD. Wherein, the third preset length threshold is less than the second preset length threshold corresponding to the first sub-queue. Finally, after receiving the confirmation information from the slave OSD: the second information, it can be determined that the data to be stored in the first or second sub-queue has been successfully written to the corresponding slave OSD.
[0088] Furthermore, a third preset length threshold can be set to be equal to a second preset length threshold used to distinguish between large and small I / O, in order to avoid issues such as large I / O where data to be stored that is not sensitive to latency occupying cache space.
[0089] Based on the same inventive concept, this application provides a CEPH data storage method. This method is applied to a monitor to elect a heterogeneous disk type by using the CRUSH algorithm as the master OSD and slave OSDs, so that the client sends the data to be stored to the elected master OSD to complete the data storage.
[0090] For ease of understanding, the following explanations cover Monitor, the CRUSH algorithm, and PG.
[0091] Mon (Monitor): One of the core components of CEPH, responsible for managing metadata.
[0092] CRUSH (Controlled Replication Under Scalable Hashing) algorithm: A pseudo-random algorithm for controlling data distribution and replication. This algorithm can achieve balanced data distribution and load in CEPH, flexibly handle cluster scaling, and support large-scale clusters.
[0093] Placement Groups (PGs) are a logical concept. A PG contains multiple OSDs. CEPH introduces this PG layer to better allocate and locate data. When storing data, CEPH divides it into objects. Each object has a unique OID, which uniquely identifies each object and stores the object's dependency on a file. Because all data in CEPH is virtualized as uniform objects, read and write efficiency is relatively high. To avoid the large addressing workload when writing a large number of objects directly to OSDs, and to prevent the inability to migrate data from a damaged OSD to a new OSD, Placement Groups (PGs) are introduced. A PG is a logical concept; in Linux systems, objects are directly visible, but PGs are not. In data addressing, it's similar to an index in a database: each object is mapped to a specific PG. Therefore, when searching for an object, you only need to find the PG to which the object belongs and then traverse that PG, without needing to traverse all objects. Furthermore, during data migration, a PG is used as the basic unit for migration; CEPH does not directly manipulate objects.
[0094] The following is a detailed description of the CEPH-based data storage method provided in this application. Please refer to it. Figure 5 When data to be stored in the client is mapped to a PG, the Monitor node can elect the primary OSD and secondary OSD disk types for the data to be stored in each PG based on the CRUSH algorithm. Here, the PG, as a logical concept, can be used as a parameter in CRUSH along with the disk type. That is, in this embodiment, an input parameter for identifying the disk type is added to CRUSH to achieve the purpose of electing heterogeneous storage media.
[0095] Therefore, this application embodiment adds input parameters to the CRUSH algorithm, including the disk types of the primary OSD and secondary OSDs. Specifically, the data to be stored is first divided into objects in the client, hashed to PGs by object names, and then the disk types of PGs, primary OSDs, and secondary OSDs are specified through the CRUSH algorithm, for example, [SSD, HDD, HDD], [SSD, SSD, HDD], or [SSD, SSD, SSD]. That is, the Monitor first determines the disk type of the primary OSD as SSD based on the parameters and type input parameters corresponding to PGs in the CRUSH algorithm; then, still through the CRUSH algorithm, the first SSD is selected as the storage disk of the primary OSD in the nodes containing SSDs. Next, still according to the aforementioned input parameters, the disk type of the secondary OSD is determined to be HDD through the CRUSH algorithm; then, in the nodes containing HDDs, the first HDD and the second HDD are selected as the storage disks of the first secondary OSD and the second secondary OSD, respectively, resulting in three copies of the storage disks (first SSD, first HDD, second HDD). Finally, the (first SSD, first HDD, second HDD) can be sent as an OSD set to the client, causing the client to send the data to be stored in the PG to the master OSD. After receiving the OSD set, the client controls the objects in the PG, i.e., the data to be stored, to be sent to the master OSD.
[0096] The number of SSD disks corresponding to the first SSD can be multiple, and the number of HDD disks corresponding to the first HDD and the second HDD can also be multiple. The number of nodes containing HDDs is not less than 2, preferably 2; then the first HDD and the second HDD are selected from different HDD-containing nodes.
[0097] Based on the same inventive concept, this application provides a device for storing data based on CEPH, which is similar to the aforementioned device. Figure 3 The method for storing data based on CEPH shown corresponds to the specific implementation of this device, which can be found in the description of the aforementioned method embodiments section. Repeated descriptions will not be repeated here. Figure 6 The main OSD in this device includes:
[0098] Receiving unit 601: Used to receive data to be stored.
[0099] The master OSD has an SSD disk type; the slave OSD has an HDD disk type.
[0100] Response unit 602: Used to store the data to be stored and send the first information to the client.
[0101] The first piece of information is used to indicate that the data to be stored is written to the main OSD.
[0102] The response unit 602 is specifically used to determine the length of the data to be stored; in response to the length of the data to be stored being less than a first preset threshold, to store and add the data to be stored to a first aggregation queue; or, in response to the length of the data to be stored being not less than the first preset threshold, to store the data to be stored.
[0103] The response unit 602 is further configured to determine the length of the data to be stored; in response to the length of the data to be stored being less than a second preset threshold, store the data to be stored and add it to the first sub-queue; or, in response to the length of the data to be stored being not less than the second preset threshold, store the data to be stored and add it to the second sub-queue.
[0104] Sending unit 603: Used to send the data to be stored to the slave OSD, so that the slave storage device stores the data to be stored.
[0105] The sending unit 603 is specifically used to determine the length of the first aggregation queue and the creation time of the first aggregation queue; in response to the length of the first aggregation queue being greater than a first preset length threshold, and / or the creation time of the first aggregation queue being greater than a first preset time threshold, the sending unit 603 sends the aggregation queue to the slave OSD, causing the slave OSD to write all the data to be stored in the aggregation queue; and receives second information; wherein the second information is used to instruct the aggregation queue to be written to the slave OSD.
[0106] The sending unit 603 is further configured to determine the length of the first sub-queue, the creation time of the first sub-queue, and the length of the second sub-queue; in response to the length of the first sub-queue being greater than a second preset length threshold, and / or the creation time of the first sub-queue being greater than a second preset time threshold, the first sub-queue is written to the slave OSD; in response to the length of the second sub-queue being greater than the third preset length threshold, the second sub-queue is written to the slave OSD; the third preset length threshold is less than the second preset length threshold.
[0107] This application provides a device for storing data based on CEPH. Specific implementation details of this device can be found in the foregoing description of the method embodiments; repeated details will not be repeated here. Figure 7 The device is used in monitors and includes:
[0108] First type unit 701: Based on the input parameters in the CRUSH algorithm, determine that the disk type of the master OSD is SSD; wherein, the input parameters include the disk types of the master OSD and the slave OSD.
[0109] First disk unit 702: Based on the CRUSH algorithm, select the first SSD as the storage disk of the main OSD in the node containing the SSD.
[0110] Second type unit 703: Based on the input parameters in the CRUSH algorithm, determine that the disk type from the OSD is HDD.
[0111] Second disk unit 704: Based on the CRUSH algorithm, in the node containing HDD, the first HDD and the second HDD are selected as the storage disks of the first slave OSD and the second slave OSD, respectively.
[0112] Collection unit 705: Sends an OSD collection to the client; wherein the OSD collection includes the first SSD, the first HDD, and the second HDD.
[0113] If the number of nodes containing HDDs is not less than 2, then the first HDD and the second HDD are selected from different nodes containing HDDs.
[0114] Based on the same inventive concept, embodiments of this application also provide a readable storage medium, including:
[0115] memory,
[0116] The memory is used to store instructions that, when executed by a processor, cause the apparatus including the readable storage medium to perform the CEPH-based data storage method as described above.
[0117] Based on the same inventive concept as the aforementioned CEPH-based data storage method, this application also provides an electronic device that can implement the functions of the aforementioned CEPH-based data storage method. Please refer to... Figure 8 The electronic device includes:
[0118] At least one processor 801 and a memory 802 connected to at least one processor 801. In this embodiment, the specific connection medium between the processor 801 and the memory 802 is not limited. Figure 8 The example shown is the connection between processor 801 and memory 802 via bus 800. Bus 800 is... Figure 8 The connections between other components are indicated by thick lines and are for illustrative purposes only, not as limiting information. The 800 bus can be divided into address bus, data bus, control bus, etc., for ease of representation. Figure 8 The term is represented by a single thick line, but this does not imply that there is only one bus or one type of bus. Alternatively, the processor 801 can also be called a controller; there is no restriction on the name.
[0119] In this embodiment, memory 802 stores instructions executable by at least one processor 801. By executing the instructions stored in memory 802, at least one processor 801 can perform the CEPH-based data storage method described above. Processor 801 can implement... Figure 6 or Figure 7 The functions of each module in the device shown.
[0120] The processor 801 is the control center of the device. It can connect to various parts of the control device through various interfaces and lines. By running or executing instructions stored in memory 802 and calling data stored in memory 802, the processor can perform various functions and process data, thereby monitoring the device as a whole.
[0121] In one possible design, processor 801 may include one or more processing units. Processor 801 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 801. In some embodiments, processor 801 and memory 802 may be implemented on the same chip; in some embodiments, they may also be implemented on separate chips.
[0122] The processor 801 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the CEPH-based data storage method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0123] Memory 802, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 802 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 802 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 802 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.
[0124] By designing and programming the processor 801, the code corresponding to the CEPH-based data storage method described in the aforementioned embodiments can be embedded into the chip, enabling the chip to execute the code during operation. Figure 3 The steps of the method for storing data based on CEPH are shown. How to design and program the processor 801 is a technique well-known to those skilled in the art and will not be described further here.
[0125] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0126] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0127] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0128] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0129] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: Universal Serial Bus flash disks, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0130] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for storing data based on CEPH, characterized in that, The method is applied to the main OSD and includes: Receive data to be stored; The system stores the data to be stored and sends a first message to the client; wherein the first message is used to indicate that the data to be stored is written to the main OSD; storing the data to be stored includes: determining the length of the data to be stored; in response to the length of the data to be stored being not less than a second preset threshold, writing the data to be stored into the local object storage engine and adding the data to be stored to a second sub-aggregation queue; the disk type of the main OSD is an SSD disk; Sending the data to be stored to the slave OSD, causing the slave OSD to store the data to be stored; wherein, sending the data to be stored to the slave OSD, causing the slave OSD to store the data to be stored, includes: determining the length of the second sub-aggregation queue; in response to the length of the second sub-aggregation queue being greater than a third preset length threshold, writing the second sub-aggregation queue to the slave OSD; the disk type of the slave OSD is an HDD disk.
2. The method as described in claim 1, characterized in that, The storage of the data to be stored includes: Determine the length of the data to be stored; In response to the fact that the length of the data to be stored is less than a first preset threshold, the data to be stored is stored and added to the first aggregation queue; or, In response to the fact that the length of the data to be stored is not less than the first preset threshold, the data to be stored is stored.
3. The method as described in claim 2, characterized in that, Sending the data to be stored to the slave OSD, so that the slave OSD stores the data to be stored, includes: Determine the length of the first aggregation queue and the creation time of the first aggregation queue; In response to the first aggregation queue having a length greater than a first preset length threshold, and / or the first aggregation queue having a creation time greater than a first preset time threshold, the aggregation queue is sent to the slave OSD, causing the slave OSD to write all the data to be stored in the aggregation queue; Receive second information; wherein the second information is used to instruct the aggregation queue to write to the OSD.
4. The method as described in claim 1, characterized in that, The storage of the data to be stored includes: In response to the fact that the length of the data to be stored is less than the second preset threshold, the data to be stored is stored and added to the first sub-queue.
5. The method as described in claim 4, characterized in that, Sending the data to be stored to the slave OSD, so that the slave OSD stores the data to be stored, includes: Determine the length of the first sub-queue and the creation time of the first sub-queue; In response to the first sub-queue's length being greater than a second preset length threshold, and / or the first sub-queue's creation time being greater than a second preset time threshold, the first sub-queue is written to the OSD; The third preset length threshold is less than the second preset length threshold.
6. A method for storing data based on CEPH, characterized in that, The method is applied to a monitor and includes: Based on the input parameters in the CRUSH algorithm, the disk type of the master OSD is determined to be SSD; wherein, the input parameters include the disk types of the master OSD and the slave OSD; Based on the CRUSH algorithm, the first SSD in the node containing the SSD is selected as the storage disk of the main OSD; Based on the input parameters in the CRUSH algorithm, the disk type from the OSD is determined to be HDD; Based on the CRUSH algorithm, in a node containing an HDD, the first HDD and the second HDD are selected as the storage disks of the first slave OSD and the second slave OSD, respectively. The system sends an OSD set to a client, wherein the OSD set includes the first SSD, the first HDD, and the second HDD, causing the client to send data to be stored to the master OSD, so that the master OSD, upon receiving the data to be stored, stores the data and sends first information to the client indicating that the data to be stored has been written to the master OSD; the system also sends the data to be stored to a slave OSD, causing the slave OSD to store the data; wherein, the master OSD storing the data to be stored includes: the master OSD determining the length of the data to be stored; in response to the length of the data to be stored being not less than a second preset threshold, writing the data to be stored into the local object storage engine and adding the data to be stored to a second sub-aggregation queue; the master OSD sending the data to be stored to the slave OSD, causing the slave OSD to store the data to be stored, includes: the master OSD determining the length of the second sub-aggregation queue; in response to the length of the second sub-aggregation queue being greater than a third preset length threshold, writing the second sub-aggregation queue into the slave OSD.
7. A device for storing data based on CEPH, characterized in that, The device is applied to the main OSD and includes: Receiving unit: Used to receive data to be stored; Response unit: used to store the data to be stored and send first information to the client; wherein, the first information is used to instruct the data to be stored to be written to the main OSD; storing the data to be stored includes: determining the length of the data to be stored; in response to the length of the data to be stored being not less than a second preset threshold, writing the data to be stored into the local object storage engine and adding the data to be stored to a second sub-aggregation queue; the disk type of the main OSD is an SSD disk; Sending unit: Sends the data to be stored to the slave OSD, causing the slave OSD to store the data to be stored; wherein, sending the data to be stored to the slave OSD, causing the slave OSD to store the data to be stored, includes: determining the length of the second sub-aggregation queue; in response to the length of the second sub-aggregation queue being greater than a third preset length threshold, writing the second sub-aggregation queue to the slave OSD; the disk type of the slave OSD is an HDD disk.
8. A readable storage medium, characterized in that, include, memory, The memory is used to store instructions that, when executed by a processor, cause a device including the readable storage medium to perform the method as described in any one of claims 1-6.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, when executing a computer program stored in the memory, implements the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Disaster tolerance configuration method and device for distributed file system and readable storage medium
CN109582509A
Control method, control device and control equipment for write operation of distributed storage system
CN111142795A
Object storage small file processing method and device, equipment and storage medium
CN111309687A