Data storage method, storage device and related equipment
By using a faster primary persistent storage medium in the storage pool to perform write caching, the data loss problem caused by BBU failure was resolved, achieving higher data storage reliability and write caching speed.
Patent Information
- Application Number
- CN202511417808.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-12-26
AI Technical Summary
The data loss risk of existing enterprise-class all-flash array memory depends on the BBU, and the memory data cannot be persisted when the BBU fails, resulting in data loss.
Two storage media with different read and write speeds are used. The first persistent storage medium with faster read and write speed is used to perform write caching, and data is quickly recovered through multi-way write caching after the storage pool is powered off, thus avoiding data loss.
It improves the reliability of data storage and the read/write speed of the write cache, reduces the risk of data loss due to BBU power failure, and improves the speed and efficiency of data recovery.
Smart Images

Figure CN121209796A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data storage, and in particular to a data storage method, a storage device and related equipment. BACKGROUND
[0002] Enterprise-level all-flash array storage is generally based on a dual-control or multi-controller hardware architecture. Data reading and writing is implemented through memory mirroring between multiple controllers to achieve high performance, and the BBU provides power to the minimum system (CPU, memory, non-volatile medium) to continue to supply power in the event of an abnormal power failure, and the memory data is persisted to the safe deposit box to achieve high reliability.
[0003] For ease of understanding, Figure 1 The architecture of the all-flash array storage in the prior art is shown, and such an architecture of the all-flash array storage at least has the following disadvantages:
[0004] Data loss risk: Reliability completely depends on the BBU. When the BBU is abnormal or has a continuous failure, the data in the memory cannot be persisted (written to the safe deposit box) and is lost. SUMMARY
[0005] The embodiments of the present application provide a data storage method, a storage device and related equipment for improving the reliability of data storage to avoid data loss risks caused by BBU failure.
[0006] The first aspect of the embodiments of the present application provides a data storage method applied to a storage device, the storage device comprising at least one control node, at least one memory and a storage pool, wherein the storage pool comprises at least two storage media with different read-write speeds, the at least one control node is in communication connection with the at least one memory and the storage pool respectively, and the method comprises:
[0007] receiving an IO request sent by a client by using the at least one control node;
[0008] if the IO request is a write request, writing data corresponding to the IO request to the at least one memory;
[0009] obtaining the data size and / or data type of the data;
[0010] writing the data from the at least one memory to a first persistent storage medium with a faster read-write speed of the two storage media according to the data size and / or the data type, so that the first persistent storage medium performs write caching on the data.
[0011] As an optional embodiment, the method further comprises:
[0012] If the storage pool is powered on again after power failure, and the write cache of the first persistent storage medium stores to-be-recovered data, a write process of the to-be-recovered data is continued; wherein the to-be-recovered data is data stored in the write cache before the storage pool is powered off.
[0013] As an optional embodiment, the method further comprises:
[0014] If the data size does not exceed a first threshold or the data type includes metadata, an identifier of the write request is obtained;
[0015] A first target write position of the data is determined according to the identifier of the write request;
[0016] The data is written from the write cache to the first target write position, wherein the first target write position is a position in the first persistent storage medium that is different from other positions in the write cache.
[0017] As an optional embodiment, the first persistent storage medium includes a first area and a second area, wherein the first area is used to execute the write cache, and the second area is used to execute persistent storage of the data;
[0018] According to the data size and / or the data type, the data is written from the at least one memory to the first persistent storage medium, so that the first persistent storage medium executes write cache for the data, comprising:
[0019] According to the data size and / or the data type, the data is written from the at least one memory to the first area, so that the first area executes write cache for the data;
[0020] The method further comprises:
[0021] An identifier of the write request is obtained;
[0022] A first target write position of the data is determined according to the identifier of the write request;
[0023] The data is written from the first area to the first target write position in the second area.
[0024] As an optional embodiment, if the write request includes a plurality of write requests, the method further comprises:
[0025] When the first target write position is determined, a sequence-preserving process is performed on the plurality of write requests to obtain a plurality of write requests after sequence-preserving processing;
[0026] merge the plurality of data corresponding to the plurality of write requests in the order of the plurality of write requests after the order preserving processing to obtain merged data;
[0027] write the merged data to the first target write position.
[0028] As an optional embodiment, if the storage device includes a plurality of control nodes, the method further includes:
[0029] obtain the first target write position of the data in the first persistent storage medium;
[0030] send a mapping relationship between the data and the first target write position to control nodes other than the at least one control node, so that the other control nodes obtain the data according to the mapping relationship.
[0031] As an optional embodiment, the two storage media further include a second persistent storage medium with a read-write speed less than that of the first persistent storage medium, and the method further includes:
[0032] if the data size exceeds a first threshold, split the data in the first persistent storage medium according to a preset storage unit to obtain a plurality of data blocks;
[0033] write the plurality of data blocks to a second target write position in the second persistent storage medium.
[0034] As an optional embodiment, the two storage media further include a second persistent storage medium with a read-write speed less than that of the first persistent storage medium, and the method further includes:
[0035] if the data size exceeds a first threshold, split the data in the first persistent storage medium according to a preset storage unit to obtain a plurality of data blocks;
[0036] calculate erasure redundancy blocks of the plurality of data blocks;
[0037] write the plurality of data blocks and the erasure redundancy blocks to a second target write position in the second persistent storage medium.
[0038] As an optional embodiment, the method further includes:
[0039] obtain the second target write position of the data in the second persistent storage medium;
[0040] send a mapping relationship between the data and the second target write position to a control node other than the at least one control node, so that the other control node obtains the data according to the mapping relationship.
[0041] The second aspect of the embodiment of the application provides a storage device, comprising at least one control node, at least one memory and a storage pool, wherein the storage pool comprises at least two storage media with different read-write speeds, the at least one control node is in communication connection with the at least one memory and the storage pool respectively, and wherein:
[0042] The at least one control node is configured to implement the data storage method provided in the first aspect of the embodiment of the application.
[0043] As an optional embodiment, the two storage media with different read-write speeds comprise a first persistent storage medium with a faster read-write speed and a second persistent storage medium with a slower read-write speed, wherein the first persistent storage medium comprises at least one of a storage level memory, a non-volatile random storage medium and a persistent memory region (PMR), and the second persistent storage medium comprises a solid state disk (SSD).
[0044] As an optional embodiment, the at least one control node comprises a processor, and the storage device further comprises a PCIE expansion slot in communication connection with the processor.
[0045] As an optional embodiment, the at least one control node is in communication connection with the at least one storage pool through an interface or a hardware chip.
[0046] As an optional embodiment, the two storage media with different read-write speeds are integrated in a same physical device, or the two storage media with different read-write speeds are two independent physical devices respectively.
[0047] The third aspect of the embodiment of the application provides a computer readable storage medium, which stores a computer program, and the computer program is configured to implement the data storage method provided in the first aspect of the embodiment of the application when executed by a processor.
[0048] The fourth aspect of the embodiment of the application provides a computer program product, which stores a computer program, and the computer program is configured to implement the data storage method provided in the first aspect of the embodiment of the application when executed by a processor.
[0049] As can be seen from the above technical solutions, the embodiment of the application has the following advantages:
[0050] Different from the prior art, in the embodiment of the application, the data of the write request is written into the first persistent storage medium with faster read-write speed in the storage pool, and the first persistent storage medium is used to perform write cache on the data, so that compared with the technical solution of performing write cache on the memory connected with the control node in the prior art, on the one hand, the risk of data loss caused by BBU power failure is avoided, and on the other hand, the first persistent storage medium with faster read-write speed in the storage pool is used to perform write cache on the data, so that the read-write speed of the write cache is further improved. BRIEF DESCRIPTION OF DRAWINGS
[0051] Figure 1 An architecture diagram of a storage device in the prior art is shown;
[0052] Figure 2 An embodiment diagram of a storage device in the embodiment of the application is shown;
[0053] Figure 3 An embodiment diagram of a data storage method in the embodiment of the application is shown;
[0054] Figure 4 An embodiment diagram of a data storage method in the embodiment of the application is shown; Figure 3 An embodiment diagram of a data storage method in the embodiment of the application is shown;
[0055] Figure 5 An embodiment diagram of a data storage method in the embodiment of the application is shown; Figure 3 An embodiment diagram of a data storage method in the embodiment of the application is shown;
[0056] Figure 6 An embodiment diagram of a storage device in the embodiment of the application is shown;
[0057] Figure 7 An embodiment diagram of a storage device in the embodiment of the application is shown;
[0058] Figure 8 An embodiment diagram of a storage device in the embodiment of the application is shown. DETAILED DESCRIPTION
[0059] The embodiment of the application provides a data storage method, a storage device and related equipment, which are used for improving the reliability of data storage to avoid the data loss risk caused by BBU failure.
[0060] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort should fall into the protection scope of the present application.
[0061] The terms "first", "second", "third", "fourth" and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a particular order or sequence. It should be understood that the data thus used can be interchanged, where appropriate, so that the embodiments described herein can be carried out in sequences other than those illustrated or described herein. Furthermore, the terms "comprise" and "have", and any variations thereof, are intended to cover non-exclusive inclusion, for example, processes, methods, systems, products, or devices that include a series of steps or units are not necessarily limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0062] For the convenience of understanding, the architecture of the storage device in the prior art will be described simply first, please refer to Figure 1 :
[0063] In Figure 1 , the existing enterprise-level all-flash array storage is generally based on a dual-control or multi-controller hardware architecture. Data reading and writing are implemented through memory mirroring between multiple controllers to achieve high performance. The BBU provides power to the minimum system (CPU, memory, non-volatile medium) in the event of abnormal power failure, and the memory data is persisted to the safe deposit box to achieve high reliability. With the gradual popularization of high-performance NVME SSD, the storage device at least has the following disadvantages: 1. Data loss risk: the reliability completely depends on the BBU. When the BBU is abnormal or has continuous failures, the data in the memory cannot be persisted (written to the safe deposit box) and is lost.
[0064] 2. The health status of the BBU needs to be continuously monitored, and needs to be replaced in time when abnormal. Usually, professional operators are needed, and the operation and maintenance is complex.
[0065] In view of this problem, the storage device in the prior art is improved in the present application. Specifically, the BBU and the safe deposit box in the control node are deleted, and two kinds of storage media with different read and write speeds are set in the storage pool. For the convenience of description, the storage medium with faster read and write speed in the embodiments of the present application is referred to as the first persistent storage medium, and the storage medium with slower read and write speed is referred to as the second persistent storage medium. For the convenience of understanding,Figure 2 The architecture of the storage device in the embodiments of the present application is shown in the following figure, in which Figure 2 The storage device includes at least one control node, at least one memory and a storage pool, the at least one control node is in communication connection with the at least one memory and the storage pool respectively, and the battery backup unit (BBU) and the safe originally preset for the at least one memory are deleted in the at least one control node, and the data storage strategy of providing write cache for the write request of the control node by using the memory in each control node is changed to avoid the data loss problem caused by the BBU failure. The data storage method of the storage device in the embodiments will be described in the following. Figure 2
[0066] For the convenience of understanding, the data storage method in the architecture will be described in the following, please refer to Figure 2 One embodiment of the data storage method in the embodiments of the present application includes: Figure 3
[0067] 301, receiving the IO request sent by the client by using the at least one control node;
[0068] For the storage device in Figure 2 The storage device can receive the IO request sent by the user client by using the at least one control node, that is, the read / write request sent by the user, the control node includes a storage controller, the control node establishes the communication connection with the user client through the wired or wireless communication network in advance, and can receive various IO requests (including read request and write request) sent by the user client in real time, and identify the request type preliminarily.
[0069] 302, if the IO request is a write request, the data corresponding to the IO request is written into the at least one memory;
[0070] If it is identified that the IO request is a write request, the data corresponding to the write request is first written into the at least one memory in communication connection with the at least one control node. Optionally, the memory in the embodiments of the present application can also be built-in in the control node, because the memory can quickly complete the temporary storage of data, and greatly shorten the preliminary response time of the above-mentioned control node to the write request of the user client. After the data corresponding to the IO request is written into the at least one memory, further processing can be performed on the data written in the memory.
[0071] 303, obtaining the data size and / or data type of the data;
[0072] Different from the prior art directly using the memory in communication connection with the control node as the write cache, the application further acquires the data size and / or data type of the data after writing the data into the memory, and executes step 304 on the data according to the data size and / or data type of the data.
[0073] 304. According to the data size and / or the data type, write the data from the at least one memory into the first persistent storage medium with faster read-write speed in the storage pool, so that the first persistent storage medium executes write cache on the data.
[0074] After acquiring the data size and / or data type of the data written into the memory, the application further writes the data from the memory into the first persistent storage medium with faster read-write speed in the storage pool, so that the first persistent storage medium executes write cache on the data.
[0075] Different from the prior art, in the embodiment of the application, the data of the write request is written into the first persistent storage medium with faster read-write speed in the storage pool, and the first persistent storage medium is used to execute write cache on the data. Therefore, compared with the technical solution of the prior art that uses the memory in communication connection with the control node to execute write cache, the embodiment of the application avoids the risk of data loss caused by BBU power failure on the one hand, and on the other hand, the embodiment of the application uses the first persistent storage medium with faster read-write speed in the storage pool to execute write cache on the data, thereby further improving the read-write speed of the write cache.
[0076] As an optional embodiment, when the storage device in Figure 2 is upgraded to the storage device including multiple control nodes in Figure 6 , the application directly writes the data into the first persistent storage medium to execute write cache, instead of using the memory to execute write cache as in the prior art, so that the performance of the storage device in the application is no longer limited by the physical channel of the memory mirror between the control nodes. Because the application only synchronizes control information between the control nodes after writing the data into the first persistent storage medium, the data amount of the memory channel is greatly reduced, the congestion of the physical channel between the control nodes caused by a large amount of IO data is avoided, and the data read-write performance of the storage device is further improved.
[0077] Further, the prior art Figure 1In order to ensure that the memory data is not lost after the system is powered off, the data in the memory needs to be directly written into the safe by using a battery backup unit (BBU) so that the existing storage device needs to read the data from the safe into the memory and then write the data from the memory to the storage medium in the storage pool after the system is powered on. Generally, only one communication path is provided between the memory and the safe, and therefore the bandwidth of the single communication path directly limits the write rate and write time of the abnormal data from the safe to the memory and / or the write rate and write time of the abnormal data to the storage medium in the storage pool, and the larger the amount of abnormal data stored in the safe, the longer the write time.
[0078] To solve the problem, the data is written into the first persistent storage medium for cache, and when the storage pool is powered on after power failure and the first persistent storage medium has data to be recovered (i.e. data that is not stored persistently), the data to be recovered can be written from the cache to other locations in the storage pool for persistent storage in a multi-path manner, so as to complete the persistent storage of the data to be recovered. Therefore, the write speed of the data to be recovered after the system is powered off and powered on is faster and more efficient than the prior art.
[0079] In order to more conveniently understand the beneficial effects of the present application, the data recovery process when the memory is powered off and powered on is described as follows:
[0080] 1. In the prior art, when the memory is powered on, the abnormal data in the safe needs to be written into the memory through the single communication channel between the safe and the memory, and the write speed and write time directly depend on the upper limit of the bandwidth of the single path. However, the present application can use multiple locations (assuming n locations, n is greater than or equal to 2) of the first persistent storage medium for cache, so that the data can be stored persistently through n paths between the n caches and the persistent storage locations when the data is recovered from the multiple caches. Obviously, the upper limit of the bandwidth of the n paths of the present application is higher than that of the single path in the prior art.
[0081] 2. In the prior art, two-step IO operations are required to perform persistent storage of the abnormal data in the safe, i.e. the first step is to write the abnormal data in the safe back to the memory, and the second step is to store the data in the memory persistently in the storage pool. However, the present application only needs one-step IO operation, i.e. directly stores the data to be recovered from the cache location of the first persistent storage medium to other locations for persistent storage in the storage pool, which obviously reduces one-step IO operation, so that the data recovery speed of the present application is faster.
[0082] Based onFigure 3 In order to facilitate dynamic management of the storage area in the first persistent storage medium, the embodiment of the application can further divide the first persistent storage medium into a first area and a second area, wherein the first area is used to perform write caching, and the second area is used to perform persistent storage of data. Therefore, when step 304 is performed, the data is written from the at least one memory to the first area in the first persistent storage medium according to the data size and / or the data type, so that the first area performs write caching on the data.
[0083] Further, after the data corresponding to the write request is written into the first area to perform write caching, the following method steps need to be performed to complete the persistent storage of the data, specifically, the method steps include:
[0084] obtaining the identifier of the write request, determining the first target write position of the data according to the identifier of the write request, and writing the data from the first area to the first target write position in the second area.
[0085] It is easy to understand that after the data is written into the first area, only the write caching of the data is completed, and in order to further realize the persistent storage of the data, the embodiment of the application further needs to obtain the identifier of the write request and determine the first target write position of the data according to the identifier of the write request, so as to perform persistent storage on the data. As an optional embodiment, the correspondence between the write request identifier and the first target write position can be pre-written in a mapping table, and after the write request identifier is obtained, the first target write position corresponding to the write request identifier is obtained from the mapping table, and the data is further written from the first area of the first persistent storage medium to the first target write position in the second area to complete the persistent write of the data.
[0086] In the embodiment of the application, the first persistent storage medium is pre-divided into a first area and a second area, and the first area is used to perform write caching, and the second area is used to perform persistent storage, so as to realize dynamic management of the first area and the second area in the data storage process. For example, when the first area is smaller than a preset threshold, part of the second area can be divided to temporarily serve as the first area to perform write caching, thereby realizing dynamic management of different areas in the first persistent storage medium.
[0087] Based on Figure 3 The embodiment is described below in detail. Figure 3
[0088] I. The data size does not exceed the first threshold or the data type is metadata
[0089] Please refer to Figure 4 ,Figure 4 As a refinement of step 304:
[0090] 401. If the data size does not exceed the first threshold or the data type includes metadata, obtaining an identification of the write request;
[0091] If the control node determines that the data size does not exceed the first threshold or the data type includes metadata, the identification of the write request is further obtained. The identification of the write request can be an identification code identifying the write request received by the control node. The first threshold can be customized according to actual needs, such as 2K, 4K or 8K, and the size of the first threshold is not specifically limited here.
[0092] Further, the data type in the embodiment of the application can be divided into metadata and business data. The metadata can be understood as the specification of the business data. The metadata does not directly record specific business information, but describes the attributes, context, structure and management information of the business data to help users understand, locate, manage and use the business data. The business data directly carries data with actual meanings such as business logic, user information and observation results.
[0093] When the data size does not exceed the first threshold or the data type is metadata, the identification of the write request is obtained, and step 402 is executed according to the identification of the write request.
[0094] 402. Determining a first target write position of the data according to the identification of the write request;
[0095] Specifically, after obtaining the identification of the write request, the first target write position of the data can also be calculated according to the identification of the write request. As an optional embodiment, the first target write position of the data can be obtained by reading a pre-set mapping table (which records the mapping relationship between the write request identification and the first target write position, or records the mapping relationship between the metadata and the first target write position). According to the identification of the write request, the first target write position of the write request data in the first persistent storage medium is read from the pre-set mapping table, and then the data is written to the first target write position to complete the persistent storage of the data.
[0096] 403. Writing the data from the write cache to the first target write position, wherein the first target write position is a position in the first persistent storage medium that is different from other positions in the write cache.
[0097] After the first target write position of the data in the first persistent storage medium is determined according to the write request identifier, the data is written from the write cache to the first target write position, so as to complete the persistent writing of the data in the persistent storage medium. It is easy to understand that the first target write position is a position in the first persistent storage medium that is different from the write cache.
[0098] Because the read and write frequencies of the small IO data and / or metadata in the storage medium are high, for this requirement, in the embodiment of the application, when the data size is less than the first threshold value and / or the data type is metadata,
[0099] The small IO data or metadata corresponding to the write request is directly written to the first target write position in the first persistent storage medium with a faster read and write speed, and the first target write position is a position in the first persistent storage medium that is different from the write cache, so as to ensure the read and write performance of the small IO data and metadata.
[0100] It should be noted that, in order to avoid the data at the write cache position and the data at the first target write position from interfering with each other, the first target write position in the first persistent storage medium is different from the write cache position, but in actual application, the first target write position and the write cache position in the first persistent storage medium can be mutually distinguished or can be used in a cross manner, as long as the persistent storage of the small IO data and metadata in the first persistent storage medium can be realized. The relative positions of the first target write position and the write cache position in the first persistent storage medium are not limited.
[0101] Further, based on Figure 4 When the control node receives a plurality of write requests, in order to avoid the out-of-order writing of data caused by the error of the order of the write requests, the embodiment of the application can also perform order-preserving processing on the plurality of write requests in the process of determining the first target write position according to the request identifier, to obtain a plurality of write requests after order-preserving processing. Then, according to the arrangement order of the plurality of write requests after order-preserving processing, a plurality of data corresponding to the plurality of write requests are merged to obtain merged data, and the merged data is written to the first target write position.
[0102] In the embodiment of the application, when a plurality of write requests are received, in order to avoid errors in the writing process of the data caused by the disorder of the order of the plurality of write requests, the application performs order-preserving processing between the plurality of write requests, to ensure that the sequence of the write requests initiated by the user is consistent with the physical order of the final writing to the storage medium, thereby avoiding the write disorder caused by factors such as hardware scheduling and concurrent competition.
[0103] As an optional embodiment, as Figure 6As shown, if the storage device includes multiple control nodes, in order to ensure that the user can obtain the data written in the first persistent storage medium through each control node, after the at least one control node writes the data into the first target write position, the at least one control node also needs to obtain the mapping relationship between the data and the first target write position, and send the mapping relationship to other control nodes except the at least one control node, so that the user can also obtain the above-mentioned data written into the first persistent storage medium through other control nodes and the mapping relationship, thereby improving the convenience of the user obtaining the above-mentioned data through each control node.
[0104] In addition, after the at least one control node sends the mapping relationship between the data and the first target write position to other control nodes, the at least one control node can also send a write success prompt information to the user client to improve the user's write experience.
[0105] It is easy to understand that the at least one control node in the embodiment of the application is a control node that sends an IO request to the user client, such as in Figure 6 , if the user client sends an IO write request to the first control node, the first control node sends a write success prompt information to the user client after completing the data storage operation of the user client.
[0106] II. The data size exceeds the first threshold
[0107] Please refer to Figure 5 , Figure 5 for Figure 3 another detailed step of step 304 in the embodiment:
[0108] 501. If the data size exceeds the first threshold, the data is split according to a preset storage unit in the first persistent storage medium to obtain a plurality of data blocks.
[0109] If the data size written into the first persistent storage medium exceeds the first threshold, the data is split according to a preset storage unit in the first persistent storage medium to obtain a plurality of data blocks.
[0110] As an optional embodiment, in the process of splitting the data according to the preset storage unit, the data size can be split and aligned according to the minimum storage unit of the second persistent storage medium to obtain a plurality of data blocks, so as to improve the read and write speed of the plurality of data blocks in the second persistent storage medium, such as when the second persistent storage medium is an SSD, the data can be split and aligned according to the logical sector of the SSD to obtain a plurality of data blocks.
[0111] Because the SSD is usually operated in the process of writing data according to the sector size (here, the sector includes but is not limited to the logical sector or the physical sector of the SSD), the data is split and aligned according to the sector size of the SSD, so that the size of each data block is consistent with the sector size of the SSD, thereby ensuring that the data of each write request covers one or more sectors, thereby improving the read-write efficiency of the data.
[0112] It should be noted that the first threshold in the embodiment of the application is any value greater than the minimum storage unit, for example, when the minimum storage unit is 512 bytes, the first threshold is 4K, 8K or 16K, etc. The size of the first threshold is not limited here.
[0113] 502, write the plurality of data blocks to a second target write position in the second persistent storage medium.
[0114] After obtaining the plurality of data blocks in step 501, the plurality of data blocks can be written to the second target write position in the second persistent storage medium, thereby completing the persistent storage of the plurality of data blocks in the storage pool.
[0115] In the embodiment of the application, when the data size is greater than the first threshold, the data is split into a plurality of data blocks, and the plurality of data blocks are written to the second persistent storage medium with slower read-write speed in the storage pool, thereby completing the persistent storage of the large data. Because the data is usually accessed less frequently when it is greater than the first threshold, the data greater than the first threshold is stored in the second persistent storage medium with slower read-write speed in the embodiment of the application, thereby realizing the reduction of storage cost without reducing the business response speed.
[0116] Based on Figure 5 In order to prevent data loss during storage, the embodiment of the application can further calculate the erasure redundancy block of the plurality of data blocks after obtaining the plurality of data blocks in step 501, and write the plurality of data blocks and the erasure redundancy block to the second target write position of the second persistent storage medium, to improve the reliability of the data during storage.
[0117] Because the erasure redundancy block splits the original data into a plurality of data blocks and a small amount of redundancy block by mathematical calculation, the redundancy block carries the key information of "recovering lost data", and its core role is to recover the original data through the remaining data blocks and the erasure redundancy block even if some data blocks are lost, thereby ensuring the reliability of the data during storage.
[0118] Based on Figure 5The embodiment sends the mapping relationship between the data and the second target write position to the other control nodes except the at least one control node after the data is written into the second target write position in the second persistent storage medium, so that the user can also obtain the data through the other control nodes, thereby improving the convenience of the user in obtaining the data through the control nodes in the actual application scenario.
[0119] Further, the at least one control node can further send a storage success prompt information to the user terminal after sending the mapping relationship between the data and the second target write position to the other control nodes except the at least one control node, so as to complete the IO request process and improve the data management efficiency.
[0120] It can be understood that the size of the serial number of each step in various embodiments of the present application does not mean the order of execution, and the execution order of each step should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0121] The data storage method in the embodiments of the present application is described in detail above, and the data storage device Figure 2 will be further described below:
[0122] Specifically in Figure 2 , the storage pool includes a first persistent storage medium with a faster read-write speed and a second persistent storage medium with a slower read-write speed, wherein the first persistent storage medium includes at least one of a storage class memory SCM, a non-volatile random access memory NVRAM and a persistent memory region PMR, and the second persistent storage medium includes a solid state disk SSD.
[0123] In actual application, the first persistent storage medium is preferably a persistent memory region PMR, because when using a storage class memory SCM or a non-volatile random access memory NVRAM, these media need to occupy an SSD position when adopting a standard solid state disk form for ease of operation and maintenance, thereby not only increasing the cost but also reducing the storage capacity density, while the persistent memory region PMR does not have the above problems.
[0124] Further, when the second persistent storage medium in the storage pool is a solid state disk (SSD), the SSD can also be a dual-port SSD supporting the NVME access protocol to further improve throughput and reduce latency to meet high performance requirements. Of course, when the storage device includes two control nodes, the dual-port can be connected to the two control nodes respectively to implement a dual-control architecture of the storage device and improve the reliability of the storage device.
[0125] As an optional embodiment, the first persistent storage medium and the second persistent storage medium in the storage pool of the present application can be integrated in the same physical device or can be two independent physical devices, for example, when the first persistent storage medium is a persistent memory region (PMR) and the second persistent storage medium is an SSD, the first persistent storage medium and the second persistent storage medium are integrated in the same physical device, and when the first persistent storage medium is a storage class memory (SCM) or a non-volatile random access memory (NVRAM) and the second persistent storage medium is an SSD, the first persistent storage medium and the second persistent storage medium are two independent physical devices.
[0126] For the convenience of understanding, Figure 7 a schematic diagram of the first persistent storage medium and the second persistent storage medium integrated in the same physical device and connected to the two control nodes through dual ports is given, Figure 8 a schematic diagram of the first persistent storage medium and the second persistent storage medium being two independent physical devices and connected to the two control nodes through dual ports is given.
[0127] In the embodiments of the present application, the first persistent storage medium and the second persistent storage medium are preferably integrated in the same physical device, thereby improving the integration of the storage pool and correspondingly improving the portability of the storage pool.
[0128] As an optional embodiment, Figure 2 When the control node in the storage pool is connected in communication with the first persistent storage medium and the second persistent storage medium in the storage pool, the control node can be connected in communication with the first persistent storage medium and the second persistent storage medium through a hardware chip or a PCIE interface.
[0129] Preferably, the embodiments of the present application connect the control node with the first persistent storage medium and the second persistent storage medium through the PCIE interface. On the one hand, with the continuous iteration of CPU specifications, the PCIE resources of the CPU are more abundant, which can be several times that of the previous generation of CPUs (for example, an AMD EPYC second-generation single CPU can provide 128 pcie4.0 lanes, and an Intel 6th-generation single CPU has up to 136 pcie5.0 lanes). Therefore, the present application uses a CPU with more abundant PCIE resources, so that the CPU directly communicates with the first persistent storage medium and the second persistent storage medium in the storage pool through the PCIE interface. This can avoid data delay caused by hardware chips and fully utilize the storage performance of different storage media in the storage pool.
[0130] As an optional embodiment, when the storage device in the control node is upgraded to a plurality of control nodes as shown in FIG. 8, the plurality of control nodes can communicate with each other through the NTB technology or the RDMA technology. Figure 2 Figure 6 As an optional embodiment, when the storage device in the control node is upgraded to a plurality of control nodes as shown in FIG. 8, the plurality of control nodes can communicate with each other through the NTB technology or the RDMA technology.
[0131] Among them, NTB is a kind of pcie communication technology, which allows two hosts or subsystems to exchange and transmit data, and is generally widely used in storage systems to mirror data between controllers. Remote direct memory access (RDMA) is a high-performance network transmission technology that allows control nodes to directly access each other's memory data without passing through CPU processing and operating system kernel. Thus, the storage device has the following advantages:
[0132] 1. Low latency: bypassing CPU and kernel, reducing data transmission path, significantly reducing communication delay.
[0133] 2. High bandwidth: supports large throughput data transmission, suitable for large-scale data center scenarios.
[0134] 3. Low CPU occupation: release CPU resources for other computing tasks, improve overall system efficiency.
[0135] Furthermore, when multiple control nodes communicate using Remote Direct Memory Access (RDMA) technology, the versatility of the storage device architecture in this application can be improved. This is because when the control nodes in this application communicate using RDMA technology, they can be decoupled from the Intel platform CPU, since only Intel CPUs support built-in NTB communication technology. This allows the control nodes in this application to run on multiple platforms such as Intel, AMD, ARM, and Hygon. Moreover, when the control nodes in this application communicate using RDMA technology, the storage device can be further decoupled from the PCIe Switch chip, because when the control nodes communicate using NTB technology, they also need the PCIe Switch chip to support NTB technology.
[0136] As an optional embodiment, for ease of understanding... Figure 2 The hardware of the control node can be expanded. In this embodiment, a PCIe expansion slot that communicates with the control node can be provided in the control node, so that users can provide a sufficient number of PCIe interfaces through the PCIe expansion slot to connect more first persistent storage media and second persistent storage media to meet higher storage capacity requirements. For ease of understanding, Figure 6 The diagram also shows the storage devices and the PCIe expansion slot.
[0137] This application also provides a computer program product on which a computer program is stored. When the computer program is executed by a processor, it is used to implement the various steps in the above method embodiments.
[0138] This application also provides a readable storage medium storing a computer program thereon, which, when executed by a processor, is used to implement the various steps in the above method embodiments.
[0139] It should be noted that the above integrated unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the entire or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0140] The above description and the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A data storage method, characterized in that, The method is applied to a storage device, the storage device including at least one control node, at least one memory, and a storage pool, wherein the storage pool includes at least two storage media with different read / write speeds, and the at least one control node is communicatively connected to the at least one memory and the storage pool, respectively. The at least one control node is used to receive I / O requests sent by the client; If the IO request is a write request, then the data corresponding to the IO request is written to the at least one memory; Obtain the data size and / or data type of the data; Based on the data size and / or the data type, the data is written from the at least one memory to a first persistent storage medium with a faster read / write speed among the two storage media, so that the first persistent storage medium performs write caching on the data.
2. The storage method according to claim 1, characterized in that, The method further includes: If the storage pool is powered off and then powered on again, and the write cache of the first persistent storage medium contains data to be recovered, then the write process for the data to be recovered is executed from the write cache; wherein, the data to be recovered is the data stored in the write cache before the storage pool was powered off.
3. The storage method according to claim 1, characterized in that, The method further includes: If the data size does not exceed the first threshold or the data type includes metadata, then obtain the identifier of the write request; The first target write location of the data is determined based on the identifier of the write request; The data is written from the write cache to the first target write location, wherein the first target write location is another location in the first persistent storage medium that is different from the write cache.
4. The storage method according to claim 1, characterized in that, The first persistent storage medium includes a first region and a second region, wherein the first region is used to perform the write caching and the second region is used to perform persistent storage of the data; The step of writing the data from the at least one memory to the first persistent storage medium according to the data size and / or the data type, so that the first persistent storage medium performs write caching on the data, includes: Based on the data size and / or the data type, the data is written from the at least one memory to the first region, such that the first region performs a write cache on the data; The method further includes: Obtain the identifier of the write request; The first target write location of the data is determined based on the identifier of the write request; The data is written from the first region to the first target write location in the second region.
5. The storage method according to claim 3 or 4, characterized in that, If multiple write requests are included, the method further includes: When determining the first target write position, the multiple write requests are processed in a pre-order manner to obtain multiple write requests after the pre-order manner is processed. According to the order of the multiple write requests after the order preservation process, the multiple data corresponding to the multiple write requests are merged to obtain the merged data; The merged data is written to the first target write location.
6. The storage method according to claim 3 or 4, characterized in that, If the storage device includes multiple control nodes, the method further includes: Obtain the first target write position of the data in the first persistent storage medium; The mapping relationship between the data and the first target write location is sent to other control nodes besides the at least one control node, so that the other control nodes can obtain the data according to the mapping relationship.
7. The storage method according to claim 1, characterized in that, The two storage media also include a second persistent storage medium with a read / write speed lower than that of the first persistent storage medium, and the method further includes: If the data size exceeds the first threshold, the data is split into multiple data blocks in the first persistent storage medium according to a preset storage unit. Write the plurality of data blocks to the second target write location in the second persistent storage medium.
8. The storage method according to claim 1, characterized in that, The two storage media also include a second persistent storage medium with a read / write speed lower than that of the first persistent storage medium, and the method further includes: If the data size exceeds the first threshold, the data is split into multiple data blocks in the first persistent storage medium according to a preset storage unit. Calculate the erasure redundancy blocks of the plurality of data blocks; The plurality of data blocks and the erasure redundancy blocks are written to the second target write location in the second persistent storage medium.
9. The storage method according to claim 7 or 8, characterized in that, The method further includes: Obtain the second target write position of the data in the second persistent storage medium; The mapping relationship between the data and the second target write location is sent to other control nodes besides the at least one control node, so that the other control nodes can obtain the data according to the mapping relationship.
10. A storage device, characterized in that, It includes at least one control node, at least one memory, and a storage pool, wherein the storage pool includes at least two storage media with different read / write speeds, and the at least one control node is communicatively connected to both the at least one memory and the storage pool, wherein: The at least one control node is used to implement the data storage method as described in any one of claims 1 to 9.
11. The storage device according to claim 10, characterized in that, The two types of storage media with different read / write speeds include a first persistent storage medium with a faster read / write speed and a second persistent storage medium with a slower read / write speed. The first persistent storage medium includes at least one of storage-class memory, non-volatile random access memory, and persistent memory region (PMR). The second persistent storage medium includes a solid-state drive (SSD).
12. The storage device according to claim 10, characterized in that, The at least one control node includes a processor, and the storage device further includes a PCIe expansion slot communicatively connected to the processor.
13. The storage device according to claim 10, characterized in that, The at least one control node is communicatively connected to the at least one storage pool via an interface or hardware chip.
14. The storage device according to claim 10, characterized in that, The two storage media with different read / write speeds are integrated into the same physical device, or the two storage media with different read / write speeds are two independent physical devices.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it is used to implement the data storage method as described in any one of claims 1 to 9.
16. A computer program product having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it is used to implement the data storage method as described in any one of claims 1 to 9.