Cache method and device of distributed storage system, computer equipment

CN116048399BActive Publication Date: 2026-09-25DAWNING INFORMATION IND (BEIJING) CO LTD +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211731302.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2026-09-25
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

[0004]上述方式虽然提升了用户写数据的响应速度,但是日志系统存储空间有限,随着持续接收到大量的写数据,后续接收的写数据长时间难以缓存,缓存效率慢降低、响应时间变长

Benefits of technology

[0071]上述分布式存储系统的缓存方法、装置、计算机设备、存储介质和计算机程序产品,对接收到的写数据进行三种不同方式的缓存,并且按照写数据的数据大小分配不同的缓存方式,使得内存、固态硬盘不会被完全消耗,因此降低了数据大小较大的写请求阻塞数据大小较小的其他写请求的概率,从而保证了各个写请求的响应时间不会过长,从总体上提升了缓存效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116048399B_ABST
    Figure CN116048399B_ABST
Patent Text Reader

Abstract

The application relates to a cache method and device of a distributed storage system, computer equipment, a storage medium and a computer program product. The method is applied to a front-end processing node of a distributed storage system, and the distributed storage system further comprises a back-end storage node, wherein the front-end processing node comprises a solid state disk. The method comprises the following steps: in the case of receiving a write request, determining the data size of write data carried by the write request, determining a target cache mode matched with the write request in each cache mode based on a pre-stored cache mode division strategy and the data size of the write data, wherein each cache mode comprises at least one of a memory cache mode, a solid state disk cache mode and a back-end cache mode. Finally, the write data is subjected to cache processing based on the target cache mode, and a write response message corresponding to the write request is fed back. The method can improve cache efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of distributed storage technology, and in particular to a caching method, apparatus, and computer device for a distributed storage system. Background Technology

[0002] In the field of distributed storage, to ensure a good user experience, distributed storage systems are generally divided into front-end processing nodes and back-end storage nodes. Front-end storage nodes receive write requests from users, process them, and then send them to the back-end storage nodes to complete the write operation and data persistence.

[0003] In related technologies, when a front-end storage node receives a write request from a user, it typically uses a log system to cache the write data sent by the user. After caching is complete, it sends a response to the user indicating that storage is complete. Then, it delivers the cached write data to the back-end storage system to complete the write data persistence, which greatly improves the response speed of the user's write data.

[0004] While the above methods improve the response speed of user writing data, the log system has limited storage space. As a large amount of writing data is continuously received, subsequent writing data is difficult to cache for a long time, resulting in slow caching efficiency and longer response time. Summary of the Invention

[0005] Therefore, it is necessary to provide a caching method, apparatus, computer device, computer-readable storage medium, and computer program product for a distributed storage system that can improve caching efficiency in response to the above-mentioned technical problems.

[0006] Firstly, this application provides a caching method for a distributed storage system.

[0007] The method is applied to a front-end processing node in a distributed storage system, which further includes a back-end storage node; the front-end processing node includes a solid-state drive; the method includes:

[0008] Upon receiving a write request, determine the size of the write data carried in the write request;

[0009] Based on the pre-stored caching method partitioning strategy and the size of the write data, a target caching method matching the write request is determined among the various caching methods; the various caching methods include at least one of memory caching, solid-state drive caching, and backend caching.

[0010] The write data is cached based on the target caching method, and a write response message corresponding to the write request is fed back.

[0011] Received write data is cached in three different ways, with different caching methods allocated according to the size of the write data. This prevents memory and SSD from being fully consumed, thus reducing data loss.

[0012] Larger write requests are less likely to block other write requests with smaller data sizes, thus ensuring that the response time of each write request is not too long, thereby improving overall cache efficiency.

[0013] In one embodiment, determining the target cache method that matches the write request among various cache methods based on a pre-stored cache method partitioning strategy and the size of the write data includes:

[0014] The correspondence between query caching methods and data size ranges is used to determine the target data size range to which the written data belongs;

[0015] 0. Based on the caching method corresponding to the target data size range, determine the target caching method corresponding to the write data.

[0016] Memory is used to cache write data with sizes in the small range, solid-state drives (SSDs) are used to cache write data with sizes in the middle range, and backend users are used to cache write data with sizes in the large range. This corresponds to the cache space size of memory, SSDs, and backends, making full use of the storage advantages of each of memory, SSDs, and backends.

[0017] In one embodiment, the step of using the caching method corresponding to the determined target data size range as the target caching method corresponding to the write data includes:

[0018] Query the current cache status parameters of the cache method corresponding to the target data range;

[0019] If the cache status parameters of the cache method corresponding to the target data range meet the condition of available 0 cache, the cache method corresponding to the target data range will be used as the target cache method for the written data.

[0020] If the current cache status parameters of the cache method corresponding to the target data range do not meet the conditions for available caching, the next priority cache method corresponding to the target data range will be selected as the target cache method for the written data, based on the priority of each cache method.

[0021] 5. When performing caching, the cache status of each caching method is considered. Based on the cache status of each caching method, the caching method for writing data is dynamically adjusted to further ensure the response efficiency of writing data.

[0022] In one embodiment, the caching methods are prioritized from highest to lowest as follows: memory caching, solid-state drive caching, and backend caching.

[0023] The step of caching the write data based on the target caching method and feeding back the write response message corresponding to the write request includes:

[0024] Query the storage address corresponding to the written data on the backend storage node, and check whether the storage addresses of each piece of data cached in the cache space with a higher priority than the target caching method overlap on the backend storage node.

[0025] If no address overlap is found, the write data is cached based on the target caching method, and a write response message corresponding to the write request is returned.

[0026] If an address overlap is found, delete the data with the same address that is cached in the cache space corresponding to the target caching method and has a higher priority than the data with the target caching method. Then, cache the data with the target caching method and send back the write response message corresponding to the write request.

[0027] For write data with overlapping addresses, invalidate the write data with overlapping addresses in the cache space with higher priority to avoid reading incorrect data when users request to read data later.

[0028] In one embodiment, the method further includes:

[0029] Upon receiving a read request for target data, the system queries the cache space corresponding to the target data in descending order of cache method priority, using the data identifier of the target data;

[0030] If the data content corresponding to the target data is found, the retrieved data content is returned in response to the read request.

[0031] Data is read according to cache priority to avoid reading incorrect data.

[0032] In one embodiment, the method according to claim 1 is characterized in that the distributed storage system includes a master front-end processing node and a plurality of slave front-end processing nodes, wherein each of the front-end processing nodes includes at least one solid-state drive; the method is applied to the master front-end processing node;

[0033] When the target caching method is a solid-state drive caching method, the step of caching the write data based on the target caching method and feeding back the write response message corresponding to the write request includes:

[0034] The write data is cached in the solid-state drive, and after the write data is cached in the solid-state drive, the write data is sent to the front-end processing node;

[0035] After determining that each of the aforementioned front-end processing nodes has cached the write data to the solid-state drive, a write response message corresponding to the write request is sent back.

[0036] Configure master and slave front-end processing nodes, and after confirming that each slave front-end processing node has cached the write data to its respective solid-state drive, send back the write response message corresponding to the write request. This ensures that if the master front-end processing node fails, other slave front-end processing nodes can still take over the normal caching of write data.

[0037] Secondly, this application also provides a caching device for a distributed storage system. The device is applied to a front-end processing node of the distributed storage system, which further includes a back-end storage node; the front-end processing node includes a solid-state drive; the device includes:

[0038] The decision module is used to determine the size of the write data carried by the write request when a write request is received.

[0039] The caching module is used to determine the target caching method that matches the write request among the various caching methods based on the pre-stored caching method partitioning strategy and the data size of the write data; the various caching methods include at least one of memory caching method, solid-state drive caching method, and backend caching method;

[0040] The feedback module is used to cache the write data based on the target caching method and to send back the write response message corresponding to the write request.

[0041] In one embodiment, the above-mentioned caching module specifically includes:

[0042] The first query unit is used to query the correspondence between the caching method and the data size range, and to determine the target data size range to which the written data belongs;

[0043] The decision unit is used to determine the target caching method corresponding to the write data based on the caching method corresponding to the target data size range.

[0044] In one embodiment, the decision-making unit specifically includes:

[0045] The query subunit is used to query the current cache status parameters of the cache method corresponding to the target data range;

[0046] The first determining subunit is used to determine the cache method corresponding to the target data range as the target cache method corresponding to the write data when the current cache status parameter of the cache method corresponding to the target data range meets the available cache conditions.

[0047] The second determining subunit is used to determine the next priority cache method of the cache method corresponding to the target data range as the target cache method when the current cache status parameter of the cache method corresponding to the target data range does not meet the available cache conditions, based on the priority of each cache method.

[0048] In one embodiment, the caching methods, from highest to lowest priority, are: memory caching, solid-state drive caching, and backend caching; the aforementioned feedback unit specifically includes:

[0049] The second query unit is used to query the storage address corresponding to the written data on the backend storage node, and whether the storage addresses of each data cached in the cache space with a higher priority than the target caching method are identical on the backend storage node.

[0050] The first caching unit is used to cache the write data based on the target caching method when no address overlap is found, and to send back the write response message corresponding to the write request.

[0051] The second caching unit is used to delete data with the same address as the write data that is cached in the cache space corresponding to the target caching method when the address overlap is found, and then cache the write data based on the target caching method and send back the write response message corresponding to the write request.

[0052] In one embodiment, the above-described apparatus further includes:

[0053] The third query unit is used to, upon receiving a read request for target data, query the data content corresponding to the target data in the cache space corresponding to the cache method in descending order of cache method priority, using the data identifier of the target data;

[0054] The response unit is used to respond to the read request and return the retrieved data content when the data content corresponding to the target data is found.

[0055] In one embodiment, the distributed storage system includes a master front-end processing node and multiple slave front-end processing nodes, wherein each front-end processing node includes at least one solid-state drive (SSD); the device is applied to the master front-end processing node; when the target caching method is SSD caching, the feedback module is specifically used for:

[0056] When the target caching method is a solid-state drive caching method, the step of caching the write data based on the target caching method and feeding back the write response message corresponding to the write request includes:

[0057] The write data is cached in the solid-state drive, and after the write data is cached in the solid-state drive, the write data is sent to the front-end processing node;

[0058] After determining that each of the aforementioned front-end processing nodes has cached the write data to the solid-state drive, a write response message corresponding to the write request is sent back.

[0059] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0060] Upon receiving a write request, determine the size of the write data carried in the write request;

[0061] Based on the pre-stored caching method partitioning strategy and the size of the write data, a target caching method matching the write request is determined among the various caching methods; the various caching methods include at least one of memory caching, solid-state drive caching, and backend caching.

[0062] The write data is cached based on the target caching method, and a write response message corresponding to the write request is fed back.

[0063] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0064] Upon receiving a write request, determine the size of the write data carried in the write request;

[0065] Based on the pre-stored caching method partitioning strategy and the size of the write data, a target caching method matching the write request is determined among the various caching methods; the various caching methods include at least one of memory caching, solid-state drive caching, and backend caching.

[0066] The write data is cached based on the target caching method, and a write response message corresponding to the write request is fed back.

[0067] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0068] Upon receiving a write request, determine the size of the write data carried in the write request;

[0069] Based on the pre-stored caching method partitioning strategy and the size of the write data, a target caching method matching the write request is determined among the various caching methods; the various caching methods include at least one of memory caching, solid-state drive caching, and backend caching.

[0070] The write data is cached based on the target caching method, and a write response message corresponding to the write request is fed back.

[0071] The aforementioned distributed storage system's caching methods, devices, computer equipment, storage media, and computer program products cache received write data in three different ways, and allocate different caching methods according to the size of the write data. This prevents memory and solid-state drives from being completely consumed, thereby reducing the probability that write requests with larger data sizes will block other write requests with smaller data sizes. This ensures that the response time of each write request is not too long, thus improving overall caching efficiency. Attached Figure Description

[0072] Figure 1 This is an application environment diagram of a caching method in a distributed storage system in one embodiment;

[0073] Figure 2 This is a flowchart illustrating a caching method in a distributed storage system according to one embodiment;

[0074] Figure 3 This is a schematic diagram of the structure of a distributed storage system in one embodiment;

[0075] Figure 4 This is a flowchart illustrating a caching method for a distributed storage system in another embodiment;

[0076] Figure 5 This is a structural block diagram of a cache device in a distributed storage system according to one embodiment;

[0077] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0078] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0079] A distributed storage system typically consists of front-end storage nodes and back-end storage nodes. Front-end storage nodes are used to interact with user terminals, while back-end storage nodes are used to store all data. Back-end storage nodes generally use inexpensive mechanical hard drives to store data.

[0080] One approach in related technologies is to deploy a logging system on the front-end storage node. This logging system primarily records write requests to shorten the processing time for user-sent write requests, enabling faster processing and return of results to the user. However, due to limited storage space in the logging system, when the front-end processing node receives a large number of write requests in a short period, the response time of the front-end processing node gradually increases, and caching efficiency slows down. Furthermore, the logging system has a relatively limited functionality; it does not provide read services. Therefore, in the read process, data cannot be directly read from the logs but must be bypassed and read from the cache or back-end storage devices.

[0081] Another approach in related technologies is to deploy solid-state drives (SSDs) on backend storage nodes and use them to cache write data. While this approach can compensate for the performance limitations of mechanical hard drives on backend storage nodes and improve the overall processing performance of the backend system, for each write operation, the frontend processing node and the backend storage node must interact before the processing result can be fed back to the user, resulting in a long processing path and high latency.

[0082] Based on this, this application proposes a caching method for a distributed storage system. This method is applied to the front-end processing nodes of the distributed storage system, which also includes back-end storage nodes. The front-end processing nodes include solid-state drives (SSDs). The method includes: upon receiving a write request, determining the size of the write data carried in the write request; based on a pre-stored caching method partitioning strategy and the size of the write data; determining a target caching method that matches the write request from among various caching methods, wherein each caching method includes at least one of memory caching, SSD caching, and back-end caching. Finally, caching the write data based on the target caching method and feeding back a write response message corresponding to the write request.

[0083] The caching method for the distributed storage system provided in this application caches the received write data in three different ways, and allocates different caching methods according to the size of the write data. This prevents memory and solid-state drives from being completely consumed, thereby reducing the probability that write requests with larger data sizes will block other write requests with smaller data sizes. This ensures that the response time of each write request is not too long, and improves the overall caching efficiency.

[0084] Furthermore, related technologies use log systems to cache data and provide read functionality. However, using memory for caching provides too little cache space, resulting in long response times for most read requests. By using the caching method for the distributed storage system provided in this application, when the distributed storage system receives a read request, it can provide read functionality on the one hand, and provide a larger cache space on the other hand, thereby reducing the response time for most read requests.

[0085] The caching method for distributed storage systems provided in this application can be applied to, for example... Figure 1 In the application environment shown, the user's terminal communicates with the front-end processing nodes in the distributed storage system via the network. The front-end processing nodes of the distributed storage system communicate with the back-end storage nodes of the distributed storage system via the network. The back-end storage nodes store the write data sent by the front-end processing nodes.

[0086] The user's terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc.

[0087] In one embodiment, such as Figure 2 As shown, a caching method for a distributed storage system is provided, which can be applied to... Figure 1 Taking the front-end processing node in the example, the following steps are included:

[0088] Step 201: Upon receiving a write request, determine the size of the write data carried in the write request.

[0089] Specifically, the front-end processing nodes of the distributed storage system are used to receive read and write requests sent by end users. When the front-end node receives a write request sent by the end user, it parses the write request, obtains the write data carried by the write request, and determines the size of the write data carried by the write request.

[0090] Step 203: Based on the pre-stored caching method partitioning strategy and the size of the write data, determine the target caching method that matches the write request among the various caching methods.

[0091] The caching methods include at least one of the following: memory caching, SSD caching, and backend caching. Memory caching caches write data in the memory of the front-end processing node; SSD caching caches write data on the SSD integrated into the front-end processing node; and backend caching sends write data directly to the backend storage node for storage. The speed of the caching methods, from highest to lowest, is: memory, SSD, and backend; the corresponding cache space, from lowest to highest, is: memory, SSD, and backend.

[0092] Specifically, the front-end processing node pre-stores the correspondence between different caching methods and the size of the written data. Among the various caching methods, it determines the target caching method that matches the write request. For example, it divides the caching methods according to the data size and determines the caching method that matches the size of the written data. Alternatively, it considers the cache status of each cache space while considering the data size, and then determines the caching method corresponding to the written data.

[0093] Step 205: Cache the written data based on the target caching method and send back the write response message corresponding to the write request.

[0094] Specifically, after the front-end processing node determines the caching method corresponding to the written data, it performs caching processing on the written data based on the target caching method, caches the written data to the corresponding cache space, and after confirming that the written data is cached in the cache space, it sends a write response message corresponding to the write request to the terminal that sent the write request, so that the terminal can confirm that the write data carried by the write request has been written.

[0095] In this embodiment, the received write data is cached in three different ways, and different caching methods are allocated according to the size of the write data. This prevents memory and solid-state drives from being completely consumed, thereby reducing the probability that a write request with a large data size will block other write requests with a smaller data size. This ensures that the response time of each write request is not too long, thus improving the overall caching efficiency.

[0096] In one embodiment, step 203 specifically includes:

[0097] Step 203A: Query the correspondence between the caching method and the data size range to determine the target data size range to which the written data belongs.

[0098] Specifically, the front-end processing node can divide different caching methods according to the data size. First, it queries the correspondence between the caching method and the data size range. Based on the size of the written data, it determines the target data size range to which the written data belongs. For example, written data with a size in the range of [0KB, 64KB] is divided into small block writes, and the corresponding caching method is memory caching; written data with a size in the range of (64KB, 1MB] is divided into medium block writes, and the corresponding caching method is solid-state drive caching; written data with a size in the range of (1MB, ∞) is called large block writes, and the corresponding caching method is back-end caching. After the front-end processing node determines the size of the written data, it determines the target data size range to which the written data belongs.

[0099] Step 203B: Determine the target caching method for writing data based on the caching method corresponding to the target data size range.

[0100] Specifically, after the front-end processing node determines the target data size range to which the written data belongs, it determines the target caching method for the written data based on the caching method corresponding to the target data size range. For example, it can determine the caching method corresponding to the target data size range based on the correspondence between the caching method and the data size range, and then use the caching method corresponding to the target data size range as the target caching method for the written data. Alternatively, after determining the caching method based on the target data size range, it can further determine whether the caching method can continue to cache data, thereby determining the target caching method for the written data.

[0101] In this embodiment, memory is used to cache write data with sizes in the small range, solid-state drives (SSDs) are used to cache write data with sizes in the middle range, and backend users are used to cache write data with sizes in the large range. This corresponds to the cache space sizes of memory, SSDs, and backends, making full use of the storage advantages of each of memory, SSDs, and backends.

[0102] In one embodiment, step 203B specifically includes:

[0103] Step B1: Query the current cache status parameters of the cache method corresponding to the target data range.

[0104] Specifically, after the front-end processing node determines the target data size range to which the written data belongs, it queries the current cache status parameters of the corresponding caching method, such as cache space occupancy ratio and cache bandwidth, to determine whether the corresponding caching method has the capacity to continue caching. If the cache space occupancy ratio is higher than the preset ratio and cache bandwidth for that caching method, it indicates that the caching method is under high load and cannot continue caching more written data, thus failing to meet the availability caching conditions. If the cache space occupancy ratio and cache bandwidth are not higher than the preset ratio and cache bandwidth for that caching method, it indicates that the caching method still has load capacity and can continue caching more written data, thus meeting the availability caching conditions.

[0105] Step B2: If the current cache status parameters of the cache method corresponding to the target data range meet the conditions for available caching, then use the cache method corresponding to the target data range as the target caching method for writing data.

[0106] Specifically, if the front-end storage node finds that the current cache status parameters of the cache method corresponding to the target data range meet the conditions for being used for caching, it means that the cache method corresponding to the target data range is currently capable of continuing to cache write data, and the cache method corresponding to the target data range is used as the target cache method for writing data.

[0107] Step B3: If the current cache status parameters of the cache method corresponding to the target data range do not meet the conditions for available caching, the next priority cache method corresponding to the target data range will be used as the target cache method for writing data.

[0108] The caching methods, ranked from highest to lowest priority, are: memory caching, solid-state drive caching, and backend caching.

[0109] Specifically, if the front-end processing node finds that the current cache status parameters of the cache method corresponding to the target data range do not meet the available cache conditions, it means that the cache method corresponding to the target data range is not capable of caching more write data. The next priority cache method corresponding to the target data range will be used as the target cache method for the write data.

[0110] In one embodiment, the front-end processing node can also query the current cache status parameters of the next priority cache method corresponding to the target data range to determine whether the next priority cache method can cache more write data. If the current cache status parameters of the next priority cache method meet the caching availability condition, the next priority cache method corresponding to the target data range is used as the target cache method corresponding to the write data. If the current cache status parameters of the next priority cache method do not meet the caching availability condition, the node continues to query the next priority cache method of the same priority level, and so on, until the lowest priority cache method is found, and the lowest priority cache method is used as the target cache method corresponding to the write data.

[0111] In this embodiment, when performing caching, the cache status corresponding to each caching method is considered simultaneously. Based on the cache status corresponding to each caching method, the caching method for writing data is dynamically adjusted to further ensure the response efficiency of writing data.

[0112] In one embodiment, the caching methods are prioritized from highest to lowest as follows: memory caching, solid-state drive caching, and backend caching. In this case, step 205 specifically includes:

[0113] Step 205A: Query the storage address corresponding to the written data on the backend storage node, and check whether there is any overlap in the storage addresses of the data cached in the cache space with a higher priority than the target caching method on the backend storage node.

[0114] Specifically, after receiving a write request from the terminal, the front-end processing node determines the write data carried in the write request and the final storage address of the write data on the back-end storage node, i.e., the address where the write data is ultimately written to disk. Before caching the write data based on the target caching method, the front-end processing node first queries the storage address of the write data on the back-end storage node and checks whether there is any overlap in the storage addresses of the cached data in the cache space with higher priority than the target caching method on the back-end storage node.

[0115] For example, if the target caching method is memory caching, since there is no caching method with higher priority than memory caching, subsequent write requests will also prioritize reading data from memory. Therefore, the front-end processing node does not need to check in the cache space corresponding to other caching methods whether there is any cached write data with the address overlapping with the write data to be cached; it can directly cache the write data. If there is write data with the address overlapping with the write data to be cached, it can be directly overwritten. If the target caching method is SSD caching, the caching method with higher priority than SSD caching is memory caching. In the cache space (memory) corresponding to memory caching, it is checked whether there is any cached write data whose storage address on the backend storage node overlaps with the storage address of the write data on the backend storage node. If the target caching method is backend caching, the caching methods with higher priority than backend caching are memory caching and SSD caching. In the cache spaces (memory and SSD) corresponding to memory caching and SSD caching, it is checked whether there is any cached write data whose storage address on the backend storage node overlaps with the storage address of the write data on the backend storage node.

[0116] Step 205B: If no address overlap is found, cache the write data based on the target caching method and send back the write response message corresponding to the write request.

[0117] Specifically, if the front-end processing node finds that the storage addresses of the data cached in the cache space corresponding to the target caching method have a higher priority than the storage addresses of the written data in the back-end storage node, and these addresses do not overlap with the storage addresses of the written data in the back-end storage node, it means that when there is a subsequent read request for the written data, the read request will not be hit in the cache space corresponding to the caching method with a higher priority than the target data range. The written data will be cached according to the target caching method, and a write response message corresponding to the write request will be sent back.

[0118] Step 205C: If overlapping addresses are found, delete the data with overlapping addresses that are cached in the cache space corresponding to the target caching method and delete the data that is cached in the cache space with higher priority than the target caching method; cache the data based on the target caching method and send back the write response message corresponding to the write request.

[0119] Specifically, if the front-end processing node finds that the storage addresses of the data cached in the cache space corresponding to the target caching method have a higher priority than the storage addresses of the write data in the back-end storage node, it means that when there is a subsequent read request for the write data, the read request will be hit in the cache space corresponding to the caching method with a higher priority than the target data range. The data cached in the cache space corresponding to the target caching method with a higher priority than the target caching method that has an address that overlaps with the write data will be deleted. Then, the write data will be cached according to the target caching method, and a write response message corresponding to the write request will be sent back.

[0120] In this embodiment, for write data with overlapping addresses, the write data with overlapping addresses in the cache space with a higher priority level will be invalidated to avoid reading incorrect data when the user requests to read data later.

[0121] In one embodiment, the method further includes:

[0122] Step 207: Upon receiving a read request for the target data, query the corresponding data content of the target data in the cache space corresponding to the cache method in descending order of cache method priority, based on the data identifier of the target data.

[0123] Specifically, after receiving a read request for target data from the terminal, the front-end processing node queries the cache space corresponding to the target data in descending order of caching priority, using the data identifier of the target data. Since the front-end processing node caches data according to steps 205A to 205C, the cache space corresponding to high-priority caching methods will not cache old written data. Therefore, when the front-end processing node queries the cache space corresponding to the target data in descending order of priority, it will not find incorrect data. Conversely, when querying the cache space corresponding to the target data in ascending order of priority, it may read incorrect data.

[0124] Step 209: If the data content corresponding to the target data is found, respond to the read request and return the found data content.

[0125] Specifically, the front-end processing node queries the cache space corresponding to the target data in descending order of caching priority. Once the target data is found, the query stops, and the retrieved data is returned in response to the read request.

[0126] For example, first, the system queries the cache space (memory) corresponding to the target data. If the target data is found, the query stops, and the system responds to the read request by returning the found data. If the target data is not found in memory, the system continues to query the cache space (SSD) corresponding to the target data. If the target data is found, the query stops, and the system responds to the read request. If the target data is not found on the SSD, the system continues to query the cache space (backend) corresponding to the target data, i.e., a read request is sent to the backend storage. If the system receives the data returned by the backend storage node, the system responds to the read request by returning the data returned by the backend storage node. If the system does not receive the data returned by the backend storage node, the system responds to the read request by returning a message indicating that the target data was not found.

[0127] In this embodiment, data is read according to cache priority to avoid reading incorrect data.

[0128] In one embodiment, the distributed storage system includes a master front-end processing node and multiple slave front-end processing nodes, wherein each front-end processing node includes at least one solid-state drive (SSD). In this case, the above method is applied to the master front-end processing node of the distributed storage system, and step 205 specifically includes:

[0129] Step X1: Cache the write data to the main SSD, and after the main SSD caches the write data, send the write data to the front-end processing node.

[0130] Specifically, the distributed storage system deploys multiple front-end processing nodes, each of which contains a solid-state drive (SSD). The master front-end processing node caches write data to its own SSD and then sends the written data to each slave front-end processing node after caching it, so that the slave front-end processing nodes can cache the write data to their own SSDs and maintain data consistency among the various front-end processing nodes.

[0131] Step X2: After confirming that each front-end processing node has cached the write data to the solid-state drive, a write response message corresponding to the write request is sent back.

[0132] Specifically, after receiving responses from each slave front-end processing node indicating that the write data has been written to the solid-state drive, the master front-end processing node determines that the write data has been cached by each slave front-end processing node. This indicates that the data cached on the solid-state drives of the master front-end processing node and the solid-state drives of each slave front-end processing node are consistent. The master front-end processing node then sends a write response message corresponding to the write request back to the terminal that sent the write request.

[0133] In this embodiment, master and slave front-end processing nodes are set up, and when it is determined that each slave front-end processing node has cached the write data to its respective solid-state drive, a write response message corresponding to the write request is fed back. This ensures that when the master front-end processing node fails, other slave front-end processing nodes can still take over the normal caching of write data.

[0134] The following is a detailed description of a specific embodiment of the caching method for the distributed storage system provided in this application.

[0135] like Figure 3 The diagram shown illustrates the structure of a multi-level cache as described in this application. User I / O refers to read or write requests sent by the user to the front-end processing node via the terminal. When the front-end processing node receives a write request from the user, it is processed by the multi-level cache distribution mechanism of the front-end processing node. When the front-end processing node receives a read request from the user, it accesses each cache space sequentially according to cache priority.

[0136] The multi-level cache distribution mechanism categorizes write requests into three types: small-block writes, medium-block writes, and large-block writes, with no overlap in size range between these types. Small-block writes are cached in the cache group's memory, maximizing memory speed without rapidly consuming cache capacity. Medium-block writes are cached on the cache group's SSDs, leveraging the SSDs' superior random write performance by mitigating the poor random write performance of backend storage devices. Large-block writes are directly submitted to the backend storage devices, ensuring their high sequential write performance without consuming cache space from other write data. Furthermore, the multi-level cache distribution mechanism considers both cache capacity and bandwidth. When the total size of write data cached on the SSD reaches a specified threshold, or when the SSD bandwidth in the cache exceeds a specified threshold, the size and priority of the medium-block write are ignored, and it is directly submitted to the backend storage device to prevent data accumulation on the SSD.

[0137] Furthermore, during the write process, new write requests may overlap with write data in the previous priority cache. If this overlap is not handled, the new write data might skip the current priority cache, resulting in a hit on older write data in the previous priority cache the next time the data is read. Therefore, when caching write data, it is necessary to first check if there is any overlapping data in the previous priority cache, and then invalidate the overlapping data in the previous priority cache to ensure that the latest write data is read the next time the data is read, thus guaranteeing data consistency.

[0138] like Figure 4The diagram shown is a flowchart illustrating a caching method for a distributed storage system according to this application.

[0139] Step 401: After receiving a write request, the front-end processing node hands over the write data carried by the write request to the multi-level cache distribution unit to determine the target caching method.

[0140] Step 402: If the multi-level cache splitter determines that the target caching method for the write data is memory caching, it will submit the write data to memory and cache it in memory. Then, it will continue to submit the data to the solid-state drive for caching and then execute step 409.

[0141] Step 403: If the multi-level cache splitter determines that the target cache method corresponding to the write data is solid-state drive cache, it determines whether there is address overlap between the write data and the write data cached in memory. If there is overlap, proceed to step 404; if there is no overlap, proceed to step 405.

[0142] Step 404: Invalidate (i.e. delete) the write data that overlaps with the write data in memory, and then proceed to step 405.

[0143] Step 405: Cache the write data to the solid-state drive, and then proceed to step 409.

[0144] Step 406: If the multi-level cache splitter determines that the target caching method corresponding to the write data is solid-state drive caching, it determines whether there is address overlap between the write data and the write data cached in memory and solid-state drive. If there is overlap, proceed to step 407; if there is no overlap, proceed to step 408.

[0145] Step 407: Invalidate (i.e. delete) the write data that overlaps with the write data in memory and solid-state drive, and then proceed to step 408.

[0146] Step 408: Send the write data to the backend storage system, and after receiving the response message from the backend storage system, execute step 409.

[0147] Step 409: Send the write response message corresponding to the write request to the user's terminal.

[0148] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0149] Based on the same inventive concept, this application also provides a caching device for a distributed storage system for implementing the caching method of the distributed storage system described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the caching device for a distributed storage system provided below can be found in the limitations of the caching method for the distributed storage system described above, and will not be repeated here.

[0150] In one embodiment, such as Figure 5 As shown, a caching device for a distributed storage system is provided. The device is applied to a front-end processing node of the distributed storage system, which also includes a back-end storage node. The front-end processing node includes a solid-state drive (SSD). The device includes:

[0151] The decision module 501 is used to determine the data size of the write data carried by the write request when a write request is received;

[0152] The caching module 503 is used to determine the target caching method that matches the write request among the various caching methods based on the pre-stored caching method partitioning strategy and the data size of the write data; the various caching methods include at least one of memory caching method, solid-state drive caching method, and backend caching method;

[0153] The feedback module 505 is used to cache the write data based on the target caching method and to send back the write response message corresponding to the write request.

[0154] In one embodiment, the cache module 503 specifically includes:

[0155] The first query unit 503A (not shown in the figure) is used to query the correspondence between the caching method and the data size range to determine the target data size range to which the written data belongs;

[0156] Decision unit 503B (not shown in the figure) is used to use the caching method corresponding to the determined target data size range as the target caching method corresponding to the write data.

[0157] In one embodiment, the decision unit 503B specifically includes:

[0158] The query subunit B1 (not shown in the figure) is used to query the current cache status of the cache method corresponding to the target data range;

[0159] The first determining subunit B2 (not shown in the figure) is used to determine the cache method corresponding to the target data range as the target cache method corresponding to the write data when the current cache status of the cache method corresponding to the target data range is cacheable.

[0160] The second determining subunit B3 (not shown in the figure) is used to determine the next priority cache method of the cache method corresponding to the target data range as the target cache method when the current cache status of the cache method corresponding to the target data range is found to be uncacheable. The cache method priorities are as follows from high to low: memory cache method, solid-state drive cache method, and backend cache method.

[0161] In one embodiment, the caching methods, from highest to lowest priority, are: memory caching, solid-state drive caching, and backend caching; the aforementioned feedback module 505 specifically includes:

[0162] The second query unit 505A (not shown in the figure) is used to query the storage address corresponding to the write data on the backend storage node, and whether the storage addresses of each data cached in the cache space with a higher priority than the target caching method are identical on the backend storage node.

[0163] The first cache unit 505B (not shown in the figure) is used to cache the write data based on the target caching method when no address overlap is found, and to send back the write response message corresponding to the write request.

[0164] The second cache unit 505C (not shown in the figure) is used to delete data with a higher priority than the target caching method that has the same address as the write data when the address overlap is found; to cache the write data based on the target caching method, and to send back the write response message corresponding to the write request.

[0165] In one embodiment, the above-described apparatus further includes:

[0166] The third query unit X1 (not shown in the figure) is used to query the data content corresponding to the target data in the cache space corresponding to the cache method in descending order of cache method priority when a read request for the target data is received.

[0167] The response unit X2 (not shown in the figure) is used to return the retrieved data content in response to the read request when the data content corresponding to the target data is retrieved.

[0168] In one embodiment, the front-end processing node includes a primary solid-state drive (SSD) and multiple secondary SSDs; when the target caching method is SSD caching, the feedback module is specifically used for:

[0169] The write data is cached in the primary solid-state drive (SSD), and after the primary SSD has cached the write data, it is sent to the secondary SSD.

[0170] After determining that the write data has been cached by each of the solid-state drives, a write response message corresponding to the write request is sent back.

[0171] The modules in the caching device of the aforementioned distributed storage system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of the computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0172] In one embodiment, a computer device is provided, which can be a front-end processing node in a distributed storage system, and its internal structure diagram can be as shown in Figure Y. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface of the computer device is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a caching method for a distributed storage system.

[0173] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0174] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0175] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0176] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0177] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0178] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0179] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0180] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A caching method for a distributed storage system, characterized in that, The method is applied to a front-end processing node in a distributed storage system, which further includes a back-end storage node; the front-end processing node includes a solid-state drive; the method includes: Upon receiving a write request, determine the size of the write data carried in the write request; The system queries the correspondence between caching methods and data size ranges to determine the target data range to which the written data belongs. If the cache space occupancy rate or cache bandwidth corresponding to the target data range does not exceed the corresponding threshold, then the caching method corresponding to the target data range is determined as the target caching method for the written data. Each caching method includes memory caching, solid-state drive caching, and backend caching. The memory caching method refers to caching the written data in the memory of the front-end processing node. The solid-state drive caching method refers to caching the written data on the solid-state drive of the front-end processing node. The backend caching method refers to sending the written data to the backend storage node for storage. In the cache space with a priority higher than the target caching method, check if there is any cached data that overlaps with the storage address of the write data on the backend storage node; if so, delete the data with the same address as the write data cached in the cache space with a priority higher than the target caching method, cache the write data based on the target caching method, and send back the write response message corresponding to the write request.

2. The method according to claim 1, characterized in that, The method further includes: Query the current cache status parameters of the cache method corresponding to the target data range; If the cache status parameters of the cache method corresponding to the target data range meet the available cache conditions, the cache method corresponding to the target data range will be used as the target cache method for the written data. If the current cache status parameters of the cache method corresponding to the target data range do not meet the conditions for available caching, the next priority cache method corresponding to the target data range will be selected as the target cache method for the written data, based on the priority of each cache method.

3. The method according to claim 1, characterized in that, The caching methods, from highest to lowest priority, are: memory caching, solid-state drive caching, and backend caching. The method further includes: Query the storage address corresponding to the written data on the backend storage node, and check whether the storage addresses of each piece of data cached in the cache space with a higher priority than the target caching method overlap on the backend storage node. If no duplicate addresses are found, the write data is cached based on the target caching method, and a write response message corresponding to the write request is sent back.

4. The method according to claim 3, characterized in that, The method further includes: Upon receiving a read request for target data, the system queries the cache space corresponding to the target data in descending order of caching priority, using the data identifier of the target data; If the data content corresponding to the target data is found, the retrieved data content is returned in response to the read request.

5. The method according to claim 1, characterized in that, The distributed storage system includes a master front-end processing node and multiple slave front-end processing nodes, wherein each of the front-end processing nodes includes at least one solid-state drive; the method is applied to the master front-end processing node; When the target caching method is a solid-state drive caching method, the step of caching the write data based on the target caching method and feeding back the write response message corresponding to the write request includes: The write data is cached in the solid-state drive, and after the write data is cached in the solid-state drive, the write data is sent to the front-end processing node; After determining that each of the aforementioned front-end processing nodes has cached the write data to the solid-state drive, a write response message corresponding to the write request is sent back.

6. A caching device for a distributed storage system, characterized in that, The device is applied to a front-end processing node in a distributed storage system, the distributed storage system further including a back-end storage node; the front-end processing node includes a solid-state drive; the device includes: The decision module is used to determine the size of the write data carried by the write request when a write request is received. A caching module is used to query the correspondence between caching methods and data size ranges to determine the target data range to which the write data belongs. If the cache space occupancy ratio or cache bandwidth corresponding to the target data range does not exceed the corresponding threshold, then the caching method corresponding to the target data range is determined as the target caching method for the write data. Each caching method includes at least one of memory caching, solid-state drive caching, and backend caching. The memory caching method refers to caching the write data in the memory of the front-end processing node. The solid-state drive caching method refers to caching the write data on the solid-state drive of the front-end processing node. The backend caching method refers to sending the write data to the backend storage node for storage. The feedback module is used to query whether there is cached data in the cache space with a priority higher than the target caching method that overlaps with the storage address of the write data on the backend storage node; if so, delete the data cached in the cache space with a priority higher than the target caching method that overlaps with the write data; perform caching processing on the write data based on the target caching method, and feed back the write response message corresponding to the write request.

7. The apparatus according to claim 6, characterized in that, The caching module includes a decision-making unit, which includes: Query the current cache status parameters of the cache method corresponding to the target data range; If the cache status parameters of the cache method corresponding to the target data range meet the available cache conditions, the cache method corresponding to the target data range will be used as the target cache method for the written data. If the current cache status parameters of the cache method corresponding to the target data range do not meet the conditions for available caching, the next priority cache method corresponding to the target data range will be selected as the target cache method for the written data, based on the priority of each cache method.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • A data caching method of a distributed storage system

    CN109947363A

  • Data processing system and method for radio astronomical data intensive scientific operation

    CN114661637A

  • Method and apparatus for data storage system

    US20170083447A1

  • Technologies for managing replica caching in a distributed storage system

    US20170251073A1