Data access method, device, storage node and storage medium

By generating the first metadata and the second metadata to manage the data to be written in persistent memory, the atomicity problem when writing data greater than 8 bytes in persistent memory is solved, and the consistency guarantee of data is achieved.

CN114647383BActive Publication Date: 2025-06-03CHONGQING UNISINSIGHT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210323902.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-29
Publication Date
2025-06-03
Estimated Expiration
2042-03-29

AI Technical Summary

Technical Problem

When writing data greater than 8 bytes into persistent memory, the atomicity of the write operation cannot be guaranteed, resulting in the data that may be incomplete after power-down.

Method used

By generating the first metadata for managing the data to be written and the second metadata for managing the write log, it is ensured that the data satisfies atomicity when written in persistent memory.

Benefits of technology

It realizes consistency guarantee for data in persistent memory, avoiding the problem of incomplete data after power failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114647383B_ABST
    Figure CN114647383B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of storage technologies, and provides a data access method, apparatus, storage node, and storage medium, which are applied to a storage node. The storage node includes a persistent memory, and the persistent memory includes a metadata area and a data area. The storage node is communicatively connected to a client. The method includes: receiving a write data request sent by the client, where the write data request includes the data length of the data to be written and the write position; generating first metadata for managing the data to be written according to the data length; generating second metadata for managing the write log of the data to be written according to the data length and the write position; after writing the data to be written into the data area, writing the first metadata into the metadata area, and writing the second metadata into the data area. The present invention can ensure the consistency of the data in the persistent memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of storage technologies, and in particular, to a data access method, apparatus, storage node, and storage medium. Background Art

[0002] Although persistent memory (PMem) can ensure that the data written to PMem will not be lost after a power failure and restart, the data written to PMem often needs to be first written to the CPU Cache and then flushed to PMem through a series of CPU instructions. Due to the hardware limitations of PMem and the CPU, writing data greater than 8 bytes to PMem and persisting it cannot guarantee the atomicity of the write operation (that is, if a power failure occurs during the process of persistently writing data, it cannot be guaranteed that the data is completely written). Therefore, how to ensure the consistency of the data persistently written to PMem is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0003] The purpose of the present invention is to provide a data access method, apparatus, storage node, and storage medium, which can ensure the consistency of the data persistently written to PMem.

[0004] To achieve the above purpose, the technical solutions adopted in the embodiments of the present invention are as follows:

[0005] In a first aspect, an embodiment of the present invention provides a data access method, which is applied to a storage node. The storage node includes persistent memory, and the persistent memory includes a metadata area and a data area. The storage node is communicatively connected to a client. The method includes: receiving a write data request sent by the client, where the write data request includes the data length and the write position of the data to be written; generating first metadata for managing the data to be written according to the data length; generating second metadata for managing the write log of the data to be written according to the data length and the write position; after writing the data to be written into the data area, writing the first metadata into the metadata area, and writing the second metadata into the data area.

[0006] Optionally, the step of generating first metadata for managing the data to be written according to the data length includes:

[0007] Calculating the number of segments according to the data length and a preset length;

[0008] Splitting the data to be written according to the number of segments to obtain at least one data segment;

[0009] Generating metadata for each data segment according to the number of the data segments and the position of each data segment in the data to be written;

[0010] Generate a reserved metadata for the data to be written, where the reserved metadata includes the value obtained by adding 1 to the number of data segments;

[0011] Use the reserved metadata and the metadata of all the data segments as the first metadata.

[0012] Optionally, the step of generating second metadata for managing the write log of the data to be written according to the data length and the write position includes:

[0013] Obtain a flag bit used to represent the successful writing of the data to be written into the data area;

[0014] Generate verification data according to the flag bit, the data length and the write position;

[0015] Use the flag bit, the data length, the write position and the verification data as the second metadata.

[0016] Optionally, the metadata area includes multiple metadata units, and the multiple metadata units are managed hierarchically. Each level corresponds to a linked list, and each linked list includes at least one management node. Each management node is used to manage metadata units with adjacent positions. The number of metadata units managed by the management nodes in the same linked list is the same, and the number of metadata units managed by the management nodes in any two linked lists is different. The step of writing the first metadata into the metadata area includes:

[0017] Determine the target level according to the data length and the preset length;

[0018] Determine a first metadata unit to be written and a second metadata unit to be written from the management nodes of the linked list corresponding to the target level, where the number of the first metadata units to be written is the number of data segments, and the number of the second metadata units to be written is 1;

[0019] Write the metadata of each data segment into each first metadata unit to be written in sequence according to the position of each data segment in the data to be written;

[0020] Write the reserved metadata into the second metadata unit to be written.

[0021] Optionally, the data area includes multiple data units, each metadata unit corresponds to a data unit, and each data segment is written into the data unit corresponding to each target metadata unit. The step of writing the second metadata into the data area includes:

[0022] Use the data unit corresponding to the second metadata unit to be written as the data unit to be written;

[0023] Write the second metadata to the data unit to be written.

[0024] Optionally, the data area includes a plurality of data units, the storage node further includes a hard disk, a disk flushing list and a recycling list, the disk flushing list includes the disk flushing positions of the data to be flushed and the data units storing the data to be flushed, the data to be flushed is the data stored in the persistent memory and to be flushed into the hard disk, the disk flushing position is the position to be written in the write data request for writing the data to be flushed, and the method further includes:

[0025] Determine valid data units and invalid data units according to the position to be written and the disk flushing position;

[0026] If the valid data units do not exist in the disk flushing list, update the valid data units to the disk flushing list, so as to flush the data in the valid data units from the persistent memory into the hard disk through the disk flushing list;

[0027] If the invalid data units exist in the disk flushing list, delete the invalid data units from the disk flushing list and insert them into the recycling list, so as to recycle the invalid data units through the recycling list.

[0028] Optionally, the step of deleting the invalid data units from the disk flushing list and inserting them into the recycling list includes:

[0029] If there are data units to be merged that meet the preset merging conditions in the recycling list, delete the data units to be merged from the recycling list;

[0030] Merge the data units to be merged with the invalid data units to obtain data units to be inserted;

[0031] Insert the data units to be inserted into the recycling list;

[0032] If the data units to be merged do not exist in the recycling list, insert the invalid data units into the recycling list.

[0033] Optionally, the method further includes:

[0034] Receive a read data request sent by the client, where the read data request includes the data length and the position to be read of the data to be read;

[0035] Read the original data from the hard disk according to the data length and the position to be read of the data to be read;

[0036] Determine whether there is the latest data corresponding to the data to be read and not stored in the hard disk in the brush disk list according to the position to be read;

[0037] Combine the original data and the latest data to obtain the data to be read.

[0038] Optionally, the storage node further includes a hard disk. When the storage node is powered off, the persistent memory stores the data to be written to the hard disk. The metadata area includes multiple metadata units, and each metadata unit includes a log identifier. The data area includes multiple data units, and each metadata unit manages one data unit. The data to be written to the hard disk exists in at least one data unit. The method further includes:

[0039] When the storage node is powered on, divide the multiple metadata units into at least one metadata unit group according to the log identifier;

[0040] Perform reconstruction processing on each metadata unit group to write the data to be written to the hard disk to the hard disk.

[0041] Optionally, each metadata unit includes the number of shards, the shard index, and the log identifier. The step of performing reconstruction processing on each metadata unit group to write the data to be written to the hard disk to the hard disk includes:

[0042] For any target metadata unit group in the at least one metadata unit group, reconstruct the target metadata unit group according to the number of shards and the shard index of the target metadata unit in the target metadata unit group to obtain the log corresponding to the target metadata unit group;

[0043] Read the status of the log from the data unit managed by the target metadata unit with the last shard index, where the status of the log is used to indicate whether the target data in the data units managed by the other target metadata units except the last shard index in the target metadata unit exists in the data area;

[0044] If the status of the log is the cache status used to indicate that the target data exists in the data area, flush the target data to the hard disk;

[0045] Flush the data in the data units managed by the metadata units in all metadata unit groups with the log status being the cache status to the hard disk, and finally write the data to be written to the hard disk to the hard disk.

[0046] Optionally, the method further includes:

[0047] If the status of the log indicates a non-cached state where the target data does not exist in the data area, the data unit of the target metadata unit is recycled;

[0048] The data units managed by the metadata units in all metadata unit groups with a non-cached log status are recycled.

[0049] In a second aspect, an embodiment of the present invention provides a data access device applied to a storage node. The storage node includes a persistent memory, the persistent memory includes a metadata area and a data area, the storage node is communicatively connected to a client, and the device includes: a receiving module, configured to receive a write data request sent by the client, where the write data request includes the data length of the data to be written and the write position; a generating module, configured to generate first metadata for managing the data to be written according to the data length; the generating module is further configured to generate second metadata for managing the log of the data to be written according to the data length and the write position; a writing module, configured to write the data to be written into the data area, then write the first metadata into the metadata area, and write the second metadata into the data area.

[0050] In a third aspect, an embodiment of the present invention provides a storage node, including a processor and a memory; the memory is used to store a program; the processor is configured to implement the data access method in the first aspect when executing the program.

[0051] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data access method in the first aspect as described above.

[0052] Compared with the prior art, the data access method, device, storage node, and storage medium provided by the embodiments of the present invention, when receiving a write data request sent by a client, generate first metadata for managing the data to be written according to the data length of the data to be written in the write data request, generate second metadata for managing the log of the data to be written according to the data length and the write position, first write the data to be written into the data area, and then write the first metadata and the second metadata into the metadata area and the data area respectively. The atomicity can be ensured when writing data to the persistent memory through the first metadata and the second metadata, thereby ensuring the consistency of the data in the persistent memory. Description of the Drawings

[0053] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0054] Figure 1 It is an example diagram of the application scenario provided by the embodiment of the present invention.

[0055] Figure 2 It is a block diagram of the storage node provided by the embodiment of the present invention.

[0056] Figure 3 It is a flow example diagram of a data access method provided by the embodiment of the present invention.

[0057] Figure 4 It is a flow example diagram of another data access method provided by the embodiment of the present invention.

[0058] Figure 5 It is a flow example diagram of another data access method provided by the embodiment of the present invention.

[0059] Figure 6 It is an example diagram of the division of the address space of the persistent memory provided by the embodiment of the present invention.

[0060] Figure 7 It is an example diagram of the hierarchical linked list provided by the embodiment of the present invention.

[0061] Figure 8 It is a flow example diagram of another data access method provided by the embodiment of the present invention.

[0062] Figure 9 It is an example diagram of the linked list update provided by the embodiment of the present invention.

[0063] Figure 10 It is a schematic diagram of the writing process of the data to be written provided by the embodiment of the present invention.

[0064] Figure 11 It is a flow example diagram of another data access method provided by the embodiment of the present invention.

[0065] Figure 12 It is a schematic diagram of various methods for determining invalid data and valid data provided by the embodiment of the present invention.

[0066] Figure 13 It is an example diagram of the insertion process of the hierarchical linked list provided by the embodiment of the present invention.

[0067] Figure 14A flowchart example of another data access method provided by an embodiment of the present invention.

[0068] Figure 15 A flowchart example of another data access method provided by an embodiment of the present invention.

[0069] Figure 16 A block diagram of a data access device provided by an embodiment of the present invention is shown.

[0070] Icons: 10 - storage node; 11 - processor; 12 - persistent memory; 13 - hard disk; 14 - bus; 20 - client; 100 - data access device; 110 - receiving module; 120 - generating module; 130 - writing module; 140 - reading module; 150 - disk flushing module; 160 - reconstruction module. Detailed implementation manners

[0071] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated herein generally may be arranged and designed in a variety of different configurations.

[0072] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0073] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not require further definition and explanation in subsequent drawings.

[0074] In the description of the present invention, it should be noted that if terms such as "upper", "lower", "inner", "outer", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings or the orientation or positional relationship when the product of the present invention is normally placed. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention.

[0075] In addition, if terms such as "first", "second", etc. are used only for distinguishing descriptions, they cannot be understood as indicating or implying relative importance.

[0076] It should be noted that, without conflict, the features in the embodiments of the present invention may be combined with each other.

[0077] Persistent memory is a non-volatile memory that can be accessed through regular memory access instructions (instead of system calls), has low latency (instead of the I / O bus), and byte-addressable (instead of block) characteristics. Byte-addressable means addressing in units of bytes rather than blocks, and non-volatile means that the data in it will not be lost after power-off. Persistent memory is usually between external memory (ordinary hard disk or solid-state drive) and memory (dynamic random access memory), and it is in the middle position in terms of capacity, performance, and price.

[0078] When using persistent memory, the technical obstacle to overcome is how to avoid data consistency problems when using persistent memory. Since the registers and caches in the processor are volatile, and large capacitors can only ensure that the data in the memory controller is written into persistent memory after power-off, the data in persistent memory may not be the latest copy of the data. The problem caused by data inconsistency between the cache and memory is called the data consistency problem. The impact of the data consistency problem ranges from data loss to system unrecoverability.

[0079] In order to avoid data consistency problems in the prior art, application programs usually explicitly call persistent instructions, and all persistent instructions must be executed sequentially and without overlap. Although this method ensures data consistency, the execution efficiency of persistent instructions is unacceptable. In cases where the completion order of persistent instructions is not emphasized, such as memory copying, the sequential constraint on persistent instructions is usually sacrificed to obtain an improvement in execution efficiency.

[0080] In view of this, the embodiments of the present invention provide a data access method, device, storage node, and storage medium, which can ensure data consistency without sacrificing the sequential constraint of persistent instructions, and will be described in detail below.

[0081] Please refer to Figure 1 , Figure 1 which is an example diagram of the application scenario provided by the embodiments of the present invention. Figure 1In this case, the storage node 10 and the client 20 are communicatively connected. The storage node 10 includes persistent memory and a hard disk. The client 20 sends a data access request (including a read data request and a write data request) to the storage node 10. When the client 20 sends a write data request, the storage node 10 can write the data to be written to the hard disk, or write it to the persistent memory, or, to ensure write performance, temporarily store the data to be written in the persistent memory and then immediately respond to the client, and then flush the data temporarily stored in the persistent memory to the hard disk. The data access method provided by the embodiment of the present invention can be applicable to at least one of the two cases where the data to be written is written to the persistent memory and the data to be written is temporarily stored in the persistent memory and then flushed to the hard disk.

[0082] In this embodiment, the storage node 10 can be any one of a single storage server, a storage array, or a server group composed of multiple storage servers, and is used to store the user data that needs to be stored for the write data request of the client 20 or the metadata for managing the user data.

[0083] The client 20 can be a general host, a server, a mobile terminal, etc. The user issues a data access request through the client 20, and the client 20 sends the data access request to the storage node 10.

[0084] The hard disk can be a mechanical hard disk such as a Serial Attached SCSI (SAS) hard disk or a Serial Advanced Technology Attachment (SATA) hard disk, or a solid state disk (SSD).

[0085] Based on Figure 1 the application scenario, the embodiment of the present invention provides Figure 1 a block diagram of the storage node 10 in Figure 2 , Figure 2 which is the block diagram of the storage node provided by the embodiment of the present invention. Figure 2 In this case, the storage node 10 includes a processor 11, persistent memory 12, a hard disk 13, and a bus 14. The processor 11, the persistent memory 12, and the hard disk 13 communicate through the bus 14.

[0086] The processor 11 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 11 or the instructions in the form of software. The above-mentioned processor 11 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0087] The memory 12 is used to store programs. For example, in the data access device of the embodiments of the present invention, the data access device includes at least one software function module that can be stored in the memory 12 in the form of software or firmware. After receiving the execution instruction, the processor 11 executes the program to implement the data access method in the embodiments of the present invention.

[0088] The memory 12 is also used to store the user data that needs to be stored in the client write data request and the metadata for managing the user data.

[0089] The memory 12 may be Figure 1 at least one of persistent memory and hard disk. Optionally, the memory 12 may be a storage device built into the processor 11 or a storage device independent of the processor 11.

[0090] The bus 13 may be an ISA bus, a PCI bus, an EISA bus, etc. Figure 2 It is represented by only one bidirectional arrow, but it does not mean that there is only one bus or one type of bus.

[0091] On the basis of Figure 1 and Figure 2 the embodiments of the present invention provide a data access method applied to the Figure 1 and Figure 2 storage node 10 in. Please refer to Figure 3 Figure 3 FIG.

[0092] Step S100, receive a write data request sent by the client, where the write data request includes the data length and the write position of the data to be written.

[0093] ​In this embodiment, the data to be written is the data that the user needs to store in the storage node 10. The storage node 10 provides an accessible storage space to the client 20. For example, the storage node 10 provides a 10GB storage space A to the client 20. The user can store data in any location of the storage space A through the client. For example, in order to write the data to be written in the first 1GB space of the storage space A, the user sends a write data request to the storage node 10 through the client 20. Among them, the write data request includes that the data length of the data to be written is 1GB, and the write position is the starting position of the storage space, that is, the 0th byte.

[0094] Step S101, generate first metadata for managing the data to be written according to the data length.

[0095] In this embodiment, the persistent memory can be managed according to a fixed length. If the data length is greater than the fixed length, the data to be written will be split into multiple data segments and written into the persistent memory. Otherwise, the data to be written can be directly written into the storage space with a fixed length. The first metadata is used to manage the data to be written. For example, the first metadata manages the data to be written by recording information such as the number of data segments into which the data to be written is split and the position of each data segment in the data to be written.

[0096] Step S102, generate second metadata for managing the write log of the data to be written according to the data length and the write position.

[0097] In this embodiment, the second metadata is used to manage the write log of the data to be written. The write log is used to record the write situation of the data to be written. The write situation may include whether the write is successful, the write position, etc. For example, the second metadata can use different flags to indicate whether the data to be written has been successfully written into the persistent memory.

[0098] After writing the data to be written into the data area, write the first metadata into the metadata area, and write the second metadata into the data area.

[0099] In this embodiment, writing the data to be written first, and then writing the first metadata and the second metadata can ensure that the data to be written is either all correctly written into the data area, that is, the data in the data area is the latest data at this time, or when power is restored after abnormal power failure, the data area is restored to the state before writing the data to be written according to the first metadata and the second metadata, that is, the data in the data area is the old data before writing the data to be written at this time, and there will be no data inconsistency situation where both the latest data and the old data exist in the data area.

[0100] The above method provided by the embodiment of the present invention can ensure atomicity when writing data to the persistent memory through the first metadata and the second metadata, and further ensure the consistency of the data in the persistent memory.

[0101] Based on Figure 3 , an embodiment of the present invention further provides a specific implementation manner for generating first metadata. Please refer to Figure 4 , Figure 4 which is a flowchart example of another data access method provided by an embodiment of the present invention. Step S101 includes the following sub-steps:

[0102] Sub-step S1010: Calculate the number of segments according to the data length and a preset length.

[0103] In this embodiment, the preset length can be set according to actual scenario requirements. For example, the preset length is set to 4KB. As a specific implementation manner, the number of segments can be calculated using the following formula:

[0104] Sub-step S1011: Split the data to be written according to the number of segments to obtain at least one data segment.

[0105] In this embodiment, it can be understood that when the data length is an integer multiple of the preset length, splitting the data to be written according to the number of segments is an average split. When the data length is not an integer multiple of the preset length, the length of the last data segment is different from that of the other data segments.

[0106] Sub-step S1012: Generate metadata for each data segment according to the number of data segments and the position of each data segment in the data to be written.

[0107] In this embodiment, the metadata for each data segment may include the number of data segments and the position of each data segment in the data to be written.

[0108] In this embodiment, in order to determine the data segments of the data to be written belonging to the same write operation, the same identifier may also be set in each data segment of the data to be written in the same write operation. The data segments with the same identifier belong to the data to be written in the same write operation. The data to be written in the write data request includes at least the following two situations: The data to be written in a write data request can be written at one time, and the data to be written in a write request needs to be written in multiple times. For the former, the data to be written is divided into multiple data segments, and the identifiers of these multiple data segments are the same. For the latter, first determine the length of each write, and then segment the data written each time. The identifiers of the data segments written in the same time are the same, and the identifiers of the data segments written in different times are different.

[0109] In this embodiment, in order to improve the reliability of the metadata of each data segment, the metadata of each data segment may further include check data to ensure the reliability of the information included in the data segment. If each data segment includes the number of data segments and the position of each data segment in the data to be written, the check data can be obtained according to the number of data segments and the position of each data segment. If each data segment includes an identifier, the number of data segments, and the position of each data segment in the data to be written, the check data can be obtained according to the identifier, the number of data segments, and the position of each data segment.

[0110] In this embodiment, in order to reduce the impact of generating check data on the write performance, addition operations can be used to obtain the check data. For example, adding the identifier, the number of data segments, and the position of each data segment in the data to be written together to obtain the corresponding check data. Of course, within an acceptable range of write performance impact, other methods can also be used to calculate the check data, such as Cyclic Redundancy Check (CRC).

[0111] Sub-step S1013: Generate a reserved metadata for the data to be written, where the reserved metadata includes the value obtained by adding 1 to the number of data segments.

[0112] In this embodiment, similar to the metadata of each data segment, in addition to including the value obtained by adding 1 to the number of data segments, the reserved metadata may also include a position, check data, and an identifier. The position of the reserved metadata is the position after the last data segment, and the identifier of the reserved metadata is the same as the identifier of any data segment.

[0113] Sub-step S1014: Use the reserved metadata and the metadata of all data segments as the first metadata.

[0114] The above method provided by the embodiment of the present invention ensures the integrity and accuracy of the first metadata for recording the data to be written by generating corresponding metadata for each data segment and then generating a corresponding reserved metadata for the data to be written as a whole.

[0115] On the basis of Figure 3 the embodiment of the present invention further provides a specific implementation manner for generating the second metadata. Please refer to Figure 5 Figure 5 which is a flowchart example of another data access method provided by the embodiment of the present invention. Step S102 includes the following sub-steps:

[0116] Sub-step S1020: Obtain a flag bit used to represent that the data to be written has been successfully written to the data area.

[0117] ​In this embodiment, different values can be set for the flag field to characterize the writing situation of the data to be written. For example, when the flag field is "cached", it indicates that the data to be written is successfully written into the data area; when the flag field is "none", it indicates that the data to be written has been flushed from the data area to the hard disk. Of course, different integer values can also be used for identification. When the value of the flag field is 1, it indicates that the data to be written is successfully written into the data area, and when it is 0, it indicates that it is flushed from the data area to the hard disk.

[0118] Sub-step S1021: Generate check data according to the flag bit, data length, and write position.

[0119] In this embodiment, in order to avoid the impact of generating check data on the writing performance, the check data can be obtained by addition operation, that is, adding the flag bit, data length, and write position together to obtain the check data. Of course, within the acceptable range of the impact on the writing performance, other methods can also be used to calculate the check data, such as Cyclic Redundancy Check (CRC).

[0120] Sub-step S1022: Use the flag bit, data length, write position, and check data as the second metadata.

[0121] The above method provided by the embodiment of the present invention can accurately identify whether the data to be written is successfully written into the data area through the flag bit, ensure the reliability of the flag bit, data length, and write position through the check data, and use the flag bit, data length, write position, and check data as the second metadata, ensuring the integrity and accuracy of writing the data to be written into the log.

[0122] It should be noted that in the application scenario where the data in the persistent memory needs to be flushed to the hard disk, in order to achieve efficient disk flushing, in addition to the second metadata, the storage node 10 also needs to generate hard disk metadata, which is used to characterize the position where the data to be written needs to be written into the hard disk. Different management methods of the hard disk space will result in different representation methods of the hard disk metadata. For example, the hard disk metadata can include: hard disk identifier, file identifier, object identifier, block identifier, version number, etc.

[0123] To more clearly illustrate the writing process of the first metadata and the second metadata, the embodiment of the present invention first introduces a specific division method of the address space of the persistent memory. Please refer to Figure 6 , Figure 6 which is the division example diagram of the address space of the persistent memory provided by the embodiment of the present invention. Figure 6 In it, the address space of the persistent memory is divided into a first reserved area, a second reserved area, a metadata area, and a data area.

[0124] The sizes of the first reserved area and the second reserved area are both 4KB, which are respectively set at the beginning and end of the persistent memory and are backups of each other. Such a setting can, on the one hand, indicate the start and end positions of the persistent memory, and on the other hand, enhance the reliability of the data therein. Taking the second reserved area as an example, this reserved area contains 3 fields: the size field is used to record the size of the entire space of the persistent memory, occupying 8B; the magic field is set to a fixed value, for example, the fixed value is 0x4A3B2C1D, occupying 4B; the CheckSum field is used to store the CRC32 checksum of the first 12 bytes (size field + magic field) of the reserved area, occupying 4B, to ensure the integrity of the reserved area data. The reserved field is a reserved field for future expansion. The first reserved area is exactly the same as the second reserved area, so it will not be elaborated here.

[0125] The metadata area is adjacent to the first reserved area and is located in the address space after the first reserved area. The address space of the metadata area is aligned in 16B sizes and is divided into multiple metadata units according to 16B sizes. A metadata contains 4 fields: the Log ID field, the count field, the index field, and the CheckSum field. The Log ID field is used to record the log ID of the current log. A log is a set of stored data written in a certain order. A log can be represented by LOG(offset, len), where offset represents the offset position where the data to be written should be written, and len represents the length of the data to be written, that is, a log can represent the data to be written written at one time. The Log ID can be the number of seconds from the current system time to January 1, 1970 plus a 6 - bit linearly increasing serial number, used to uniquely represent a log. If it crosses seconds, the serial number starts counting from 0 again.

[0126] The count field represents the number of data segments + 1 into which the data to be written in one write is split. For example, if the data to be written is 256KB and is split into 64 4KB data segments, then the value of the count field is 65, and the count field occupies 2B.

[0127] The index field represents the relative position of this data segment in the data to be written. For example, if the data to be written is divided into 64 data segments, then the value of the count field in the metadata unit of the first data segment is 65, and the value of the index field is 1. The index values in the metadata units of the remaining data segments increase sequentially, and the index field occupies 2B.

[0128] The CheckSum field is used to store the checksum of the first 12B (i.e., log id + count + index) in this metadata unit, which occupies 4B. Since the metadata is read and written frequently, the checksum is calculated by adding the previous three fields instead of using CRC32 because CRC32 is time-consuming.

[0129] Whether it is the CheckSum field in the metadata area or the CheckSum field in the data area, the number of bytes occupied is 4B, which is less than 8B. On a 64-bit computer, only when 8-byte data is stored in persistent memory can atomicity be guaranteed. Therefore, data consistency is ensured.

[0130] The data area is adjacent to the metadata area and is located in the address space after the metadata area. The address space of the data area is aligned in 4KB increments and is divided into multiple data units of 4KB size. The data units correspond one-to-one with the metadata units. Each metadata unit manages the corresponding data unit. The last 4KB of each data unit is used to store hard disk metadata and secondary metadata, and the remaining 4KB is used to store the data segments of the data. Figure 6 The hard disk metadata in it includes a hard disk identification field, a hard disk location field, a file identification, an object identification, a block identification, and a version number. The secondary metadata includes a write position, a data length, a status flag (corresponding to the aforementioned flag bit), and check data.

[0131] For Figure 6 Regarding the partitioning method of the address space in it, given the size of the persistent memory sizepmem, the starting addresses of the first reserved area, the metadata area, the data area, and the second reserved area can be determined, and the starting address of any metadata unit can be further obtained. Since the metadata units and data units correspond one-to-one, the starting address of any data unit can also be obtained to write the first metadata, the second metadata, and the data to be written into their respective address spaces. The method for determining the starting address will be specifically described below.

[0132] The first reserved area is the first 4KB of the address space of the persistent memory. Therefore, the starting address space addrsuper1 of the first reserved area is: addrsuper1 = 0.

[0133] The starting address space of the metadata area is: addrmetabase = addrsuper1 + sizesuper1.

[0134] To calculate the starting address of the data area, first calculate the total number of metadata units cntmeta, which is also the total number of data units cntdata. The calculation method for the total number of metadata units is:

[0135] cntmeta = cntdata = (sizepmem - sizesuper1 * 2) / (sizedata + sizemeta), where sizedata is the length of the data unit and sizemeta is the length of the metadata unit. Thus, the starting address of the data area is:

[0136] Addrdatabase = 0 + sizesuper1 + ceiling(cntmeta * sizemeta), where ceiling means aligning the obtained sum result to 4K, corresponding to the partitioning of the persistent memory address space Figure 6 and the pad alignment in it.

[0137] After determining the starting addresses of the metadata area and the data area, the positions of any metadata unit and any data unit can be calculated by multiplication.

[0138] Similarly, the starting address addrsuper2 of the second reserved area is:

[0139] addrsuper2 = Addrdatabase + cntdata * sizedata.

[0140] Combined with Figure 6 the example diagram of the space partitioning of the persistent memory shown, the embodiments of the present invention specifically describe how to determine and write the metadata units and data units corresponding to the first metadata, the second metadata, and the data to be written.

[0141] To more efficiently determine the metadata units and data units corresponding to the first metadata, the second metadata, and the data to be written, the embodiments of the present invention manage the metadata units in a hierarchical management manner. Each level corresponds to a linked list, and each linked list includes at least one management node. Each management node is used to manage adjacent metadata units. The number of metadata units managed by the management nodes in the same linked list is the same, and the number of metadata units managed by the management nodes in any two linked lists is different. The level of the linked list can be determined according to the maximum length of the data written at one time and the length of the data unit. For example, if the maximum length of the data written at one time is set to 256KB and the length of the data unit is 4KB, then the level of the linked list is: 1 - 256KB / 4KB, that is, 1 - 64KB. Please refer to Figure 7 , Figure 7 which is the example diagram of the hierarchical linked list provided by the embodiments of the present invention, Figure 7 in which, there are a total of 64 levels of linked lists, linked list 1 - linked list 64. The number of metadata units managed by one management node in each linked list is respectively: 1, 2, 3, 4,..., 64.

[0142] Based on Figure 3 , and at the same time combined with Figure 6Spatial partitioning and Figure 7 The example of the hierarchical linked list in Figure 7 illustrates the process of writing the first metadata into the metadata area. Please refer to Figure 8 , Figure 8 FIG. Figure 8 is a flowchart example of another data access method provided by an embodiment of the present invention. Step S103 includes the following sub-steps for writing the first metadata:

[0143] Sub-step S103-10: Determine the target level according to the data length and the preset length.

[0144] In this embodiment, according to the data length and the preset length, the number of data segments, that is, the number of segments, can be obtained. According to the number of segments, the initial level is determined. The initial level = the number of segments + 1. If there is a management node in the linked list corresponding to the initial level, the initial level is determined as the target level. If there is no management node in the linked list corresponding to the initial level, starting from the initial level, the level of the first linked list with a management node found is determined as the target level. For example, the number of segments is 5, the initial level is 5 + 1 = 6. There is no management node in linked list 6. Starting from linked list 6, the first linked list with a management node found is linked list 10, so the level 10 of linked list 10 is the target level.

[0145] Sub-step S103-11: Determine the first metadata unit to be written and the second metadata unit to be written from the management nodes of the linked list corresponding to the target level. Among them, the number of the first metadata units to be written is the number of data segments, and the number of the second metadata units to be written is 1.

[0146] In this embodiment, if there are multiple management nodes in the linked list corresponding to the target level, any one of the management nodes is selected, and then the first metadata unit to be written and the second metadata unit to be written are determined from this management node. The number of the first metadata units to be written is the number of data segments, that is, the number of segments.

[0147] Sub-step S103-12: Write the metadata of each data segment into each first metadata unit to be written in sequence according to the position of each data segment in the data to be written.

[0148] In this embodiment, the first metadata unit to be written and the second metadata unit to be written are contiguous in address. When there are multiple first metadata units to be written, the multiple first metadata units to be written are also contiguous in address. According to the position of each data segment in the data to be written, the metadata of each data segment is written into the first data unit to be written in sequence, that is, the metadata of the first data segment is written into the first first metadata unit to be written, the metadata of the second data segment is written into the second first metadata unit to be written, and so on.

[0149] Sub-step S103-13: Write the reserved metadata into the second metadata unit to be written.

[0150] In this embodiment, the second metadata unit to be written is adjacent to the last first metadata unit to be written.

[0151] Continuing to refer to Figure 8 , step S103 further includes the following sub-steps to write the second metadata:

[0152] Sub-step S103-20: Use the data unit corresponding to the second metadata unit to be written as the data unit to be written.

[0153] Sub-step S103-21: Write the second metadata into the data unit to be written.

[0154] It should be noted that the data segments are written into the data units corresponding to their metadata units in sequence according to their positions in the data to be written. For example, the first data segment is written into the data unit corresponding to the metadata unit of the first data segment. The first metadata and the second metadata are written simultaneously and are written only after all the data segments written in one go of the data to be written are written into the corresponding data units, thereby ensuring the data consistency of the data segments in the case of abnormal power-off.

[0155] It should also be noted that since the first metadata unit to be written and the second metadata unit to be written have been occupied, they need to be deleted from the corresponding management nodes in the corresponding linked list. If the sum of the number of the first metadata unit to be written and the second metadata unit to be written is equal to the number of metadata units managed by the management node, the corresponding management node is directly deleted. If the sum of the two is less than the number of metadata units managed by the management node, in addition to deleting them from the corresponding management nodes in the corresponding linked list, at this time, since the number of metadata units managed by the management node will change, in order to meet the hierarchical management of the linked list, the metadata unit needs to be inserted into the linked list at the corresponding level according to the number of metadata units in the changed management node. Please refer to Figure 9 , Figure 9 is an example diagram of the linked list update provided by the embodiment of the present invention. There are 5 first metadata units to be written and 1 second metadata unit to be written. Linked list 6 is empty and linked list 7 is not empty. Then, 6 metadata units are allocated from the management node in linked list 7, and there is still 1 metadata unit left in linked list 7. This remaining metadata unit is inserted from linked list 7 into linked list 1. It should be noted that Figure 9 only linked lists 1, 6, and 7 related to the example are drawn in , and the other linked lists are represented by ellipses, which does not mean that the other linked lists do not exist.

[0156] It should also be noted that, as a specific implementation, initially there can be only one level of linked list including management nodes. For example, there is only one management node in linked list 64, which manages all metadata units, and the linked lists of the remaining levels are empty. As management nodes are allocated, the management nodes of this level of linked list are gradually split and inserted into the corresponding level of linked list. Those skilled in the art can obtain the specific process without creative efforts according to the allocation process described above.

[0157] The embodiments of the present invention are also based on Figure 6 to provide a schematic diagram of the write process of data to be written. Please refer to Figure 10 , Figure 10 which is a schematic diagram of the write process of data to be written provided by the embodiments of the present invention. Figure 10 In [the figure], the data length of the data to be written is 254KB. Since 245KB is not an integer multiple of 4KB, it is necessary to read data from the hard disk and pad it to an integer multiple of 4KB, which is 256KB. Refer to the example in the diagonal box in the figure. 256KB is split into 64 4KBs, occupying 65 metadata units in the metadata area and 65 data units in the data area. Figure 10 Only the field values of the metadata units of the first data segment, the metadata units reserved for metadata, and the data units corresponding to the first data segment are exemplified in [the figure]. The other data segments are similar and will not be elaborated here.

[0158] In this embodiment, for the application scenario where the data to be written is temporarily stored in persistent memory and finally stored in the hard disk, it is necessary to flush (or store) the data to be written in persistent memory to the hard disk. Once the data to be written is successfully flushed to the hard disk, in order to timely release the space in persistent memory occupied by the data to be written, it is necessary to release the metadata units and the corresponding data units occupied by the data to be written for use when writing other data to be written later. To achieve more efficient disk flushing and release, the embodiments of the present invention introduce a disk flushing list and a reclaim list. When the data to be written is written to persistent memory, the data to be written and the data units storing the data to be written are also added to the disk flushing list. After the data in the disk flushing list is flushed to the hard disk, the data units corresponding to the flushed data are added to the reclaim list for reclamation. As a specific implementation, the disk flushing process and the reclaim process can be periodically executed by two independent threads respectively. The embodiments of the present invention also provide a specific implementation of the disk flushing process. Please refer to Figure 11 , Figure 11 which is a flowchart example of another data access method provided by the embodiments of the present invention. This method further includes the following steps:

[0159] Step S200, determine valid data units and invalid data units according to the write position and the disk flushing position.

[0160] In this embodiment, the valid data unit is the data unit that needs to be written to the hard disk, and the invalid data unit is the data unit that does not need to be written to the hard disk. For example, if the first write request writes data A at position 1, then A is written to the persistent memory. Before A is written to the hard disk, if write request 2 writes data B at position 1, then when actually writing to the disk, data A does not need to be written to the hard disk, and only the latest data B needs to be written to the hard disk.

[0161] In this embodiment, the disk write list may include multiple data to be written, and each data to be written corresponds to a disk write position. The position to be written needs to be compared one by one with each disk write position in the disk write list to determine all valid data units and invalid data units. The method of comparing the position to be written with any disk write position is the same. Only the comparison between the position to be written and any disk write position will be described below.

[0162] According to the differences between the position to be written and the disk write position, it can be divided into the following three cases, and each case includes three different situations. Each situation of each case will be described in detail below.

[0163] For the convenience of description, the position to be written and the disk write position are represented by login and logcmp respectively. Both the position to be written and the disk write position include their respective start offset beginOffset and end offset endOffset. To facilitate distinguishing the writing order of the data to be written and the data to be written to the disk, both the data to be written and the data to be written to the disk carry their respective version numbers. The larger the version number, the later the corresponding data is written and the newer the data is. According to the start offset beginOffset and end offset endOffset of the position to be written and the disk write position, the valid data and invalid data therein can be determined. The data unit corresponding to the valid data is the valid data unit, and the data unit corresponding to the invalid data is the invalid data unit. The specific implementation methods for determining the invalid data and valid data will be described in detail by case and by situation below.

[0164] Case 1: The start offset of login is less than the start offset of logcmp

[0165] Case 1.1: The end offset of login is less than the end offset of logcmp

[0166] In this case, it is necessary to compare the version numbers of login and logcmp. If the version number of login is less than the version number of logcmp, then it means that the data of logcmp in login is invalid data, and the rest of the data is valid data; otherwise, assign the end offset of login to the start offset of logcmp to indicate that the data of login in logcmp is valid.

[0167] In this embodiment, the embodiments of the present invention illustrate each situation and case with diagrams. Please refer to Figure 12 , Figure 12 which is a schematic diagram of various methods for determining invalid data and valid data provided by the embodiments of the present invention. For the sake of easy description, Figure 12 in all the figures, the white part in the rectangular box represents valid data, and the shaded diagonal part represents invalid data. The schematic diagram of valid data and invalid data in case 1.1 is as shown in Figure 12 (a).

[0168] Case 1.2: The end offset of login is equal to the end offset of logcmp

[0169] In this case, it is necessary to compare the version numbers of login and logcmp. If the version number of login is less than the version number of logcmp, it means that the data of logcmp in login is invalid data, and the start offset of logcmp is assigned to the end offset of login to represent it, and the rest of the data is valid data; otherwise, it means that all the data in logcmp is invalid data. The schematic diagram of valid data and invalid data in case 1.2 is as shown in Figure 12 (b).

[0170] Case 1.3: The end offset of login is greater than the end offset of logcmp

[0171] In this case, it is necessary to compare the version numbers of login and logcmp. If the version number of login is less than the version number of logcmp, it means that the data of logcmp in login is invalid data, and the start offset of logcmp is assigned to the end offset of login to represent it; otherwise, it means that all the data in logcmp is invalid data. The schematic diagram of valid data and invalid data in case 1.3 is as shown in Figure 12 (c).

[0172] Case 2: The start offset of login is less than the start offset of logcmp

[0173] Case 2.1: The end offset of login is less than the end offset of logcmp

[0174] In this case, it is necessary to compare the version numbers of login and logcmp. If the version number of login is less than the version number of logcmp, it means that all the data in login is invalid data; otherwise, the end offset of login is assigned to the start offset of logcmp to indicate the valid data of logcmp. The schematic diagram of valid data and invalid data in case 2.1 is as shown in Figure 12 (d).

[0175] Case 2.2: The end offset of login is equal to the end offset of logcmp

[0176] In this case, it is necessary to compare the version numbers of login and logcmp. If the version number of login is less than the version number of logcmp, it means that the data in the entire login is invalid data and the data in the entire logcmp is valid data. Otherwise, it means that the data in the entire logcmp is invalid data and the data in the entire login is valid data. The schematic diagrams of valid and invalid data in Case 2.2 are as Figure 12 shown in (e).

[0177] Case 2.3: The end offset of login is greater than the end offset of logcmp

[0178] In this case, it is necessary to compare the version numbers of login and logcmp. If the version number of login is less than the version number of logcmp, it means that the data of logcmp in login is invalid data and the rest of the data is valid data. Assign the end offset of logcmp to the start offset of login to represent it. Otherwise, it means that the data in the entire logcmp is invalid data and the rest of the data is valid data. The schematic diagrams of valid and invalid data in Case 2.3 are as Figure 12 shown in (f).

[0179] Case 3: The start offset of login is greater than the start offset of logcmp

[0180] Case 3.1: The end offset of log in is less than the end offset of logcmp

[0181] In this case, it is necessary to compare the version numbers of login and logcmp. If the version number of login is less than the version number of logcmp, it means that the data in the entire login is invalid data and the rest of the data is valid data. Otherwise, it means that the data of logcmp in login is invalid data and the rest of the data is valid data. The schematic diagrams of valid and invalid data in Case 3.1 are as Figure 12 shown in (g).

[0182] Case 3.2: The end offset of login is equal to the end offset of logcmp

[0183] In this case, it is necessary to compare the version numbers of login and logcmp. If the version number of login is less than that of logcmp, it means that the data in the entire login is invalid data; otherwise, it means that the data of logcmp in login is invalid data, and the start offset of login is assigned to the end offset of logcmp to represent it. The schematic diagrams of valid and invalid data in Case 3.2 are as Figure 12 (g) shows.

[0184] Case 3.3: The end offset of login is greater than the end offset of logcmp

[0185] In this case, it is necessary to compare the version numbers of login and logcmp. If the version number of login is less than that of logcmp, it means that the data of logcmp in log in is invalid data, and the rest of the data is valid data. The end offset of logcmp is assigned to the start offset of login to represent it; otherwise, it means that the data of login in logcmp is invalid data, and the rest of the data is valid data. The start offset of login is assigned to the end offset of logcmp to represent it. The schematic diagrams of valid and invalid data in Case 3.3 are as Figure 12 (h) shows.

[0186] It should be noted that as a specific implementation, for invalid data, the flag bit in the second metadata corresponding to the invalid data needs to be set from "cached" to "none" to recycle the data unit corresponding to the invalid data.

[0187] Step S201, if the valid data unit does not exist in the disk flushing list, update the valid data unit to the disk flushing list to flush the data in the valid data unit from the persistent memory to the hard disk through the disk flushing list.

[0188] It should be noted that in the above 9 cases, if the existing valid data units in the disk flushing list change, the first metadata and the second metadata of the changed valid data units need to be updated; if a part of the data unit corresponding to the data to be written becomes invalid, the first metadata and the second metadata of the data to be written need to be updated correspondingly; and the invalid data unit is inserted into the recycling list to recycle it through the recycling list; if all the data units corresponding to the data to be written become invalid, it needs to be inserted into the recycling list to recycle the invalid data units through the recycling list.

[0189] Step S202, if an invalid data unit exists in the disk writing list, delete the invalid data unit from the disk writing list and insert it into the recycling list, so as to recycle the invalid data unit through the recycling list.

[0190] In this embodiment, in order to recycle data units in a timely and effective manner, and at the same time prevent the persistent memory space from being overly fragmented, the embodiment of the present invention further provides a specific implementation manner for inserting into the recycling list:

[0191] First, if there are mergeable data units in the recycling list that meet the preset merge conditions, delete the mergeable data units from the recycling list.

[0192] In this embodiment, meeting the preset merge conditions may mean that the addresses of the data units in the recycling list and the invalid data units are continuous. The addresses being continuous can mean that the end address of the invalid data unit and the start address of the data unit in the recycling list are continuous, or the start address of the invalid data unit and the end address of the data unit in the recycling list are continuous, or the start address of the invalid data unit and the end address of one data unit are continuous, and the end address of the invalid data unit and the start address of another data unit are continuous.

[0193] In this embodiment, in order to quickly find the mergeable data units that meet the preset merge conditions, the red - black tree method can be used to manage the data units in the persistent memory.

[0194] Secondly, merge the mergeable data units with the invalid data units to obtain the data units to be inserted.

[0195] Thirdly, insert the data units to be inserted into the recycling list.

[0196] Finally, if there are no mergeable data units in the recycling list, insert the invalid data units into the recycling list.

[0197] It can be understood that when recycling the data units in the recycling list, since the data units and the metadata units are in one - to - one correspondence, it is also necessary to update the information in the metadata units corresponding to the data units accordingly. For example, set the values of each field in the metadata unit to 0 or invalid values.

[0198] It should be noted that, as an implementation, the recycle list can be a list independent of the rank linked list. The process of recycling the recycle list is actually a process of inserting the metadata units corresponding to the data units in the recycle list into the rank linked list. As another implementation, the recycle list can be the rank linked list itself. Inserting the data unit to be inserted into the recycle list is the process of inserting the metadata unit corresponding to the data unit to be inserted into the rank linked list. When a new metadata unit is inserted into the rank linked list, it may cause changes to the nodes in the rank linked list. Please refer to Figure 13 , Figure 13 which is an example diagram of the insertion process of the rank linked list provided by an embodiment of the present invention. Figure 13 In it, the metadata unit 3 to be inserted, the linked list 2 includes metadata units 1 and 2, and the linked list 3 includes metadata units 4, 5, and 6. Since both metadata unit 2 and metadata unit 4 are consecutive with metadata unit 3, the metadata units 1 - 6 can be merged. After merging, the metadata units 1 - 6 are used as a management node and inserted into the linked list 6.

[0199] The above method provided by this embodiment can avoid the disk flushing of invalid data, reduce the amount of data written to the hard disk, thereby reducing the impact on the write performance. At the same time, it timely releases the data units corresponding to the invalid data to improve the utilization rate of the persistent memory. When inserting into the recycle list, it tries to merge the data units that need to be recycled, effectively reducing the fragmentation of the persistent memory space.

[0200] In this embodiment, when reading the data stored in the storage node 10, the data to be read may be stored in the persistent memory, may be stored in the hard disk, or may be partly in the persistent memory and partly in the hard disk. To correctly read the data, an embodiment of the present invention also provides a specific implementation method for reading data. Please refer to Figure 14 , Figure 14 which is a flow example diagram of another data access method provided by an embodiment of the present invention. This method further includes the following steps:

[0201] Step S300: Receive a read data request sent by the client. Among them, the read data request includes the data length and the read position of the data to be read.

[0202] Step S301: Read the original data from the hard disk according to the data length and the read position of the data to be read.

[0203] Step S302: Determine whether there is the latest data corresponding to the data to be read and not stored in the hard disk in the disk flushing list according to the read position.

[0204] Step S303: Combine the original data and the latest data to obtain the data to be read.

[0205] In this embodiment, the original data and the latest data may overlap or not. When they overlap, the overlapping part of the original data is replaced with the latest data to obtain the data to be read. When they do not overlap, the original data and the latest data are concatenated to obtain the data to be read.

[0206] In this embodiment, if there is no latest data in the brush disk list that corresponds to the data to be read and has not been stored in the hard disk, it indicates that the original data read from the hard disk is already the latest data. At this time, the original data is the data to be read.

[0207] In this embodiment, in the actual operating environment, storage node 10 may suddenly lose power. After the power loss, if the data being written is only half-written, when storage node 10 restarts, reading the position of the data being written at the time of power loss may result in data that is neither the latest nor the second latest. At this time, data inconsistency will occur. To avoid data inconsistency, the embodiment of the present invention also provides a process for reconstructing data after power-on. Please refer to Figure 15 , Figure 15 which is a flowchart example of another data access method provided by the embodiment of the present invention. The method includes the following steps:

[0208] Step S400, when the storage node is powered on, divide multiple metadata units into at least one metadata unit group according to the log identifier.

[0209] In this embodiment, the log identifier may be an identification field stored in the metadata unit. For the data to be written to the disk for the same write operation, the log identifiers in the corresponding metadata units are the same.

[0210] In this embodiment, when storage node 10 is powered off, the persistent memory stores the data to be written to the hard disk. The data to be written to the disk exists in at least one data unit and corresponds to at least one metadata unit.

[0211] Step S401, perform reconstruction processing on each metadata unit group to write the data to be written to the disk into the hard disk.

[0212] In this embodiment, reconstruction can be performed according to the information recorded in the metadata units of each metadata unit group.

[0213] Before specific reconstruction, to ensure data reliability, as a specific implementation, the information in the first reserved area or the second reserved area can be read first. By comparing the CheckSum field, the magic field, and the size field, it is determined whether the information is complete. If it is not complete, it is considered the first initialization of the persistent memory, and the entire persistent memory is initialized.

[0214] During initialization, traverse each metadata unit, assign fields other than the CheckSum field to 0, calculate the checksum value of the metadata unit and write it into the CheckSum field. Then, take the entire metadata area as a shard, and generate a hierarchical list and the corresponding red-black tree according to the recycling process in the foregoing embodiment. Finally, persist the first 8 bytes (log identifier field) of the metadata unit, and after completion, persist the last 8 bytes (count field, index field, and CheckSum field) of the metadata unit.

[0215] If the information in the first reserved area and the second reserved area is complete, compare the size field to see if it is consistent with the current size of the persistent memory. If it is inconsistent, an error is considered to have occurred. If it is consistent, start the specific reconstruction process.

[0216] As a specific implementation, the specific reconstruction process can be:

[0217] First, for any target metadata unit group in at least one metadata unit group, reconstruct the target metadata unit group according to the number of shards and the shard index of the target metadata unit in the target metadata unit group to obtain the log corresponding to the target metadata unit group.

[0218] In this embodiment, as a specific implementation, the number of shards and the shard index can be obtained from the count field and the index field in the metadata unit respectively. Thus, the metadata units can be sorted according to the value of the index field, and then the corresponding sorted data units can be obtained.

[0219] Second, read the status of the log from the data unit managed by the target metadata unit with the last shard index, where the status of the log is used to indicate whether the target data in the data units managed by the target metadata units other than the last shard index exists in the data area.

[0220] In this embodiment, since the data unit stored by the last target metadata unit manages the second metadata, which includes a flag bit, a data length, a write position to be written, and check data, the flag bit, the data length, and the write position to be written can be verified through the check data, and subsequent processing can be performed after the verification passes.

[0221] Third, if the status of the log is the cache status used to indicate that the target data exists in the data area, flush the target data to the hard disk.

[0222] Finally, flush the data in the data units managed by the metadata units in all metadata unit groups with the log status of the cache status to the hard disk, and finally write the data to be written to the disk to the hard disk.

[0223] In this embodiment, if the status of the log is a non-cached status indicating that the target data does not exist in the data area, at this time, the data unit of the target metadata unit needs to be recycled, and the specific processing is as follows:

[0224] If the status of the log is a non-cached status indicating that the target data does not exist in the data area, then the data unit of the target metadata unit is recycled.

[0225] Recycle the data units managed by the metadata units in all metadata unit groups with the log status being the non-cached status.

[0226] In this embodiment, the recycling of data units has been described in step S202 of the foregoing embodiment, and will not be elaborated here.

[0227] In order to execute the corresponding steps in the above embodiments and each possible implementation manner, an implementation manner of a data access device 100 is given below. Please refer to Figure 16 , Figure 16 which shows a block diagram of the data access device 100 provided in the embodiment of the present invention. It should be noted that the basic principle and the technical effects generated by the data access device 100 provided in this embodiment are the same as those in the above embodiments. For the sake of brief description, some parts of this embodiment are not mentioned.

[0228] The data access device 100 includes a receiving module 110, a generating module 120, a writing module 130, a reading module 140, a disk flushing module 150, and a reconstruction module 160.

[0229] The receiving module 110 is configured to receive a write data request sent by a client. Among them, the write data request includes the data length and the write position of the data to be written.

[0230] Optionally, the receiving module 110 is further configured to receive a read data request sent by a client. Among them, the read data request includes the data length and the read position of the data to be read.

[0231] The generating module 120 is configured to generate first metadata for managing the data to be written according to the data length.

[0232] Optionally, the generating module 120 is specifically configured to: calculate the number of segments according to the data length and a preset length; split the data to be written according to the number of segments to obtain at least one data segment; generate metadata for each data segment according to the number of data segments and the position of each data segment in the data to be written; generate a reserved metadata for the data to be written, where the reserved metadata includes the value after adding 1 to the number of data segments; and use the reserved metadata and the metadata of all data segments as the first metadata.

[0233] The generating module 120 is further configured to generate second metadata for managing the writing of the data to be written into the log according to the data length and the writing position.

[0234] Optionally, the generating module 120 is specifically further configured to: obtain a flag bit for characterizing the successful writing of the data to be written into the data area; generate verification data according to the flag bit, the data length, and the writing position; and use the flag bit, the data length, the writing position, and the verification data as the second metadata.

[0235] The writing module 130 is configured to, after writing the data to be written into the data area, write the first metadata into the metadata area and write the second metadata into the data area.

[0236] Optionally, the metadata area includes multiple metadata units, and the multiple metadata units are hierarchically managed. Each level corresponds to a linked list, and each linked list includes at least one management node. Each management node is used to manage the metadata units with adjacent positions. The number of metadata units managed by the management nodes in the same linked list is the same, and the number of metadata units managed by the management nodes in any two linked lists is different. The writing module 130 is specifically configured to: determine a target level according to the data length and a preset length; determine a first metadata unit to be written and a second metadata unit to be written from the management nodes of the linked list corresponding to the target level, where the number of the first metadata units to be written is the number of data segments, and the number of the second metadata units to be written is 1; write the metadata of each data segment into each first metadata unit to be written in sequence according to the position of each data segment in the data to be written; and write the reserved metadata into the second metadata unit to be written.

[0237] Optionally, the data area includes multiple data units, each metadata unit corresponds to a data unit, and each data segment is written into the data unit corresponding to each target metadata unit. When the writing module 130 writes the second metadata into the data area, it is specifically configured to: use the data unit corresponding to the second metadata unit to be written as the data unit to be written; and write the second metadata into the data unit to be written.

[0238] Optionally, the reading module 140 is configured to: read the original data from the hard disk according to the data length and the reading position of the data to be read; determine whether there is the latest data corresponding to the data to be read and not stored in the hard disk in the disk flushing list according to the reading position; and combine the original data and the latest data to obtain the data to be read.

[0239] Optionally, the data area includes a plurality of data units, and the storage node further includes a hard disk, a disk flushing list, and a recycling list. The disk flushing list includes the disk flushing positions of the data to be flushed and the data units storing the data to be flushed. The data to be flushed is the data stored in the persistent memory and to be flushed into the hard disk. The disk flushing module 150 is configured to: determine valid data units and invalid data units according to the write position and the disk flushing positions; if the valid data units are not present in the disk flushing list, update the valid data units to the disk flushing list so as to flush the data in the valid data units from the persistent memory into the hard disk through the disk flushing list; if the invalid data units are present in the disk flushing list, delete the invalid data units from the disk flushing list and insert them into the recycling list so as to recycle the invalid data units through the recycling list.

[0240] Optionally, the disk flushing module 150 is specifically configured to: if there are data units to be merged that meet the preset merging conditions in the recycling list, delete the data units to be merged from the recycling list; merge the data units to be merged with the invalid data units to obtain data units to be inserted; insert the data units to be inserted into the recycling list; if there are no data units to be merged in the recycling list, insert the invalid data units into the recycling list.

[0241] Optionally, the storage node further includes a hard disk. When the storage node is powered off, the persistent memory stores data to be written to the hard disk, which is called data to be disk-written. The metadata area includes a plurality of metadata units, and each metadata unit includes a log identifier. The data area includes a plurality of data units, and each metadata unit manages one data unit. The data to be disk-written exists in at least one data unit. The reconstruction module 160 is configured to: when the storage node is powered on, divide the plurality of metadata units into at least one metadata unit group according to the log identifiers; perform reconstruction processing on each metadata unit group to write the data to be disk-written into the hard disk.

[0242] Optionally, each metadata unit includes the number of shards, the shard index, and the log identifier. The reconstruction module 160 is specifically configured to: for any target metadata unit group in at least one metadata unit group, reconstruct the target metadata unit group according to the number of shards and the shard index of the target metadata units in the target metadata unit group to obtain the log corresponding to the target metadata unit group; read the status of the log from the data unit managed by the target metadata unit with the last shard index, where the status of the log is used to indicate whether the target data in the data units managed by the target metadata units other than the last shard index in the target metadata unit exists in the data area; if the status of the log is the cached status indicating that the target data exists in the data area, flush the target data into the hard disk; flush the data in the data units managed by the metadata units in all metadata unit groups with the log status being the cached status into the hard disk, and finally write the data to be disk-written into the hard disk.

[0243] Optionally, the reconstruction module 160 is further specifically configured to: if the status of the log indicates a non-cached state where the target data does not exist in the data area, recycle the data unit of the target metadata unit; recycle the data units managed by the metadata units in all metadata unit groups whose log status is in the non-cached state.

[0244] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the data access method as described above is implemented.

[0245] In summary, an embodiment of the present invention provides a data access method, apparatus, storage node, and storage medium, which are applied to a storage node. The storage node includes a persistent memory, the persistent memory includes a metadata area and a data area, and the storage node is communicatively connected to a client. The method includes: receiving a write data request sent by the client, where the write data request includes the data length of the data to be written and the write position; generating first metadata for managing the data to be written according to the data length; generating second metadata for managing the write log of the data to be written according to the data length and the write position; after writing the data to be written into the data area, writing the first metadata into the metadata area and writing the second metadata into the data area. Compared with the prior art, the embodiment of the present invention can ensure atomicity when writing data to the persistent memory through the first metadata and the second metadata, thereby ensuring the consistency of the data in the persistent memory.

[0246] The above is only a specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A data access method, characterized in that, it is applied to a storage node, the storage node includes persistent memory, the persistent memory includes a metadata area and a data area, the storage node is communicatively connected to a client, and the method includes: Receiving a write data request sent by the client, where the write data request includes the data length and the write position of the data to be written; Generating first metadata for managing the data to be written according to the data length, the data to be written includes multiple data segments, the first metadata includes the metadata of each data segment and reserved metadata, the reserved metadata includes the value after adding 1 to the number of data segments, and the position of the reserved metadata is after the last data segment among the multiple data segments; Generating second metadata for managing the write log of the data to be written according to the data length and the write position; After writing the data to be written into the data area, writing the first metadata into the metadata area and writing the second metadata into the data area.

2. The data access method according to claim 1, characterized in that, the step of generating first metadata for managing the data to be written according to the data length includes: Calculating the number of segments according to the data length and a preset length; Dividing the data to be written according to the number of segments to obtain at least one data segment; Generating the metadata of each data segment according to the number of data segments and the position of each data segment in the data to be written; Generating a reserved metadata for the data to be written; Taking the reserved metadata and the metadata of all data segments as the first metadata.

3. The data access method according to claim 1, characterized in that, the step of generating second metadata for managing the write log of the data to be written according to the data length and the write position includes: Obtaining a flag bit for indicating successful writing of the data to be written into the data area; Generating verification data according to the flag bit, the data length and the write position; Taking the flag bit, the data length, the write position and the verification data as the second metadata.

4. The data access method according to claim 2, characterized in that, the metadata area includes multiple metadata units, the multiple metadata units are hierarchically managed, each level corresponds to a linked list, each linked list includes at least one management node, each management node is used to manage metadata units with adjacent positions, the number of metadata units managed by management nodes in the same linked list is the same, and the number of metadata units managed by management nodes in any two linked lists is different. The step of writing the first metadata into the metadata area includes: Determining a target level according to the data length and the preset length; Determining a first metadata unit to be written and a second metadata unit to be written from the management nodes of the linked list corresponding to the target level, where the number of the first metadata units to be written is the number of data segments, and the number of the second metadata units to be written is 1; According to the position of each said data segment in the data to be written, write the metadata of each said data segment into each said first metadata unit to be written in sequence; Write the reserved metadata into the second metadata unit to be written.

5. The data access method according to claim 4, wherein, the data area includes a plurality of data units, each metadata unit corresponds to a data unit, and each said data segment is written into the data unit corresponding to each said metadata unit. The step of writing the second metadata into the data area includes: Regarding the data unit corresponding to the second metadata unit to be written as the data unit to be written; Write the second metadata into the data unit to be written.

6. The data access method according to claim 1, wherein, the data area includes a plurality of data units, the storage node further includes a hard disk, a disk flushing list and a recycling list. The disk flushing list includes the disk flushing position of the data to be flushed and the data unit storing the data to be flushed. The data to be flushed is the data stored in the persistent memory and to be flushed into the hard disk. The disk flushing position is the position to be written in the write data request for writing the data to be flushed. The method further includes: Determine valid data units and invalid data units according to the position to be written and the disk flushing position; If the valid data units do not exist in the disk flushing list, update the valid data units to the disk flushing list, so as to flush the data in the valid data units from the persistent memory into the hard disk through the disk flushing list; If the invalid data units exist in the disk flushing list, delete the invalid data units from the disk flushing list and insert them into the recycling list, so as to recycle the invalid data units through the recycling list.

7. The data access method according to claim 6, wherein, The step of deleting the invalid data units from the disk flushing list and inserting them into the recycling list includes: If there are data units to be merged that meet the preset merging conditions in the recycling list, delete the data units to be merged from the recycling list; Merge the data units to be merged with the invalid data units to obtain data units to be inserted; Insert the data units to be inserted into the recycling list; If the data units to be merged do not exist in the recycling list, insert the invalid data units into the recycling list.

8. The data access method according to claim 6, wherein, the method further includes: Receive a read data request sent by the client, wherein the read data request includes the data length and the position to be read of the data to be read; Read the original data from the hard disk according to the data length and the position to be read of the data to be read; Determine whether there is the latest data corresponding to the data to be read and not stored in the hard disk in the disk flushing list according to the position to be read; Combine the original data and the latest data to obtain the data to be read.

9. The data access method according to claim 1, wherein, The storage node further includes a hard disk. When the storage node is powered off, the persistent memory stores the data to be written to the hard disk, which is called data to be disk-written. The metadata area includes a plurality of metadata units, and each metadata unit includes a log identifier. The data area includes a plurality of data units, and each metadata unit manages one of the data units. The data to be disk-written exists in at least one of the data units. The method further includes: When the storage node is powered on, divide the plurality of metadata units into at least one metadata unit group according to the log identifier; Perform a reconstruction process on each metadata unit group to write the data to be disk-written to the hard disk.

10. The data access method according to claim 9, wherein, Each metadata unit includes the number of shards, a shard index, and a log identifier. The step of performing a reconstruction process on each metadata unit group to write the data to be disk-written to the hard disk includes: For any target metadata unit group in the at least one metadata unit group, reconstruct the target metadata unit group according to the number of shards and the shard index of the target metadata unit in the target metadata unit group to obtain the log corresponding to the target metadata unit group; Read the status of the log from the data unit managed by the target metadata unit with the last shard index, where the status of the log is used to indicate whether the target data in the data units managed by the target metadata units other than the last shard index exists in the data area; If the status of the log is a cached status indicating that the target data exists in the data area, flush the target data to the hard disk; Flush the data in the data units managed by the metadata units in all the metadata unit groups with the log status being the cached status to the hard disk, and finally write the data to be disk-written to the hard disk.

11. The data access method according to claim 10, wherein, The method further includes: If the status of the log is a non-cached status indicating that the target data does not exist in the data area, recycle the data unit of the target metadata unit; Recycle the data units managed by the metadata units in all the metadata unit groups with the log status being the non-cached status.

12. A data access device, wherein, Applied to a storage node, the storage node includes persistent memory, the persistent memory includes a metadata area and a data area, the storage node is communicatively connected to a client, and the device includes: A receiving module, configured to receive a write data request sent by the client, where the write data request includes the data length and the write position of the data to be written; A generation module, configured to generate first metadata for managing the data to be written according to the data length, where the data to be written includes a plurality of data segments, the first metadata includes metadata of each data segment and reserved metadata, the reserved metadata includes a value obtained by adding 1 to the number of data segments, and the position of the reserved metadata is after the last data segment among the plurality of data segments; The generation module is further configured to generate second metadata for managing a write log of the data to be written according to the data length and the write position; A write module, configured to, after writing the data to be written into the data area, write the first metadata into the metadata area and write the second metadata into the data area.

13. A storage node, characterized in that it includes a processor and a memory; the memory is used for storing a program; the processor is configured to, when executing the program, implement the data access method according to any one of claims 1-11.

14. A computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, it implements the data access method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Array reconstruction method and device based on metadata

    CN104598171A

  • Distributed storage CEPH based erasure correction code overwriting method

    CN105930103A

  • I / O self-adaption-based file system log mode

    CN105956090A