Method for realizing data compression in RAID (Redundant Array of Independent Disks) card and RAID control chip
By combining IO merging and data compression technology within the RAID card with RAID redundancy, the high cost and high power consumption issues of SSD compression solutions are solved, achieving more efficient data compression and redundancy protection, simplifying host configuration, and reducing CPU load.
Patent Information
- Application Number
- CN202511041667.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-31
AI Technical Summary
Existing SSD compression solutions result in high SSD costs and increased power consumption. They also fail to achieve redundancy protection and data compression coordination among multiple SSDs. Host-side compression solutions consume CPU resources and are highly complex. Furthermore, OS data cannot be restored after compression, causing the system to fail to boot.
By using IO merging, data compression, and IO redirection technologies within the RAID card, combined with RAID redundancy, online compression is achieved. The compression function is offloaded to the CPU and above the SSD, simplifying host configuration.
It reduces SSD wear and power consumption, improves compression performance, reduces wear on SSD chips, simplifies host configuration, reduces CPU load, and achieves better redundancy protection and compression performance.
Smart Images

Figure CN120872247A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data compression technology, specifically to a method for implementing data compression within a RAID card and a RAID control chip. Background Technology
[0002] With the development of storage media technology, SSD solutions are increasingly being adopted for server boot devices. Mainstream SSDs are categorized into SATA and NVMe based on their access interface protocols. Regardless of the type, the primary storage medium for SSDs is NAND flash. It is well known that write operations to NAND flash memory chips wear down their lifespan. Therefore, the FTL (Folded Layer Transmission) module in SSD firmware design primarily aims to reduce write amplification, balance wear between chips, and increase the lifespan of SSD chips and data retention time. However, the FTL module cannot solve all NAND flash wear problems. To reduce the amount of data written to NAND flash, some manufacturers implement hardware compression algorithms within the SSD. These algorithms operate in two modes: one is online data compression, where data compression is performed during the host write I / O process. The advantage of this approach is less wear on the chips, but the disadvantages are that it requires memory for I / O caching, the data compression ratio is not high, and the hardware cost is relatively high. The other approach is background data compression, where data is written to NAND flash memory first. The data to be moved is compressed during background garbage collection (GC) in the FTL (Flash Memory Transfer Layer). The advantage of this approach is that it doesn't require memory for data caching, and the data compression ratio can be relatively high. The disadvantage is that the initial data written to the host needs to be stored in the NAND flash, increasing the amount of data written. In short, compression on SSDs requires an additional hardware compression engine, posing significant challenges to SSD cost and power consumption control. Furthermore, this approach cannot achieve data protection across SSDs or handle SSD failures. For the host system, RAID technology is used to prevent single-disk failures. If data distribution across multiple disks uses replication or striping, it can lead to wasted space for user data (replication) or data skipping within a single SSD (striping), significantly reducing the effectiveness of single-disk data compression.
[0003] On the other hand, SSD-based host software is also implementing flash-awareness storage solutions to reduce write pressure on the SSD. One approach is to implement open channels in the host software, moving the FTL layer to the host driver. The main idea is to optimize I / O write addresses by aggregating write I / O and converting random write I / O into sequential large I / O, reducing the number of random write I / Os to the SSD. This reduces fragmented data on the NAND flash, thereby reducing the probability of SSD GC and maintaining a better internal data distribution. Another approach is to consider reducing the amount of data written. Through data reduction technology, the content that needs to be written is merged on the host side, reducing the total amount of data written to the SSD and thus extending the SSD's lifespan. Since host machines have more abundant hardware resources than SSDs, implementing these solutions is technically easier. However, due to the flexible and complex usage environment of host machines, the solution of using SSDs for migration is subject to interference from the host machine's software environment, which may have a very large impact and make it impossible to implement. To implement deduplication and compression on the host machine's system disk, if a software solution is adopted, it is necessary not only to modify the OS block device driver, but also to consider modifying the BIOS / UEFI block device driver to ensure that the compressed OS image can be correctly loaded into memory. Secondly, during OS operation, deduplication and compression consume significant CPU resources, which may affect the performance of applications.
[0004] In summary, current SSD in-disk compression solutions increase the complexity of SSD implementation, leading to higher SSD costs, increased power consumption and heat generation, and an inability to achieve redundancy protection and coordinated data compression across multiple SSDs. On the other hand, implementing compression solutions on the host side requires utilizing the host's CPU computing resources, which can cause contention for system computing power. If OS data needs to be compressed, the BIOS / UEFI also needs to implement corresponding functions; otherwise, the compressed OS data cannot be restored, causing the system to fail to boot. This further increases the complexity of BIOS / UEFI implementation. Summary of the Invention
[0005] To overcome the aforementioned technical problems in the prior art, this invention provides a method for implementing data compression within a RAID card and a RAID control chip. By employing IO merging, data compression, and IO redirection techniques on the RAID card, online compression is achieved. Furthermore, RAID redundancy is implemented on the hardware side, solving the high power consumption problem of existing SSD compression solutions. This results in better reliability, closer resemblance to the original user data content, and superior compression performance. Simultaneously, by offloading the compression function to a location below the CPU and above the SSD, the online compression capability is better utilized, while also reducing wear on the SSD chips, simplifying host configuration, and making the system more environmentally friendly.
[0006] To achieve the above objectives, this invention provides a method for implementing data compression within a RAID card, comprising the following steps: S1: Receiving multiple write IO commands from the host and obtaining the logical block address and length of each IO; S2: Performing a merging operation on the multiple IOs to generate a merged IO block; S3: Performing a data compression operation on the merged IO block through a data compression engine to obtain compressed data; S4: Establishing a mapping relationship between the logical block address and the physical block address of the compressed data, and storing it in a logical-to-physical mapping table; S5: Writing the compressed data to a solid-state drive through the RAID engine, and simultaneously persistently storing the metadata of the logical-to-physical mapping table; S6: Sending a reclamation command or a write-zero command to the solid-state drive to reclaim unused address space after compression.
[0007] Preferably, the operating conditions of S2 are: the logical block addresses of multiple IOs are consecutive; the total length of the merged IO block is greater than or equal to a preset threshold; and the starting address and length of the merged IO block both satisfy 4KB boundary alignment.
[0008] Preferably, the structure of the logical-to-physical mapping table in S4 includes a RAID address field, a compression flag field, and a data length field. The total address length of the RAID address field, the compression flag field, and the data length field is 64 bits. The RAID address field is in 4KB units. The compression flag field is used to identify whether the data is compressed. The data length field is used to record the actual length of the compressed data, which is padded to 4KB when the data is downloaded to disk.
[0009] Preferably, S5 specifically includes: the logical-to-physical mapping table occupies a preset capacity of total storage space, and the metadata is protected by RAID 1 data mirroring; the metadata is stored separately from user data, and the persistent metadata space is not visible to the host.
[0010] Preferably, the logical-to-physical mapping table is stored in 4KB pages in the metadata persistence space. The addressing method includes: obtaining the logical block address of the access request; dividing the logical block address by 512 to calculate the page number of the metadata corresponding to the logical block address; reading the corresponding logical-to-physical mapping table page number data from the metadata persistence RAID group according to the page number; taking the remainder of the logical block address by 512 to calculate the offset of the physical block address corresponding to the logical block address in the specified logical-to-physical mapping table page number; reading the corresponding physical block address; and issuing a read / write request to the RAID group where the user data is located according to the physical block address.
[0011] Preferably, S6 specifically includes: unused address space after compression is a data hole, and the processing rule for the data hole is: when the data hole is ≥4KB, a recycling command or a write zero command is sent to notify the solid-state drive to reclaim the space; when the data hole is <4KB, zeros are added and then written to the solid-state drive.
[0012] Accordingly, the present invention also provides a RAID control chip, the RAID control chip comprising: a front-end module for receiving host IO commands and in-band management commands; a merging module for performing merging operations on IOs with consecutive logical block addresses; a compression engine for performing hardware compression or decompression on merged IO blocks; an address remapping module for performing mapping and addressing of logical block addresses and physical block addresses on data before and after compression; a RAID engine module for processing RAID read / write data, configuring RAID, supporting SSD hot-swapping and RAID 1 automatic disk replacement reconstruction, and performing redundancy protection algorithms on compressed data; a metadata mapping module for persisting, restoring, adding, deleting, and searching logical-to-physical mapping entries; and a back-end module for issuing read / write commands to the solid-state drive.
[0013] Preferably, the mapping metadata module maintains a logical-to-physical mapping table. The logical address is represented by the index of the logical-to-physical mapping table. The physical address includes a RAID address, a compression flag, and a data length. The sum of the lengths of the RAID address, the compression flag, and the data length is 64 bits. The RAID address is in 4KB units. The compression flag is used to identify whether the data to be downloaded to the disk is compressed. The data length is used to identify the length of the compressed data, which is padded to 4KB when downloaded to the disk.
[0014] Preferably, the persistent space of the logical-to-physical mapping table of the RAID control chip is allocated from each hard disk space by a preset capacity for storage, metadata is protected by RAID1 data mirroring, and user data is protected by RAID10 / 5 / 6 levels.
[0015] The present invention has at least the following technical effects through the technical solution provided by the present invention: This invention achieves online compression by performing IO merging, data compression, and IO redirection on the RAID card. Compared to existing SSD compression solutions, it not only has lower overall power consumption but also implements RAID redundancy on the hardware side, resulting in better reliability. Furthermore, compared to compressing SSDs within a RAID configuration, because compression is performed on top of the RAID, it more closely approximates the original user data content, leading to better compression performance. Simultaneously, by implementing both redundancy protection and compression, the burden of compression calculations on the host CPU is reduced. By offloading the compression function below the CPU and above the SSD, the online compression capability is better utilized, wear on SSD chips is reduced, and host configuration is simplified, making it more environmentally friendly. Moreover, the data compression method within the RAID card provided by this invention reduces the total amount of data written to the SSD, and the physical address mapping scheme uses in-situ mapping. This method is simple to implement, and the method of reclaiming 4K write holes by reclaiming or writing zeros reduces the space allocation of internal physical addresses in the SSD, which is beneficial to the lifespan of the NAND flash. Attached Figure Description
[0016] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating a method for implementing data compression within a RAID card, as provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the addressing method from persistent logic to physical mapping table in the method provided in the embodiments of the present invention; Figure 3 This is a schematic diagram of the data merging, data compression, and address remapping process during host write I / O in the method provided by the embodiments of the present invention; Figure 4 This is a schematic diagram of the host read I / O process in the method provided by the embodiments of the present invention; Figure 5 This is a logic framework diagram of the RAID control chip provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the write I / O processing flow of the RAID controller chip provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of the RAID controller chip read I / O processing flow provided in an embodiment of the present invention. Detailed Implementation
[0017] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.
[0018] In this invention, the terms "system" and "network" are used interchangeably. "Multiple" refers to two or more; therefore, in this invention, "multiple" can also be understood as "at least two." "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / ", unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, it should be understood that in the description of this invention, terms such as "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or order.
[0019] Please refer to Figure 1 This invention provides a method for implementing data compression within a RAID card, comprising the following steps: S1: Receive multiple write IO commands issued by the host and obtain the logical block address and length of each IO; S2: Perform a merge operation on multiple I / O operations to generate a merged I / O block; S3: Perform data compression operation on the merged IO block through the data compression engine to obtain compressed data; S4: Establish a mapping relationship between logical block addresses and physical block addresses of the compressed data, and store it in a logical-to-physical mapping table; S5: Writes compressed data to the solid-state drive through the RAID engine, and persistently stores the metadata of the logical-to-physical mapping table. S6: Send a reclamation command or write zero command to the solid-state drive to reclaim unused address space after compression.
[0020] In this embodiment of the invention, a method for implementing data compression within a RAID card is provided. For S1, when the host issues a write IO command request, the method receives the write IO command request, obtains the logical block address and length of each IO, and then executes S2. Specifically, multiple IOs are selected to perform a merging operation, generating a merged IO block. The condition for performing the merging operation on multiple IOs is that the logical block addresses of the IOs to be merged are consecutive, and the total length of the generated merged IO block after performing the merging operation on multiple IOs with consecutive logical block addresses must be greater than or equal to a preset threshold. In this embodiment of the invention, the preset threshold is set to 12KB, that is, the length of the merged IO block obtained after merging is ≥12KB. Furthermore, the starting address and length of the merged IO block both satisfy the 4KB boundary alignment, meaning the data length of the merged IO block must be an integer multiple of 4KB, and the starting position of the merged IO block must also be divisible by 4KB. Since solid-state drives (hereinafter referred to as SSDs) and hard drives operate in units of physical sectors, and the sector size of modern storage devices is usually 4KB, if the data is not aligned, reading and writing will span two sectors, requiring additional operations on multiple sectors, which will significantly reduce performance. Moreover, if the starting address and length of the merged IO block both satisfy the 4KB boundary alignment, it can also improve the subsequent compression efficiency. After alignment, the data can be compressed directly, avoiding the latency and resource waste caused by splitting and reassembling. After obtaining the merged IO block, continue to execute S3.
[0021] Furthermore, for S3, after obtaining merged I / O blocks with a data length ≥ 12KB, the data compression engine performs data compression on the merged I / O blocks to obtain compressed data. This compressed data is then written to the SSD using a RAID algorithm to provide data protection. To reduce the complexity of the write operation, a RAID overwrite method is used here. If the data length of the merged I / O block is < 12KB, the compression operation is skipped, and the data is written directly. Simultaneously, in operation S4, a mapping relationship is established between logical block addresses and compressed data physical block addresses, and stored in a logical-to-physical mapping table to record the actual physical address where the compressed data is stored. A linear lookup table is used to implement the logical-to-physical mapping table; the linear table only stores physical addresses, and logical addresses are represented by linear table indices. Specifically, the structure of the logical-to-physical mapping table in S4 includes a RAID address field, a compression flag field, and a data length field. The total address length of the RAID address field, compression flag field, and data length field is 64 bits. The specific address length occupied by each field is determined according to the actual situation. For example… The RAID address field has a length of 58 bits, the compression flag field has a length of 1 bit, and the data length field has a length of 5 bits. The RAID address field is in 4KB units. The compression flag field is used to indicate whether the data is compressed, and it is set to "1" if compressed and "0" if not compressed. The data length field is used to record the actual length of the compressed data. When the data is downloaded to disk, it is padded with 4KB, which can support the real-time compression of 32×4KB=128KB of data to disk. Then, S5 is executed.
[0022] Furthermore, for S5, the compressed data obtained from S4 is written to the solid-state drive via a RAID engine, while the logical-to-physical mapping table metadata is persistently stored. Specifically, the persistent space for the logical-to-physical mapping table is allocated from each hard drive space with a preset capacity for storage, and RAID is used. Metadata is protected using data mirroring, while user data is protected using conventional RAID levels such as RAID 10 / 5 / 6. The persistent metadata space is invisible to the host. In one implementation, the logical-to-physical mapping table is stored in 4KB pages within the persistent metadata space, thus achieving a 0.2% space usage for the logical-to-physical mapping table. Specifically, each logical-to-physical mapping entry occupies 8 bytes (58 bits + 1 bit + 5 bits = 64 bits = 8 bytes), and each 4KB page stores 512 entries. The total number of entries = storage capacity / 4KB × 512, calculating the metadata percentage as 8 bytes / 4KB = 0.2%. The 4KB pages forcibly isolate metadata from user data. Metadata is stored in a separate RAID group (i.e., dedicated SSD space), inaccessible to the host, thus achieving separate metadata storage. The 4KB pages are adapted for RAID 1 writes. RAID 1 mirrors data in 4KB blocks, with the page size aligned to the RAID block to avoid cross-block writes, thereby protecting metadata using RAID 1 data mirroring. Figure 2 As shown, the addressing method specifically includes: obtaining the logical block address of the access request ( Figure 2 The logical block address is represented as LBA, and the logical block address is parsed. The logical block address is then divided by 512 to calculate the page number of the metadata corresponding to that logical block address. Figure 2 The page number is represented as "page". Based on the page number, the corresponding logical-to-physical mapping table page number is read from the metadata persistent RAID group. Figure 2 The logical-to-physical mapping table page number is represented as an L2P Page; then, the logical block address is modulo 512 to calculate the offset of the physical block address corresponding to that logical block address in the specified logical-to-physical mapping table page number, and the corresponding physical block address is read out. Figure 2 The physical block address is represented as PA). Based on the physical block address, a read / write request is sent to the RAID group where the user data is located; then S6 is executed.
[0023] Furthermore, for S6, for unused address space after compression, a reclamation command or a write-zero command is sent to the SSD to reclaim the space; specifically, unused address space after compression is called a data hole, and the processing rules for data holes are as follows: when the data hole is ≥4KB, a reclamation command or a write-zero command is sent to notify the SSD to reclaim the space; when the data hole is <4KB, zeros are added before writing to the SSD.
[0024] In one implementation, please refer to Figure 3 This describes the data merging, data compression, and address remapping process during host write I / O in this invention. Taking the host issuing six 4KB small I / O write requests as an example, the access logical address ( Figure 3 The logical addresses (represented as LBAs) are 2, 6, 8, 1, 3, and 7. Through the IO merging process, the three small IOs with logical addresses 1, 2, and 3 are merged into a large IO with a logical address of 1 and a size of 12KB, and marked as 12KB@1. Simultaneously, the IOs with logical addresses 6, 7, and 8 are merged into a large IO of 12KB@6. Since logical addresses 4 and 5 have not been written to by the host, they are set to -1 and point to a page of all zeros. If the length of the compressed data is not less than the length of the original data, the original data is downloaded to disk. Further, the data compression engine compresses the large IO of 12KB@1 into a small IO of 4KB@1, and the large IO of 12KB@6 into a small IO of 6KB@6. Before the data is downloaded to the RAID engine, the IOs corresponding to the compressed data need to undergo address remapping, that is, mapping from logical block address to physical block address. After address remapping, the logical address (1,2,3) corresponding to IO 12KB@1 is mapped to the physical address (1,1,1). Figure 3 The physical address is represented as PBA, and its length is 4KB. The logical address (6,7,8) corresponding to IO 12KB@6 is mapped to the physical address (6,6,6), which has a length of 6KB. Before being written to the disk, it needs to be made up to 8KB, and the 2KB of idle space is padded with zeros. For the complete 4KB aligned addresses 2, 3, 4, 5, 7, and 8 in the RAID address space, since no data has been written, a write zero or garbage collection command can be issued to the RAID. The RAID will split the command into the corresponding SSD disk address, so that the SSD flash converter can release the internal logical address to physical address address mapping and trigger internal garbage collection at an appropriate time to reduce the occupation of NAND flash space.
[0025] In one implementation, please refer to Figure 4 This describes the host's I / O read process in this invention. Specifically, when the host reads I / O data, it first needs to look up the logical-to-physical mapping table (hereinafter, logical address is abbreviated as LBA, and physical address as PBA) to find the address and length of the compressed data in the RAID space. After reading the data, it performs a decompression operation. Then, based on the address and data length of the read request, it extracts the data content returned to the host from the decompressed data. Taking the host reading 4KB of data at address 7 as an example, ... Figure 4The solid arrow starting with 4KB@7 indicates the start of a read request. The read / write address remapping module queries the PBA=6 corresponding to LBA=7, which has a length of 6KB (PBA6, PBA7, block devices require read and write access to be aligned to 4KB, but the alignment here is less than 8KB). The RAID engine reads the data from these two addresses in the RAID address space and returns it to the compression engine via the dashed arrow in the figure. The compression engine analyzes that the effective data in the 8KB data is 6KB in size and decompresses it into 12KB data. At this time, the data covers the range of LBA6, LBA7, and LBA8 and is then handed over to the read / write merging module. The read / write merging module extracts the data corresponding to LBA=7 into 4KB@7
[1101] and returns it to the host.
[0026] In the method for implementing data compression within a RAID card provided by this invention, to prevent the loss of compressed data written to the SSD due to a sudden power failure, an enterprise-grade SSD with power-loss protection is typically used to prevent data loss caused by abnormal power failure. Meanwhile, since the logical-to-physical mapping table (L2P) metadata and the compressed user data are not written to disk in the same I / O operation, there is a risk of inconsistency between the metadata and user data. Therefore, the following two solutions can be adopted to address this problem: Option 1: Metadata is stored centrally, and the backend uses append-only write to write user data to avoid overwriting old data. At the same time, after the metadata is successfully written to disk, a write-back success response is sent to the host to ensure ACID properties of write I / O transactions. This approach requires dynamic backend physical space allocation, which requires a backend bitmap to record the allocation status of physical space. The physical space allocation bitmap also needs to be written to disk in real time to ensure that physical space is not leaked and that valid data is not overwritten by newly written data.
[0027] Option 2: Store metadata and user data together; specifically, an SSD supporting VSS functionality is required, which supports writing L2P data into the DIF metadata; during the host write process, the compressed data and L2P metadata are written to disk together; if there are write holes, the RAID card needs to fill in the write zero operation, sending a write zero command along with the L2P metadata; the SSD must have the ability to process write zero commands and save L2P metadata, and this command operation should not occupy the NAND user data space, reserving metadata space for separate operation.
[0028] Based on the same inventive concept, please refer to Figure 5This invention provides a RAID control chip for implementing data compression within a RAID card. The RAID control chip includes: a front-end module for receiving host I / O commands and in-band management commands; a merging module for performing merging operations on I / Os with consecutive logical block addresses; a compression engine for performing hardware compression or decompression on merged I / O blocks; an address remapping module for mapping and addressing logical block addresses and physical block addresses of data before and after compression; a RAID engine module for processing RAID read / write data, configuring RAID, supporting SSD hot-swapping and RAID 1 automatic disk replacement reconstruction, and performing redundancy protection algorithms on compressed data; a metadata mapping module for persisting, restoring, adding, deleting, and searching logical-to-physical mapping entries; and a back-end module for issuing read / write commands to the solid-state drives. The RAID control chip also includes a conventional data buffer module and a metadata read cache module.
[0029] In the RAID control chip provided by this invention, the mapping metadata module maintains a logical-to-physical mapping table. The logical address is represented by the index of this logical-to-physical mapping table. The physical address includes the RAID address, compression flag, and data length. The total length of the RAID address, compression flag, and data length is 64 bits. The specific address length occupied by each field is determined according to the actual situation. For example, if the length of the RAID address is 58 bits and the length of the compression flag is 1 bit, then the length of the data length is 5 bits. The length of the RAID address is in 4KB units. The compression flag length is used to identify whether the data to be downloaded to the disk is compressed. The data length is used to identify the length of the compressed data, which is padded to 4KB when downloaded to the disk. Furthermore, the persistent space of the logical-to-physical mapping table is allocated from each hard disk space with a preset capacity for storage. The metadata is protected using RAID1 data mirroring, and user data is protected using RAID10 / 5 / 6 levels. In one embodiment, the preset capacity can be 0.2% of the total capacity.
[0030] In the RAID control chip provided by this invention, such as Figure 6 As shown, the write I / O processing steps of this RAID controller chip are as follows: Figure 6 Steps 1-15 in the diagram are as follows: 1. The host sends a write I / O command to the front-end module; 2. The merging module retrieves commands, parses the addresses and lengths of multiple write I / O commands, and performs a merging operation on I / O commands with consecutive addresses; 3. The merging module initiates an efficient data transmission mechanism (hereinafter referred to as HDMA request) to the data buffer module. 4. The data buffer module initiates HDMA data fetching and writing to the host and caches it in the data buffer module, and sends an HDMA completion response back to the merging module; 5. The merging module forwards the merged I / O to the address remapping module; 6. The address remapping module sends a compression request to the compression engine. The compression engine compresses the data in the data buffer module and returns the compressed data to the address remapping module. 7. The address remapping module initiates the writing of compressed data to the RAID engine module; 8. The RAID engine module obtains the compressed data from the data buffer module; 9. The RAID engine module sends a write compressed data request to the backend module, and the backend module writes the compressed data to the solid-state drive. 10. The address remapping module sends a write metadata request to the mapping metadata module; 11. The mapping metadata module sends a write metadata request to the RAID engine module; 12. The RAID engine module writes metadata to the backend module, and the backend module writes metadata to the solid-state drive. 13. After the metadata is written, the mapping metadata module writes the completion response back to the front-end module; 14. The front-end module writes the completed response back to the host; 15. The metadata mapping module invalidates the corresponding cached metadata in the cache reading module.
[0031] In the RAID control chip provided by this invention, such as Figure 7 As shown, the processing steps for the read I / O of this RAID controller chip are as follows: Figure 7 Steps 1-15 in the diagram are as follows: 1. The host sends a read I / O command to the front-end module; 2. The merging module reads IO requests; 3. The merging module forwards read I / O to the address remapping module; 4. The address remapping module queries the metadata read cache module for the mapping metadata based on the logical address of the read IO. If a match is found, it jumps to sequence number 8. 5. If no match is found, the address remapping module sends a request to the mapping metadata module to read the mapping data; 6. The mapping metadata module sends a read metadata request to the RAID engine module; 7. The RAID engine module sends a read data request to the backend module, and the backend module reads data from the solid-state drive and responds to the RAID engine module. 8. The address remapping module sends a request to the RAID engine module to read the compressed data based on the physical address of the metadata; 9. The RAID engine module sends a request to the backend module to read the compressed data, and the backend module reads the data from the solid-state drive. 10. The RAID engine module places the read data into the data buffer module; 11. The address remapping module sends a decompression command to the compression engine, and the compression engine decompresses the data in the data buffer module; 12. The data buffer module sends the decompressed data to the host via HDMA; 13. The address remapping module sends data back to the front-end module to complete the data transmission; 14. The front-end module reads the completed response back to the host; 15. The mapping metadata module locks the mapping metadata read this time to the metadata read cache module.
[0032] This invention provides a method for data compression within a RAID card and a RAID control chip. By employing I / O merging, data compression, and I / O redirection techniques on the RAID card, online compression is achieved. Compared to existing SSD compression solutions, this method not only offers lower overall performance and power consumption but also implements RAID redundancy on the hardware side, resulting in better reliability. Compared to compressing SSDs for RAID, compression is performed on top of the RAID, more closely approximating the original user data content and achieving better compression results. Simultaneously, because the RAID control chip simultaneously implements redundancy protection and compression functions, it reduces the burden of compression calculations on the host CPU, offloading the compression function below the CPU and above the SSD. This allows for better utilization of online compression capabilities while reducing wear on SSD chips and simplifying host configuration, making it more environmentally friendly. Finally, the compression method provided by this invention reduces the total amount of data written to the SSD. The physical address mapping scheme uses in-situ mapping, simplifying implementation. Furthermore, methods such as zero-writing or reclamation commands to reclaim 4KB write holes reduce the space allocation of internal physical addresses within the SSD, significantly extending the lifespan of the NAND flash.
[0033] The optional embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the scope of the technical concept of the embodiments of the present invention, various simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the protection scope of the embodiments of the present invention.
[0034] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the embodiments of the present invention will not describe the various possible combinations separately.
[0035] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a microcontroller, chip, or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0036] Furthermore, various different implementations of the present invention can be combined arbitrarily, as long as they do not violate the spirit of the present invention, they should also be regarded as the content disclosed in the present invention.
Claims
1. A method for implementing data compression within a RAID card, characterized in that, Includes the following steps: S1: Receive multiple write IO commands issued by the host and obtain the logical block address and length of each IO; S2: Perform a merge operation on multiple I / O operations to generate a merged I / O block; S3: Perform data compression operation on the merged IO block through the data compression engine to obtain compressed data; S4: Establish a mapping relationship between logical block addresses and physical block addresses of the compressed data, and store it in a logical-to-physical mapping table; S5: Writes compressed data to the solid-state drive through the RAID engine, and persistently stores the metadata of the logical-to-physical mapping table. S6: Send a reclamation command or write zero command to the solid-state drive to reclaim unused address space after compression.
2. The method for implementing data compression within a RAID card according to claim 1, characterized in that, The operating condition for S2 is that the logical block addresses of multiple IOs are consecutive; The total length of the merged I / O blocks is greater than or equal to a preset threshold. The starting address and length of the merged IO block both satisfy 4KB boundary alignment.
3. The method for implementing data compression within a RAID card according to claim 1, characterized in that, The structure of the logical-to-physical mapping table in S4 includes a RAID address field, a compression flag field, and a data length field. The total address length of the RAID address field, the compression flag field, and the data length field is 64 bits. The RAID address field is in 4KB units; The compression flag field is used to identify whether the data is compressed; The data length field is used to record the actual length of the compressed data, which is padded to 4KB when the data is downloaded to disk.
4. The method for implementing data compression within a RAID card according to claim 1, characterized in that, S5 specifically includes: The logical-to-physical mapping table occupies a preset capacity of total storage space, and metadata is protected through RAID 1 data mirroring. The metadata is stored separately from user data, and the persistent metadata space is not visible to the host.
5. A method for implementing data compression within a RAID card according to claim 4, characterized in that, The logical-to-physical mapping table is stored in 4KB pages in the metadata persistence space, and the addressing methods include: Obtain the logical block address of the access request, divide the logical block address by 512 to calculate the page number of the metadata corresponding to the logical block address, and read the corresponding logical-to-physical mapping table page number data from the metadata persistent RAID group according to the page number; Take the remainder of the logical block address with 512, calculate the offset of the physical block address corresponding to the logical block address in the specified logical-to-physical mapping table page number, and read out the corresponding physical block address. Based on the physical block address, a read / write request is sent to the RAID group where the user data is located.
6. The method for implementing data compression within a RAID card according to claim 1, characterized in that, S6 specifically includes: Unused address space after compression is called a data hole. The processing rules for the data holes are as follows: When the data hole is ≥4KB, send a recycling command or write zero command to notify the solid-state drive to reclaim space; When the data gap is less than 4KB, it is padded with zeros before being written to the solid-state drive.
7. A RAID control chip implementing any one of claims 1-6, characterized in that, The RAID control chip includes: Front-end module: Used to receive host I / O commands and in-band management commands; Merging module: Used to perform merging operations on I / Os with consecutive logical block addresses; Compression engine: Used to perform hardware compression or decompression on merged I / O blocks; Address remapping module: used to perform logical block address and physical block address mapping and addressing on the data before and after compression; RAID engine module: Used to process RAID read and write data, configure RAID, support SSD hot-swapping and RAID 1 automatic reconstruction function after disk replacement, and perform redundancy protection algorithm for compressed data; The metadata mapping module is used to persist, restore, and add / delete search logic to physical table entries; Backend module: Used to issue read and write commands to the solid-state drive.
8. The RAID control chip according to claim 7, characterized in that, The mapping metadata module maintains a logical-to-physical mapping table. The logical address is represented by the index of the logical-to-physical mapping table. The physical address includes the RAID address, compression tag, and data length. The total length of the RAID address, the compression tag, and the data length is 64 bits. The RAID addresses are in 4KB units; The compression marker is used to indicate whether the data on the lower disk is compressed; The data length is used to identify the length of the compressed data, which is padded to 4KB when the data is downloaded to disk.
9. The RAID control chip according to claim 8, characterized in that, The RAID controller chip's logical-to-physical mapping table persistent space is allocated from each hard disk space with a preset capacity for storage. Metadata is protected using RAID 1 data mirroring, while user data is protected using RAID 10 / 5 / 6 levels.
Citation Information
Cited By
Data storage method of GDDR memory, GPU, equipment and medium
CN121300713A
Data storage method of gddr memory, gpu, device and medium
CN121300713B
Data storage method and data storage device based on block storage engine
CN121957509A
A data storage method and data storage device based on a block storage engine
CN121957509B