Disk array data writing method and device, electronic equipment and storage medium
By generating and merging triple parity blocks in low-write scenarios, the write amplification and low efficiency issues of RAID 5 and RAID 6 in low-write scenarios are solved, achieving efficient data writing and multi-disk fault tolerance, and improving the performance and reliability of the storage system.
Patent Information
- Application Number
- CN202511524851.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-23
AI Technical Summary
Traditional RAID 5 and RAID 6 architectures suffer from write amplification, complex and inefficient parity calculations, and are intolerant of multiple disk failures in low-write scenarios, thus failing to effectively support the storage requirements for high reliability and high performance.
After receiving a write request from the host, if the write mode is determined to be lowercase, the new data and old data are obtained, incremental calculation is performed to generate check increment data, and the updated triple check block is generated. Finally, the data and check block are persisted to the physical disk respectively.
Reduce read/write overhead in low-write scenarios, reduce write amplification effect, improve verification calculation and data writing efficiency, ensure data integrity in the event of multiple disk failures, and improve storage system performance and reliability.
Smart Images

Figure CN120994144A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data storage, and particularly relates to a disk array data writing method and device, electronic equipment and a storage medium. BACKGROUND
[0002] As a core fault-tolerant mechanism of a storage system, RAID technology is widely used in enterprise-level data centers and key business systems. With the increase of single-disk capacity and the expansion of storage array size, the traditional RAID 5 and RAID 6 architecture faces significant limitations, such as insufficient data security protection in scenarios with extremely high reliability requirements. The traditional RAID architecture is difficult to support the dual requirements of performance and reliability of the storage system, especially in PB-level data application scenarios such as distributed storage and hyper-converged architecture, the performance bottleneck and the defect of insufficient fault tolerance are more prominent, and a new RAID architecture supporting triple verification is urgently needed to solve the performance problems caused by write amplification and the multi-disk fault tolerance problem. SUMMARY
[0003] The present application provides a disk array data writing method and device, electronic equipment and a storage medium to at least solve the problem of write amplification caused by the triple verification RAID architecture in the related art.
[0004] The present application provides a disk array data writing method, comprising: receiving a write request of a host, determining a write mode of corresponding data based on the write request, and the write mode comprising full write, large write and small write; if the write mode corresponding to the write request is small write, obtaining new data and old data of at least one target data block corresponding to the write request; performing incremental calculation based on the new data and the old data, generating verification incremental data of the at least one target data block, and writing the data corresponding to the write request into the at least one target data block for persistent processing; after the incremental calculation of the at least one target data block is completed, merging the verification incremental data of the at least one target data block, and generating updated triple verification blocks based on the merging result and the old data; persisting the triple verification blocks to physical disks respectively, so as to write the data corresponding to the write request into the physical disks in the disk array.
[0005] The present application also provides a disk array data writing device, comprising: a determination unit configured to receive a write request of a host, determine a write mode of corresponding data based on the write request, and the write mode comprising full write, large write and small write; an obtaining unit configured to, if the write mode corresponding to the write request is small write, obtain new data and old data of at least one target data block corresponding to the write request; The first generation unit is configured to perform incremental calculation based on the new data and the old data, generate the check incremental data of the at least one target data block, and write the data corresponding to the write request into the at least one target data block for persistent processing. The second generation unit is configured to merge the check incremental data of the blocks, and generate updated triple-check subblocks based on the merging result and the old data. The write unit is configured to persist the triple-check subblocks to the physical disks respectively, so as to write the data corresponding to the write request into the physical disks in the disk array.
[0006] The application further provides an electronic device, including a memory configured to store a computer program, and a processor configured to execute the computer program to implement the steps of any of the above methods.
[0007] The application further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0008] The application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0009] The application provides a disk array data write method and device, an electronic device and a storage medium. After receiving a host write request, the write mode of full write, large write or small write is determined first. For the small write scenario, the new data and the old data of the target data block are acquired first to perform incremental calculation to generate check incremental data. After the incremental calculation of the target data block is completed, the check incremental data is merged, and the updated triple-check subblocks are generated in combination with the old data. Finally, the target data block data and the triple-check subblocks are persisted respectively. Therefore, the problems of write amplification, complex and low-efficiency check calculation, and difficulty in tolerating multiple disk failures in the prior art RAID 5 and RAID 6 in the small write scenario can be solved. The technical effects of reducing the read-write overhead in the small write scenario, reducing the write amplification effect, improving the check calculation and data write efficiency, guaranteeing the data integrity in multiple disk failures through triple-checking, and improving the performance and reliability of the storage system are achieved.
[0010] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0011] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them: Figure 1A flowchart of a disk array data writing method provided by an embodiment of the present application is shown in the figure. Figure 2 A structural diagram of a disk array data writing device provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0012] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help the understanding of the present disclosure. These should be considered in the context of the overall description and should not be considered to limit the scope of the present disclosure. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to make the description clear and concise, the description of well-known functions and structures is omitted in the following description.
[0013] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0014] In conjunction with the specific application environment architecture or specific hardware architecture on which the disk array data writing method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0015] An embodiment of the present application provides a disk array data writing method, Figure 1 A flowchart of a disk array data writing method provided by an embodiment of the present application is shown in the figure.
[0016] As Figure 1 shown, the method comprises the following steps: Step 101, receiving a write request of a host, determining a write mode of corresponding data based on the write request, the write mode comprising full write, large write and small write.
[0017] In the embodiment of the present application, during the operation of the RAID-TP storage system, first, a write request from a host is received, the write request comprising relevant information of data to be written by the host, such as data volume, target storage location identifier, etc. After receiving the write request, the data corresponding to the write request is analyzed to determine the adaptive data write mode, wherein the write mode specifically comprises three types of full write, large write and small write.
[0018] Here, "full write" refers to a host write request involving data that can completely cover all data blocks in a stripe in the storage system, without the need for additional reading of existing old data in the stripe to complete data write-related processing; "uppercase" is that after analyzing the data corresponding to the host write request, the data range does not completely cover all data blocks in a stripe, but by reading the old data in the stripe that is not covered, the host write data can be combined to form complete data that can cover the entire stripe, and then the processing is completed according to the logic similar to full write; "lowercase" is that after analyzing the data corresponding to the host write request, the data range only involves part of the data blocks in a stripe, and the processing is completed without the need to read the old data in the uncovered part to combine complete stripe data.
[0019] By first receiving the host write request and determining the corresponding write mode, the storage system can subsequently use an adaptive process to handle data write, avoiding the problem of low processing efficiency caused by using a unified write process, laying the foundation for subsequent efficient completion of data write and verification processing, and helping to improve the overall write performance and adaptability of the storage system.
[0020] Step 102, if the write request corresponds to a lowercase write mode, obtaining new data and old data of at least one target data block corresponding to the write request.
[0021] In the embodiments of the present application, after determining that the write mode corresponding to the host write request is lowercase, the storage system will further locate at least one target data block associated with the write request, which is the specific data storage unit in the storage system that needs to be updated by the data of this write request. Subsequently, the storage system needs to obtain the new data and old data of the target data block respectively: wherein the new data refers to the latest data of the target data block to be written by the host through this write request, which will be used to overwrite the original data content in the target data block; the old data refers to the original data stored in the target data block before receiving this write request, which is an important basis for subsequent data update and verification calculation.
[0022] When obtaining new data, the storage system receives the write data issued by the host from the upper module, and performs preliminary processing according to the preset data storage format to ensure that the new data conforms to the storage specification of the target data block; when obtaining old data, the storage system sends a data read request to the corresponding storage module according to the physical storage location information of the target data block, and retrieves the historical storage data of the target data block from the storage medium to complete the acquisition of the old data.
[0023] By accurately obtaining the new data and the old data of the target data block, necessary data support is provided for subsequent data incremental calculation and verification update in the small write scenario, avoiding calculation errors caused by incomplete data acquisition, and also laying a foundation for subsequent reduction of write overhead and optimization of write performance, which helps to improve the data processing accuracy and efficiency of the storage system in the small write scenario.
[0024] In step 103, incremental calculation is performed based on the new data and the old data to generate the verification incremental data of at least one target data block, and the data corresponding to the write request is written into at least one target data block for persistent processing.
[0025] In the embodiments of the present application, after obtaining the new data and the old data of at least one target data block in the small write scenario, the storage system will perform incremental calculation operation based on the two types of data to generate the verification incremental data corresponding to the target data block. The core of the incremental calculation here is to obtain the difference data between the new data and the old data through logical operation, and the difference data is the verification incremental data, which is the key basis for subsequent update of the verification chunk. Compared with directly recalculating the verification value for the complete data, the calculation amount can be greatly reduced.
[0026] After completing the generation of the verification incremental data, the storage system will simultaneously write the new data corresponding to the write request into at least one target data block and start the persistent processing procedure. Persistent processing refers to the operation of stably storing the new data in the target data block to the storage medium, ensuring that the written new data will not be lost even in the case of subsequent temporary failure of the storage system, and guaranteeing the stability of data storage.
[0027] By generating the verification incremental data through incremental calculation, the redundant operation of repeated operation on complete data in traditional verification calculation is effectively avoided, and the consumption of computing resources is reduced. At the same time, the new data is written into the target data block in time and the persistent processing is completed, ensuring the immediacy and safety of the data. This process lays a foundation for efficient update of the verification chunk in the future, which helps to improve the overall processing efficiency of the storage system in the small write scenario and reduce unnecessary resource overhead.
[0028] In step 104, after the incremental calculation of at least one target data block is completed, the verification incremental data of at least one target data block is merged, and the updated triple verification chunk is generated based on the merging result and the old data.
[0029] In the embodiments of the present application, after the incremental calculation of each of the at least one target data block is completed and the corresponding check incremental data is generated, the storage system initiates a merging operation of the check incremental data. The merging operation refers to integrating the check incremental data generated by the plurality of target data blocks according to a preset logical rule, eliminating the redundant information among the plurality of check incremental data, and forming a unified merging result. The merging result can collectively reflect the overall impact of the data update of all target data blocks on the check sub-blocks, and avoid repeated calculation and resource waste caused by processing each group of check incremental data separately.
[0030] After obtaining the merging result of the check incremental data, the storage system calls a preset check calculation logic to cooperatively operate the merging result and the old check data stored in the storage medium, and generates an updated triple check sub-block through incremental update of the old check data. The triple check sub-block contains three independent check information, which can provide higher level data fault tolerance protection for the storage system, and can further improve the security of data in the multi-disk failure scenario compared with the traditional double check architecture.
[0031] Through the merging processing of the check incremental data of the plurality of target data blocks, the data processing amount in the check calculation process is greatly reduced, and the computing resource overhead is reduced. Meanwhile, based on the merging result and the old data, the triple check sub-block is generated, which realizes efficient update of the check sub-block under the premise of ensuring the accuracy of the check, and lays a foundation for subsequent data persistence and improvement of the fault tolerance capability of the storage system.
[0032] In step 105, the triple check sub-blocks are respectively persisted to the physical disks to write the data corresponding to the write request into the physical disks in the disk array.
[0033] In the embodiments of the present application, after the updated triple check sub-block is generated, the storage system will initiate a persistence process of the triple check sub-block to write the three independent check sub-blocks into the corresponding physical disks in the disk array. The "persistence" refers to stably storing the check information of the triple check sub-block to the non-volatile storage area of the physical disk, ensuring that the check information will not be lost due to temporary power failure, device fluctuation and the like of the storage system, and ensuring the long-term reliability of the check data. In the specific execution process, the storage system will determine the physical disk position corresponding to each check sub-block according to the hardware configuration and the sub-block storage rule of the disk array, initiate a write instruction to the target physical disk through the underlying storage interface, write the check sub-block data into the disk storage unit in a preset format, and monitor the data transmission state and disk response in the write process to ensure that each check sub-block can be completely and accurately stored to the corresponding physical disk.
[0034] When the triple-check chunk is completely persisted, the data corresponding to the host write request is completely written into the physical disk of the disk array in combination with the new data of the target data block that has been persisted. This process completes the persistence of the triple-check chunk separately and sequentially, which not only guarantees the storage consistency of the check data and the target data block data, but also strengthens the fault tolerance of the disk array in response to multiple disk failures by relying on the independent storage characteristics of the triple-check chunk. The beneficial effects are that the data check failure problem caused by the delayed or incomplete storage of the check chunk is avoided, and the bottom-layer guarantee for the data security of the storage system is provided by solidifying the check data at the physical disk level, thereby further improving the reliability and stability of the overall data storage.
[0035] The application provides a disk array data write method. After receiving a host write request, the write mode of full write, large write or small write is determined. For the small write scenario, the new data and the old data of a target data block are first acquired to generate check incremental data through incremental calculation. After the incremental calculation of the target data block is completed, the check incremental data is merged, and updated triple-check chunks are generated in combination with the old data. Finally, the target data block data and the triple-check chunks are respectively persisted. Therefore, the problems of write amplification, complex and low-efficiency check calculation, and difficulty in tolerating multiple disk failures of RAID 5 and RAID 6 in the prior art can be solved, the read-write overhead in the small write scenario is reduced, the write amplification effect is reduced, the check calculation and data write efficiency are improved, the data integrity is guaranteed in the case of multiple disk failures through triple-checking, and the technical effects of improving the performance and reliability of the storage system are achieved.
[0036] Under the technical solution framework disclosed in the foregoing embodiments, in order to achieve better resource allocation for data write, the application embodiments further provide an implementation of resource allocation, which comprises: dynamically allocating computing and caching resources for a stripe corresponding to a write request before acquiring new data and old data of at least one target data block corresponding to the write request; creating a stripe lock and synchronizing the state information of each node in a failure domain through a lock communication mechanism.
[0037] Further, in the application embodiments, the dynamic allocation of computing and caching resources for the stripe corresponding to the write request is embodied as: estimating and reserving volatile memory pages for data temporary storage according to the total number of data blocks and check blocks in the stripe write unit; allocating a plurality of input-output control blocks for each data block and each check block; reserving memory resources for creating data block control structures and check block control structures.
[0038] Specifically, under the technical solution framework disclosed in the foregoing embodiments, to achieve efficient resource allocation for data writing, the embodiments of the present application further optimize the resource management link for the data writing process in the small write scenario. Specifically, before obtaining the new data and the old data of at least one target data block corresponding to the writing request, dynamic allocation of computing and caching resources needs to be completed for the stripe associated with the writing request, and a stripe lock is created at the same time, and the state information of each node in the fault domain is synchronized by means of the lock communication mechanism. The stripe is the basic unit of data division in the storage system, which is composed of multiple data blocks and check blocks, and the sub-blocks in the same stripe are distributed on different physical disks. The stripe lock is used to avoid data conflicts caused by simultaneous operation of multiple writing requests on the same stripe, and the lock communication mechanism can ensure that all nodes in the fault domain are aware of the current occupancy and processing state of the stripe in real time, thereby ensuring the consistency of resource allocation.
[0039] When dynamically allocating computing and caching resources for the stripe, specific operations need to be performed in combination with the actual data size and processing requirements of the stripe: first, according to the total number of data blocks and check blocks in the stripe write unit, a volatile memory page is estimated and reserved, which is used to temporarily store the new data, the old data of the data blocks, and the intermediate data generated in the subsequent calculation process, to ensure that there is no storage resource shortage during data processing; second, multiple input / output control blocks (IOBs) are allocated for each data block and each check block. The input / output control block is a key component that connects data and hardware interfaces, and is used to manage the transmission logic of data between memory and physical disks; finally, special memory resources are reserved to create data block control structures (such as IPK, responsible for managing IO operations of a single data block, containing information such as offset and length of data read / write) and check block control structures (such as PSIO, used to manage the calculation and disk landing process of the check block), to provide structural support for efficient processing of subsequent data blocks and check blocks.
[0040] Through the above resource allocation method, accurate reservation and synchronization of resources can be completed before data writing, avoiding processing delays caused by resource shortage or node state asynchronization, effectively improving the smoothness and stability of data writing in the small write scenario, and laying a reliable resource foundation for subsequent new data and old data acquisition, check increment calculation, etc.
[0041] Under the technical solution framework disclosed in step 102, the technical solution in step 102 is embodied as: saving the new data corresponding to the writing request to the input / output buffer area indicated by the sub-block input / output management structure; for the old data that needs to be read, asynchronously reading the disk data through the virtualization layer module and saving it to the input / output buffer area indicated by the stripe cache management structure.
[0042] Specifically, under the technical solution framework of determining the writing mode as lowercase, and needing to obtain the new data and the old data of at least one target data block in step 102, the specific implementation process needs to rely on specific data management structures and modules in the storage system to complete the storage and reading of data. Among them, for the new data corresponding to the write request, it needs to be saved to the input / output buffer (IOB, used to temporarily store data in the data transmission process, realizing the data connection between the memory and the physical disk) indicated by the input / output management structure (IPK, which is used to manage the IO operation of a single data block, including the offset, length and other key information of data reading and writing) of the block. Through the accurate positioning of the data storage location by IPK, it is ensured that the new data can be temporarily stored according to the preset path, and preparation is made for subsequent data processing and persistence.
[0043] And for the old data that needs to be read, since it is stored in the physical disk, it needs to initiate an asynchronous reading operation with the help of the virtualization layer (VL) module - asynchronous reading can avoid blocking other processes due to waiting for disk data to return, improving the overall processing efficiency, and the virtualization layer module is responsible for coordinating the interaction logic between the storage system and the physical disk, ensuring that the data reading request can be accurately issued to the corresponding disk. The read old data will finally be saved to the input / output buffer indicated by the stripe cache management structure (SDE, which is used to manage the read cache and other information of the stripe, and the stripe is the basic unit of data division, containing multiple blocks distributed in different disks). Through the management of the stripe-level cache by SDE, the centralized temporary storage and rapid calling of the old data are realized, and data support is provided for the comparison and incremental calculation of the subsequent new data.
[0044] Through the above specific implementation, the ordered storage and efficient reading of new data and old data can be realized, avoiding data storage confusion or reading delay affecting the subsequent process, and at the same time relying on the synergistic effect of IPK, SDE and virtualization layer module, the accuracy and timeliness of data acquisition are guaranteed, laying a reliable data foundation for the incremental calculation link of step 102, and further improving the efficiency and stability of data processing in the lowercase scenario.
[0045] Under the technical solution framework disclosed in step 103, the technical solution of performing incremental calculation in step 103 is further specified as: performing exclusive or operation on the new data and the old data of at least one target data block to obtain the check incremental data; saving the check incremental data in the management structure of the first check block, and the management structure of the first check block has a mapping relationship with the management structures of the second check block and the third check block.
[0046] Specifically, under the technical scheme framework of step 103 of performing incremental calculation based on new data and old data to generate check incremental data, the specific implementation of incremental calculation needs to be realized through the cooperation of specific operation logic and data management structure. Specifically, for the new data (i.e., the latest data to be written by the host) and the old data (i.e., the original storage data to be updated in the target data block) of the at least one target data block that has been acquired, an exclusive-OR operation is used to complete the incremental calculation operation. The exclusive-OR operation can efficiently capture the difference information between the new data and the old data, and the result obtained through the operation is the check incremental data. This data can accurately reflect the impact of the update of the target data block on the check sub-block, and compared with the traditional full-check calculation, the calculation complexity and resource consumption are greatly reduced.
[0047] After obtaining the check incremental data, it needs to be stored in the management structure of the first check sub-block, wherein a mapping relationship is pre-established among the management structure of the first check sub-block, the management structure of the second check sub-block, and the management structure of the third check sub-block. The check sub-block management structure (such as the IPK structure corresponding to the check block, used to manage the IO operation and data storage information of the check sub-block) is the core of ensuring the ordered management of the check data. Through the establishment of the mapping relationship among the three, the management structures of the second check sub-block and the third check sub-block do not need to store the check incremental data separately, but can obtain the data from the management structure of the first check sub-block through mapping association, avoiding the repeated storage of the check incremental data, saving the memory resources, and at the same time ensuring that the subsequent triple check sub-block update can be calculated based on the unified check incremental data, and the consistency of the check data is ensured.
[0048] Through the above specific implementation of incremental calculation, not only the efficient generation of check incremental data is realized by means of exclusive-OR operation, but also the data storage logic is optimized by relying on the mapping relationship among the check sub-block management structures, effectively reducing the calculation and storage resource overhead, providing accurate and efficient data support for the subsequent merging of check incremental data and the update of triple check sub-blocks, and further improving the processing efficiency of the storage system in the small write scenario.
[0049] Under the technical scheme framework disclosed in step 104, the merging of the check incremental data of the at least one target data block in step 104 is embodied as: creating an exclusive-OR calculation structure, and sequentially merging the check incremental data of the at least one target data block; taking the check sub-block management structure generated by the first target data block as the reference structure, and taking the check sub-block management structure generated by the subsequent target data block as the subsequent structure to participate in the merging calculation.
[0050] Specifically, under the technical scheme framework of merging the check delta data of the at least one target data block in step 104, the specific implementation of the merging operation needs to rely on the cooperation of a specific computing structure and a block management structure. First, an exclusive or computing structure (XOR structure for short, used to uniformly manage the merging computing logic of the check delta data and record the data source information and intermediate results in the computing process) needs to be created, and the check delta data generated by each target data block is sequentially included in the merging process in a preset order through the structure, so that each group of check delta data can be accurately identified and processed, and the problem of data omission or repeated calculation in the merging process is avoided.
[0051] In the selection of the structure of the merging computation, the check block management structure generated by the first target data block is taken as the reference structure (baseIpk, which stores the check delta data corresponding to the first target data block and related IO management information, and provides an initial data reference for subsequent merging computation), and the check block management structures generated by the subsequent target data blocks are taken as the subsequent structures (NextIpk) to gradually participate in the merging computation. In the specific merging process, the exclusive or computing structure will first take the check delta data in the reference structure as the basis, and then integrate the check delta data in the subsequent structures into the calculation result of the reference structure through exclusive or operation. Through this progressive merging mode, a unified merging result reflecting the data update influence of all target data blocks is gradually formed, and it is ensured that the merging result can comprehensively cover the check delta information of all target data blocks, thereby providing complete and accurate data basis for generating the updated triple check block based on the merging result and the old data.
[0052] Through the above merging mode, the ordered integration of the check delta data is realized by means of the exclusive or computing structure, and the integrity and accuracy of the merging result are ensured by means of the progressive merging logic of the reference structure and the subsequent structure, thereby avoiding the calculation redundancy caused by the dispersed processing of multiple groups of check delta data, effectively reducing the resource consumption in the merging process, laying a reliable foundation for the efficient generation of the subsequent triple check block, and further improving the efficiency of the check data processing in the small write scenario.
[0053] Under the technical scheme framework disclosed in step 104, generating the updated triple check block based on the merging result and the old data in step 104 is embodied as follows: the old check data and the merged check delta data are calculated through exclusive or operation to obtain new check data; the new check data is saved in the stripe cache management structure, and the context state of the acceleration processing unit is updated to a check data ready state.
[0054] Specifically, under the technical scheme framework of step 104 of generating updated triple-check chunks based on the merging result and old data, the specific implementation process needs to be completed through the cooperation of specific operation logic and data management structure. First, for the old check data (i.e., the historical check data before the update of the triple-check chunks, including the old data corresponding to the P, Q, and R triple-check chunks) stored in the storage system, an exclusive-OR operation is performed with the check incremental data obtained by the previous merging. The exclusive-OR operation can efficiently utilize the logical relationship between the two, quickly derive new check data reflecting the check state after the data update, and the new check data is the core content of the updated triple-check chunks. Compared with the traditional full-recompute check data method, the calculation amount and time consumption are greatly reduced.
[0055] After obtaining the new check data, it needs to be saved in the stripe cache management structure (i.e., SDE, which is used to manage the read cache information of the stripe, and can realize the temporary storage and quick calling of the new check data, and prepare for the subsequent persistent operation). At the same time, the context state (i.e., ApuContext, which is used to record the running state of the APU, the check block size, and other key information) of the acceleration processing unit (i.e., APU, which is a hardware unit for improving data calculation and processing efficiency) needs to be updated to the check data ready state (such as setting ApuContext->events = Pr|Qr|Rr). This state updating operation can real-time feedback the preparation of the new check data, ensure that the subsequent triple-check chunk persistent process can be started in time, and avoid process blocking caused by different state information.
[0056] Through the above specific implementation, the efficient generation of new check data is realized by means of exclusive-OR operation, the reliable storage of new check data is guaranteed by relying on the stripe cache management structure, and the smoothness of process connection is ensured by updating the context state of the acceleration processing unit, which lays an accurate and timely data and state foundation for the subsequent persistence of triple-check chunks, and further improves the efficiency and stability of triple-check chunk update in the small write scenario.
[0057] Under the technical scheme framework disclosed in the foregoing embodiments, the embodiments of the present application are further specified as: in response to the completion of the incremental calculation of the target data block, performing data landing processing on the target data block.
[0058] Specifically, under the technical solution framework disclosed in the foregoing embodiments, the subsequent processing operation of the target data block is further refined for the data processing flow in the small write scenario. Specifically, after the target data block completes the incremental calculation (that is, the process of generating the check incremental data based on the exclusive OR operation of the new data and the old data), the storage system will immediately perform data landing processing on the target data block. The "data landing processing" here refers to the operation of writing the new data (that is, the data to be updated transmitted by the host through the write request) obtained in the target data block from the temporary storage area (such as the input / output buffer) to the physical disk, ensuring that the new data can be stably stored in the physical storage medium of the disk array, avoiding data loss due to temporary system failure or resource fluctuation.
[0059] When performing the data landing processing, the storage system will rely on the preset block input / output management structure (such as IPK, used to manage the IO operation of a single data block, including the offset, length and other key information of data read / write) to explicitly determine the physical disk location and storage format corresponding to the target data block, and initiate a write request to the corresponding physical disk through the underlying storage interface. At the same time, the system will monitor the data transmission state in the write process in real time, including data integrity check, disk response feedback, etc., to ensure that the new data can be completely and accurately written to the physical disk. It is worth noting that this data landing processing does not need to wait for the completion of the subsequent check incremental data merging and triple check block update flow, and can be independently and parallelly executed, thereby reducing the waiting delay of the overall data write.
[0060] By immediately performing the data landing processing after the target data block completes the incremental calculation, the physical storage solidification of the new data can be quickly realized, the safety and timeliness of data storage are improved, the parallel processing logic is used to reduce process blocking, the overall write efficiency of the storage system in the small write scenario is further optimized, sufficient time is reserved for the promotion of the subsequent check-related flow, and the smoothness of the data write whole flow is ensured.
[0061] Under the technical solution framework disclosed in the foregoing embodiments, if the write mode determined in step 101 is full write, the full write operation can be performed by using but not limited to the following mode: based on the write request, data of all data blocks in the stripe is obtained; after the data of all data blocks is obtained, the first check block, the second check block and the third check block are calculated and generated based on the data cached by the data blocks; after the calculation of the check blocks is initiated, the landing operation of all data blocks is performed; and after the calculation of the check blocks is completed, each check block is respectively persisted to the corresponding physical disk.
[0062] Further, in the embodiments of the present application, the first, second and third check sub-blocks are generated based on the data cached in the data blocks, which is embodied as: memory resources and acceleration processing unit resources are respectively applied for the first, second and third check sub-blocks; the first, second and third check sub-blocks are sequentially calculated, and the calculation results are temporarily stored in the respective associated memory pages.
[0063] Further, in the embodiments of the present application, each check sub-block is respectively persisted to the corresponding physical disk, including: after all the check sub-blocks are calculated, the write disk operation of each check sub-block is sequentially initiated; after the write disk of each check sub-block is successful, the state of the check sub-block in the stripe acceleration processing unit context is updated, and the memory and acceleration processing unit resources occupied by the check sub-block are released.
[0064] Specifically, under the technical solution framework disclosed in the foregoing embodiments, when it is determined in step 101 that the write mode is full write, the data write operation needs to be performed according to the process adapted to the full write scenario. Specifically, first, based on the write request of the host, the stripe corresponding to the request is located, and the to-be-written data of all data blocks in the stripe is obtained - here, the stripe is a basic storage unit in the storage system that contains multiple data blocks and check blocks, and in the full write scenario, all data blocks in the stripe need to be overwritten, so it is necessary to ensure that the to-be-written data of all data blocks is completely obtained. After the data of all data blocks is obtained, instead of relying on the old data and old check blocks, the first, second and third check sub-blocks are generated based on the to-be-written data temporarily stored in the data block cache, so as to build a triple check mechanism and provide higher level fault tolerance protection for data. After initiating the check sub-block calculation process, without waiting for the completion of the calculation of the check sub-blocks, the landing operation of all data blocks can be synchronously performed, that is, the data in the data block cache is written to the corresponding physical disk through the underlying storage interface, so as to realize the rapid storage of the data blocks; after the calculation of the check sub-blocks is completed, the first, second and third check sub-blocks are respectively persisted to the corresponding physical disks in the disk array, and the complete data write in the full write scenario is completed.
[0065] When the triple check sub-blocks are generated based on the data cached in the data blocks, independent memory resources and acceleration processing unit (APU) resources need to be respectively applied for the first, second and third check sub-blocks - the memory resources are used to temporarily store the calculation results and related context information of the check sub-blocks, and the acceleration processing unit resources are used to improve the efficiency of the check calculation, so as to avoid the influence of the calculation delay on the overall write performance. After the resource application is completed, the first, second and third check sub-blocks are sequentially calculated according to a preset order, and after the calculation of each check sub-block is completed, the calculation result of the check sub-block is temporarily stored in the memory page associated with the check sub-block, so as to ensure that the calculation results of the check sub-blocks are independently stored and do not interfere with each other, and to provide accurate data basis for subsequent persistence operation.
[0066] In the process of persisting each check chunk to the corresponding physical disk, after all check chunks are calculated, write disk operation is initiated to the physical disk corresponding to each check chunk in a preset order to ensure that the write disk process is orderly. When a single check chunk is successfully written, the state of the check chunk in the ApuContext (a structure used to record the running state of APUs in the stripe and check block information) on the corresponding stripe acceleration processing unit needs to be updated in time, for example, updating the state to "write disk complete (Pw / Qw / Rw)", and releasing the memory resources and acceleration processing unit resources previously occupied by the check chunk to avoid resource waste caused by long-term occupation of resources and to ensure efficient recycling of storage system resources.
[0067] Through the design of the above full-write process, on the one hand, the parallel execution of data block landing and check chunk calculation reduces the overall write delay; on the other hand, through independent resource allocation and orderly state management, the accuracy of triple check chunk calculation and storage is ensured, which not only improves the write efficiency in full-write scenarios, but also relies on triple check to strengthen data reliability, effectively adapting to storage scenarios with high requirements for write performance and data security.
[0068] Under the technical solution framework disclosed in the foregoing embodiments, if the write mode determined in step 101 is large write, the embodiments of the present application can also use but are not limited to the following methods for full-write operation: combining the old data on the member disk and the data of the host write request to form a full stripe; and multiplexing the full-write process to land the data chunk and the check chunk in parallel.
[0069] Specifically, under the technical solution framework disclosed in the foregoing embodiments, when the write mode determined in step 101 is large write, the data write operation needs to be performed according to a specific process adapted to the large write scenario. Specifically, the core feature of the large write scenario is that the data range corresponding to the host write request does not completely cover all data blocks in the stripe, so first the old data corresponding to the data blocks not covered by the host write request needs to be read through the underlying interaction logic of the storage system - these old data are stored in the member disk (i.e. the physical disk constituting the RAID array) of the disk array, and the data transmission link needs to be established between the virtualization layer module and the member disk to ensure that the old data can be completely and accurately transferred from the physical disk to the system cache area.
[0070] After the old data reading of the member disk is completed, the old data is combined with the new data delivered by the host write request according to the data block distribution rule of the strip to form full-strip data that can completely cover the entire strip. The "full strip" here refers to the data amount that matches the storage capacity of all data blocks in the strip, which meets the basic condition for subsequent full-write process. The combination process needs to map the old data and the new data to different data blocks in the strip according to the strip block mapping relationship, so as to ensure that the data distribution logic in the strip is consistent with the full-write scenario.
[0071] After the full-strip data is constructed, the data processing flow in the previous full-write scenario is directly reused, that is, based on the combined full-strip data, the first, second and third check blocks are first calculated and generated, and then the parallel disk writing operation of the data blocks and the check blocks is started. Parallel disk writing refers to the synchronous execution of the writing of the data blocks (including the combined old data and new data) to the corresponding member disk and the writing of the check blocks to the specified physical disk. It is not necessary to wait for the writing of one type of block to be completed before starting the writing of another type of block. Through parallel processing, the overall writing time is reduced. At the same time, relying on the mature resource scheduling and state management logic in the full-write process, the accuracy and stability of the disk writing of the data blocks and the check blocks are ensured.
[0072] Through the design of the above full-write process, it is not necessary to develop a completely new processing logic for the full-write scenario. Only by reading the old data and combining the full-strip, the full-write process is reused, which not only reduces the system design complexity, but also continues the high performance advantage of the full-write scenario by means of parallel disk writing, and at the same time, relying on the three check blocks, the data reliability is guaranteed, which effectively adapts to the storage scenario where the host write data does not cover the full strip but has a higher requirement for the write efficiency.
[0073] It should be noted that the embodiments of the present disclosure can include a plurality of steps, which are numbered for the convenience of description, but these numbers do not limit the execution time slots and execution order between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not limit this.
[0074] Corresponding to the above-mentioned disk array data writing method, the present disclosure also proposes a disk array data writing device. Since the device embodiments of the present disclosure correspond to the above-mentioned method embodiments, for the details not disclosed in the device embodiments, the above-mentioned method embodiments can be referred to, and the present disclosure will not be described in detail.
[0075] Figure 2 A structural schematic diagram of a disk array data writing device provided by an embodiment of the present disclosure is shown in Figure 2 as shown, comprising: A determination unit 21 is configured to receive a write request of a host, determine a write mode of corresponding data based on the write request, and the write mode includes full-write, full-write and small-write. The acquisition unit 22 is configured to acquire new data and old data of at least one target data block corresponding to the write request if the write mode corresponding to the write request is lowercase; The first generation unit 23 is configured to perform incremental calculation based on the new data and the old data, generate check incremental data of the at least one target data block, and write data corresponding to the write request into the at least one target data block for persistent processing; The second generation unit 24 is configured to merge the check incremental data of the blocks, and generate updated triple-check subblocks based on a merging result and the old data; The write unit 25 is configured to persist the triple-check subblocks to physical disks respectively, so as to write data corresponding to the write request into the physical disks in the disk array.
[0076] It should be noted that the foregoing explanation and description of the method embodiments are also applicable to the device of the present embodiment, and the principle is the same, which will not be limited in the present embodiment.
[0077] The features of the embodiments corresponding to the disk array data writing device can be referred to the related description of the embodiments corresponding to the disk array data writing method, which will not be described herein.
[0078] Embodiments of the present application also provide an electronic device, including a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-mentioned disk array data writing method embodiments.
[0079] Embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to perform the steps in any of the above-mentioned disk array data writing method embodiments when running.
[0080] In an exemplary embodiment, the above-mentioned computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0081] Embodiments of the present application also provide a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned disk array data writing method embodiments.
[0082] The embodiment of the present application further provides another computer program product, comprising a nonvolatile computer readable storage medium, the nonvolatile computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps in any of the above-mentioned disk array data writing method embodiments.
[0083] Those skilled in the art will further appreciate that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, various components have been described above generally in terms of their functionality, which could be implemented in either hardware or software. The particular implementation of the functional aspects in either hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled persons can implement the described functionality in varying ways for each particular application, but such implementation should not be construed to limit the scope of the application.
[0084] The above provides a kind of disk array data writing method and device, electronic equipment and storage medium provided by the present application in detail.The principle and implementation of the present application are described in the specific examples in this paper, the above example is only used to help understand the method and its core idea of the present application.It should be pointed out that, for the ordinary skilled person in the art, without departing from the principle of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for writing data to a disk array, characterized in that, The method comprises the following steps: receiving a write request of a host, determining a write mode of corresponding data based on the write request, the write mode comprising full write, large write and small write; if the write mode corresponding to the write request is small write, obtaining new data and old data of at least one target data block corresponding to the write request; performing incremental calculation based on the new data and the old data, generating check incremental data of the at least one target data block, and writing data corresponding to the write request into the at least one target data block for persistent processing; after the incremental calculation of the at least one target data block is completed, merging the check incremental data of the at least one target data block, and generating updated triple check blocks based on the merging result and the old data; persisting the triple check blocks to physical disks respectively, so as to write the data corresponding to the write request into the physical disks in a disk array.
2. The disk array data write method of claim 1, wherein, Before obtaining the new data and the old data of the at least one target data block corresponding to the write request, the method further comprises the following steps: dynamically allocating computing and caching resources for the stripe corresponding to the write request; creating a stripe lock and synchronizing the state information of each node in the fault domain through a lock communication mechanism.
3. The disk array data write method of claim 2, wherein, The step of dynamically allocating computing and caching resources for the stripe corresponding to the write request comprises the following steps: estimating and reserving volatile memory pages for data temporary storage according to the total number of data blocks and check blocks in the stripe write unit; allocating a plurality of input / output control blocks for each data block and each check block; reserving memory resources for creating data block control structures and check block control structures.
4. The method of claim 1, wherein, The step of obtaining the new data and the old data of the at least one target data block corresponding to the write request comprises the following steps: saving the new data corresponding to the write request to an input / output buffer area indicated by a block input / output management structure; for the old data that needs to be read, asynchronously reading disk data through a virtualization layer module and saving the disk data to an input / output buffer area indicated by a stripe cache management structure.
5. The method of claim 4, wherein, The step of performing incremental calculation based on the new data and the old data, and generating check incremental data of the at least one target data block comprises the following steps: performing exclusive or operation on the new data and the old data of the at least one target data block to obtain the check incremental data; saving the check incremental data in a first check block management structure, and the first check block management structure has a mapping relationship with second check block management structures and third check block management structures.
6. The method of claim 1, wherein, The step of merging the check incremental data of the at least one target data block comprises the following steps: creating an exclusive or calculation structure to sequentially merge the check incremental data of the at least one target data block; taking the check block management structure generated by the first target data block as a reference structure, and taking the check block management structures generated by subsequent target data blocks as subsequent structures to participate in the merging calculation.
7. The method of claim 1, wherein, The step of generating updated triple check blocks based on the merging result and the old data comprises the following steps: performing exclusive or operation on the old check data and the merged check incremental data to obtain new check data; saving the new check data in a stripe cache management structure, and updating the context state of an acceleration processing unit to a check data ready state.
8. The method of claim 1, wherein, The method further comprises the following steps: In response to the target data block completing the incremental calculation, performing data flushing processing on the target data block.
9. The method of claim 1, wherein, When the write mode is full write, further comprising: Based on the write request, obtaining data of all data blocks in the stripe; After the data of all data blocks is obtained, based on the data cached by the data blocks, first, second and third check chunks are calculated and generated; After initiating the check chunk calculation, performing the flushing operation of all data blocks; After the check chunk calculation is completed, each check chunk is respectively persisted to the corresponding physical disk.
10. The method of claim 9, wherein, The calculation of the first, second and third check chunks based on the data cached by the data blocks comprises: Applying memory resources and acceleration processing unit resources for the first, second and third check chunks respectively; The first, second and third check chunks are sequentially calculated, and the calculation results are temporarily stored in the associated memory page.
11. The disk array data writing method of claim 10, wherein, The respective persistence of each check chunk to the corresponding physical disk comprises: After all check chunks are calculated, sequentially initiate the disk writing operation of each check chunk; After each check chunk is successfully written to the disk, update its state in the stripe acceleration processing unit context, and release the memory and acceleration processing unit resources it occupies.
12. The method of claim 11, wherein, When the write mode is large write, further comprising: Combining the old data on the member disk and the data of the host write request to form a full stripe; Reuse the full write process to parallel flush the data chunk and the check chunk.
13. A disk array data writing apparatus, characterized by comprising: Comprising: A determination unit configured to receive a write request of a host, determine a write mode of corresponding data based on the write request, the write mode comprising full write, large write and small write; An obtaining unit configured to, if the write mode corresponding to the write request is small write, obtain new data and old data of at least one target data block corresponding to the write request; A first generating unit configured to perform incremental calculation based on the new data and the old data, generate check incremental data of the at least one target data block, and write data corresponding to the write request to the at least one target data block for persistent processing; A second generating unit configured to merge the check incremental data of the block, and generate updated triple check chunks based on the merging result and the old data; A write unit configured to persist the triple check chunks to physical disks respectively, so as to write data corresponding to the write request to the physical disks in the disk array.
14. An electronic device, comprising: Comprising: At least one processor; And A memory connected in communication with the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the disk array data write method of any one of claims 1-12.
15. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the disk array data write method according to any one of claims 1-12. The computer instructions are used to enable the computer to perform the disk array data write method according to any one of claims 1-12.
Citation Information
Patent Citations
Batch write check-based redundant-array-of-independent-disk method
CN106293990A
RAID5 verification method for performing data verification by array disk
CN115237342A
RAID-based write data cache acceleration method
CN115686366A
Incremental updating method and device of memory storage system, equipment, medium and product
CN115981875A
Hot data management method, system and device and computer readable storage medium
CN116382568A