Erasure code strip updating method and device based on multiple queues and electronic equipment
By adopting a multi-queue-based erasure code strip update method in an all-flash cluster, dynamically adjusting the queue allocation and strip generation strategies of data blocks, the problem of low SSD disk life is solved, and the effect of extending the service life of SSD and ensuring data read and write efficiency is achieved.
Patent Information
- Application Number
- CN202510101745.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
The existing erasure code strip update solution cannot effectively solve the problem of low life of SSD disks in all-flash clusters. It is mainly due to the increase in write operations caused by frequent data updates, which leads to increased wear and tear of SSD disks.
The multi-queue-based erasure code strip update method is adopted. By obtaining the basic metadata information of the data block, the queue allocation and strip generation strategies of the data block are dynamically adjusted, the write amplification effect is reduced, the write load is dispersed, and the SSD wear is reduced.
It effectively extends the service life of the SSD, ensures the data read and write efficiency after update operation, adapts to different storage needs and access scenarios, and ensures the long-term and stable operation of the system.
Smart Images

Figure CN120010783A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer storage technology, and more specifically, to a method, device and electronic device for updating erasure code stripes based on multiple queues. Background Art
[0002] With the rapid development of information technology, massive amounts of data are constantly being generated, and how to reliably store this data has become a challenge. Erasure coding technology is a reliability assurance technology that is widely used in various large data centers. By processing data in blocks and calculating additional check blocks, the original data can still be restored through decoding when some data is lost, providing high availability assurance for data storage. In a solid-state drive (SSD) cluster, erasure coding reduces data redundancy with its unique data sharding and redundancy mechanism. While ensuring data security and reliability, it improves SSD storage utilization, reduces overall storage costs, and drives all-flash cluster storage towards high efficiency and economy.
[0003] However, the life of SSD disks is directly affected by the number of write operations. SSD disks with frequent write operations have higher wear and are more prone to failure. In a write-intensive all-flash cluster, since erasure coding technology relies on check blocks to ensure data integrity and recoverability, frequent data updates will trigger frequent check block calculation and update operations, introducing a large number of additional write operations, aggravating SSD disk wear, and thus significantly reducing the life of SSD disks in the all-flash cluster.
[0004] Most of the existing optimization solutions for erasure code stripe updates are suitable for HDD clusters, focusing on how to increase the speed of stripe updates and reduce the data transmission flow during stripe updates, but not on the number of writes, which is a design indicator that directly affects the life of SSD disks. Therefore, many existing mature erasure code update optimization solutions are not suitable for all-flash clusters. Summary of the invention
[0005] In view of the defects of the prior art, the purpose of the present application is to provide a multi-queue based erasure code stripe update method, device and electronic device, aiming to solve the problem that the existing erasure code stripe update scheme cannot solve the low lifespan of SSD disks in an all-flash cluster.
[0006] To achieve the above objectives, in a first aspect, the present application provides a multi-queue based erasure code stripe update method, comprising: Get the basic metadata information of the data blocks in the original stripe; Based on the basic metadata information, the data block is placed into different preset encoding queues, where the different preset encoding queues include a high-frequency update encoding queue, a low-frequency update encoding queue, and a no-update encoding queue; Based on the data blocks in different preset encoding queues, different preset generation strategies are used to generate stripes.
[0007] This application uses a specific data block metadata management method to dynamically adjust the queue allocation and stripe generation strategy of the data block, reduce the write amplification effect during the erasure code stripe update process, effectively disperse the write load, reduce the wear of the SSD caused by the update operation, extend the service life of the SSD, and ensure the data reading and writing efficiency after the update operation.
[0008] According to a multi-queue based erasure code stripe updating method provided by the present application, the data blocks are placed into different preset coding queues based on the basic metadata information, including: In a set time window, the update frequency and locality characteristics of each data block are calculated through the basic metadata information, and position coding is performed based on the update frequency and locality characteristics; The position weight of the data block is calculated based on the position coding, and the data block is placed into a corresponding preset coding queue based on the position weight.
[0009] According to a multi-queue based erasure code stripe updating method provided by the present application, the position weight of the data block is calculated based on the position code, including: Within a set time window, determine the data block with the highest update frequency among all data blocks; For the remaining data blocks, a dot product is performed between the position code and the data block with the highest update frequency to obtain a position weight of the current data block relative to the data block with the highest update frequency.
[0010] According to a multi-queue based erasure code stripe updating method provided by the present application, the data block is placed into a corresponding preset coding queue based on the position weight, including: Multiply the position weight of each data block by its respective update frequency to obtain a comprehensive weight; The data blocks with comprehensive weights greater than the preset threshold and the data blocks with the highest update frequency are placed in the high-frequency update coding queue, the other data blocks are placed in the low-frequency update coding queue, and the data blocks not captured within the set time window are placed in the no-update coding queue.
[0011] According to a multi-queue based erasure code stripe updating method provided by the present application, the stripes are generated using different preset generation strategies based on data blocks in different preset coding queues, including: For the high-frequency update encoding queue, stripes are sequentially generated according to the existing data block sequence in the queue and the currently set encoding. If an update operation occurs during or after stripe generation, the corresponding encoding calculation is updated in the cache layer; Based on the health status information of each hard disk, the new stripe generated by the high-frequency update encoding queue is transferred to a hard disk with good health for storage.
[0012] According to a multi-queue based erasure code stripe updating method provided by the present application, the method further includes: Dynamically adjust the queue to which the data block belongs, and regenerate and store stripes based on the current cluster load.
[0013] This application dynamically adjusts the storage strategy according to file access patterns and update conditions to adapt to different storage requirements and access scenarios, ensuring long-term stable operation of the system.
[0014] In a second aspect, the present application provides a multi-queue based erasure code stripe updating device, comprising: An acquisition module is used to obtain basic metadata information of data blocks in the original stripe; An adjustment module, configured to place the data block into different preset encoding queues based on the basic metadata information, wherein the different preset encoding queues include a high-frequency update encoding queue, a low-frequency update encoding queue, and a non-update encoding queue; The generation module is used to generate stripes using different preset generation strategies based on data blocks in different preset encoding queues.
[0015] In a third aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is used to execute the multi-queue based erasure code stripe update method described in the first aspect or any possible implementation of the first aspect.
[0016] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the multi-queue based erasure code stripe update method described in the first aspect or any possible implementation of the first aspect.
[0017] In a fifth aspect, the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the multi-queue based erasure code stripe update method described in the first aspect or any possible implementation of the first aspect.
[0018] It can be understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0019] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the prior art: Through specific data block metadata management methods, the queue allocation and stripe generation strategies of data blocks are dynamically adjusted to reduce the write amplification effect during the erasure code stripe update process, effectively disperse the write load, reduce the wear of the SSD caused by the update operation, extend the service life of the SSD, and ensure the data reading and writing efficiency after the update operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 It is a flowchart of a multi-queue based erasure code stripe updating method provided in an embodiment of the present application; Figure 2 It is a schematic diagram of the stripe generation process provided by an embodiment of the present application; Figure 3 It is an overall architecture diagram of a multi-queue based erasure code stripe update system provided in an embodiment of the present application; Figure 4 It is a structural diagram of a multi-queue based erasure code stripe updating device provided in an embodiment of the present application; Figure 5 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0023] The term "and / or" in this article is a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The symbol " / " in this article indicates that the associated objects are in an or relationship, for example, A / B means A or B.
[0024] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0025] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. For example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.
[0026] Next, combine Figure 1-Figure 3 The erasure code stripe updating method based on multiple queues provided in the embodiment of the present application is introduced.
[0027] Figure 1 is a flow chart of a method for updating erasure codes based on multiple queues provided in an embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps: Step 100, obtaining basic metadata information of data blocks in the original stripe; Optionally, when data is first written, the data is divided into blocks in the order in which the data stream is written and a corresponding stripe, ie, an original stripe, is created. At this time, basic metadata information of the data blocks in the original stripe can be obtained.
[0028] Optionally, the metadata of each data block includes information such as the logical address, timestamp, operation type (read / write / update) of the data block, thereby further obtaining information such as the update frequency and read locality of each data block.
[0029] The update frequency of a data block indicates the number of times a data block is updated within a certain period of time, reflecting the update frequency of the data block; while read locality refers to the access mode of the data block, that is, whether there are concentrated read requests for the data block, that is, which data blocks may be accessed in the near time when accessing the current data block.
[0030] Optionally, after obtaining the basic metadata information, a dynamically updated hash table may be maintained in memory to record the basic metadata information of each data block.
[0031] Step 110, based on the basic metadata information, the data blocks are placed into different preset encoding queues, where the different preset encoding queues include a high-frequency update encoding queue, a low-frequency update encoding queue, and a non-update encoding queue; Based on the basic metadata information of each data block, the update frequency and read locality of the data block can be obtained. Based on this information, the data blocks can be divided into three categories and placed in a high-frequency update encoding queue, a low-frequency update encoding queue, and a no-update encoding queue respectively.
[0032] Step 120: Based on the data blocks in different preset encoding queues, different preset generation strategies are used to generate stripes.
[0033] For data blocks in different encoding queues, the data blocks in each queue will generate corresponding stripes according to a certain strategy.
[0034] The high-frequency update encoding queue includes data blocks that are included in the current time window and have a high update frequency. It is mainly used to store data blocks that need to be updated frequently, and an efficient stripe generation and storage strategy is adopted for these data blocks.
[0035] The low-frequency update encoding queue includes data blocks with moderate update frequency within the current time window. The data blocks with a certain update frequency placed in this queue are grouped into the same stripe to reduce the frequency of recalculation of encoding blocks caused by updates.
[0036] The no-update encoding queue contains data blocks that have not been updated in the current time window. Since these data blocks have not been updated for a long time, these no-update data blocks are fully used to generate new stripes to avoid mixing no-update data blocks with high-frequency update data blocks to form stripes, causing no-update data blocks to be repeatedly read into memory for recalculation of encoding blocks, consuming computing resources and memory space.
[0037] The present application provides an erasure code stripe update method based on multiple queues. Through a specific data block metadata management method, the queue allocation and stripe generation strategy of the data block are dynamically adjusted to reduce the write amplification effect during the erasure code stripe update process, effectively disperse the write load, reduce the wear of the SSD caused by the update operation, extend the service life of the SSD, and ensure the data reading and writing efficiency after the update operation.
[0038] In some embodiments, step 110 specifically includes: Step 1101, within a set time window, obtain the update frequency and locality characteristics of each data block through basic metadata information calculation, and perform position coding based on the update frequency and locality characteristics; Step 1102, calculating the position weight of the data block based on the position coding, and placing the data block into a corresponding preset coding queue based on the position weight.
[0039] Optionally, within a set time window, the metadata in the hash table can be used to count the operation type and corresponding number of each data block, and the data block with two or more write operations is considered to have been updated, and the corresponding number of writes is recorded as the update frequency of the block within the time window. At the same time, within the set locality window, the timestamp and operation type of each data block are counted through the metadata in the hash table, and the data blocks with the same type of operation near the timestamp (such as multiple blocks read at the same time) are considered to have locality.
[0040] Optionally, after obtaining the locality characteristics and update frequency of the data block, the above two characteristics may be written into the hash table.
[0041] Perform positional encoding (PE) on a data block within a set time window to mark the position characteristics of each data block in the current time window, put blocks with similar update frequencies into the same queue while trying to preserve the locality between data blocks.
[0042] The specific encoding calculation formula for position encoding is as follows:
[0043]
[0044] Among them, pos represents the position of the data block in the sequence written in the time window, expressed as a numerical value; window is the size of the set time window, corresponding to the number of data blocks in the currently set time window; size is the number of encoding bits currently used, which is a variable that can be set according to the window size; i is a variable that changes according to the value of the encoding bit size, and is used to participate in the encoding calculation to complete the encoding.
[0045] By the above method, the position code of each data block can be obtained. The position code of each data block can be expressed as:
[0046]
[0047] After obtaining the position code of each data block, the position weight of the data block can be calculated based on the position code, and the data block can be placed in a corresponding preset coding queue based on the position weight.
[0048] In some embodiments, calculating the position weight of the data block based on the position code in step 1102 specifically includes: Step 11021, determining the data block with the highest update frequency among all data blocks within the set time window; Step 11022: for the remaining data blocks, perform a dot product between the position code and the data block with the highest update frequency to obtain the position weight of the current data block relative to the data block with the highest update frequency.
[0049] Specifically, first, find the data block with the highest frequency in the set time window, then obtain the position code of the data block, multiply the position code of other data blocks by the position code of the data block with the highest update frequency, and obtain the relative position weights of all other data blocks to the data block with the highest frequency.
[0050] In some embodiments, placing the data block into the corresponding preset encoding queue based on the position weight in step 1102 specifically includes: Step 11023, multiplying the position weight of each data block by its respective update frequency to obtain a comprehensive weight; Step 11024, put the data blocks with comprehensive weights greater than the preset threshold and the data blocks with the highest update frequency into the high-frequency update encoding queue, put other data blocks into the low-frequency update encoding queue, and put the data blocks not captured within the set time window into the no-update encoding queue.
[0051] Multiply the relative position weight with the frequency of each update to obtain a weight that combines the position information and the update frequency. Sort the weights, put the data blocks with high weights and the data blocks with the highest frequency into the high-frequency update encoding queue, put other data blocks in the time window into the low-frequency update encoding queue, and finally put the data blocks not captured in the time window into the no-update encoding queue.
[0052] Optionally, the preset threshold can be freely set according to actual needs.
[0053] In some embodiments, step 120 specifically includes: Step 1201, for a high-frequency update encoding queue, sequentially generate stripes according to the existing data block sequence in the queue and the currently set encoding. If an update operation occurs during or after stripe generation, the corresponding encoding calculation is updated in the cache layer; Step 1202: Based on the health status information of each hard disk, the new stripe generated by the high-frequency update encoding queue is transferred to a hard disk with good health for storage.
[0054] For high-frequency update encoding queues, low-frequency update encoding queues, and no-update encoding queues, the stripes of the queues will be generated sequentially according to the order of the existing data blocks in the queue and the currently set encoding, that is, in a first-in-first-out order.
[0055] Optionally, for a non-update encoding queue, if all data blocks in an existing stripe still belong to the non-update queue, the stripe is not re-divided and generated.
[0056] Optionally, for the low-frequency update encoding queue, a repartition threshold is set to avoid the generation of low-frequency update queue stripes in which some update blocks account for a low proportion.
[0057] For example, if the repartition threshold is set to 80%, that is, when more than 80% of the data blocks in a stripe still belong to the no-update queue, even if multiple data blocks in the stripe are divided into the low-frequency update queue according to metadata characteristics, the data blocks of the stripe will not be repartitioned to form a new stripe to avoid frequent encoding disk reading and recalculation.
[0058] Optionally, for high-frequency update encoding queues, since the newly generated stripes will be temporarily stored in the cache and will not be written to the disk immediately, if an update operation occurs during or after the stripe generation process, the corresponding encoding calculation will continue to be updated in the cache layer, reducing the frequency of writing to the SSD and the frequency of reading SSD data, thereby reducing the wear of the SSD and ensuring that high-frequency write data can be updated quickly and maintain storage stability.
[0059] Optionally, for new stripes generated by high-frequency update encoding queues, the SMART and other disk health status information of the current cluster SSD disk can be combined to transfer the new stripes generated by the frequent update queue to disks with better health for storage, thereby reducing the wear on the SSD disk caused by frequent update reads and writes in the future.
[0060] In some embodiments, the method further comprises: Dynamically adjust the queue to which the data block belongs, and regenerate and store stripes based on the current cluster load.
[0061] As the access and update patterns of data blocks change, the data block metadata in the hash table will be updated in real time. The queue to which the data block belongs can be dynamically adjusted based on these changes, and the stripes can be regenerated and stored based on the current cluster load.
[0062] This mechanism ensures that the system can always adapt to changing storage needs and access patterns, ensuring that data block management is always in the best state. Through dynamic monitoring and adjustment, the system can continuously optimize the utilization of storage resources, avoid excessive write amplification caused by frequent updates, and thus extend the service life of the SSD.
[0063] Figure 2 is a schematic diagram of the stripe generation process provided by the embodiment of the present application, such as Figure 2 As shown, in one embodiment of the present application, the cluster adopts (6, 4) RS encoding. There are 16 data blocks (i.e., window=16) currently entering the time window, namely ADD_0, ADD_1, ADD_2, ADD_3, ADD_4, ADD_5, ADD_6, ADD_7, ADD_8, ADD_9, ADD_10, ADD_11. In the original stripe, ADD_0, ADD_1, ADD_2, ADD_3 generate corresponding check blocks to form stripe 1, ADD_4, ADD_5, ADD_6, ADD_7 form stripe 2, and ADD_8, ADD_9, ADD_10, ADD_11 form stripe 3.
[0064] The embodiment uses a time window to capture a data block write record in a trace. According to statistics, the update frequency of data block ADD_9 in the time window is the highest, which is 3; while the update frequency of data blocks ADD_6, ADD_7 and ADD_11 is 2; the update frequency of data blocks ADD_2, ADD_3, ADD_5 and ADD_8 is 1. Among them, since the update frequency of ADD_0, ADD_1, ADD_4, and ADD_10 is 0, they directly enter the no-update encoding queue.
[0065] If we simply arrange the data blocks from small to large based on the update rate, and evenly divide the remaining data blocks into the remaining low-frequency update encoding queue and high-frequency update encoding queue, the data blocks entering the low-frequency update encoding queue should be ADD_2, ADD_5, ADD_3, ADD_8, and the data blocks entering the high-frequency update encoding queue should be ADD_9, ADD_6, ADD_7, ADD_11. However, this queue grouping scheme does not take the locality between data blocks into account, so each block will be position-encoded to explore the locality between data blocks while considering the update frequency.
[0066] For the convenience of coding examples in the embodiment, the value of size is set to 4, that is, the position code is a 4-bit code.
[0067] Next, the embodiment uses position coding to position code the data blocks in the current window in combination with the update frequency and locality of each data block. Take data block ADD_9, data block ADD_11 and data block ADD_3, and display their position coding calculation results. The relative position of data block ADD_9 in the current time window is marked as 0, 2, and 3. According to the position coding calculation formula, the three position codes of ADD_9 are obtained as [sin0, cos0, sin0 / 16, cos0 / 16], [sin2, cos2, sin2 / 16, cos2 / 16], and [sin3, cos3, sin3 / 4, cos3 / 4]. Similarly, the two position codes of data block ADD_11 can be calculated as [sin1, cos1, 1sin11 / 16, cos11 / 16] and [sin12, cos12, sin12 / 16, cos12 / 16]. For ADD_3, its position encoding is [sin1, cos1, 1sin1 / 16, cos1 / 16].
[0068] A smaller sliding window location window is set in the time window to explore the local association between data blocks. Since data block ADD_9 is the data block with the highest update frequency in the current location window, the position weight is calculated according to the position code to find data blocks with a similar update frequency and local association with ADD_9. Taking data blocks ADD_11 and ADD_3 in the current location window as examples, by calculating the position weight, it is determined which block should enter the frequent update queue together with ADD_9. The position weight of ADD_11 relative to ADD_9 is [sin1, cos1, 1sin11 / 16, cos11 / 16] dot multiplication [sin3, cos3, sin3 / 4, cos3 / 4], the result is 0.732082528, and then multiplied by the update frequency of ADD_11 2, the final position weight is 1.464. For ADD_3, its position encoding is [sin1, cos1, 1sin1 / 16, cos1 / 16], and the dot product result with ADD_9 is 1.538349817, which is then multiplied by ADD_3's own update frequency 1, and the final weight is 1.538.
[0069] From the above calculation and analysis, it can be obtained that even though data block ADD_11 has a higher update frequency than data block ADD_3, since data block ADD_3 has a better spatiotemporal correlation with ADD_9, the data block that finally follows data block ADD_9 into the frequent update queue is ADD_3, not ADD_11.
[0070] Combined with the current encoding parameters (each stripe generates 2 check blocks from 4 data blocks), for ease of understanding, the embodiment chooses to select 4 data blocks from the time window to enter the high-frequency update encoding queue. Combined with the above classification method, the position weights of all data blocks in the current window are calculated, and the blocks that finally enter the high-frequency update encoding queue are ADD_3, ADD_9, ADD_6, and ADD_7. In the weight calculation, data blocks ADD_2, ADD_5, ADD_11, and ADD_8 with lower weight values enter the low-frequency update encoding queue. In actual applications, how many data blocks are selected to enter the high-frequency update encoding queue based on the position weight calculation results can be set by yourself.
[0071] Figure 3 is an overall architecture diagram of a multi-queue erasure code stripe update system provided in an embodiment of the present application, such as Figure 3 As shown, in one embodiment of the present application, a multi-queue-based erasure code stripe update system is used to implement the multi-queue-based erasure code stripe update method provided by the present application, and the system specifically includes the following components: 1. Data block metadata management module This module is responsible for recording metadata information such as the logical address, update frequency, and read locality of each data block in real time. The metadata of the data block is updated regularly and passed to the queue management module for subsequent queue classification and stripe optimization. This module is also responsible for monitoring the access mode of the data block and providing data support for dynamically adjusting the queue classification and stripe generation strategy.
[0072] 2. Queue management module It includes no update queue, low frequency update queue and frequent update queue. This module dynamically allocates data blocks to corresponding queues according to the metadata of data blocks, and adjusts the queue division criteria according to system load and file access mode. The queue management module cooperates with the stripe generation module to generate adaptive storage stripes according to the characteristics of the queue and regularly optimize the storage strategy.
[0073] 3. Strip generation module This module generates optimized stripes for data blocks in different queues according to the characteristics of the queues and the currently set encoding parameters. At the same time, the stripe generation module collaborates with the performance optimization module to continuously optimize the stripe generation strategy according to the access mode and storage requirements of the data blocks, and re-layout the storage of the newly generated stripe data blocks to ensure efficient use of system storage space and rapid data recovery.
[0074] 4. Performance optimization module This module is responsible for monitoring the performance of the entire storage system, analyzing the access frequency of data blocks, read locality, and SSD load. Based on this information, the system automatically adjusts the queue division criteria, stripe generation strategy, and other storage parameters to ensure that the system can adapt to changing storage needs. The core goal of the performance optimization module is to balance storage load, reduce write amplification effects, and improve the efficiency of file updates and data recovery.
[0075] Figure 4 is a schematic diagram of the structure of the erasure code stripe updating device based on multiple queues provided in an embodiment of the present application, such as Figure 4 As shown, the device includes an acquisition module 410, an adjustment module 420 and a generation module 430, wherein: An acquisition module 410 is used to acquire basic metadata information of data blocks in the original stripe; An adjustment module 420, configured to place data blocks into different preset encoding queues based on basic metadata information, wherein the different preset encoding queues include a high-frequency update encoding queue, a low-frequency update encoding queue, and a non-update encoding queue; The generating module 430 is configured to generate stripes using different preset generating strategies based on data blocks in different preset encoding queues.
[0076] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the device are similar to those described in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method, which will not be repeated here.
[0077] Based on the method in the above embodiment, Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5 As shown, an embodiment of the present application provides an electronic device, which may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the erasure code stripe update method based on multiple queues in the above embodiment.
[0078] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the erasure code stripe update method based on multiple queues described in various embodiments of the present application.
[0079] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the multi-queue based erasure code stripe update method in the above embodiment.
[0080] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the erasure code stripe update method based on multiple queues in the above embodiment.
[0081] It is understandable that the processor in the embodiment of the present application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0082] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0083] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.
[0084] It should be understood that the various numerical numbers involved in the embodiments of the present application are only used for the convenience of description and are not used to limit the scope of the embodiments of the present application.
[0085] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A multi-queue based erasure code stripe update method, characterized in that: include: Get the basic metadata information of the data blocks in the original stripe; Based on the basic metadata information, the data block is placed into different preset encoding queues, where the different preset encoding queues include a high-frequency update encoding queue, a low-frequency update encoding queue, and a no-update encoding queue; Based on the data blocks in different preset encoding queues, different preset generation strategies are used to generate stripes.
2. The erasure code stripe update method based on multiple queues according to claim 1, characterized in that: The step of placing the data blocks into different preset encoding queues based on the basic metadata information includes: In a set time window, the update frequency and locality characteristics of each data block are calculated through the basic metadata information, and position coding is performed based on the update frequency and locality characteristics; The position weight of the data block is calculated based on the position coding, and the data block is placed into a corresponding preset coding queue based on the position weight.
3. The erasure code stripe update method based on multiple queues according to claim 2 is characterized in that: The calculating the position weight of the data block based on the position code comprises: Within a set time window, determine the data block with the highest update frequency among all data blocks; For the remaining data blocks, a dot product is performed between the position code and the data block with the highest update frequency to obtain a position weight of the current data block relative to the data block with the highest update frequency.
4. The erasure code stripe update method based on multiple queues according to claim 2 is characterized in that: The step of placing the data block into a corresponding preset encoding queue based on the position weight includes: Multiply the position weight of each data block by its respective update frequency to obtain a comprehensive weight; The data blocks with comprehensive weights greater than the preset threshold and the data blocks with the highest update frequency are placed in the high-frequency update coding queue, the other data blocks are placed in the low-frequency update coding queue, and the data blocks not captured within the set time window are placed in the no-update coding queue.
5. The method for updating erasure codes stripes based on multiple queues according to claim 1, characterized in that: The step of generating stripes based on data blocks in different preset encoding queues using different preset generation strategies respectively includes: For the high-frequency update encoding queue, stripes are sequentially generated according to the existing data block sequence in the queue and the currently set encoding. If an update operation occurs during or after stripe generation, the corresponding encoding calculation is updated in the cache layer; Based on the health status information of each hard disk, the new stripe generated by the high-frequency update encoding queue is transferred to a hard disk with good health for storage.
6. The erasure code stripe update method based on multiple queues according to claim 1, characterized in that: The method further comprises: Dynamically adjust the queue to which the data block belongs, and regenerate and store stripes based on the current cluster load.
7. A multi-queue based erasure code stripe update device, characterized in that: include: An acquisition module is used to obtain basic metadata information of data blocks in the original stripe; An adjustment module, configured to place the data block into different preset encoding queues based on the basic metadata information, wherein the different preset encoding queues include a high-frequency update encoding queue, a low-frequency update encoding queue, and a non-update encoding queue; The generation module is used to generate stripes using different preset generation strategies based on data blocks in different preset encoding queues.
8. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the multi-queue based erasure code stripe update method as described in any one of claims 1-6.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor, the processor executes the erasure code stripe updating method based on multiple queues as described in any one of claims 1-6.
10. A computer program product, characterized in that When the computer program product runs on a processor, the processor executes the erasure code stripe updating method based on multiple queues as described in any one of claims 1-6.