Storage system array generation method and electronic device

By using a sliding window algorithm and a flexible virtual cell partitioning method, the problem of low resource utilization caused by the requirement that the total number of hard drives in a RAID10 storage system must be even is solved, thus achieving efficient and flexible use of storage resources and redundancy protection.

CN121349378BActive Publication Date: 2026-04-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, RAID10 storage systems require a total number of hard drives to be even, resulting in lower flexibility and utilization of storage resources.

Method used

By setting a sliding window to traverse the storage unit array, the target window is accurately determined. Based on the parity of the number of storage units, virtual units are flexibly divided to form a storage redundancy array, ensuring the smooth construction of the redundancy array and improving the flexibility and space utilization of the hard drive.

Benefits of technology

It enables the efficient construction of redundant storage arrays in both odd and even number of hard drives, improving the flexibility and utilization of storage resources, and reducing the waste of hard drive resources and the risk of data loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121349378B_ABST
    Figure CN121349378B_ABST
Patent Text Reader

Abstract

This application discloses a storage system array generation method and electronic device, relating to the field of computer storage technology. By setting a sliding window to traverse the array, the target window is accurately determined, and suitable storage units are selected. For an even number of storage units, two are paired to form a redundant group that mirrors each other, and then striped to form an array. For an odd number of storage units, multiple virtual sub-units are determined using reference values. The storage system's redundant array is formed based on these multiple virtual sub-units. The virtual units are flexibly divided according to their parity, ensuring the redundant array can be successfully assembled, achieving data redundancy protection, improving hard drive flexibility and space utilization, and solving problems such as low flexibility and utilization of storage resources in related technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer storage technology, and in particular to a method for generating storage system arrays and electronic devices. Background Technology

[0002] RAID 10 (Redundant Array of Independent Disks Level 10) pairs physical hard drives to form multiple independent RAID 1 (Redundant Array of Independent Disks Level 1) mirror groups. Within each mirror group, data is synchronously written to both drives to create data copies. Then, RAID 0 (Redundant Array of Independent Disks Level 0) striping is performed, splitting the data into multiple segments and writing them sequentially and alternately to different mirror groups, allowing multiple mirror groups to process data read and write tasks in parallel. However, because this technology relies on pairwise RAID 1 mirror groups, the total number of hard drives must be even, resulting in lower flexibility and utilization of storage resources in the storage system. Summary of the Invention

[0003] This application provides a method for generating storage system arrays and an electronic device to at least solve the problems of low flexibility in the use of storage resources and low resource utilization in related technologies.

[0004] This application provides a method for generating a storage system array. The storage system includes multiple storage units. The method includes: obtaining a first capacity value of the storage units; generating a first array based on the first capacity value; traversing the first array based on a pre-set sliding window, obtaining at least one of a first number of storage units, a first reference value, and a first total capacity value within the current window during the traversal; determining a target window from the multiple windows based on at least one of the first number of storage units, the first reference value, and the first total capacity value; dividing a first virtual unit from the storage units within the target window based on the first reference value; if the number of storage units is even, forming a storage redundancy array of the storage system based on the multiple first virtual units; if the number of storage units is odd, determining a second reference value based on the first reference value; dividing the first virtual unit into multiple virtual sub-units based on the second reference value; and forming a storage redundancy array of the storage system based on the divided multiple virtual sub-units.

[0005] This application also provides a storage system array generation apparatus. The storage system includes multiple storage units, comprising: an acquisition module, configured to acquire a first capacity value of the storage units and generate a first array based on the first capacity value; a determination module, configured to traverse the first array based on a pre-set sliding window, and during the traversal process acquire at least one of a first number of storage units, a first reference value, and a first total capacity value within the current window, and determine a target window from the multiple windows based on at least one of the first number of storage units, the first reference value, and the first total capacity value; and a partitioning module, configured to partition a first virtual unit from the storage units within the target window based on the first reference value, wherein if the number of storage units is even, a storage redundancy array of the storage system is formed based on the multiple first virtual units, and if the number of storage units is odd, a second reference value is determined based on the first reference value, and the first virtual unit is divided into multiple virtual sub-units based on the second reference value, and a storage redundancy array of the storage system is formed based on the partitioned multiple virtual sub-units.

[0006] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of the above-described memory system array generation method when executing the computer program.

[0007] This application utilizes a sliding window to traverse an array, precisely determining the target window and selecting suitable storage units. For an even number of storage units, pairwise redundant groups are formed, which are then striped to create an array. For an odd number of storage units, multiple virtual sub-units are determined using reference values. These virtual sub-units are then used to create a redundant storage array for the storage system. The virtual units are flexibly divided according to their parity, ensuring the successful construction of the redundant array, achieving data redundancy protection, and improving hard drive flexibility and space utilization. Therefore, it addresses the issues of low storage resource utilization and flexibility in related technologies. Attached Figure Description

[0008] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 A flowchart illustrating a method for generating a storage system array, as provided in this application embodiment;

[0010] Figure 2 An example diagram of a sliding window algorithm provided in this application embodiment;

[0011] Figure 3A schematic diagram of a sliding window for filtering storage units is provided in one embodiment of this application;

[0012] Figure 4 A schematic diagram of a sliding window for filtering storage units is provided for another embodiment of this application;

[0013] Figure 5 A schematic diagram illustrating a remaining capacity selection rule provided in an embodiment of this application;

[0014] Figure 6 This is a schematic diagram of a RAID 10 organization form provided in an embodiment of this application;

[0015] Figure 7 A schematic diagram illustrating an example of excess capacity in a RAID 10 array disk, provided as an embodiment of this application;

[0016] Figure 8 A schematic diagram of a RAID 10 configuration for selecting excess capacity, provided as an embodiment of this application;

[0017] Figure 9 A schematic diagram illustrating the reordering of disks based on remaining capacity, provided as an embodiment of this application;

[0018] Figure 10 This is a distribution diagram of a storage system array generation device provided in an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0020] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0021] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] The specific application environment architecture or specific hardware architecture on which the execution of the storage system array generation method depends is described here.

[0023] The embodiments of this application provide a method for generating a storage system array. The method is described in detail below in conjunction with the execution flow of the storage system array generation method.

[0024] Specifically, Figure 1 This is a flowchart illustrating the storage system array generation method provided in an embodiment of this application.

[0025] like Figure 1 As shown in the flowchart, the method for generating the storage system array includes the following steps:

[0026] In step S101, the first capacity value of the storage unit is obtained, and a first array is generated based on the first capacity value.

[0027] The storage system includes multiple storage units, which are basic units with independent storage functions and can also be called hard disks; the first capacity value is the numerical value of the available storage capacity provided by the storage unit to the user; the first array is an ordered set of array data structures generated based on the first capacity value.

[0028] It is understandable that extracting the first capacity value of the storage unit can standardize and quantify the capacity information. Generating the first array based on the first capacity value can transform the abstract capacity value into a structured data carrier that can be operated by programs or systems, providing a standardized computing basis for subsequent storage management processing.

[0029] Specifically, data in the storage system is stored in units of storage cells. By obtaining the first capacity value of the storage cell, a first array and the identifier of the storage cell can be generated to achieve granular splitting of the first capacity value. For example, if the first capacity value is 1024MB, an array of 1024 elements [1,1,...,1] is generated with 1MB as the smallest unit, and each element represents an allocatable basic block.

[0030] Further, in an embodiment of this application, generating a first array based on a first capacity value includes: obtaining a sorting rule, wherein the sorting rule includes an ascending order rule and a descending order rule; if the sorting rule is an ascending order rule, then the first capacity value is sorted in ascending order; if the sorting rule is a descending order rule, then the first capacity value is sorted in descending order; and generating a first array based on the sorted first capacity value.

[0031] Among them, sorting rules are the overall standards or logic that determine the order of elements in a set of data; ascending order rules arrange elements from smaller / earlier / lower values ​​to larger / later / higher values ​​according to the sorting rules; descending order rules arrange elements from larger / later / higher values ​​to smaller / earlier / lower values ​​according to the sorting rules.

[0032] It is understandable that sorting the first capacity value according to the sorting rules of ascending or descending order to generate the first array can upgrade the capacity information in the array from unordered structure to ordered structure, thereby realizing ordered retrieval of capacity information, improving data processing efficiency, adapting to the strategic needs of storage resource scheduling, and enhancing business flexibility.

[0033] Specifically, in an unsorted array, the capacity values ​​are distributed randomly, while sorting in ascending / descending order allows the array elements to present a clear order. If you need to find the smallest available capacity block, the first element of the ascending array is the target, without having to traverse the entire array; if you need to find the largest available capacity block, the first element of the descending array satisfies the requirement, saving the cost of comparing each element one by one.

[0034] This orderliness allows subsequent search and filtering operations to be optimized based on sequential characteristics, reducing the time complexity from unordered traversal to ordered search, especially in arrays split from large values, where the efficiency improvement is significant.

[0035] For example, sorting all storage units by capacity from smallest to largest would result in a sorted array: sorted_disk=[s1,s2,...,s...]. n (s1≤s2≤...≤s) n The sorting process places storage units with similar capacities next to each other, making it easier to filter them later.

[0036] This application embodiment extracts the first capacity value of the storage unit, which can standardize and quantify the capacity information. A first array is generated based on the first capacity value, which can transform the abstract capacity value into a structured data carrier that can be operated by a program or system, providing a standardized computing basis for subsequent storage management processing.

[0037] In step S102, the first array is traversed based on a pre-set sliding window. During the traversal, at least one of the first number of storage units, the first reference value, and the first total capacity value in the current window is obtained. Based on at least one of the first number of storage units, the first reference value, and the first total capacity value, the target window is determined from multiple windows.

[0038] Here, the sliding window is a movable subsequence range set on the first array for batch extraction of local data; the first quantity is the total number of storage units contained in the current sliding window, i.e., the number of elements in the window; the first reference value is a benchmark reference index calculated based on the characteristics of the storage units in the current window; the first total capacity value is the sum of the capacities of all storage units in the current sliding window; and the target window is a specific window that meets preset conditions and is selected from all generated windows based on the first quantity, the first reference value, and the first total capacity value during the process of traversing the first array through the sliding window.

[0039] Understandably, the sliding window divides the full data into multiple windows through local traversal of a fixed / dynamic range. Each window corresponds to a local combination. By extracting the first quantity, first total capacity, and first reference value within the window, windows that do not meet the conditions can be quickly eliminated. Finally, the target window is determined from the remaining windows, achieving efficient global adaptation filtering, improving the efficiency of full data processing, and ensuring the accurate adaptation of the target window.

[0040] Specifically, a storage system may have dozens or even hundreds of storage units. If four storage units are selected to build a RAID10, the number of combinations will grow exponentially by analyzing each combination. For example, there are over 3.92 million combinations if four out of 100 storage units are selected. By using a sliding window traversal, the analysis scope can be limited to continuous local combinations, which greatly reduces the computational cost. The target window is determined by the first quantity, the first total capacity value, and the first reference value. This ensures that the traversal logic of the sliding window can adapt to changes in real time, ensuring that the selected window fully meets the requirements. The target window is then re-selected to ensure business continuity.

[0041] Furthermore, in the embodiments of this application, traversing the first array based on a pre-set sliding window includes: obtaining the target size, sliding step size, and sliding direction of the sliding window; determining the starting position of the sliding window in the first array according to the sliding direction; and moving the sliding window of the target size according to the sliding step size and sliding direction, starting from the starting position of the first array.

[0042] The target size is the number of storage units covered by the sliding window in one pass, i.e., the number of array elements contained in the window; the sliding step is the number of storage units that the sliding window skips backward or forward relative to the current position each time it moves; the sliding direction is the direction in which the sliding window moves on the first array, i.e., whether the window moves from the beginning to the end of the array or from the end to the beginning; and the starting position is the initial position at which the sliding window begins its first traversal.

[0043] Understandably, by determining the starting position of the sliding window in the first array based on the sliding direction, and moving the sliding window of the target size from the starting position of the first array according to the sliding step and sliding direction, the traversal logic can be flexibly adjusted according to specific needs, significantly reducing computational costs and improving the efficiency of storage resource analysis.

[0044] Specifically, the target size of the sliding window is m, ensuring that each window contains exactly m storage units. A window must contain m storage units; calculating the minimum capacity using the leftmost element is wasteful. The sliding step size is 1, moving 1 position at a time to ensure that adjacent windows differ by only 1 storage unit. For example, if the previous window is [i, i+m-1] and the next window is [i+1, i+m], the entire array is traversed for all combinations of m consecutive storage units. The sliding direction is from the beginning to the end of the array, ensuring sequential traversal. For each window [s[i], s[i+1], ..., s[i+m-1]], all possible minimum capacity starting windows are completely covered. The starting position is the first element of the array, i.e., index 0. The window range is from index i=0 to i=nm, a total of n-m+1 windows, ensuring traversal starts from the window with the smallest minimum capacity, without missing any possible candidate windows.

[0045] For example, RAID10 requires an even number of storage units, so the target size can be set to 4 to ensure that each window covers exactly the number of units that might meet the requirements. If detecting abnormal capacity fluctuations in two consecutive storage units, setting the target size to 2 can focus on small-scale local features. When a quick scan of all storage units is needed, the step size can be set to equal the target size to reduce the number of traversals. When fine-grained investigation is needed, the step size can be set to the maximum overlap to ensure that no consecutive unit combinations are missed. If the first array is sorted in descending order of capacity, forward sliding can prioritize the analysis of high-capacity combinations, while reverse sliding prioritizes the analysis of low-capacity combinations, so that the traversal order is aligned with business priorities.

[0046] Further, in embodiments of this application, determining a target window from multiple windows based on at least one of a first number of storage units, a first reference value, and a first total capacity value includes: calculating a required capacity value based on the first number of storage units and the first reference value; calculating a remaining capacity value within the current window based on the first total capacity value and the required capacity value; and determining the target window from multiple windows based on the remaining capacity value.

[0047] The required capacity value is the total capacity needed by the current window to meet business needs; the remaining capacity value is the total available capacity of the current window after the required capacity value has been met.

[0048] Understandably, the required capacity value is first calculated based on the first number of storage units and the first reference value. Then, the remaining capacity value in the current window is calculated based on the first total capacity value and the required capacity value. Finally, the target window is determined from multiple windows based on the remaining capacity value. By calculating the remaining capacity value, it is determined whether the window is sufficient, and the actual available capacity of the window can be accurately assessed. When the remaining capacity values ​​of multiple windows can meet the business needs, the target window is selected based on the remaining capacity value to achieve optimal resource utilization.

[0049] Specifically, such as Figure 2 As shown, the first total capacity value reflects the theoretical maximum capacity, which is sum_cap=s[i]+s[i+1]+...+s[i+m-1]. The first reference value is a benchmark reference index calculated based on the characteristics of the storage units in the current window, which is min_cap=s[i]. The required capacity value is calculated based on the first number of storage units and the first reference value. However, the actual usable capacity is the remaining part after deducting the necessary reservations, which needs to exclude invalid windows with high total capacity but low actual usable capacity. Therefore, the remaining capacity value in the current window is calculated as waste=sum_cap-min_cap×m based on the first total capacity value and the required capacity value. After traversing all windows, the window with the smallest remaining capacity value can be selected to reduce capacity waste.

[0050] In this embodiment, the sliding window divides the full data into multiple windows through local traversal of a fixed / dynamic range. Each window corresponds to a local combination. By extracting the first quantity, first total capacity, and first reference value within the window, windows that do not meet the conditions can be quickly eliminated. Finally, the target window is determined from the remaining windows, achieving efficient global adaptation filtering, improving the efficiency of full data processing, and ensuring the accurate adaptation of the target window.

[0051] In step S103, a first virtual unit is divided from the storage units in the target window according to the first reference value. If the number of storage units is even, a storage redundancy array of the storage system is formed based on multiple first virtual units. If the number of storage units is odd, a second reference value is determined based on the first reference value. The first virtual unit is divided into multiple virtual sub-units based on the second reference value. A storage redundancy array of the storage system is formed based on the multiple virtual sub-units after division.

[0052] The first virtual unit is a logical storage unit divided from the physical storage units of the target window based on the first reference value; the storage redundancy array is a storage structure composed of multiple storage units; the second reference value is a new reference value adjusted based on the first reference value when the number of storage units in the target window is odd; the virtual subunit is a smaller logical storage unit that further subdivides the first virtual unit based on the second reference value, where the second reference value is smaller than the first reference value.

[0053] Understandably, by flexibly adapting to the odd-even difference in the number of storage units, for an even number of storage units, an array can be formed based on the first virtual unit, which can efficiently generate a storage redundancy array and build a stable redundancy relationship by taking advantage of the symmetry of quantity. For an odd number of storage units, virtual sub-units are divided by a second reference value, breaking the even-number dependency restriction. The first virtual unit of a certain unit is split into multiple sub-units that adapt to other units, forming a mirror or stripe relationship with the sub-units of the remaining units. This can effectively avoid some resources being idle due to an odd number of units. The unified division standard for building a storage redundancy array ensures that storage resource utilization can be maximized and data reliability can be guaranteed regardless of whether the number of units is even or odd. This can reduce the minimum number of storage units required to form a storage redundancy array and improve the flexibility and resource utilization of the storage system.

[0054] Specifically, when the number of storage units is even, the embodiments of this application conform to the logic of using mirroring and striping in RAID10, without the need for additional unit splitting or adjustment. From a practical application perspective, the pairing relationship of even-numbered units is naturally clear. Every two storage units can form an independent mirror, and multiple mirrors can be integrated into a redundant array through striping. The whole process does not require complex logic conversion, and the data synchronization and fault recovery efficiency is extremely high.

[0055] However, when the number of storage units is odd, the RAID10 architecture in related technologies faces the dilemma of idle resources or redundancy failure. For example, out of 5 storage units, only 1 can be discarded, and only 4 can be used to form an array, resulting in a 20% waste of storage resources. Alternatively, forcibly using 5 disks to build an array can cause chaotic redundancy relationships, increasing the risk of data loss in case of failure. The embodiments of this application can solve the problem of not being able to use an odd number of storage units to form a RAID10 array by using a second reference value and the logic of splitting virtual sub-units. This can reduce the minimum number of storage units required to form a redundant storage array and can form a redundant storage array when the number of storage units is odd, effectively improving the flexibility and resource utilization of the storage system.

[0056] Furthermore, in the embodiments of this application, the storage redundancy array of the storage system is composed of multiple first virtual units, including: generating multiple first mirror groups based on the multiple first virtual units, wherein the first virtual units within the first mirror groups are mirror units of each other, and the mirror units synchronously store the same data; generating the storage redundancy array of the storage system based on the multiple first mirror groups, and recording the first metadata information of the storage redundancy array, wherein the multiple first mirror groups within the storage redundancy array allow parallel processing of data read and write tasks, and the first metadata information includes a first mapping relationship between storage units and first virtual units, and the starting address of the first virtual units.

[0057] The first mirror group is a logical group composed of multiple first virtual units used to achieve data mirroring redundancy; the mirror unit is a unit within the first mirror group that mirrors each other and synchronously stores the same data; the first metadata information is a dataset describing the storage redundancy array structure, mapping and other key attributes; the first mapping relationship is the correspondence between physical storage units and first virtual units, used to abstract the storage capacity of physical hardware into logical virtual units; the starting address is the logical starting position of the first virtual unit in the storage system.

[0058] It is understandable that the first virtual units within the first mirror group are mirror units of each other and store the same data synchronously. Multiple first mirror groups are generated based on multiple first virtual units, and a storage redundancy array of the storage system is generated. The first metadata information of the storage redundancy array is recorded. The mirror units ensure complete data consistency through a synchronization mechanism. Multiple first mirror groups form a storage redundancy array. Through parallel design, the overall system performance is improved, the storage resources are managed and scalable, and the operation and maintenance complexity is reduced.

[0059] Specifically, the minimum window size (min_cap) is used as the standard, such as... Figure 3 As shown, the selected disk capacity is divided into two parts: the first part consists of disks with equal capacity, and the second part consists of disks with unequal capacity. Each disk in the first part is virtualized as member disks member_disk0~member_disk1, and the mapping relationship between each disk and the member disks is recorded, as follows: Figure 3 The metadata information is shown below:

[0060] member_disk0:disk1,start_lba=0,capacity=min_cap;

[0061] member_disk1:disk2,start_lba=0,capacity=min_cap;

[0062] member_disk2:disk3,start_lba=0,capacity=min_cap;

[0063] disk1, disk2, and disk3 form a target window of size m=3. Since the array is in ascending order, disk1 is the first reference value. These three storage units are divided into equal-capacity partitions, that is, only the parts that are equal to the first reference value are taken. These partitioned logical units are the first virtual units. For example, the first reference value part of disk1 is virtualized as member_disk0, the corresponding part of disk2 is virtualized as member_disk1, and the corresponding part of disk3 is virtualized as member_disk2.

[0064] The first virtual unit is grouped into a first mirror group. For example, member_disk0 and member_disk1 form a first mirror group. member_disk0 and member_disk1 in the group are mirror units of each other. Data written to member_disk0 will be synchronized to member_disk1 in real time. The two store the exact same data. Multiple first mirror groups are combined to form a storage redundancy array, such as RAID10.

[0065] Furthermore, in the embodiments of this application, the storage redundancy array of the storage system is composed of multiple virtual sub-units after division, including: generating multiple second mirror groups based on the multiple virtual sub-units, wherein the virtual sub-units in the second mirror groups are mirror units of each other, and the mirror units synchronously store the same data; generating the storage redundancy array of the storage system based on the multiple second mirror groups, and recording the second metadata information of the storage redundancy array, wherein the multiple second mirror groups in the storage redundancy array allow parallel processing of data read and write tasks, and the second metadata information includes the second mapping relationship between storage units and virtual sub-units, and the starting address of the virtual sub-units.

[0066] The second mirror group is a group of multiple virtual sub-units. All virtual sub-units in the group are mirror units of each other, and the data is stored in complete synchronization to ensure that one piece of data has multiple backups. The second metadata information is the information that records the configuration of the storage redundancy array. The second mapping relationship is the associated data that clarifies the correspondence between physical storage units and virtual sub-units.

[0067] It is understandable that the second mirror group is generated from multiple virtual sub-units. The virtual sub-units within the second mirror group are mirror units of each other and store the same data synchronously. The storage redundancy array of the storage system is generated from multiple second mirror groups to record the second metadata information of the storage redundancy array. The virtual sub-units within the group are mirror units of each other, and the data is completely synchronized. The existence of mirror units allows the system to read multiple copies in parallel when reading data, which can improve read performance, ensure data reliability, realize the manageability and scalability of storage resources, and reduce the complexity of operation and maintenance.

[0068] Specifically, such as Figure 4 As shown, if m is even, skip the following steps; if m is odd, virtualize each member disk into two member disks based on half the minimum window capacity, i.e., min_cap / 2, as follows: Figure 4 As shown, the corresponding mapping metadata needs to be adjusted as follows:

[0069] member_disk0:disk1,start_lba=0,capacity=min_cap / 2;

[0070] member_disk1:disk2,start_lba=0,capacity=min_cap / 2;

[0071] member_disk2:disk3,start_lba=0,capacity=min_cap / 2;

[0072] member_disk3:disk1,start_lba=min_cap / 2,capacity=min_cap / 2;

[0073] member_disk4:disk2,start_lba=min_cap / 2,capacity=min_cap / 2;

[0074] member_disk5:disk3,start_lba=min_cap / 2,capacity=min_cap / 2;

[0075] The diagram shows that each of the original physical disks disk1, disk2, and disk3 is virtualized into two member disks. For example, disk1 is split into member_disk0 and member_disk3. member_disk0 is a virtual sub-unit disk1 with an initial LBA of 0 and a capacity of min_cap / 2. member_disk3 is another virtual sub-unit disk1 with an initial LBA of min_cap / 2 and a capacity of min_cap / 2.

[0076] Figure 4Virtual member disks of the same color form a mirror relationship, meaning that the virtual sub-units within the second mirror group are mirror images of each other and store the same data synchronously. Second mirror group 1 consists of member_disk0 (upper half of disk1), member_disk1 (upper half of disk2), and member_disk2 (upper half of disk3), achieving mutual mirroring and synchronous storage of the same data among the three. Second mirror group 2 consists of member_disk3 (lower half of disk1), member_disk4 (lower half of disk2), and member_disk5 (lower half of disk3), achieving mutual mirroring and synchronous storage of the same data among the three.

[0077] Figure 4 The configuration information of member_disk reflects the second metadata information, including the second mapping relationship between storage units and virtual sub-units, and the starting address of the virtual sub-unit. For example, it is clear that each virtual sub-unit of member_disk0 corresponds to the physical storage unit of disk1. The start_lba of each virtual sub-unit is the starting address information of the metadata record. For example, start_lba=0 for member_disk0 and start_lba=min_cap / 2 for member_disk3.

[0078] Furthermore, in the embodiments of this application, after dividing the first virtual unit from the storage units within the target window according to the first reference value, the method further includes: dividing the second virtual unit from the storage units within the target window according to the first reference value, wherein the sum of the capacity values ​​of the second virtual unit and the first virtual unit is the first capacity value of the storage unit; generating a second array according to the second capacity value of the second virtual unit; performing a capacity reuse process on the second array to generate a storage redundancy array of the storage system; the capacity reuse process includes: sorting the second array; determining a third reference value from the sorted second array; dividing the third virtual unit and the fourth virtual unit from the virtual units in the second array based on the third reference value; forming a storage redundancy array of the storage system based on multiple third virtual units; updating the second array according to the third capacity value of the fourth virtual unit; and performing a capacity reuse process on the updated second array.

[0079] The second virtual unit is a virtual storage unit partitioned from the storage units within the target window based on a first reference value; the first capacity value is the total storage unit capacity, which is the sum of the capacity values ​​of the first and second virtual units; the second array is an array set composed of the second capacity values ​​of multiple second virtual units; the third reference value is a capacity baseline value determined from the sorting results of the second array; the third virtual unit is a virtual unit partitioned from the second virtual units based on the third reference value, with a capacity equal to the third reference value; the fourth virtual unit is a virtual unit partitioned from the second virtual units based on the third reference value, with a capacity equal to the third reference value; the third capacity value is...

[0080] The remaining capacity portion after dividing the third virtual unit from the second virtual unit.

[0081] Understandably, by dividing the storage units into second, third, and fourth virtual units and cyclically updating the array, the remaining capacity of the storage units can be accurately mined, avoiding the fragmentation and waste of capacity caused by traditional partitioning methods. Based on reference values, suitable virtual units are selected to form a redundant array, making the allocation of redundant resources more targeted and effectively resisting the risk of data loss. The cyclically executed capacity reuse process can adjust the redundant configuration in real time according to changes in the capacity of the virtual units, without the need to re-plan the overall storage architecture, so that storage resources can be fully utilized, adapting to changes in storage pressure and demand at different stages, and improving the fault tolerance and data reliability of the storage system.

[0082] Specifically, such as Figure 5 As shown in the diagram, disk1, disk2, disk3, disk4, disk5, and disk6 are physical storage units. A second virtual unit is partitioned from the storage units within the target window based on the first reference value, as shown by the yellow, green, and other colored areas in the diagram. The remaining portion is the first virtual unit. The sum of the capacities of the two virtual units constitutes the first capacity value of the storage unit.

[0083] like Figure 6As shown, the capacity reuse process sorts the second array, arranging the disks in ascending order of remaining capacity. For example, sort_0 to sort_6 correspond to disk1 to disk7 respectively. The colors from red to dark blue represent the increasing remaining capacity. The relationship between the remaining capacities is analyzed, and the minimum remaining capacity min_cap_round_1, i.e., the capacity of the red area of ​​disk1, is used as the benchmark to define the capacity threshold that can be used to build a redundant array. This ensures that the capacity allocated from each disk is evenly available. Based on the third reference value, the third virtual unit that makes up the redundant array and the fourth virtual unit that updates the array are divided. The third virtual unit is the part of disk1 to disk7 that has the same capacity as min_cap_round_1, that is, each disk allocates space with the same capacity as the red area. These allocated spaces are combined to form a new RAID redundant array. The fourth virtual unit is the remaining capacity after each disk is allocated, such as the remaining part of the yellow area of ​​disk2, the remaining part of the green area of ​​disk3, etc., which corresponds to the operation of updating the second array.

[0084] Figure 6 The RAID structure, which is composed of equal-capacity portions divided in the disk, is used to form a storage redundancy array based on multiple third virtual units. For example, RAID10 is composed of multiple sets of mirrored stripes. The remaining fourth virtual units of each disk are reassembled into a new second array, and the process of sorting, determining reference values, and updating is performed again.

[0085] Furthermore, in embodiments of this application, before the second array performs a capacity reuse process to generate a storage redundant array of the storage system, the process further includes: obtaining a third number of second virtual cells and a third total capacity value; if the third number is greater than a preset second number threshold and the third total capacity value is greater than a preset second capacity threshold, then the capacity reuse process is performed on the second array; if the second number is less than or equal to a preset second number threshold, or the third total capacity value is less than or equal to a preset second capacity threshold, then the capacity reuse process is not performed on the second array.

[0086] The third quantity is the total number of second virtual units participating in the capacity reuse process; the third total capacity value is the sum of the capacities of all second virtual units; and the second quantity threshold is the minimum number of virtual units that are allowed to execute the capacity reuse process.

[0087] Understandably, by obtaining the third quantity and the third total capacity value, and comparing the third quantity with the second quantity threshold, if the third quantity is greater than the preset second quantity threshold and the total capacity is greater than the preset second capacity threshold, the capacity reuse process is executed on the second array. By setting the threshold, it can be ensured that only the remaining capacity that meets the quantity and has sufficient total capacity will enter the reuse process, thereby avoiding the generation of invalid redundant arrays from the source, reducing system ineffective overhead, and improving resource utilization efficiency.

[0088] Specifically, as shown in Figure 7, the disks are sorted in ascending order of remaining capacity. In the capacity reuse process, the second array is sorted. The disks corresponding to sort_0 to sort_6, disk1 to disk7, are sorted in ascending order of remaining capacity. For example, disk1 has the least remaining capacity in the red area, and disk7 has the most remaining capacity in the dark blue area. Each cell in the figure represents 5G. The minimum remaining capacity of disk1, the 5G capacity threshold, and the minimum number of 4 disks are used to determine the third reference value from the sorted second array. The remaining capacity of disk1 is used as the implicit benchmark. At the same time, the capacity benchmark that can be used to build a redundant array is defined by the restrictions of remaining capacity ≥ 5G and at least 4 disks.

[0089] Figure 7 In this process, a portion matching the minimum capacity is allocated from the remaining capacity of each disk to construct a RAID10 array. Based on a third reference value, a third and fourth virtual unit are created. The third virtual unit is a portion from disks 1 to 7 that matches the remaining capacity of disk 1. In other words, each disk allocates space with the same capacity as the red area. These allocated spaces form a new RAID10 redundant array. The fourth virtual unit is the remaining capacity after each disk is allocated, such as the remaining portion of the yellow area of ​​disk 2 and the remaining portion of the green area of ​​disk 3. This remaining capacity is then incorporated back into the array for reuse in the next round.

[0090] Further, in embodiments of this application, generating a second array based on the second capacity value of the second virtual unit includes: obtaining a third capacity threshold of the second array; filtering out second virtual units whose second capacity value is less than the third capacity threshold; and generating a second array based on the second capacity values ​​of the remaining second virtual units.

[0091] The third capacity threshold is the capacity judgment criterion for selecting second virtual units that have the value of being reused.

[0092] Understandably, the construction of redundant storage arrays needs to meet minimum capacity coordinability. The third capacity threshold is essentially the minimum capacity threshold for redundant arrays. Only virtual units with remaining capacity greater than this threshold can be divided into capacity blocks that can be collaboratively built in subsequent processes. This ensures that the final generated redundant array can truly achieve the goals of data backup and performance optimization, ensuring the effectiveness and reliability of the redundant array and improving the accuracy of storage capacity utilization.

[0093] Specifically, suppose a storage system has disk0 to disk45 hard disks, and the remaining capacity of each storage unit, that is, the second capacity value of the second virtual unit, are as follows: disk0: 3GB, disk1: 6GB, disk2: 8GB, disk3: 4GB, disk4: 7GB.

[0094] A third capacity threshold of 5GB is pre-set, meaning that only virtual units with a remaining capacity of ≥5GB have enough space to participate in the subsequent construction of redundant arrays. The remaining capacity of each storage unit is checked one by one: disk0 (3GB) <5GB, so disk0 is filtered; disk3 (4GB) <5GB, so disk3 is filtered; disk1 (6GB) ≥5GB, so disk1 is retained; disk2 (8GB) ≥5GB, so disk2 is retained; disk4 (7GB) ≥5GB, so disk4 is retained. The remaining capacity of the retained disk1, disk2, and disk4 is combined into a new second array. Subsequent capacity reuse processes such as sorting, dividing virtual units, and constructing redundant arrays are executed based on the remaining capacity of these three storage units.

[0095] Furthermore, in the embodiments of this application, the storage redundancy array of the storage system is composed of multiple third virtual units, including: generating multiple third mirror groups based on the multiple third virtual units, wherein the third virtual units within the third mirror groups are mirror units of each other, and the mirror units synchronously store the same data; generating the storage redundancy array of the storage system based on the multiple third mirror groups, and recording the third metadata information of the storage redundancy array, wherein the multiple third mirror groups within the storage redundancy array allow parallel processing of data read and write tasks, and the third metadata information includes the third mapping relationship between the storage unit and the third virtual unit, and the starting address of the third virtual unit.

[0096] The third mirror group is a data redundancy group composed of multiple third virtual units; the third metadata information is a set of configuration information used in the storage system to manage the redundant array and virtual units; the third mapping relationship is an address lookup table of physical storage and logical virtual units in the third metadata information.

[0097] It is understandable that by using multiple third virtual units to generate multiple third mirror groups, the third virtual units within the third mirror groups can synchronously store the same data, and a storage redundancy array of the storage system can be generated. The third metadata information of the storage redundancy array can be recorded. The redundancy coverage can be expanded through the redundancy array composed of multiple third mirror groups, so that the mirror units can synchronously store the same data and process data read and write tasks in parallel. This makes resource allocation more flexible and improves storage utilization.

[0098] Specifically, such as Figure 8 As shown, a single storage space block of the same color is a third virtual unit, such as the red area of ​​disk1 and the blue area of ​​disk5 in the figure. Each independent color block is a third virtual unit. RAID1 composed of storage spaces of the same color is a third mirror group. If there is a red disk1 and a red area of ​​another disk, then these two red virtual units form a third mirror group. RAID0 composed of RAID1 of different colors is a storage redundancy array composed of multiple third mirror groups. As shown in the figure, RAID1 of different colors such as red RAID1 and blue RAID1 together form RAID0, forming a redundant array that processes data read and write in parallel with multiple mirror groups. The metadata `mem_frg_disk0:disk1`, `start_lba=min_cap`, and `capacity=min_cap_round_1` clearly define the correspondence between physical storage units and the third virtual unit, meaning that the segment space of the physical disk belongs to the virtual unit relationship. The starting address of the third virtual unit, `start_lba=min_cap`, records the starting position of the third virtual unit in the physical disk. The fourth virtual unit, `disk_1:last_lba=min_cap+min_cap_round_1` in the metadata, is the starting address of the fourth virtual unit after the third virtual unit is partitioned.

[0099] For example, the metadata records are as follows:

[0100] mem_frg_disk0:disk1,start_lba=min_cap,capacity=min_cap_round_1;

[0101] mem_frg_disk1:disk5,start_lba=min_cap,capacity=min_cap_round_1;

[0102] mem_frg_disk2:disk6,start_lba=min_cap,capacity=min_cap_round_1;

[0103] mem_frg_disk3:disk7,start_lba=min_cap,capacity=min_cap_round_1;

[0104] The starting address for updating the remaining capacity of the above four disks is:

[0105] disk_1:last_lba=min_cap+min_cap_round_1;

[0106] disk_5:last_lba=min_cap+min_cap_round_1;

[0107] disk_6:last_lba=min_cap+min_cap_round_1;

[0108] disk_7:last_lba=min_cap+min_cap_round_1.

[0109] Furthermore, in embodiments of this application, before performing the capacity reuse process on the updated second array, the method further includes: obtaining the second number of virtual cells and the second total capacity value in the updated second array; if the second number is greater than a preset first number threshold and the second total capacity value is greater than a preset first capacity threshold, then performing the capacity reuse process on the updated second array; if the second number is less than or equal to the preset first number threshold, or the second total capacity value is less than or equal to the preset first capacity threshold, then stopping the capacity reuse process on the updated second array.

[0110] Wherein, the first quantity threshold is a pre-set minimum number of virtual units that allow the second quantity to continue executing the capacity reuse process; the first capacity threshold is a pre-set minimum total capacity value that allows the second total capacity to continue executing the capacity reuse process.

[0111] It is understandable that by comparing the updated second quantity in the second array with the first quantity threshold, and comparing the second total capacity value with the first capacity threshold, the capacity reuse process can be performed on the updated second array or the capacity reuse process can be stopped. This can reserve sufficient remaining capacity to meet the basic requirements for building an effective redundant array, ensure the effectiveness of the redundant array, protect data security and performance, and improve capacity utilization.

[0112] Specifically, such as Figure 9As shown, the first quantity threshold is the minimum number of virtual units allowed to continue the capacity reuse process. As clearly shown, the loop stops when the number of available disks is less than 4. If the number of remaining disks is less than 4, reuse will not be performed even if there is remaining capacity. The first capacity threshold is the minimum total capacity value allowed to continue the process. As shown, the loop stops when the remaining capacity is less than 5GB. If the remaining capacity of a single disk is less than 5GB, even if the number is sufficient, an effective redundant array cannot be built due to capacity fragmentation. Disks with remaining space are reordered, and the steps are repeated until the remaining capacity of all disks is less than the threshold or the number of available disks is less than 4. Then, the capacity reuse process is performed on the updated second array. The change in disk sorting from the left to the right color blocks in the diagram is a visualization of the updated second array, stopping when the remaining capacity is less than 5GB or the number of available disks is less than 4.

[0113] By flexibly adapting to the parity of the number of storage units, it ensures that regardless of whether the number of units is even or odd, for an even number of storage units, the array is composed of the first virtual unit, which can efficiently realize the traditional redundancy architecture and build a stable redundancy relationship by taking advantage of the symmetry of the number; for an odd number of storage units, the virtual sub-units are divided by the second reference value, breaking the even number dependency restriction. The first virtual unit of a certain unit is split into sub-units that adapt to other units, forming a mirror or stripe relationship with the sub-units of the remaining units. This can effectively avoid some resources being idle due to the odd number, unify the division standard to build a storage redundancy array, maximize the utilization of storage resources and ensure data reliability.

[0114] This application's embodiments flexibly adapt to the parity of the number of storage units, ensuring that regardless of whether the number of units is even or odd, for an even number of storage units, an array composed of first virtual units can efficiently realize a traditional redundant architecture, leveraging the advantage of quantity symmetry to build a stable redundancy relationship; for an odd number of storage units, virtual sub-units are divided by a second reference value, breaking the even-number dependency restriction, splitting the first virtual unit of a certain unit into sub-units adapted to other units, forming a mirror or stripe relationship with the sub-units of the remaining units, which can effectively avoid some resources being idle due to an odd number, unifying the division standard to build a storage redundant array, maximizing storage resource utilization and ensuring data reliability.

[0115] To better understand the solution of this application, the storage system array generation method or execution flow of this application is described below through a specific embodiment, as follows:

[0116] First, select the most suitable disk group from the available disks.

[0117] The backend of a storage system typically connects to a storage unit expansion cabinet via protocols such as NVME, SAS, and RDMA. Each expansion cabinet contains dozens of storage units of varying capacities and models. Based on the customer's redundancy requirements, the storage system selects a different number of storage units to form different levels of RAID and then provides the user with available storage capacity.

[0118] This assumes the storage backend has n storage units, numbered from 0 to n-1, and m disks need to be selected to form a RAID 10 array, where 3 <= m <= n. The disk selection is performed as follows:

[0119] 1. Sort storage units by capacity: Sort all storage units in ascending order of capacity. The sorted array is sorted_disk=[s1,s2,...,s...]. n (s1≤s2≤...≤s) n The purpose of sorting is to make "storage units with similar capacities adjacent" to facilitate subsequent filtering. Here, we will not limit the total capacity requirement of the customer for the time being, but this can be used as a direction for future improvement.

[0120] 2. Use a sliding window of size m to traverse the sorted array. The core logic of the sliding window is: in the m storage units within the window, the leftmost element s[i] is the "minimum capacity" of the current window (because the array is already sorted). It is necessary to calculate the temporary "waste" of each window and select the window with the smallest temporary "waste".

[0121] Window range: from index i=0 to i=nm (a total of n-m+1 windows);

[0122] For each window [s[i], s[i+1], ..., s[i+m-1]], calculate two key metrics:

[0123] a. Minimum window size: min_cap = s[i];

[0124] b. Total window capacity: sum_cap = s[i] + s[i+1] + ... + s[i+m-1];

[0125] c. Temporary wasted amount of window space: waste = sum_cap - min_cap × m.

[0126] II. Handling of odd-numbered disks.

[0127] If m is even, skip step two in the following steps.

[0128] 1. Using the minimum window capacity `min_cap` as the standard, as shown in the diagram below, divide the selected disk capacity into two parts. The first part consists of disks with equal capacities, while the second part contains disks of varying capacities. Each disk in the first part is then virtualized as member disks `member_disk0` to `member_disk1`, and the mapping relationship between each disk and its member disks is recorded. Figure 3 For example, the metadata information is as follows:

[0129] member_disk0:disk1,start_lba=0,capacity=min_cap;

[0130] member_disk1:disk2,start_lba=0,capacity=min_cap;

[0131] member_disk2:disk3,start_lba=0,capacity=min_cap.

[0132] 2. If m is an odd number, then each member disk will be virtualized into two member disks, based on half the minimum window capacity (min_cap / 2), as follows: Figure 4 As shown, the corresponding mapping metadata needs to be adjusted as follows:

[0133] member_disk0:disk1,start_lba=0,capacity=min_cap / 2;

[0134] member_disk1:disk2,start_lba=0,capacity=min_cap / 2;

[0135] member_disk2:disk3,start_lba=0,capacity=min_cap / 2;

[0136] member_disk3:disk1,start_lba=min_cap / 2,capacity=min_cap / 2;

[0137] member_disk4:disk2,start_lba=min_cap / 2,capacity=min_cap / 2;

[0138] member_disk5:disk3,start_lba=min_cap / 2,capacity=min_cap / 2.

[0139] 3. For example Figure 6As shown, the storage units in this part can ultimately organize data according to the following mechanism: storage spaces of the same color form RAID1, and the same data is saved on two different disks to ensure redundancy backup; storage spaces of different colors form RAID0, which ensures RAID0.

[0140] 4. Record the starting address of the remaining capacity of each member disk in the form of metadata, as follows:

[0141] disk_1:last_lba=min_cap;

[0142] disk_2:last_lba=min_cap;

[0143] disk_3:last_lba=min_cap;

[0144] The above solutions address the issue that the number of disks in a RAID 10 array must be even, and reduce the minimum number of disks in a RAID 10 array from 4 to 3. Due to the serialized I / O processing characteristics of mechanical hard drives, this strategy works best for solid-state drives.

[0145] 3. Reuse of excess capacity.

[0146] 1. For example Figure 7 The diagram shows the remaining space of each disk. Here, it's assumed that the RAID 10 array consists of 8 disks, with disk0 being the smallest in terms of storage capacity. Based on the data organization strategies in 2.1.1 and 2.1.2, its excess capacity is 0. The disks are sorted in ascending order of remaining capacity, with two restrictions: first, disks with remaining capacity less than a certain threshold (e.g., 5GB, each cell in the diagram represents 5GB); second, when the number of disks eligible for sorting is less than 4, the reuse of excess capacity ends. This means that the minimum requirement of 4 disks for a standard RAID 10 array must be met.

[0147] 2. For example Figure 5 The diagram shows a RAID 10 array composed of disks with the smallest remaining capacity (min_cap_round_1) as the standard. The three disks with the smallest remaining capacity and the largest remaining capacity are selected, and each is allocated a storage space of min_cap_round_1 size to form a RAID 10 array. Figure 8 As shown, storage spaces of the same color form RAID 1, ensuring redundancy by storing the same data on two different disks. Storage spaces of different colors form RAID 0, ensuring redundancy. Metadata is recorded as follows:

[0148] mem_frg_disk0:disk1,start_lba=min_cap,capacity=min_cap_round_1;

[0149] mem_frg_disk1:disk5,start_lba=min_cap,capacity=min_cap_round_1;

[0150] mem_frg_disk2:disk6,start_lba=min_cap,capacity=min_cap_round_1;

[0151] mem_frg_disk3:disk7,start_lba=min_cap,capacity=min_cap_round_1;

[0152] The starting address for updating the remaining capacity of the above four disks is:

[0153] disk_1:last_lba=min_cap+min_cap_round_1;

[0154] disk_5:last_lba=min_cap+min_cap_round_1;

[0155] disk_6:last_lba=min_cap+min_cap_round_1;

[0156] disk_7:last_lba=min_cap+min_cap_round_1.

[0157] 3. For example Figure 9 As shown, disks with unused remaining space are reordered, and the above steps are repeated until all disks have less than the initial threshold, such as 5GB or fewer than 4 usable disks. This allows for maximizing the use of excess capacity when forming a RAID 10 array.

[0158] In summary, the storage system array generation method proposed in this application uses a sliding window to traverse the array, accurately determines the target window, and filters out suitable storage units. For an even number of storage units, it performs pairwise formation of redundant groups that are mirror images of each other, and then stripes them to form an array. For an odd number of storage units, it determines multiple virtual sub-units based on reference values, and forms a storage redundancy array of the storage system based on the multiple virtual sub-units. The virtual units are flexibly divided according to the parity of the storage units to ensure that the redundant array can be successfully built, realize data redundancy protection, improve storage stability, and enhance the flexibility and space utilization of storage units.

[0159] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0160] Embodiments of this application also provide a storage system array generation apparatus. Figure 10 This is a schematic diagram of the structure of the storage system array generation device provided in the embodiments of this application.

[0161] like Figure 10 As shown, the storage system array generation device 1000 includes: an acquisition module 1001, a determination module 1002, and a first partitioning module 1003.

[0162] The acquisition module 1001 is used to acquire a first capacity value of a storage unit and generate a first array based on the first capacity value; the determination module 1002 is used to traverse the first array based on a pre-set sliding window, and acquire at least one of the first number, first reference value, and first total capacity value of storage units in the current window during the traversal process, and determine a target window from multiple windows based on at least one of the first number, first reference value, and first total capacity value of storage units; the first partitioning module 1003 is used to partition a first virtual unit from the storage units in the target window based on the first reference value. If the number of storage units is even, a storage redundancy array of the storage system is formed based on multiple first virtual units. If the number of storage units is odd, a second reference value is determined based on the first reference value, and the first virtual unit is divided into multiple virtual sub-units based on the second reference value. A storage redundancy array of the storage system is formed based on the multiple virtual sub-units after partitioning.

[0163] Furthermore, in this embodiment of the application, the acquisition module 1001 is further configured to: acquire sorting rules, wherein the sorting rules include ascending order rules and descending order rules; if the sorting rule is an ascending order rule, then sort the first capacity value in ascending order; if the sorting rule is a descending order rule, then sort the first capacity value in descending order; and generate a first array based on the sorted first capacity value.

[0164] Furthermore, in this embodiment of the application, the determining module 1002 is further configured to: obtain the target size, sliding step size and sliding direction of the sliding window; determine the starting position of the sliding window in the first array according to the sliding direction; and move the sliding window of the target size according to the sliding step size and sliding direction starting from the starting position of the first array.

[0165] Furthermore, in this embodiment of the application, the determining module 1002 is further configured to: calculate the required capacity value based on the first number of storage units and the first reference value; calculate the remaining capacity value in the current window based on the first total capacity value and the required capacity value; and determine the target window from multiple windows based on the remaining capacity value.

[0166] Furthermore, in this embodiment, the first partitioning module 1003 is further configured to: generate multiple first mirror groups based on multiple first virtual units, wherein the first virtual units within the first mirror groups are mirror units of each other, and the mirror units synchronously store the same data; generate a storage redundancy array of the storage system based on the multiple first mirror groups, and record the first metadata information of the storage redundancy array, wherein the multiple first mirror groups within the storage redundancy array allow parallel processing of data read and write tasks, and the first metadata information includes a first mapping relationship between storage units and first virtual units, and the starting address of the first virtual units.

[0167] Furthermore, in this embodiment, the first partitioning module 1003 is further configured to: generate multiple second mirror groups based on multiple virtual sub-units, wherein the virtual sub-units within the second mirror groups are mirror units of each other, and the mirror units synchronously store the same data; generate a storage redundancy array of the storage system based on the multiple second mirror groups, and record the second metadata information of the storage redundancy array, wherein the multiple second mirror groups within the storage redundancy array allow parallel processing of data read and write tasks, and the second metadata information includes a second mapping relationship between storage units and virtual sub-units, and the starting address of the virtual sub-units.

[0168] Furthermore, in this embodiment of the application, the storage system array generation device 1000 further includes: a second partitioning module, configured to partition a first virtual unit from the storage units in the target window according to a first reference value, and then partition a second virtual unit from the storage units in the target window according to the first reference value, wherein the sum of the capacity values ​​of the second virtual unit and the first virtual unit is the first capacity value of the storage unit; generate a second array according to the second capacity value of the second virtual unit, and perform a capacity reuse process on the second array to generate a storage redundancy array of the storage system, wherein the capacity reuse process includes: sorting the second array, determining a third reference value from the sorted second array, partitioning a third virtual unit and a fourth virtual unit from the virtual units in the second array based on the third reference value, forming a storage redundancy array of the storage system according to multiple third virtual units, updating the second array according to the third capacity value of the fourth virtual unit, and performing a capacity reuse process on the updated second array.

[0169] Furthermore, in this embodiment, the first partitioning module 1003 is further configured to: generate multiple third mirror groups based on multiple third virtual units, wherein the third virtual units within the third mirror groups are mirror units of each other, and the mirror units synchronously store the same data; generate a storage redundancy array of the storage system based on the multiple third mirror groups, and record the third metadata information of the storage redundancy array, wherein the multiple third mirror groups within the storage redundancy array allow parallel processing of data read and write tasks, and the third metadata information includes the third mapping relationship between the storage unit and the third virtual unit, and the starting address of the third virtual unit.

[0170] Furthermore, in this embodiment of the application, the storage system array generation device 1000 further includes: a third partitioning module, configured to obtain a second number of virtual cells and a second total capacity value in the updated second array before performing a capacity reuse process on the updated second array; if the second number is greater than a preset first number threshold and the second total capacity value is greater than a preset first capacity threshold, then perform a capacity reuse process on the updated second array; if the second number is less than or equal to the preset first number threshold, or the second total capacity value is less than or equal to the preset first capacity threshold, then stop performing a capacity reuse process on the updated second array.

[0171] Furthermore, in this embodiment of the application, the storage system array generation apparatus 1000 further includes: an execution module, configured to perform a capacity reuse process on the second array to generate a storage redundant array of the storage system, and before doing so, obtain a third number and a third total capacity value of the second virtual cells; if the third number is greater than a preset second number threshold and the third total capacity value is greater than a preset second capacity threshold, then perform a capacity reuse process on the second array; if the second number is less than or equal to a preset second number threshold, or the third total capacity value is less than or equal to a preset second capacity threshold, then do not perform a capacity reuse process on the second array.

[0172] Furthermore, in this embodiment, the first partitioning module 1003 is further configured to: obtain a third capacity threshold of the second array; filter out second virtual units whose second capacity value is less than the third capacity threshold; and generate a second array based on the second capacity value of the remaining second virtual units.

[0173] In summary, the storage system array generation device proposed in this application uses a sliding window to traverse the array, accurately determines the target window, and filters out suitable storage units. For an even number of storage units, it performs pairwise formation of redundant groups that are mirror images of each other, and then stripes them to form an array. For an odd number of storage units, it determines multiple virtual sub-units based on reference values, and forms a storage redundancy array of the storage system based on the multiple virtual sub-units after division. The virtual units are flexibly divided according to the parity of the storage units to ensure that the redundant array can be successfully built, realize data redundancy protection, improve storage stability, and enhance the flexibility and space utilization of storage units.

[0174] For a description of the features in the embodiments corresponding to the above-mentioned storage system array generation apparatus, please refer to the relevant descriptions in the embodiments corresponding to the storage system array generation method, which will not be repeated here.

[0175] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in the above-described embodiments of the storage system array generation method.

[0176] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0177] The foregoing has provided a detailed description of a storage system array generation method and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for generating a storage system array, characterized in that, The storage system includes multiple storage units, wherein the method includes: Obtain the first capacity value of the storage unit, and generate a first array based on the first capacity value; The first array is traversed based on a pre-set sliding window. During the traversal, the first number of storage units, the first reference value, and the first total capacity value in the current window are obtained. Based on the first number of storage units, the first reference value, and the first total capacity value, the target window is determined from multiple windows. A first virtual unit is divided from the storage units within the target window based on the first reference value. If the number of storage units is even, a storage redundancy array of the storage system is formed based on multiple first virtual units. If the number of storage units is odd, a second reference value is determined based on the first reference value. The first virtual unit is divided into multiple virtual sub-units based on the second reference value. A storage redundancy array of the storage system is formed based on the multiple virtual sub-units after division. The first reference value is the minimum capacity value within the current window, and the second reference value is less than the first reference value.

2. The storage system array generation method according to claim 1, characterized in that, The step of generating the first array based on the first capacity value includes: Obtain the sorting rules, wherein the sorting rules include ascending order rules and descending order rules; If the sorting rule is an ascending order rule, then the first capacity value is sorted in ascending order; If the sorting rule is a descending order rule, then the first capacity value is sorted in descending order; The first array is generated based on the sorted first capacity values.

3. The storage system array generation method according to claim 1, characterized in that, The step of traversing the first array based on a pre-set sliding window includes: Obtain the target size, sliding step size, and sliding direction of the sliding window; The starting position of the sliding window in the first array is determined based on the sliding direction; Starting from the beginning position of the first array, the sliding window of the target size is moved according to the sliding step size and the sliding direction.

4. The storage system array generation method according to claim 1, characterized in that, The step of determining a target window from multiple windows based on at least one of the first number of storage units, the first reference value, and the first total capacity value includes: The required capacity value is calculated based on the first number of storage units and the first reference value; Calculate the remaining capacity value within the current window based on the first total capacity value and the required capacity value; The target window is determined from multiple windows based on the remaining capacity value.

5. The method for generating a storage system array according to claim 1, characterized in that, The storage redundancy array of the storage system, composed of a plurality of the first virtual units, includes: Multiple first mirror groups are generated based on multiple first virtual units, wherein the first virtual units within the first mirror group are mirror units of each other, and the mirror units synchronously store the same data. A storage redundancy array of the storage system is generated based on multiple first mirror groups, and the first metadata information of the storage redundancy array is recorded. The multiple first mirror groups in the storage redundancy array allow parallel processing of data read and write tasks. The first metadata information includes a first mapping relationship between the storage unit and the first virtual unit, and the starting address of the first virtual unit.

6. The method for generating a storage system array according to claim 1, characterized in that, The step of assembling the storage redundancy array of the storage system based on the divided virtual sub-units includes: Multiple second mirror groups are generated based on the multiple virtual sub-units, wherein the virtual sub-units in the second mirror group are mirror units of each other, and the mirror units synchronously store the same data. The storage system generates a storage redundancy array based on multiple second mirror groups, and records the second metadata information of the storage redundancy array. The multiple second mirror groups within the storage redundancy array allow parallel processing of data read and write tasks. The second metadata information includes a second mapping relationship between the storage unit and the virtual sub-unit, and the starting address of the virtual sub-unit.

7. The method for generating a storage system array according to claim 1, characterized in that, After dividing the first virtual unit from the storage unit within the target window according to the first reference value, the method further includes: A second virtual unit is divided from the storage units within the target window based on the first reference value, and the sum of the capacity values ​​of the second virtual unit and the first virtual unit is the first capacity value of the storage unit. A second array is generated based on the second capacity value of the second virtual unit, and a capacity reuse process is performed on the second array to generate a storage redundancy array of the storage system. The capacity reuse process includes: sorting the second array, determining a third reference value from the sorted second array, dividing a third virtual unit and a fourth virtual unit from the virtual units in the second array based on the third reference value, forming a storage redundancy array of the storage system based on multiple third virtual units, updating the second array based on the third capacity value of the fourth virtual unit, and performing a capacity reuse process on the updated second array.

8. The method for generating a storage system array according to claim 7, characterized in that, The storage redundancy array of the storage system composed of multiple third virtual units includes: Multiple third mirror groups are generated based on multiple third virtual units, wherein the third virtual units within the third mirror group are mirror units of each other, and the mirror units synchronously store the same data. The storage redundant array of the storage system is generated based on multiple third mirror groups, and the third metadata information of the storage redundant array is recorded. The multiple third mirror groups in the storage redundant array allow parallel processing of data read and write tasks. The third metadata information includes the third mapping relationship between the storage unit and the third virtual unit, and the starting address of the third virtual unit.

9. The method for generating a storage system array according to claim 7, characterized in that, Before performing the capacity reuse process on the updated second array, the following steps are also included: Get the second number of virtual cells and the second total capacity value in the updated second array; If the second quantity is greater than the preset first quantity threshold and the second total capacity value is greater than the preset first capacity threshold, then the capacity reuse process is executed on the updated second array. If the second quantity is less than or equal to a preset first quantity threshold, or if the second total capacity value is less than or equal to a preset first capacity threshold, then the capacity reuse process for the updated second array is stopped.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the storage system array generation method as described in any one of claims 1 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Method and apparatus for creating redundancy array in disc RAID

    CN101620518A

  • Storage system condition indicator and method

    CN101872319A