A distributed storage pool high water level reconstruction strategy optimization method, device and equipment
By simulating data distribution after a failure and generating a water level reconstruction strategy, resource allocation was optimized, solving the problem of data loss when a hard drive fails under high water level conditions in the storage pool, and improving system stability and data reliability.
Patent Information
- Application Number
- CN202412000325.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-12-31
AI Technical Summary
Existing technologies lack precise reconstruction control in high-water level scenarios of storage pools when hard drives fail, making it unable to effectively handle multiple hard drive failures, resulting in data loss and unwritable business processes.
By acquiring the status information of object storage devices and storage pool information, the system simulates the data distribution after a failure, generates a water level reconstruction strategy, optimizes resource allocation and reconstruction process, and ensures the accuracy and efficiency of data migration.
It improves the reliability of data storage and the stability of the system, reduces the risk of data loss, optimizes resource utilization, shortens fault response time, and maintains high system availability.
Smart Images

Figure CN119806919B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of storage, in particular to a distributed storage pool high water level reconstruction strategy optimization method, device and equipment. BACKGROUND
[0002] In the storage pool high water level scenario, when a hard disk fails, the data on the hard disk is reconstructed to other hard disks of the node where the hard disk is located, accompanied by the reconstructed data and business, and the other hard disks of the node are full, resulting in read-only business. To solve this problem, a known method is to set a water level threshold when a fault occurs, such as 85%, and when the storage pool reaches this water level, the hard disk that fails is not allowed to be kicked out of the storage pool to trigger reconstruction. However, this scheme only looks at the storage pool water level, and the control is not accurate enough, and it cannot face the multi-hard disk failure scenario, such as the current storage pool water level being 84%, at this time multiple hard disks fail, and reconstruction is allowed to be triggered, but in fact it cannot be reconstructed completely, and the water level of other hard disks will rise to the threshold, making the business read-only and not writable, and unable to guarantee that the failed data can be reconstructed completely, resulting in data loss. SUMMARY
[0003] Therefore, the present application provides a distributed storage pool high water level reconstruction strategy optimization method, device and equipment to solve the problem of inaccurate data reconstruction control and inability to face multi-hard disk failure.
[0004] In a first aspect, the present application provides a distributed storage pool high water level reconstruction strategy optimization method, which comprises:
[0005] Obtaining historical distribution information of an object storage device data set and a placement group in a storage pool, the object storage device data set comprising state information corresponding to a plurality of object storage devices respectively, and storage pool information of the plurality of object storage devices, the plurality of object storage devices comprising an object storage device that has failed and an object storage device in a normal state;
[0006] Based on the state information of each object storage device and the storage pool information, simulating to determine the current distribution information of the placement group in the storage pool if the object storage device that has failed is kicked out, and determining whether to reconstruct the placement group corresponding to the object storage device that has failed;
[0007] If reconstruction is performed, determining the current water level information of each object storage device in a normal state according to the reconstruction preset bandwidth, the current distribution information, and the storage pool information;
[0008] Generating a water level reconstruction strategy according to the current water level information, wherein the water level reconstruction strategy is used to indicate whether to kick out the object storage device that has failed.
[0009] The application provides a distributed storage pool high water level reconstruction strategy optimization method, which has the following advantages.
[0010] In the method, the placement group is reconstructed in time when the object storage device fails, so that the data is ensured not to be lost due to the failure, and the reliability of data storage is improved. Therefore, by automatically identifying the failed device and simulating the reconstruction process, the uncertainty and potential cost in the actual reconstruction process can be reduced. The risk caused by improper data migration in the reconstruction process is reduced. The need for reconstruction and how to reconstruct are quickly determined, which can shorten the time for the system to respond to the failure and improve the overall performance. The stability and fault tolerance of the system are enhanced by effectively managing and replacing the failed device. Furthermore, by reasonably allocating and adjusting the storage resources, the utilization rate of the storage space in the storage pool can be maximized. According to the current water level information, the system can dynamically allocate resources to ensure the balance of the storage pool. Finally, according to the water level reconstruction strategy, it is determined whether the object storage device that has failed is kicked out, and the high availability of the system is maintained by timely reconstruction and optimization of resource allocation.
[0011] In an optional embodiment, the storage pool information includes water level information of the storage pool, and the number of placement groups in the storage pool.
[0012] Based on the state information of each object storage device and the storage pool information, the current distribution information of the placement groups in the storage pool after the object storage device that has failed is kicked out is simulated and determined, and it is determined whether the placement group corresponding to the object storage device that has failed is reconstructed, specifically including:
[0013] The state information of each object storage device, the water level information of the storage pool, and the number of placement groups in the storage pool are input into a data distribution algorithm to simulate the current distribution information of the placement groups in the storage pool after the object storage device that has failed is kicked out.
[0014] The current distribution information is compared with the historical distribution information.
[0015] When it is determined that the current distribution information and the historical distribution information are inconsistent, it is determined that the placement group corresponding to the object storage device that has failed is reconstructed.
[0016] Specifically, the water level information of the storage pool (i.e., the usage of the storage resource) is used to evaluate the necessity of reconstruction and optimize the allocation of the storage resource. The number of placement groups in the storage pool helps to more accurately simulate and optimize data distribution. The state information of the object storage device is used to distinguish the object storage device that has failed from the object storage device that has not failed. By inputting these information and the number of placement groups in the storage pool into the data distribution algorithm, the current distribution information of the placement group in the storage pool after the object storage device that has failed is kicked out can be simulated. By comparing the historical distribution information with the current distribution information, it can be ensured that the data remains consistent after the failure occurs, thereby improving the reliability of the data. By simulating the distribution after the failed device is kicked out, the possible data distribution problem can be predicted in advance, and the reaction time when the actual failure occurs can be reduced. The simulation result helps the decision maker to determine whether the placement group needs to be reconstructed, thereby avoiding unnecessary reconstruction operation. Based on the historical and current data, the system can intelligently adjust the storage pool configuration, improve the resource utilization efficiency, determine the best reconstruction time, reduce the data interruption and system downtime, and generate the water level reconstruction strategy according to the current water level information, which helps to accurately control data migration and avoid resource waste. The automated process and intelligent decision-making reduce the operation and maintenance workload and improve the operation and maintenance efficiency. By optimizing the storage resource utilization and reducing unnecessary reconstruction, the operation and maintenance cost is controlled.
[0017] In an optional embodiment, the storage pool information further includes the redundancy type of the storage pool and the data storage amount of each placement group corresponding to each object storage device in the storage pool;
[0018] If reconstruction is performed, the current water level information of each object storage device in a normal state is determined according to the reconstruction preset bandwidth, the current distribution information, and the storage pool information, specifically including:
[0019] The total capacity to be reconstructed is determined according to the redundancy type of the storage pool, the current distribution information, and the data storage amount of each placement group corresponding to the object storage device that has failed in the storage pool;
[0020] The current water level information of each object storage device in a normal state is determined according to the reconstruction preset bandwidth, the total capacity to be reconstructed, the current distribution information, the number of placement groups in the storage pool, and the data storage amount of each placement group corresponding to the object storage device in a normal state in the storage pool.
[0021] Specifically, different redundancy types require different methods to determine the total capacity to be reconstructed. The current information of placement groups can identify the number of placement groups stored in failed object storage devices and the number of placement groups stored in non-failed object storage devices, thereby determining the data storage volume of the placement groups that need to be migrated, and ultimately the total capacity to be reconstructed. Then, based on the data storage volume of the placement groups corresponding to each object storage device, storage resources are allocated and optimized more accurately, avoiding data overload or resource waste. Utilizing the pre-set reconstruction bandwidth and current distribution information, the system can quickly calculate the bandwidth and resources required for reconstruction, thereby improving reconstruction efficiency. By determining the current water level information of each normal object storage device, the system can better balance the load, avoiding performance degradation or failures due to overload. Accurate water level information and redundancy strategies help accelerate data recovery and reduce the impact of failures on business operations. Based on the actual data storage volume of each device, the system can allocate resources more rationally, improving the overall performance of the storage pool. By combining pre-set bandwidth and actual data storage volume, the system can avoid bandwidth overload and ensure a smooth data reconstruction process. During reconstruction, accurate water level information helps to quickly locate problematic devices, improving the efficiency of fault handling. Therefore, it is necessary to determine the current status of each object storage device in a normal state. By optimizing the reconstruction strategy and resource allocation, system stability is enhanced, and system fluctuations caused by improper reconstruction are reduced.
[0022] In one optional implementation, the total capacity to be reconstructed is determined based on the redundancy type of the storage pool, current distribution information, and the data storage volume of the placement group corresponding to the failed object storage devices in the storage pool, including:
[0023] Extract the first mapping relationship from the current distribution information. The first mapping relationship is the first mapping relationship re-established between the placement group to which the data stored on the failed object storage device belongs and the object storage device in normal condition.
[0024] Based on the redundancy type of the storage pool and the data storage volume of the placement group corresponding to the failed object storage device, determine the raw capacity of the placement group corresponding to the failed object storage device on the normal object storage device that has established a first mapping relationship with it.
[0025] The total capacity to be reconstructed is determined based on the raw capacity of the placement group corresponding to the failed object storage device on the normal object storage device that has established a first mapping relationship with itself.
[0026] Specifically, based on the first mapping relationship, the system determines which normal object storage device can be assigned to the placement group to which the data stored on the failed object storage device belongs; that is, it extracts the first mapping relationship. Then, based on the redundancy type of the storage pool and the data storage volume of the placement group corresponding to the failed object storage device, the system determines the raw capacity of the placement group corresponding to the failed object storage device on the normal object storage device with which it has established the first mapping relationship. Finally, based on the raw capacity of the placement group corresponding to the failed object storage device on the normal object storage device with which it has established the first mapping relationship, the system determines the total capacity to be reconstructed. The total reconstructed capacity is the sum of all raw capacities. In the above method, by extracting the first mapping relationship—that is, the mapping between the placement group to which the data stored on the failed object storage device belongs and the normal object storage device—the system can accurately calculate the total capacity to be reconstructed, avoiding insufficient or excessive capacity. Based on the redundancy type of the storage pool and the data storage volume of the failed placement group, the system can optimize the use of redundant resources, ensuring data security after reconstruction. By determining the raw capacity—that is, the space on the normal object storage device that can be used to store failed data—the system can utilize existing resources more effectively. Determining the total capacity to be refactored allows for more efficient planning of the refactoring process, reducing the time and resources required and accelerating data recovery. Accurate calculation of the total capacity to be refactored improves the operational efficiency of the storage pool and the reliability of the system, while also reducing costs and enhancing the user experience.
[0027] In one optional implementation, based on the preset reconstruction bandwidth, the total capacity to be reconstructed, the current distribution information, the number of placement groups in the storage pool, and the data storage volume of the placement group corresponding to the object storage device in normal state in the storage pool, the current water level information of each object storage device in normal state is determined, including:
[0028] Based on the pre-set bandwidth and total capacity to be reconstructed, predict the incremental business data during the reconstruction period;
[0029] The average data increment for each placement group is obtained based on the number of placement groups in the storage pool and the business data increment during the reconstruction period.
[0030] Based on the current distribution information of placement groups in the storage pool, the redundancy type of the storage pool, the number of placement groups in the storage pool, the average data increment of each placement group, and the data storage volume of the placement group corresponding to the object storage device in normal state in the storage pool, the current water level information of each object storage device in normal state is determined.
[0031] Specifically, by predicting the incremental business data during the reconstruction period, bandwidth resources can be planned and adjusted in advance, improving reconstruction efficiency and reducing reconstruction time. Based on the average data increment of the placement group, storage resources can be rationally allocated to avoid overloading certain storage devices while ensuring that other devices are not idle. By monitoring the current water level information of each storage device, potential storage bottlenecks can be identified and addressed in a timely manner, enhancing system stability. Accurate water level information helps to quickly locate and recover data on faulty devices, accelerating fault recovery. This reduces system downtime caused by reconstruction or fault recovery, improving user experience.
[0032] In one optional implementation, the predicted increase in service data during the reconstruction period, based on the preset reconstruction bandwidth and the total capacity to be reconstructed, includes:
[0033] Based on the total capacity to be reconstructed and the preset bandwidth for reconstruction, predict the data reconstruction time;
[0034] Based on the incremental business data during the failure period of the failed object storage device and the time of the failure, obtain the business data growth rate;
[0035] By leveraging the data restructuring time and the growth rate of business data, predict the incremental business data during the restructuring period.
[0036] In an alternative implementation, when the water level reconstruction strategy indicates that a failed object storage device cannot be removed, the method further includes:
[0037] Save the water level information for each object storage device in the storage pool;
[0038] Periodically monitor changes in the water level information of each object storage device;
[0039] When the water level information of the first object storage device drops below a preset water level threshold based on the changes, it is re-determined whether the object storage device that has failed meets the conditions for being kicked out. Here, the first object storage device is any object storage device in the storage pool.
[0040] In an alternative implementation, when the water level reconstruction strategy indicates that a failed object storage device is being removed, the method further includes:
[0041] After performing the operation of kicking out the failed object storage device, determine whether to trigger a refactoring operation;
[0042] Identify the danger level corresponding to the current water level information of each object storage device that is in a normal state;
[0043] Once a refactoring operation is triggered, obtain the actual business increments generated during the refactoring period;
[0044] Based on the current water level information of each object storage device in normal condition, the corresponding danger level, data reconstruction time, actual business increment, and predicted business data increment during the reconstruction period, the reconstruction control strategy is determined.
[0045] In a second aspect, the present invention provides a distributed storage pool high-water level reconstruction strategy optimization device, the device comprising:
[0046] The data acquisition module is used to acquire the object storage device dataset and the historical distribution information of the placement groups in the storage pool. The object storage device dataset includes status information corresponding to multiple object storage devices, as well as storage pool information to which the multiple object storage devices belong. The multiple object storage devices include object storage devices that have failed and object storage devices that are in normal condition.
[0047] The processing module is used to simulate and determine the current distribution information of the placement group in the storage pool after the failed object storage device is kicked out, based on the status information of each object storage device and the storage pool information, and to determine whether to reconstruct the placement group corresponding to the failed object storage device.
[0048] The water level acquisition module is used to determine the current water level information of each object storage device in normal state based on the preset reconstruction bandwidth, the current distribution information, and the storage pool information if reconstruction is performed.
[0049] The processing module is further configured to generate a water level reconstruction strategy based on the current water level information, wherein the water level reconstruction strategy is used to indicate whether to kick out the object storage device that has failed.
[0050] The high-water level reconstruction strategy optimization device for distributed storage pools provided by this invention has the following advantages:
[0051] In the methods described above, timely reconstruction of the placement group in the event of object storage device failure ensures that data is not lost due to the failure, thus improving data storage reliability. Therefore, by automatically identifying faulty devices and simulating the reconstruction process, and through pre-simulation and strategy generation, uncertainties and potential costs in the actual reconstruction process can be reduced. Risks caused by improper data migration during reconstruction are also reduced. Quickly determining whether and how to reconstruct shortens system response time and improves overall performance. Effective management and replacement of faulty devices enhance system stability and fault tolerance. Furthermore, reasonable allocation and adjustment of storage resources maximizes the utilization of storage space in the storage pool. Based on current water level information, the system can dynamically allocate resources to ensure the balance of the storage pool. Finally, based on the water level reconstruction strategy, it is determined whether to remove faulty object storage devices, and high system availability is maintained through timely reconstruction and optimized resource allocation.
[0052] Thirdly, the present invention provides a computer device, including: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the distributed storage pool high water level reconstruction strategy optimization method of the first aspect or any corresponding embodiment described above.
[0053] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the distributed storage pool high-water level reconstruction strategy optimization method of the first aspect or any corresponding embodiment described above.
[0054] Fifthly, the present invention provides a computer program product, including computer instructions, which are used to cause a computer to execute the distributed storage pool high-water level reconstruction strategy optimization method of the first aspect or any corresponding embodiment described above. Attached Figure Description
[0055] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of the present invention, the drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0056] Figure 1 This is a flowchart illustrating a method for optimizing a high-water level reconstruction strategy for a distributed storage pool, as provided in an embodiment of the present invention.
[0057] Figure 2This is a flowchart illustrating another method for optimizing the high-water level reconstruction strategy of a distributed storage pool provided in an embodiment of the present invention.
[0058] Figure 3 This is a flowchart illustrating another method for optimizing a high-water level reconstruction strategy for a distributed storage pool, as provided in an embodiment of the present invention.
[0059] Figure 4 This is a flowchart illustrating another method for optimizing a high-water level reconstruction strategy for a distributed storage pool, as provided in an embodiment of the present invention.
[0060] Figure 5 This is a schematic diagram of the overall process of a distributed storage pool high-water level reconstruction strategy optimization method provided in an embodiment of the present invention;
[0061] Figure 6 This is a structural block diagram of a distributed storage pool high-water level reconstruction strategy optimization device provided in an embodiment of the present invention;
[0062] Figure 7 This is a schematic diagram of the hardware structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0064] Before introducing the embodiments of this application, the following content will be introduced first:
[0065] Distributed storage typically consists of several server nodes (each with several disks). Distributed storage software assembles these server nodes into a distributed storage cluster, managing the nodes and providing storage services. The software manages the storage resources (disks) on the servers by creating storage pools. A storage pool can be understood as a logical container for storing data objects and related metadata; it can be viewed as a group of relatively independent data storage spaces. A storage pool contains multiple server nodes, and their hard drives are managed by an object storage device service, typically one hard drive per object storage device service. In addition, the storage pool contains several placement groups. A placement group is a logical concept, mapped to object storage devices through a data distribution algorithm, resulting in two storage pool-level data redundancy methods: replication and erasure. Therefore, a placement group exists across multiple object storage devices. When business data needs to be written to the storage pool, the client will cut the data into fixed sizes. This fixed-size data is called an object. Objects and placement groups are associated with each other through a consistent hash algorithm. The corresponding placement group can be calculated through the object name. Placement groups exist on multiple object storage devices. In this way, the object is written to the disk through the object storage device service.
[0066] A Placement Group (PG) is a logical concept used to distribute data objects across different Object Storage Daesel (OSD) devices in a storage system. Through a data distribution algorithm, PGs are mapped to OSDs to achieve balanced data distribution and redundant storage. Specifically, the data distribution algorithm maps data objects to different PGs according to certain rules. Each PG has a unique identifier and is associated with a set of OSDs. When a data object needs to be stored, it is assigned to a specific PG, and then the data is stored on the corresponding OSD according to the mapping between the PG and OSD. This approach achieves balanced data distribution, avoids overloading certain OSDs, and improves data reliability and fault tolerance. By storing data on multiple OSDs, redundant data backup can be achieved; if one OSD fails, data can be recovered from other OSDs.
[0067] In practical applications, after prolonged use, the water level in the distributed storage pool will gradually increase, and the hard drive failure rate will also rise. When a hard drive fails, during the failover rebalancing process, the failed drive will be removed from the storage pool, and its data will be reconstructed and distributed to other hard drives, ensuring data security and sufficient redundancy. Generally, the storage pool sets certain thresholds; for example, when the water level or hard drive water level reaches 95%, the storage pool is marked as full, and at this point, read operations are allowed only, not write operations.
[0068] However, in the above method, there is a scenario where a hard drive fails in a high-water level storage pool. After the hard drive is kicked out, the data on it is reconstructed to other hard drives on the same node. Along with the reconstructed data and business, the other hard drives on that node are filled up, resulting in read-only and unwritable business operations.
[0069] In related technologies, to avoid the above situation, a fault threshold is set, such as 85%. When the storage pool reaches this level, the faulty hard drive is not allowed to be removed from the storage pool to trigger reconstruction. However, this solution only considers the storage pool level, which is not precise enough and cannot handle scenarios with multiple hard drive failures. For example, if the current storage pool level is 84%, and multiple hard drives fail, reconstruction is allowed, but it cannot be completed in practice, and the level of other hard drives will rise to the threshold, making the service read-only and not writable.
[0070] This invention provides an optimization method for high-water mark reconstruction strategy of distributed storage pool. By accurately predicting the reconstructed data, incomplete data reconstruction is avoided, thereby improving the security of data reconstruction.
[0071] According to an embodiment of the present invention, an embodiment of a method for optimizing a high-water level reconfiguration strategy for a distributed storage pool is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0072] This embodiment provides an optimization method for high-water mark reconstruction strategy of distributed storage pool, which can be used in electronic devices such as clients and servers. Figure 1 This is a flowchart of a distributed storage pool high-water level reconstruction strategy optimization method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps:
[0073] Step S101: Obtain the historical distribution information of the object storage device dataset and the placement groups in the storage pool.
[0074] Specifically, the object storage device dataset includes status information corresponding to multiple object storage devices, as well as storage pool information to which the multiple object storage devices belong. The multiple object storage devices include object storage devices that have failed and object storage devices that are in normal condition.
[0075] Step S102: Based on the status information of each object storage device and the storage pool information, simulate and determine the current distribution information of the placement group in the storage pool after the failed object storage device is kicked out, and determine whether to reconstruct the placement group corresponding to the failed object storage device.
[0076] Specifically, based on the data distribution algorithm, the placement group distribution information after the failed object storage device is kicked out is simulated by using the status information of the failed object storage device, the status information of the object storage device in normal state, and the storage pool information.
[0077] Then, it is confirmed whether the placement group distribution information after the faulty object storage device is kicked out is consistent with the historical distribution information. If the data is inconsistent, the placement group corresponding to the faulty object storage device is reconstructed.
[0078] Step S103: If reconstruction is to be performed, determine the current water level information of each object storage device in normal state based on the reconstruction preset bandwidth, current distribution information, and storage pool information.
[0079] Step S104: Generate a water level reconstruction strategy based on the current water level information.
[0080] Among them, the water level reconstruction strategy is used to indicate whether to kick out object storage devices that have failed.
[0081] Specifically, once the current water level of an OSD reaches a certain height, read operations are allowed, but write operations are prohibited. Therefore, if simulations determine that reconstruction is necessary, the current water level of each normally functioning object storage device needs to be determined based on the preset reconstruction bandwidth, current distribution information, and storage pool information. The current water level of each normally functioning OSD has a significant impact on the storage pool's water level. This is because it's necessary to consider whether the sum of the water levels of the remaining normally functioning OSDs will reach the upper limit of the storage pool's water level if a failed OSD is removed, resulting in read-only access and no write access. Moreover, there is also a risk when the water level of a single OSD reaches the preset upper limit. Therefore, a reconstruction strategy needs to be generated based on the current water level of each normally functioning OSD to indicate whether to remove the failed object storage device.
[0082] The preset bandwidth for reconstruction represents the current data reconstruction speed. Based on this speed, the amount of data storage to be added during reconstruction can be determined. Based on the current distribution and storage pool information, the water level of the object storage device in a normal state before the increase in data storage can be determined. Then, based on the increased data storage and the water level of the object storage device in a normal state before the increase, the current water level of the object storage device in a normal state after reconstruction can be determined.
[0083] If the current water level information of an object storage device in normal condition does not reach the preset threshold standard, the faulty object storage device is not allowed to be kicked out. If the current water level information of an object storage device in normal condition does not exceed the preset threshold, the faulty object storage device is kicked out of operation.
[0084] The distributed storage pool high-water mark reconfiguration strategy optimization method provided in this embodiment ensures data integrity and improves data storage reliability by promptly reconfiguring placement groups when object storage devices fail. Therefore, by automatically identifying faulty devices and simulating the reconfiguration process, pre-simulation and strategy generation reduce uncertainties and potential costs in the actual reconfiguration process. It also reduces risks caused by improper data migration during reconfiguration. Quickly determining whether and how to reconfigure shortens system response time and improves overall performance. Effective management and replacement of faulty devices enhance system stability and fault tolerance. Furthermore, reasonable allocation and adjustment of storage resources maximizes storage space utilization in the storage pool. Based on current water mark information, the system dynamically allocates resources to ensure storage pool balance. Finally, based on the water mark reconfiguration strategy, it determines whether to remove faulty object storage devices and maintains high system availability through timely reconfiguration and optimized resource allocation.
[0085] This embodiment provides an optimization method for high-water level reconstruction strategy of distributed storage pool, which can be used in the aforementioned mobile terminals, such as mobile phones and tablets. Figure 2 This is a flowchart illustrating another method for optimizing the high-water level reconstruction strategy of a distributed storage pool provided in an embodiment of the present invention, as shown below. Figure 2 As shown, based on the aforementioned embodiments, the storage pool information includes the water level information of the storage pool and the number of groups placed in the storage pool;
[0086] Based on the status information of each object storage device and the storage pool information, the simulation determines the current distribution information of the placement group in the storage pool after the failed object storage device is kicked out, and determines whether to reconstruct the placement group corresponding to the failed object storage device. The specific steps include the following:
[0087] Step S201: Input the status information of each object storage device, the water level information of the storage pool, and the number of placement groups in the storage pool into the data distribution algorithm to simulate the current distribution information of placement groups in the storage pool after the failed object storage device is kicked out.
[0088] The specific data distribution algorithm can directly simulate the corresponding results, that is, simulate the current distribution information of the storage pool group after the failed object storage device is kicked out.
[0089] Step S202: Compare the current distribution information with the historical distribution information.
[0090] Specifically, the distribution information of storage devices that have not been kicked out of the faulty object is compared with the distribution information of storage devices that have been kicked out of the faulty object.
[0091] Step S203: When it is determined that the current distribution information and the historical distribution information are inconsistent, it is determined to reconstruct the placement group corresponding to the object storage device that has failed.
[0092] If the current distribution information is inconsistent with the historical distribution information, a capacity balancing operation is triggered to reconstruct the placement group corresponding to the object storage device that has failed.
[0093] This invention provides a method for optimizing a high-water level reconfiguration strategy for a distributed storage pool. The storage pool's water level information (i.e., the usage of storage resources) is used to assess the necessity of reconfiguration and optimize storage resource allocation. The number of placement groups in the storage pool helps to more accurately simulate and optimize data distribution. The status information of object storage devices is used to distinguish between failed and intact object storage devices. This information, along with the number of placement groups in the storage pool, is input into a data separation algorithm. This algorithm can simulate the current separation information of placement groups in the storage pool after a failed object storage device is removed. By comparing historical and current distribution information, it ensures that data remains consistent after a failure, thereby improving data reliability. By simulating the distribution after a failed device is removed, potential data distribution problems can be predicted in advance, reducing response time when actual failures occur. The simulation results help decision-makers determine whether placement groups need to be reconfigured, thus avoiding unnecessary reconfiguration operations. Based on historical and current data, the system can intelligently adjust the storage pool configuration, improving resource utilization efficiency. It also determines the optimal reconfiguration timing, reducing data interruption and system downtime. Water level reconstruction strategies generated based on current water level information help to precisely control data migration and avoid resource waste. Automated processes and intelligent decision-making reduce operational workload and improve operational efficiency. By optimizing storage resource utilization and reducing unnecessary reconstructions, operational costs are controlled.
[0094] This embodiment provides an optimization method for high-water level reconstruction strategy of distributed storage pool, which can be used in the aforementioned mobile terminals, such as mobile phones and tablets. Figure 3 This is a flowchart illustrating another method for optimizing a high-water level reconstruction strategy in a distributed storage pool, as provided in this embodiment of the invention. Figure 3 As shown.
[0095] Specifically, the storage pool information also includes the redundancy type of the storage pool, as well as the data storage volume of the placement group corresponding to each object storage device in the storage pool.
[0096] Based on any of the foregoing embodiments, if reconstruction is performed, the current water level information of each object storage device in a normal state is determined according to the reconstruction preset bandwidth, current distribution information, and storage pool information, specifically including:
[0097] Step 301: Determine the total capacity to be reconstructed based on the redundancy type of the storage pool, the current distribution information, and the data storage volume of the placement group corresponding to the failed object storage device in the storage pool.
[0098] In an optional example, the total capacity to be refactored can be determined by the following steps:
[0099] Step a1: Extract the first mapping relationship from the current distribution information.
[0100] Specifically, the first mapping relationship is a newly established mapping relationship between the placement group to which the data stored on the failed object storage device belongs and the object storage device in normal condition.
[0101] Step a2: Based on the redundancy type of the storage pool and the data storage volume of the placement group corresponding to the failed object storage device, determine the raw capacity of the placement group corresponding to the failed object storage device on the normal object storage device that has established a first mapping relationship with it.
[0102] Specifically, the redundancy type of the storage pool is obtained, and the method for obtaining the raw capacity on the object storage device is determined based on the redundancy type. If the redundancy type of the storage pool is the replica type, the raw capacity of the placement group on the object storage device is the data storage amount of the placement group corresponding to the object storage device that has failed.
[0103] If the redundancy type of the storage pool is erasure K+M type, the raw capacity of the placement group on the object storage device is the data storage amount of the placement group corresponding to the failed object storage device divided by the total number of fragments into which the original data is divided and the number of verification fragments.
[0104] In a specific example, the raw capacity of the placement group on the normal state object storage device with which it establishes the first mapping relationship has two conversion methods based on the redundancy type of the storage pool. The first is that when the redundancy type of the storage pool is replica, the raw capacity of the placement group on the object storage device is the current data volume of the placement group. The second is that when the redundancy type of the storage pool is erasure K+M, the raw capacity of the placement group on the object storage device = the data volume of the placement group / (K+M), where K is the number of fragments into which the original data is divided, and M is the number of verification fragments.
[0105] Step a3: Determine the total capacity to be reconstructed based on the raw capacity of the placement group corresponding to the failed object storage device on the normal object storage device that has established a first mapping relationship with itself.
[0106] In a specific example, if the object storage device members corresponding to placement group PG1.0 include (osd.0, osd.2, osd.4), meaning the data in the placement group is stored (or backed up) on different OSD devices, such as osd.0, osd.2, and osd.4, and then osd.4 fails, the simulation will kick osd.4 out. After the simulation determines that placement group PG1.0 needs to be rebuilt, the data storage of the placement group in osd.4 needs to be allocated to the non-failed member osd.6. Based on the redundancy type of the storage pool and the data storage volume of osd.4 corresponding to the placement group, this will constitute a portion of the raw capacity on osd.6. Of course, if other OSD devices fail, and the corresponding data storage volume of the placement group is also allocated to osd.6, then the raw capacity also needs to be determined in the same way. Finally, the sum of all the raw capacity that needs to be transferred to osd.6 is the raw capacity of the placement group corresponding to the failed object storage device on the normal object storage device with which it has established the first mapping relationship. Here, osd.6 refers to the object storage device in normal condition that has established the first mapping relationship with the placement group corresponding to the failed object storage device. Therefore, the purpose of the first mapping relationship is to find the object storage device in the placement group corresponding to the failed object storage device that needs to migrate data.
[0107] Step S302: Based on the preset reconstruction bandwidth, total capacity to be reconstructed, current distribution information, number of placement groups in the storage pool, and data storage volume of the placement group corresponding to the object storage device in normal state in the storage pool, determine the current water level information of each object storage device in normal state.
[0108] In an optional example, see [link to implementation details of the method steps]. Figure 4 As shown, it includes the following steps:
[0109] Step S401: Based on the preset bandwidth for reconstruction and the total capacity to be reconstructed, predict the incremental business data during the reconstruction period.
[0110] Specifically, based on the total capacity to be reconstructed and the preset reconstruction bandwidth, the data reconstruction time is predicted. That is, the data reconstruction time is calculated by dividing the total capacity to be reconstructed by the preset reconstruction bandwidth. Based on the increase in business data during the failure period of the failed object storage device and the failure time, the business data growth rate is obtained. This growth rate is calculated by dividing the quantity by the time. Using the data reconstruction time and the business data growth rate, the increase in business data during the reconstruction period is predicted. Multiplying the data reconstruction time by the business data growth rate yields the increase in business data during the reconstruction period.
[0111] Step S402: Based on the number of placement groups in the storage pool and the business data increment during the reconstruction, obtain the average data increment for each placement group.
[0112] Specifically, based on the number of placement groups in the storage pool, the incremental business data during the reconstruction period is evenly distributed to each placement group to simulate the average data increment of each placement group.
[0113] Step S403: Based on the current distribution information of placement groups in the storage pool, the redundancy type of the storage pool, the number of placement groups in the storage pool, the average data increment of each placement group, and the stored capacity information of each object storage device in normal state, determine the current water level information of each object storage device in normal state.
[0114] Specifically, based on the current distribution information of placement groups in the storage pool and the number of placement groups in the storage pool, the average data increment of each placement group is added to the stored capacity information of each object storage device in normal state. The stored capacity of the object storage device in normal state plus the average data increment of each placement group is the current water level information of the object storage device in normal state.
[0115] Specifically, in the above method steps, the current water level information of the object storage device in a normal state is determined based on the total amount to be reconstructed and the incremental business data. By optimizing the water level reconstruction strategy, data can be more evenly distributed in the storage pool, thereby improving data availability and reliability. Accurate prediction of the water level increase caused by reconstruction after a hard drive failure in a high-water level scenario, and prediction of the incremental business data during reconstruction, allows for setting a danger level based on the difference from the preset value. This enables control over whether automatic reconstruction needs to be triggered, and how to adjust the reconstruction speed after triggering reconstruction, preventing incomplete data reconstruction from affecting business read / write operations. This improves product usability, and the optimized strategy helps to better manage redundant data, ensuring that the system can still quickly recover data when some storage nodes fail.
[0116] This embodiment provides a method for optimizing a high-water level reconstruction strategy for a distributed storage pool, which can be used in the aforementioned mobile terminals, such as mobile phones and tablets. Based on the method corresponding to any of the foregoing embodiments, in an optional example, when it is determined, based on the current water level information of each object storage device in a normal state, that a faulty object storage device cannot be removed, the method further includes the following steps, for example:
[0117] Step b1: Save the water level information for each object storage device in the storage pool.
[0118] Step b2 involves periodically detecting changes in the water level information of each object storage device.
[0119] Step b3: When the water level information of the first object storage device drops below a preset water level threshold based on the changes, re-determine whether the object storage device that has failed meets the conditions for being kicked out.
[0120] The first object storage device is any object storage device in the storage pool.
[0121] Specifically, the water level information of each object storage device is periodically judged according to the above process. For example, the judgment is recalculated every five minutes according to the above judgment process. If the faulty object storage device is not allowed to be kicked out, the current water level in the storage pool involved in the faulty object storage device is saved, and the water level is detected based on the timer function. If the water level drops by more than 2%, the judgment is re-performed until the kick-out condition is met, and then the recorded object storage device water level information is cleared.
[0122] In layman's terms, this means that there are faulty object storage devices that have previously failed but did not meet the conditions for removal at the time. In subsequent processes, these devices need to be repeatedly judged according to a preset cycle until the removal conditions are met, at which point the judgment stops and the faulty object storage device is removed.
[0123] In a specific example, if the simulated data level on an object storage device exceeds a preset threshold of 93%, the faulty object storage device is not allowed to be removed, and this is recorded. If the simulated data level on an object storage device does not exceed the preset threshold of 93%, a danger level is set based on the difference between the simulated data level and the preset threshold. When the difference between the data level of all object storage devices and the preset threshold is greater than 20%, the danger level is safe. When the difference between the data level of all object storage devices and the preset threshold is between 5% and 20%, the danger level is warning. When the difference between the data level of all object storage devices and the preset threshold is less than 5%, the danger level is dangerous.
[0124] However, in practical applications, it is necessary to preset the range of hazard levels according to the actual situation and set the preset values based on the actual situation.
[0125] In an optional example, when determining that a faulty object storage device is being removed based on the current water level information of each normal object storage device, the method includes the following steps:
[0126] Step c1: After performing the operation of kicking out the failed object storage device, determine whether to trigger a reconstruction operation.
[0127] Step c2: Identify the danger level corresponding to the current water level information of each object storage device in normal condition.
[0128] Step c3: When it is determined that the refactoring operation is triggered, obtain the actual business increment generated during the refactoring time.
[0129] Specifically, the actual business increment generated during the reconstruction period is obtained in order to control the reconstruction speed in the later stages.
[0130] Based on the current water level information of each object storage device in normal condition, the corresponding danger level, data reconstruction time, actual business increment, and predicted business data increment during the reconstruction period, the reconstruction control strategy is determined.
[0131] Specifically, if kicking out a trigger for refactoring is allowed and it is at a dangerous level, then it is also necessary to monitor data writing in the timer function and adjust different refactoring strategies based on the estimated refactoring time at the time of prediction and the business increment during the refactoring period.
[0132] In a specific example: if the reconstruction time is within 50% of the estimated time and the actual business increment is less than 50% of the estimated business increment, no adjustment is made. If the reconstruction time is within 50% of the estimated time and the actual business increment is greater than 50% of the estimated business increment, the business growth rate is calculated based on the current actual business increment and the actual reconstruction time. The remaining business increment is obtained by subtracting the actual business increment from the estimated business increment. The remaining business increment is divided by the business growth rate to obtain the remaining reconstruction time. The reconstruction speed is obtained by dividing the current remaining reconstruction data volume by the remaining reconstruction time. If the current reconstruction speed is greater than this calculated value, it remains unchanged; otherwise, it is adjusted to this calculated value.
[0133] This embodiment provides an optimization method for high-water level reconstruction strategy of distributed storage pool, which can be used in the aforementioned mobile terminals, such as mobile phones and tablets. Figure 5 This is a schematic diagram of the overall process of a distributed storage pool high-water level reconstruction strategy optimization method provided by an embodiment of the present invention, as shown below. Figure 5 As shown, the details are as follows:
[0134] Step 1: Obtain the OSD involved in this failure and record the storage pool information involved in the newly failed OSD.
[0135] Specifically, when an OSD fails, the current object storage device dataset of the cluster is copied first. Since the object storage device data in the cluster can change, the subsequent simulation process is based on the copied object storage device dataset, which avoids the problem of incorrect simulation results caused by changes in object storage device data. The OSD that failed in this case is obtained from the copied object storage device dataset.
[0136] In a specific example, the storage pool information involved by all OSDs in the copied object storage device dataset is obtained, including OSDs that have failed and OSDs that have not failed. The failed OSDs in the copied object storage device dataset are marked as OUT, and the storage pool information involved after kicking out the failed OSD is simulated.
[0137] Step 2: Simulate and calculate the placement group distribution information after the faulty OSD is kicked out.
[0138] Specifically, based on the storage pool information involved after the faulty OSD is kicked out, the placement group distribution information of the faulty OSD after being kicked out is calculated using a data distribution algorithm.
[0139] Step 3: Calculate the current data volume and total capacity to be reconstructed of the OSD based on the placement group distribution information and the placement group data volume.
[0140] Specifically, the redundancy type of the storage pool is obtained, and the method for obtaining the raw capacity on the object storage device is determined based on the redundancy type. If the redundancy type of the storage pool is replica type, the raw capacity of the placement group on the object storage device is the data storage amount of the placement group corresponding to the failed object storage device. If the redundancy type of the storage pool is erasure K+M type, the raw capacity of the placement group on the object storage device is the data storage amount of the placement group corresponding to the failed object storage device divided by the total number of fragments into which the original data is divided and the number of verification fragments.
[0141] Iterate through all placement groups in the storage pool to obtain the raw capacity of each placement group on the object storage device, and then sum the raw capacities of all placement groups in the storage pool on the object storage device to determine the total capacity to be reconstructed.
[0142] For example, if the members of a placement group are OSD1, OSD2, and OSD4, and OSD4 fails, the members of the placement group will then be OSD1 and OSD2. After simulation, the members will be OSD1, OSD2, and OSD6. In this case, the placement group needs to reconstruct the data onto OSD6. Therefore, the current level of OSD6 is the amount of data already stored in OSD6 plus the amount of data to be reconstructed.
[0143] Step 4: Simulate and calculate the business increment during the reconstruction period and add the business increment during the reconstruction period to the OSD data volume to obtain the current water level information of each object storage device in normal state.
[0144] The time required for data reconstruction is calculated by dividing the total capacity to be reconstructed by the preset reconstruction bandwidth. Then, the increase in business data during the object storage device failure period is divided by the failure time to obtain the business data growth rate. This growth rate is then multiplied by the time required for data reconstruction to obtain the increase in business data during the reconstruction period. Based on the current distribution information of the placement groups, the increase in business data during the reconstruction period is evenly distributed to each placement group, simulating the average data increase for each placement group, thus determining the current status of each object storage device in normal condition.
[0145] Step 5: Determine whether the current water level information of each object storage device in normal state exceeds the threshold.
[0146] If the current water level information of each object storage device in normal condition does not exceed the threshold, the faulty OSD can be removed; if the current water level information of each object storage device in normal condition exceeds the threshold, the faulty OSD cannot be removed.
[0147] In a specific example, when the faulty hard drive of a node is a solid-state drive (SSD), if the number of faulty OSDs exceeds 30%, it is determined whether the faulty OSD should be kicked out.
[0148] Step 6: If the faulty OSD is allowed to be kicked out, set the danger level.
[0149] If a faulty OSD is allowed to be removed and its danger level is dangerous, the data reconstruction speed will be adjusted according to the business increment.
[0150] Step 7: If the faulty OSD is not allowed to be kicked out, the OSDs that failed to be kicked out before the fault is detected and recalculated periodically or based on the water level change to determine whether they can be kicked out.
[0151] In an optional example, a storage pool water level reconstruction model can be constructed based on ARIMA (Autoregressive Integral Moving Average) as the time series forecasting model. See the following for details:
[0152] Collect historical data, including storage utilization, read / write load, and node status.
[0153] Specifically, it monitors and records the space utilization of each storage device or node, and obtains the read and write request volume of each storage unit.
[0154] Historical data was used to train the storage pool water level reconstruction model.
[0155] Parameters such as storage utilization, read / write load, and node status are preprocessed, and the preprocessed data is used to train the storage pool water level reconstruction model. The preprocessing includes removing outliers, filling in missing values, and normalizing the data.
[0156] The storage pool water level reconstruction model is used to predict whether a faulty storage device can be reconstructed.
[0157] By predicting future load conditions and dynamically adjusting accordingly, the load across nodes is balanced, hotspot issues are reduced, and overall system performance is improved. This predictive and adaptive storage pool refactoring optimization method effectively manages the storage pool's state, enhancing overall system performance and reliability.
[0158] Specifically, in the above method steps, the business growth rate is calculated based on the current actual business increment and the actual reconstruction time. The remaining business increment is obtained by subtracting the actual business increment from the estimated business increment. The remaining business increment is divided by the business growth rate to obtain the remaining reconstruction time. The reconstruction speed is obtained by dividing the current remaining reconstruction data volume by the remaining reconstruction time. The reconstruction speed is adjusted based on the business growth rate and reconstruction time. A risk level is set based on the difference from the preset value, thereby controlling whether automatic reconstruction needs to be triggered and how to adjust the reconstruction speed after triggering reconstruction, preventing incomplete data reconstruction from affecting business read and write operations. This improves the usability of the product.
[0159] This embodiment also provides a distributed storage pool high-water level reconstruction strategy optimization device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0160] This embodiment provides a distributed storage pool high-water level reconstruction strategy optimization device, such as... Figure 6 As shown, it includes:
[0161] The data acquisition module 601 is used to acquire the object storage device dataset and the historical distribution information of the placement groups in the storage pool. The object storage device dataset includes status information corresponding to multiple object storage devices, as well as storage pool information to which the multiple object storage devices belong. The multiple object storage devices include object storage devices that have failed and object storage devices that are in normal condition.
[0162] The processing module 602 is used to simulate and determine the current distribution information of the placement group in the storage pool after the failed object storage device is kicked out, based on the status information of each of the object storage devices and the storage pool information, and to determine whether to reconstruct the placement group corresponding to the failed object storage device.
[0163] The water level acquisition module 603 is used to determine the current water level information of each object storage device in normal state based on the reconstruction preset bandwidth, the current distribution information, and the storage pool information if reconstruction is performed.
[0164] The processing module 602 is further configured to generate a water level reconstruction strategy based on the current water level information, wherein the water level reconstruction strategy is used to indicate whether to kick out the object storage device that has failed.
[0165] In an optional example, the storage pool information includes the water level information of the storage pool and the number of groups placed in the storage pool; the processing module 602 is specifically used for:
[0166] The status information of each object storage device, the water level information of the storage pool, and the number of placement groups in the storage pool are input into the data distribution algorithm to simulate the current distribution information of the placement groups in the storage pool after the failed object storage device is kicked out.
[0167] The current distribution information is compared with the historical distribution information;
[0168] When it is determined that the current distribution information and the historical distribution information are inconsistent, it is determined that the placement group corresponding to the object storage device that has failed should be reconstructed.
[0169] In an optional example, the storage pool information may also include the redundancy type of the storage pool and the data storage amount of the placement group corresponding to each object storage device in the storage pool;
[0170] The water level acquisition module 603 is used to determine the total capacity to be reconstructed based on the redundancy type of the storage pool, the current distribution information, and the data storage volume of the placement group corresponding to the object storage device that has failed in the storage pool.
[0171] Based on the preset reconstruction bandwidth, the total capacity to be reconstructed, the current distribution information, the number of placement groups in the storage pool, and the data storage volume of the placement group corresponding to the object storage device in normal state in the storage pool, the current water level information of each object storage device in normal state is determined.
[0172] In an optional example, the processing module 602 is specifically used to extract a first mapping relationship from the current distribution information, the first mapping relationship being a newly established first mapping relationship between the placement group to which the data stored by the failed object storage device belongs and the object storage device in normal condition;
[0173] Based on the redundancy type of the storage pool and the data storage volume of the placement group corresponding to the failed object storage device, determine the raw capacity of the placement group corresponding to the failed object storage device on the normal object storage device that has established a first mapping relationship with itself.
[0174] The total capacity to be reconstructed is determined based on the raw capacity of the placement group corresponding to the faulty object storage device on the normal object storage device that has established a first mapping relationship with itself.
[0175] In an optional example, the water level acquisition module 603 is specifically used to predict the incremental business data during the reconstruction period based on the preset reconstruction bandwidth and the total capacity to be reconstructed;
[0176] Based on the number of placement groups in the storage pool and the increase in business data during the reconstruction, the average data increment of each placement group is obtained;
[0177] Based on the current distribution information of the placement groups in the storage pool, the redundancy type of the storage pool, the number of placement groups in the storage pool, the average data increment of each placement group, and the data storage volume of the placement group corresponding to the object storage device in the storage pool that is in a normal state, the current water level information of each object storage device in a normal state is determined.
[0178] In an optional example, the water level acquisition module 603 is specifically used to predict the data reconstruction time based on the total capacity to be reconstructed and the preset reconstruction bandwidth;
[0179] Based on the incremental business data during the failure period of the object storage device that has failed and the time of the failure, the growth rate of business data is obtained.
[0180] Using the data reconstruction time and the growth rate of the business data, predict the increment of the business data during the reconstruction period.
[0181] In an optional example, the processing module 602 is further configured to save the water level information of each of the object storage devices in the storage pool when the water level reconstruction strategy indicates that the failed object storage device cannot be kicked out;
[0182] Periodically detect changes in the water level information of each of the aforementioned object storage devices;
[0183] When the water level information of the first object storage device drops below a preset water level threshold based on the changes, it is re-determined whether the object storage device that has failed meets the conditions for being kicked out, wherein the first object storage device is any object storage device in the storage pool.
[0184] In an optional example, the processing module 602 is further configured to determine whether to trigger a reconstruction operation after performing the operation of kicking out the faulty object storage device when the water level reconstruction strategy indicates that the faulty object storage device is kicked out;
[0185] Identify the danger level corresponding to the current water level information of each object storage device that is in a normal state;
[0186] When a refactoring operation is determined to be triggered, the actual business increment generated within the refactoring time period is obtained;
[0187] Based on the current water level information of each object storage device in normal condition, the corresponding danger level, the data reconstruction time, the actual business increment, and the predicted business data increment during the reconstruction period, a reconstruction control strategy is determined.
[0188] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0189] In this embodiment, the distributed storage pool high-water level reconstruction strategy optimization device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit), a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0190] This embodiment provides a distributed storage pool high-water level reconstruction strategy optimization device. In the above method, by promptly reconstructing the placement group when an object storage device fails, data loss due to failure can be ensured, improving data storage reliability. Therefore, by automatically identifying faulty devices and simulating the reconstruction process, and through pre-simulation and strategy generation, uncertainties and potential costs in the actual reconstruction process can be reduced. Risks caused by improper data migration during reconstruction are also reduced. Quickly determining whether and how to reconstruct can shorten system response time and improve overall performance. Effective management and replacement of faulty devices enhance system stability and fault tolerance. Furthermore, by rationally allocating and adjusting storage resources, the utilization rate of storage space in the storage pool can be maximized. Based on the current water level information, the system can dynamically allocate resources to ensure the balance of the storage pool. Finally, based on the water level reconstruction strategy, it is determined whether to remove faulty object storage devices, and high availability of the system is maintained through timely reconstruction and optimized resource allocation.
[0191] This invention also provides a computer device having the above-described features. Figure 6 The distributed storage pool high-water level reconstruction strategy optimization device is shown.
[0192] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 7 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 7 Take a processor 10 as an example.
[0193] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware integrated circuit. The aforementioned hardware integrated circuit may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.
[0194] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.
[0195] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0196] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0197] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.
[0198] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0199] A portion of this invention can be applied to computer program products, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer. Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and all such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A method for optimizing a distributed storage pool high water mark reconfiguration strategy, characterized in that, The method comprises: acquiring object storage device dataset and historical distribution information of placement groups in a storage pool, the object storage device dataset comprising state information corresponding to a plurality of object storage devices respectively, and storage pool information to which the plurality of object storage devices belong, the plurality of object storage devices comprising an object storage device having occurred a fault and an object storage device in a normal state; based on the state information of each object storage device and the storage pool information, simulating to determine current distribution information of placement groups in the storage pool if the object storage device having occurred a fault is kicked out; determining whether to reconstruct a placement group corresponding to the object storage device having occurred a fault according to the current distribution information and the historical distribution information; if reconstruction is performed, determining current water level information of each object storage device in a normal state according to a reconstruction preset bandwidth, the current distribution information, and the storage pool information; generating a water level reconstruction strategy according to the current water level information, wherein the water level reconstruction strategy is used to indicate whether to kick out the object storage device having occurred a fault.
2. The method of claim 1, wherein, The storage pool information comprises water level information of the storage pool and a number of placement groups in the storage pool; The simulation to determine the current distribution information of placement groups in the storage pool if the object storage device having occurred a fault is kicked out based on the state information of each object storage device and the storage pool information, and the determination of whether to reconstruct the placement group corresponding to the object storage device having occurred a fault, specifically comprises: inputting the state information of each object storage device, the water level information of the storage pool, and the number of placement groups in the storage pool into a data distribution algorithm to simulate the current distribution information of placement groups in the storage pool if the object storage device having occurred a fault is kicked out; comparing the current distribution information with the historical distribution information; when it is determined that the current distribution information and the historical distribution information are inconsistent, determining to reconstruct the placement group corresponding to the object storage device having occurred a fault.
3. The method of claim 2, wherein, The storage pool information further comprises a redundancy type of the storage pool and a placement group data storage amount corresponding to each object storage device in the storage pool; The determination of the current water level information of each object storage device in a normal state according to the reconstruction preset bandwidth, the current distribution information, and the storage pool information if reconstruction is performed, specifically comprises: determining a total reconstruction capacity according to the redundancy type of the storage pool, the current distribution information, and the placement group data storage amount corresponding to the object storage device having occurred a fault in the storage pool; determining the current water level information of each object storage device in a normal state according to the reconstruction preset bandwidth, the total reconstruction capacity, the current distribution information, the number of placement groups in the storage pool, and the placement group data storage amount corresponding to the object storage device in a normal state in the storage pool.
4. The method of claim 3, wherein, The method further comprises: extracting a first mapping relationship from the current distribution information, the first mapping relationship being a first mapping relationship re-established between a placement group to which data stored by the failed object storage device belongs and an object storage device in a normal state; determining, according to the redundancy type of the storage pool and the data storage amount of the placement group corresponding to the failed object storage device, a raw capacity of the placement group corresponding to the failed object storage device on the object storage device in the normal state with which the first mapping relationship is established; determining the total capacity to be reconstructed based on the raw capacity of the placement group corresponding to the failed object storage device on the object storage device in the normal state with which the first mapping relationship is established.
5. The method of claim 3 or 4, wherein, The method further comprises: predicting a service data increment during the reconstruction according to the reconstruction preset bandwidth and the total capacity to be reconstructed; obtaining an average data increment of each placement group based on the number of placement groups in the storage pool and the service data increment during the reconstruction; determining the current water level information of each object storage device in the normal state based on the current distribution information of the placement groups in the storage pool, the redundancy type of the storage pool, the number of placement groups in the storage pool, the average data increment of each placement group, and the data storage amount of the placement group corresponding to the object storage device in the normal state.
6. The method of claim 5, wherein, The method further comprises: predicting a data reconstruction time based on the total capacity to be reconstructed and the reconstruction preset bandwidth; obtaining a service data growth rate based on a service data increment during a failure of the failed object storage device and a failure occurrence time; predicting the service data increment during the reconstruction by using the data reconstruction time and the service data growth rate.
7. The method of claim 1-4, wherein, When the water level reconstruction strategy indicates that the failed object storage device cannot be kicked out, the method further comprises: saving the water level information of each object storage device in the storage pool; periodically detecting changes in the water level information of each object storage device; when it is determined according to the changes that the water level information of a first object storage device decreases by a preset water level threshold, re-determining whether the failed object storage device meets the kicked-out condition, wherein the first object storage device is any object storage device in the storage pool.
8. The method of claim 6, wherein, When the water level reconstruction strategy indicates that the failed object storage device is kicked out, the method further comprises: After performing the operation of kicking out the failed object storage device, it is determined whether to trigger a reconstruction operation; Identify the risk level corresponding to the current water level information of each object storage device in normal state; When it is determined to trigger the reconstruction operation, the actual service increment generated within the reconstruction time is obtained; According to the risk level corresponding to the current water level information of each object storage device in normal state, the data reconstruction time, the actual service increment, and the predicted service data increment during the reconstruction, a reconstruction control strategy is determined.
9. A distributed storage pool high water mark reconstitution strategy optimization apparatus, characterized in that, The device comprises: A data acquisition module is configured to acquire historical distribution information of object storage device data sets and placement groups in a storage pool, wherein the object storage device data sets comprise state information corresponding to a plurality of object storage devices respectively, and storage pool information to which the plurality of object storage devices belong, and the plurality of object storage devices comprise a failed object storage device and an object storage device in normal state; A processing module is configured to simulate and determine current distribution information of placement groups in the storage pool after the failed object storage device is kicked out based on the state information of each object storage device and the storage pool information, and determine whether to reconstruct the placement group corresponding to the failed object storage device according to the current distribution information and the historical distribution information; A water level acquisition module is configured to determine current water level information of each object storage device in normal state according to reconstruction preset bandwidth, the current distribution information, and the storage pool information if reconstruction is performed; The processing module is further configured to generate water level reconstruction strategy according to the current water level information, wherein the water level reconstruction strategy is used to indicate whether to kick out the failed object storage device.
10. A computer device, comprising: Comprise: A memory and a processor are communicatively connected between each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the distributed storage pool high water level reconstruction strategy optimization method of any one of claims 1 to 8.
11. A computer readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, and the computer instructions are used to make the computer execute the distributed storage pool high water level reconstruction strategy optimization method of any one of claims 1 to 8.
Citation Information
Patent Citations
Method, device and equipment for controlling data reconstruction and readable medium
CN113687798A
Method and device for checking validity of data reconstruction
CN117555735A