A disk scheduling method

By collecting historical access patterns and real-time traffic monitoring of disk array cards, the disk status is dynamically adjusted, solving the problems of unbalanced bandwidth allocation and response delay in traditional RAID technology, achieving flexible and adaptive disk scheduling, and improving the performance and reliability of the storage system.

CN120428922BActive Publication Date: 2025-09-16INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510855776.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-16
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

When faced with a mismatch between high-speed disk performance and bandwidth, traditional RAID technology suffers from rigid static configuration and insufficient dynamic scheduling, leading to unbalanced bandwidth allocation and a surge in response delays, impacting business continuity and system reliability.

Method used

By collecting the historical input/output access patterns of disk array cards and using pre-trained disk scheduling models to generate scheduling strategies that adapt to business rules, the interface traffic is monitored in real time and bandwidth utilization is calculated. The dormant and active states of the disks are dynamically adjusted to achieve flexible and adaptive bandwidth allocation.

Benefits of technology

It effectively solves the problems of unbalanced bandwidth allocation and sudden load response delay in traditional technologies, ensures business continuity and system reliability, and improves the performance and resource utilization of storage systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120428922B_ABST
    Figure CN120428922B_ABST
Patent Text Reader

Abstract

This application discloses a disk scheduling method, which relates to the field of storage system technology. The method involves collecting historical input / output access patterns of disk array cards, which include access information for disk input / output requests during historical time periods, and using a pre-trained disk scheduling model to generate a disk scheduling strategy that adapts to business rules. The method also monitors the interface traffic of the disk array cards in real time, calculates the bandwidth utilization of the disk array cards based on the interface traffic, and dynamically adjusts the number of active and dormant disks based on the strategy. In this way, the method flexibly and adaptively allocates the number of disks and bandwidth, effectively resolving problems such as bandwidth imbalance and sudden load response delays caused by rigid configurations in traditional technologies, thereby ensuring business continuity and system reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of storage systems, and in particular to a disk scheduling method, a disk scheduling device, an electronic device, a storage medium, and a program product. Background Art

[0002] With the booming development of data-intensive applications such as artificial intelligence training, cloud computing, and real-time data analytics, the demand for high-throughput and low-latency storage systems is exploding. Redundant Array of Independent Disks (RAID) technology, leveraging parallel input / output and redundancy mechanisms across multiple disks, strikes a balance between improved storage performance and data security, making it a mainstream solution for enterprise-class storage. However, with the widespread adoption of high-speed storage media, traditional RAID architectures are gradually exposing significant shortcomings. Mainstream RAID cards, limited by the Peripheral Component Interconnect Express (PCIe) interface specifications and memory buffering capabilities, struggle to fully handle the aggregate bandwidth of multiple high-speed disks, hindering the full utilization of storage performance.

[0003] Furthermore, traditional RAID groups employ a fixed number of disks and bandwidth allocation strategies, which cannot be dynamically adjusted based on real-time load. Bandwidth competition between RAID synchronization tasks and business input / output (I / O) has long been a problem, and related technologies rely solely on fixed thresholds for speed limiting, lacking adaptive adjustment capabilities. This makes storage systems prone to bandwidth imbalances and surges in response latency when faced with sudden loads or high-concurrency requests, seriously impacting business continuity and system reliability. In summary, traditional RAID technology faces multiple technical bottlenecks when addressing the mismatch between high-speed disk performance and bandwidth, including rigid static configuration, insufficient dynamic scheduling, and low energy efficiency. Innovative optimization solutions are urgently needed to overcome this dilemma. Summary of the Invention

[0004] The present application provides a disk scheduling method, device, electronic device, storage medium and program product to at least solve the problem of unbalanced bandwidth allocation caused by static configuration in related technologies.

[0005] The present application provides a disk scheduling method, comprising: obtaining historical input / output access patterns of a disk array card, the historical input / output access patterns including access information of disk input / output requests within a historical time period; determining a disk scheduling strategy based on the historical input / output access patterns and a pre-trained disk scheduling model; monitoring interface traffic of the disk array card and calculating bandwidth utilization of the disk array card based on the interface traffic; and adjusting the number of disks in a dormant state and an active state based on the bandwidth utilization and the disk scheduling strategy.

[0006] The present application also provides a disk scheduling device, comprising:

[0007] An acquisition module is used to acquire a historical input / output access pattern of the disk array card, where the historical input / output access pattern includes access information of disk input / output requests within a historical time period;

[0008] The policy generation module is used to determine the disk scheduling policy based on historical input / output access patterns and pre-trained disk scheduling models;

[0009] The acquisition module is also used to monitor the interface traffic of the disk array card and calculate the bandwidth utilization of the disk array card based on the interface traffic;

[0010] The adjustment module is used to adjust the number of disks in the dormant state and the active state according to bandwidth utilization and disk scheduling policy.

[0011] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned disk scheduling methods when executing the computer program.

[0012] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned disk scheduling methods are implemented.

[0013] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above disk scheduling methods when executed by a processor.

[0014] This application collects historical input / output access patterns of disk array cards, which cover access information for disk input / output requests over historical time periods, and uses a pre-trained disk scheduling model to generate a disk scheduling strategy tailored to business patterns. It also monitors the interface traffic of the disk array cards in real time, calculates the RAID card bandwidth utilization based on the interface traffic, and dynamically adjusts the number of active and dormant disks based on the strategy. This allows for flexible and adaptive disk number and bandwidth allocation, effectively resolving bandwidth imbalances and sudden load response delays caused by rigid configurations in traditional technologies, ensuring business continuity and system reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0016] Figure 1 A schematic diagram of the hardware architecture relied upon for executing a disk scheduling method provided in an embodiment of the present application;

[0017] Figure 2 A flowchart of a disk scheduling method provided in an embodiment of the present application;

[0018] Figure 3 A timing diagram of the RAID card bandwidth dynamic optimization process;

[0019] Figure 4 A schematic diagram of the structure of a disk scheduling device provided in an embodiment of the present application;

[0020] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0023] In order to more clearly illustrate the embodiments of the present application, the following briefly introduces the technical terms used in the embodiments:

[0024] Redundant Array of Independent Disks (RAID) is a collection of technologies that virtualizes multiple physical hard drives into a single logical volume. It improves storage performance, reliability, and availability through data redundancy and parallel I / O operations.

[0025] RAID levels represent different data organization and redundancy schemes within RAID technology. Each level employs a different redundancy strategy, resulting in varying performance, reliability, and cost. For example, RAID 0 uses striping to improve read and write performance but lacks redundancy; RAID 1 uses mirroring for data redundancy, achieving 50% space utilization; and RAID 5 combines striping with distributed parity, balancing performance, security, and cost. Different RAID levels offer unique characteristics in terms of performance, reliability, and cost, allowing for trade-offs to be made to meet diverse application requirements, such as data protection and improved read and write speeds.

[0026] The RAID card coordinates data input and output requests from the disks. When writing data, the host (CPU / memory) sends the data to the RAID card's cache. The RAID card then splits or mirrors the data across multiple disks, depending on the RAID level (e.g., RAID 0 / 1 / 5). When reading data, the RAID card reads the data from multiple disks in parallel, combines it, and returns it to the host.

[0027] In RAID technology, striping divides data into equal-sized blocks. These blocks are distributed across multiple physical hard drives, allowing data reads and writes to proceed in parallel, improving I / O speed and achieving I / O parallelization. Different RAID levels handle striping differently, and properly setting the stripe size can also optimize performance.

[0028] Non-Volatile Memory Host Controller Interface (NVME) is a high-speed storage protocol standard for solid-state drives (SSDs). Designed specifically for PCIe-based storage devices, it fully utilizes the performance of SSDs.

[0029] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0030] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the disk scheduling method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0031] The hardware architecture on which the disk scheduling method is executed depends on includes: RAID controller and disk array, high-performance bus interface, disk monitoring module, hierarchical scheduling system, distributed cache system, real-time operating system, network storage architecture, metadata management service, etc.

[0032] The RAID controller supports dynamic disk management and has sufficient cache to store dynamic scheduling metadata. The disk array supports mixed deployment of different storage media types and provides hot-swappable capabilities to enable dynamic switching of disk status. A PCIe 4.0 / 5.0 bus connects the RAID controller to the host system, ensuring high-bandwidth data transmission and preventing the bus from becoming a performance bottleneck. Real-time disk health indicators such as temperature, erase / write counts, and read / write latency provide data support for scheduling decisions. A real-time I / O monitoring component is deployed to collect key metrics such as disk bandwidth, queue depth, and response time. A pre-trained Long Short-Term Memory (LSTM) prediction module or reinforcement learning model is integrated to generate scheduling policies based on historical I / O patterns and real-time load. Kernel modules or drivers are used to control disk activation / deactivation and dynamically adjust striping configurations. Hotspot data is cached to reduce disk access frequency, while support for write-mode improves write performance. A low-latency kernel or dedicated storage operating system is used to ensure fast scheduling decisions and deterministic responses to I / O requests. Distributed metadata services are used to record information such as disk status and stripe mapping relationships to ensure data consistency during the scheduling process.

[0033] The above architecture provides a low-latency, highly reliable execution environment for dynamic disk scheduling through the collaboration of hardware acceleration, software definition, and intelligent scheduling.

[0034] like Figure 1 As shown, Figure 1A schematic diagram of the hardware architecture on which the disk scheduling method provided in an embodiment of the present application relies includes a disk array card bandwidth sensor 101, a dynamic disk scheduling engine 102, a disk activation / sleep circuit 103, a disk input / output monitoring unit 104, a data migration background thread 105, and a write-time redirection engine 106.

[0035] Among them, the disk array card bandwidth sensor 101 monitors the bandwidth usage of the RAID card in real time, analyzes and evaluates the performance indicators of the storage system, switches the bandwidth or generates an early warning when the bandwidth usage is abnormal, and can also provide optimization suggestions. The dynamic disk scheduling engine 102 is used to dynamically adjust the number of dormant / active disks according to the bandwidth usage to achieve the optimal match between performance and resources. The disk activation / sleep circuit 103 is used to control the dormant / activated disks according to the adjustment of the number of dormant / active disks by the dynamic disk scheduling engine 102. The disk input / output monitoring unit 104 is used to monitor the disk bandwidth. The data migration background thread 105 is used to migrate cold data from the dormant disk. The write-time redirection engine 106 is used to write new business data to the active disk without affecting business continuity.

[0036] An embodiment of the present application provides a disk scheduling method, and the method is described in detail in conjunction with the execution process of the disk scheduling method.

[0037] like Figure 2 As shown, Figure 2 A flowchart of a disk scheduling method provided in an embodiment of the present application is provided. The method includes the following steps S201 to S204:

[0038] S201: Obtain historical input / output access patterns of a disk array card.

[0039] The historical I / O access pattern is the access information of disk I / O requests made by the disk array card within a historical period of time, including but not limited to the disk address of the I / O request, the request arrival timestamp, the data block size, and the operation type (read / write). The historical I / O access pattern has a time series characteristic.

[0040] S202: Determine a disk scheduling strategy based on historical input / output access patterns and a pre-trained disk scheduling model.

[0041] The pre-trained disk scheduling model is built based on a long short-term memory (LSTM) prediction module. Through recurrent calculations, the pre-trained disk scheduling model identifies frequently accessed disk locations, periodic peaks in input / output requests, and temporal correlation patterns between read and write operations. Disk scheduling policies are then determined based on these identification results.

[0042] The disk scheduling strategy includes but is not limited to: if the probability that the first disk will not be accessed in the future time period is greater than a first preset probability value, then adjusting the first disk to a dormant state; if the probability that the second disk will be accessed in the future time period is greater than a second preset probability value, then adjusting the second disk to an active state, wherein the first and second disks are any disks among the multiple disks corresponding to the disk array card; sorting input / output requests according to the disk positions of the input / output requests in the future time period; adjusting the stripe size according to the load type of the disk array card; matching the association pattern of read and write operations in the time dimension and performing stripe migration.

[0043] The above embodiment learns historical I / O access patterns through a disk scheduling model to dynamically adapt to business loads and adjust disk resources, avoiding over-configuration of the number of disks and meeting peak demands with fewer disks through a dynamic activation mechanism.

[0044] S203: Monitor the interface traffic of the disk array card, and calculate the bandwidth utilization of the disk array card according to the interface traffic.

[0045] In some embodiments, the interface traffic of the disk array card is first monitored, and then the bandwidth utilization is calculated based on the interface traffic and the maximum theoretical bandwidth of the disk array card, wherein the interface traffic represents the number of bytes actually transmitted per unit time.

[0046] Specifically, an embedded sensor chip can monitor the PCIe / SATA interface traffic of the RAID card and calculate bandwidth utilization based on the maximum theoretical bandwidth of the disk array card. Embedded sensor chips, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), can accelerate calculations.

[0047] The above embodiment can realize accurate quantification and dynamic control of storage system transmission efficiency by monitoring disk array card interface traffic and calculating bandwidth utilization.

[0048] In some embodiments, the available bandwidth of the disk array card is calculated based on the interface traffic and its level pattern of the disk array card, and the total disk bandwidth requirement is simultaneously obtained. A first bandwidth threshold is determined based on the total disk bandwidth requirement and a first preset coefficient. If the available bandwidth of the disk array card is less than the first bandwidth threshold, an early warning strategy is generated. The early warning strategy includes, but is not limited to, at least one of storing the early warning event in the system log, generating an early warning pop-up window, and pushing an early warning notification.

[0049] For example, let's calculate the available bandwidth of a RAID card with 8 PCIe 4.0 lanes and RAID 0, representing four NVMe SSDs with a single drive bandwidth of 3 GB / s. The theoretical RAID bandwidth is: 3 GB / s x 4 = 12 GB / s. Limited by the 16 GB / s bandwidth of PCIe 4.0 x 8, the actual available bandwidth is approximately 12 GB / s.

[0050] The above embodiment calculates the available bandwidth of the disk array card through interface traffic and RAID level mode, and generates a dynamic warning strategy based on the total disk bandwidth demand and preset coefficients, which can achieve accurate prediction and automatic response to storage system bandwidth bottlenecks.

[0051] Optionally, obtaining the total disk bandwidth requirement includes monitoring the input / output rates of multiple disks corresponding to the disk array card, calculating the bandwidth requirement of each disk based on the I / O rate and the number of I / O operations, and then calculating the total disk bandwidth requirement based on the bandwidth requirements of each disk. By monitoring the I / O rate of each disk in the disk array and calculating the bandwidth requirement based on the number of operations, fine-grained quantification of storage resource consumption and dynamic load analysis can be achieved.

[0052] This embodiment of the application provides an optional implementation method. When calculating the total disk bandwidth requirement, the disk array card level mode is combined. For disk array cards of different level modes, the total disk bandwidth requirement is calculated in different ways. For example, for RAID 0, the bandwidth requirements of each corresponding disk are directly accumulated to obtain the total disk bandwidth requirement. For RAID 5, the checksum overhead coefficient is calculated based on the corresponding number of disks, and then the total disk bandwidth requirement is calculated by combining the checksum overhead coefficient and the bandwidth requirements of each disk. Dynamic calculation of the total disk bandwidth requirement provides an accurate basis for resource scheduling.

[0053] A first bandwidth threshold is determined based on the total disk bandwidth requirement and a first preset coefficient. The first preset coefficient may be 0.9, which is not specifically limited in this application. The total disk bandwidth requirement is denoted as Y, and the first bandwidth threshold may be 0.9Y. The available bandwidth of the disk array card is then compared to see whether it is less than the first bandwidth threshold.

[0054] If the available bandwidth of the disk array card falls below the first bandwidth threshold, indicating that the available bandwidth of the disk array card does not meet the disk requirements, an early warning strategy is generated, including but not limited to storing the warning event in the system log with a timestamp, calculating and recording the bandwidth difference; generating a visual early warning pop-up window, such as a red warning pop-up window on the management node, indicating that the RAID card has a bandwidth bottleneck; and pushing early warning notifications, such as sending a Simple Network Management Protocol (SNMP) trap to the operation and maintenance platform via the baseboard management controller (BMC), intelligent platform management interface (IPMI), or email. This multi-dimensional early warning strategy enables a hierarchical response to faults.

[0055] In other embodiments, a second bandwidth threshold is determined based on the total required bandwidth of the disk and a second preset coefficient, where the second preset coefficient is less than the first preset coefficient. When the available bandwidth of the disk array card is less than the first bandwidth threshold and greater than or equal to the second bandwidth threshold, the stripe size of the disk array card is adjusted, or the interface channel of the disk array card is upgraded.

[0056] The second preset coefficient can be 0.8, and this application does not impose any specific restrictions on this. According to the total required bandwidth Y of the disk and the second preset coefficient 0.8, the second bandwidth threshold is calculated to be 0.8Y. When the available bandwidth of the disk array card is less than the first bandwidth threshold 0.9Y and greater than or equal to the second bandwidth threshold 0.8Y, it means that the available bandwidth of the disk array card is slightly mismatched with the total required bandwidth of the disk. In this case, the stripe size of the disk array can be adjusted to optimize the stripe strategy and avoid the cost overhead caused by directly upgrading the hardware; or, the interface channel of the disk array card can be upgraded to achieve capacity expansion. The interface channel upgrade strategy achieves a fundamental solution to the bottleneck.

[0057] The above embodiment combines the total required disk bandwidth with dual preset coefficients to build a hierarchical response mechanism, dynamically adjusting the stripe size or upgrading the interface channel when the available bandwidth is in different threshold ranges, thereby achieving refined management and step-by-step optimization of the storage system bandwidth resources.

[0058] If the available bandwidth of the disk array card is less than the second bandwidth threshold, it indicates a serious mismatch between the available bandwidth of the disk array card and the total required bandwidth of the disks. The disk array card's level mode can be adjusted, for example, by downgrading RAID5 to RAID0 to sacrifice redundancy in exchange for increased bandwidth. Alternatively, the target disks whose load exceeds the threshold can be identified, and then scheduling of the target disks can be reduced, dynamically reducing their scheduling weights to isolate the highly loaded disks. For example, reducing the stripe allocation ratio of the disk from 25% to 10% can be used to migrate traffic to other less loaded disks.

[0059] By adjusting the RAID level mode when available bandwidth is severely insufficient or isolating high-load disks, the storage system can achieve elastic degradation and load balancing under extreme bottlenecks. This dual mechanism enables hierarchical handling in extreme scenarios: RAID level adjustment unlocks bandwidth potential at the architectural level, while disk isolation eliminates local hotspots at the load distribution level. This elastic handling strategy ensures business continuity while achieving efficient utilization of storage resources. By dynamically downgrading RAID levels or isolating faulty disks, the system can continue to provide services despite hardware resource constraints, improving resource utilization and reducing the risk of business interruption caused by bandwidth bottlenecks.

[0060] The above embodiment can match the corresponding optimization strategy according to the severity of the bottleneck through the differentiated threshold setting system to avoid overreaction or underresponse.

[0061] S204: Adjust the number of disks in the dormant state and the active state according to the bandwidth utilization and the disk scheduling policy.

[0062] Combined with intelligent disk scheduling strategies, the number of active and dormant disks can be dynamically adjusted. For example, during peak business I / O periods, synchronization tasks can be reduced in resource usage, with more dormant disks reserved for business operations. During idle business hours, more disks can be activated to accelerate synchronization tasks, synergizing the bandwidth of the two tasks and dynamically adjusting the bandwidth based on real-time conditions.

[0063] If bandwidth utilization is greater than or equal to the first utilization threshold, it indicates abnormal bandwidth usage and the storage system is approaching a performance bottleneck, requiring a bandwidth policy change. In this case, the number of disks to be activated is calculated based on the disk scheduling policy. The scheduling engine then selects disks from multiple disks that correspond to the number of disks to be activated. Disks with high health and a long remaining lifespan are selected. This control then transitions the pending disks to an active state.

[0064] For example, assume the first utilization threshold is 80%. When bandwidth utilization reaches 80%, dynamic disk activation proactively addresses bandwidth bottlenecks. When bandwidth utilization exceeds 80%, the RAID 10 array activates two spare SSDs, increasing total bandwidth from 20 GB / s to 28 GB / s. This increases read and write throughput during peak hours by 35%, and reduces response latency from 6 ms to 3 ms.

[0065] The above embodiment achieves elastic expansion of bandwidth resources. The disk scheduling policy calculates the number of disks to be activated. The scheduling engine selects the optimal candidate disks from the disk pool based on multi-dimensional indicators and activates them in advance when utilization approaches the threshold. This eliminates the risk of bandwidth bottlenecks and prevents performance fluctuations caused by blind activation.

[0066] In some embodiments, after the disk to be activated is controlled to enter the active state, the data to be written is divided into multiple data blocks in response to received write requests, and then these multiple data blocks are allocated to the active disks based on preset rules. The preset rules can be dynamically adjusted based on dimensions such as disk health, real-time load, and media type. After the disk to be activated enters the active state, the write data is divided into data blocks and allocated to the active disks according to the preset rules, achieving a refined improvement in the storage system's write performance and load balancing.

[0067] For example, Figure 3 As shown, Figure 3 This is a timing diagram of the RAID card bandwidth dynamic optimization process. The application layer initiates read and write requests to the scheduling engine. The scheduling engine obtains the current bandwidth status from the RAID card and passes it to the policy decision maker. The policy decision maker uses this information to determine whether a threshold (such as 80%) has been triggered and outputs the number of disks to be expanded. Based on this number of disks, the scheduling engine activates the disks and allocates stripes through the disk scheduler. The disk scheduler then writes the new business data to the active disks in the disk group. After the disk operation is completed, it returns an operation completion signal. This dynamic scheduling solves the problem that the fixed bandwidth of traditional RAID cards cannot adapt to sudden business bursts, and can expand disk resources in real time based on business load.

[0068] The above embodiment divides the data to be written into multiple data blocks and leverages the parallel processing capabilities of active disks to enable write bandwidth to scale linearly with the number of disks. This improves write performance while ensuring load balancing and data reliability, providing core support for the elastic expansion of large-scale storage systems.

[0069] When bandwidth utilization is greater than or equal to the second utilization threshold, the disk scheduling policy calculates the number of disks to be put into hibernation. Disks are then selected based on the number and disk heat, further controlling the hibernation of these disks. The second utilization threshold is greater than the first utilization threshold. The hibernation policy is triggered when bandwidth utilization is in the medium-to-high load range to avoid performance jitter caused by frequent hibernation during low load periods. Disk heat includes access frequency, read / write throughput, and temperature.

[0070] Continuing with the previous example, assume that the second utilization threshold is 90%, the RAID card bandwidth is 32 GB / s, there are eight active NVMe SSDs with a total bandwidth of 56 GB / s, and the calculated number of dormant disks is three. In this case, three low-power NVMe SSDs are selected and put into dormancy.

[0071] The above embodiment calculates the number of disks to be put into hibernation based on the disk scheduling strategy when the bandwidth utilization is higher than the second utilization threshold and the second utilization threshold is greater than the first utilization threshold, and screens the hibernation targets in combination with the disk heat, thereby achieving refined resource control and energy efficiency optimization of the storage system in high-load scenarios.

[0072] In some embodiments, after the control disk enters a dormant state, received write requests are forwarded to an active disk other than the dormant disk, and cold data from the dormant disk is migrated to the active disk. A background thread can be initiated to migrate the data. New data is written to the active disk via redirect-on-write, ensuring business continuity and transparency during the data migration process. Cold data includes files that have not been accessed in seven days.

[0073] Write requests are automatically forwarded to active disks, avoiding I / O interruptions caused by disk hibernation. Simultaneously, a background thread migrates cold data from dormant disks in a non-blocking manner. During the migration process, new data is written directly to active disks through write-time redirection, ensuring that the business I / O path always points to available storage resources.

[0074] When the bandwidth utilization is less than the first utilization threshold, it indicates that the storage load is low and the current operating state of the disk is maintained.

[0075] In some embodiments, after dynamically adjusting the number of disks in a dormant or active state, the order in which the I / O requests are processed is determined in response to received I / O requests, and the active disks are then controlled to process the I / O requests according to this order. Alternatively, the stripe size of the disk array card corresponding to the I / O request is determined in response to received I / O requests, and the active disks are then controlled to store the data corresponding to the I / O request according to this stripe size.

[0076] After dynamically adjusting the number of disks, the system determines the processing order and stripe size based on I / O request characteristics, enabling refined control of storage performance and resource adaptation. Optimizing the processing order of I / O requests significantly reduces response latency. Furthermore, dynamically determining the stripe size based on request characteristics maximizes parallel read and write efficiency. This collaborative mechanism not only avoids resource waste under fixed configurations, but also optimizes the underlying storage links based on real-time load characteristics, improving the overall response consistency of the storage system and enhancing resource utilization.

[0077] In summary, the disk scheduling method provided in the embodiments of the present application collects historical I / O access patterns of disk array cards, uses a pre-trained disk scheduling model, analyzes historical data, and formulates a dynamic disk scheduling strategy. This allows the number of disks and bandwidth allocation to be flexibly adjusted as business changes, rather than being fixed. The method also monitors the interface traffic of the disk array cards, calculates the bandwidth utilization of the RAID card based on the interface traffic, and adjusts the number of active / dormant disks in combination with the dynamic disk scheduling strategy. From the three dimensions of data-driven strategy generation, real-time load dynamic adaptation, and bandwidth competition collaborative processing, the number of disks and bandwidth allocation are transformed from "fixed and rigid" to "flexible and adaptive," effectively resolving problems such as bandwidth imbalance and sudden load response delays caused by policy rigidity in traditional technologies, thereby ensuring business continuity and system reliability.

[0078] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0079] The embodiment of the present application also provides a disk scheduling device, such as Figure 4 As shown, the device includes:

[0080] An acquisition module 401 is configured to acquire a historical input / output access pattern of a disk array card, wherein the historical input / output access pattern includes access information of disk input / output requests within a historical time period.

[0081] A policy generation module 402 is configured to determine a disk scheduling policy based on historical input / output access patterns and a pre-trained disk scheduling model;

[0082] The monitoring module 403 is used to monitor the interface traffic of the disk array card and calculate the bandwidth utilization of the disk array card according to the interface traffic;

[0083] The adjustment module 404 is configured to adjust the number of disks in the dormant state and the active state according to bandwidth utilization and disk scheduling policy.

[0084] As an optional implementation provided in an embodiment of the present application, the monitoring module 403 is further used to: calculate the available bandwidth of the disk array card based on the interface traffic and the level mode of the disk array card; obtain the total disk bandwidth requirement; determine a first bandwidth threshold based on the total disk bandwidth requirement and a first preset coefficient; and generate an early warning strategy when the available bandwidth of the disk array card is less than the first bandwidth threshold; wherein the early warning strategy includes at least one of storing the early warning event in the system log, generating an early warning pop-up window, and pushing an early warning notification.

[0085] As an optional implementation provided in an embodiment of the present application, the monitoring module 403 is specifically used to: monitor the input / output rates of multiple disks corresponding to the disk array card; calculate the bandwidth requirements of each disk based on the input / output rate and the number of input / output operations of each disk; and calculate the total disk bandwidth requirement based on the bandwidth requirements of each disk.

[0086] As an optional implementation provided in an embodiment of the present application, the adjustment module 404 is further configured to: determine a second bandwidth threshold based on a total disk bandwidth requirement and a second preset coefficient; the second preset coefficient is less than the first preset coefficient; if the available bandwidth of the disk array card is less than the first bandwidth threshold and greater than or equal to the second bandwidth threshold, adjust the stripe size of the disk array card or upgrade the interface channel of the disk array card; if the available bandwidth of the disk array card is less than the second bandwidth threshold, adjust the level mode of the disk array card.

[0087] As an optional implementation provided in an embodiment of the present application, the adjustment module 404 is further configured to: when the available bandwidth of the disk array card is less than the second bandwidth threshold, determine a target disk whose load exceeds the load threshold; and reduce scheduling of the target disk.

[0088] As an optional implementation provided in an embodiment of the present application, the adjustment module 404 is specifically used to: when the bandwidth utilization is greater than or equal to the first utilization threshold, calculate the number of disks to be activated according to the disk scheduling policy; select the disks to be activated corresponding to the number of disks to be activated from multiple disks through the scheduling engine; and control the disks to be activated to enter the active state.

[0089] As an optional implementation provided in an embodiment of the present application, the adjustment module 404 is also used to: after controlling the disk to be activated to enter the active state, in response to a received write request, divide the data to be written into multiple data blocks; and allocate multiple data blocks to the disk to be activated in the active state based on preset rules.

[0090] As an optional implementation provided in an embodiment of the present application, the adjustment module 404 is also used to: when the bandwidth utilization is greater than or equal to the second utilization threshold, calculate the number of disks to be put into hibernation according to the disk scheduling strategy; wherein the second utilization threshold is greater than the first utilization threshold; select the disk to be put into hibernation according to the number of disks to be put into hibernation and the disk heat; and control the disk to be put into hibernation.

[0091] As an optional implementation provided in an embodiment of the present application, the adjustment module 404 is also used to: after controlling the disk to be dormant to enter a dormant state, forward the received write request to an active disk other than the disk to be dormant; and migrate the cold data of the disk to be dormant to the active disk.

[0092] As an optional implementation provided in an embodiment of the present application, the device also includes a processing module, which is used to: respond to received input / output requests, determine the processing order of the input / output requests; control the active disk to process the input / output requests according to the processing order; or, respond to received input / output requests, determine the stripe size of the disk array card corresponding to the input / output request; control the active disk to store data corresponding to the input / output request according to the stripe size.

[0093] For the description of the features in the embodiment corresponding to the disk scheduling device, please refer to the relevant description of the embodiment corresponding to the disk scheduling method, and no further details will be given here.

[0094] The embodiment of the present application also provides an electronic device, such as Figure 5 As shown, it includes a memory 501 and a processor 502. The memory 501 stores a computer program, and the processor 502 is configured to run the computer program to execute the steps in any of the above disk scheduling method embodiments.

[0095] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above disk scheduling method embodiments when running.

[0096] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0097] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above disk scheduling method embodiments are implemented.

[0098] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned disk scheduling method embodiments are implemented.

[0099] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0100] The above describes in detail the disk scheduling method, device, electronic device, storage medium, and program product provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core concept of the present application. It should be noted that for ordinary technicians in this technical field, without departing from the principles of the present application, various improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A disk scheduling method, characterized in that: include: Acquire a historical input / output access pattern of the disk array card, wherein the historical input / output access pattern includes access information of disk input / output requests within a historical time period; Determining a disk scheduling strategy based on the historical input / output access pattern and a pre-trained disk scheduling model; Monitoring the interface traffic of the disk array card, and calculating the bandwidth utilization of the disk array card according to the interface traffic; Adjusting the number of disks in a dormant state and an active state according to the bandwidth utilization and the disk scheduling policy; The method further includes: calculating the available bandwidth of the disk array card based on the interface traffic and the level mode of the disk array card; obtaining a total disk bandwidth requirement; determining a second bandwidth threshold based on the total disk bandwidth requirement and a second preset coefficient; and adjusting the level mode of the disk array card when the available bandwidth is less than the second bandwidth threshold.

2. The method according to claim 1, characterized in that The method further comprises: Determining a first bandwidth threshold according to the total disk bandwidth requirement and a first preset coefficient; wherein the first preset coefficient is greater than the second preset coefficient; generating an early warning strategy when the available bandwidth of the disk array card is less than the first bandwidth threshold; The warning strategy includes at least one of storing the warning event in the system log, generating a warning pop-up window, and pushing a warning notification.

3. The method according to claim 2, characterized in that Obtaining the total disk bandwidth requirement includes: Monitoring the input / output rates of the plurality of disks corresponding to the disk array card; Calculate the bandwidth requirements of each disk based on its input / output rate and number of input / output operations. The total disk bandwidth requirement is calculated based on the bandwidth requirements of the various disks.

4. The method according to claim 1, wherein The method further comprises: When the available bandwidth of the disk array card is less than the first bandwidth threshold and greater than or equal to the second bandwidth threshold, the stripe size of the disk array card is adjusted, or the interface channel of the disk array card is upgraded.

5. The method according to claim 4, characterized in that The method further comprises: When the available bandwidth of the disk array card is less than the second bandwidth threshold, determining a target disk whose load exceeds a load threshold; Reduce scheduling for the target disk.

6. The method according to claim 1, characterized in that The adjusting the number of disks in a dormant state and an active state according to the bandwidth utilization and the disk scheduling policy includes: When the bandwidth utilization is greater than or equal to a first utilization threshold, calculating the number of disks to be activated according to the disk scheduling policy; Selecting, by a scheduling engine, disks to be activated corresponding to the number of disks to be activated from a plurality of disks; Control the disk to be activated to enter an active state.

7. The method according to claim 6, characterized in that The method further comprises: After controlling the disk to be activated to enter an active state, in response to a received write request, dividing the data to be written into a plurality of data blocks; Allocate multiple data blocks to the disk to be activated that is in an active state based on a preset rule.

8. The method according to claim 6, characterized in that The method further comprises: When the bandwidth utilization is greater than or equal to a second utilization threshold, the number of disks to be put into hibernation is calculated according to the disk scheduling policy; wherein the second utilization threshold is greater than the first utilization threshold; Selecting a disk to be hibernated based on the number of disks to be hibernated and the disk temperatures; Control the hard disk to enter a dormant state.

9. The method according to claim 8, characterized in that The method further comprises: After controlling the disk to be dormant to enter a dormant state, forwarding the received write request to an active disk other than the disk to be dormant; Migrate the cold data of the disk to be dormant to the active disk.

10. The method according to claim 1, characterized in that The method further comprises: In response to received input / output requests, determine a processing order of the input / output requests; and control active disks to process the input / output requests according to the processing order; Alternatively, in response to the received input / output request, a stripe size of the disk array card corresponding to the input / output request is determined; and an active disk is controlled to store data corresponding to the input / output request according to the stripe size.

Citation Information

Patent Citations

  • Low-energy-consumption data cold magnetic storage method and device based on machine learning

    CN113190173A

  • Disk scheduling strategy setting method, system and equipment and storage medium

    CN118733266A