Resource scheduling method and device for storage equipment, equipment and storage medium
By obtaining the failure risk information and load characteristics of storage devices and dynamically adjusting resource allocation, the problem that traditional storage device resource scheduling strategies are unable to cope with complex scenarios is solved, and the stability of storage devices and resource utilization are improved.
Patent Information
- Application Number
- CN202511186800.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-22
AI Technical Summary
The resource scheduling strategies of traditional storage devices are difficult to cope with complex scenarios, resulting in performance and stability being affected and the inability to perform dynamic resource allocation.
By acquiring monitoring data related to storage device failure risks, using a preset risk prediction model to process the monitoring data to determine failure risk information, processing input and output trajectory data to obtain image features, and identifying load characteristics, the allocation of input and output path resources can be dynamically adjusted.
It achieves stable operation of storage devices and improved resource utilization, and ensures efficient operation of equipment by preventing failures in advance and allocating resources reasonably.
Smart Images

Figure CN120670177A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a resource scheduling method, apparatus, device, and storage medium for a storage device. Background Art
[0002] Traditional storage device resource scheduling relies on fixed policies, making it difficult to cope with complex scenarios. The inability to dynamically allocate resources when faced with varying loads and failure risks impacts the performance and stability of storage devices. Therefore, new resource scheduling solutions are urgently needed. Summary of the Invention
[0003] In view of the above problems, the present application provides a resource scheduling method, apparatus, device and storage medium for storage devices that improve resource utilization.
[0004] According to a first aspect of the present application, a resource scheduling method for a storage device is provided, comprising: obtaining monitoring data related to the failure risk of the storage device; processing the monitoring data using a preset risk prediction model to determine multiple failure risk information of the storage device; processing input and output trajectory data of the storage device to obtain image features of the input and output trajectory data, the input and output trajectory data representing the usage of physical paths and / or logical paths involved in the storage device during data transmission; identifying load features in the image features; and adjusting the allocation of input and output path resources of the storage device based on the load features and the multiple failure risk information, the input and output path resources including physical resources and / or logical resources used for data transmission.
[0005] A second aspect of the present application provides a resource scheduling apparatus for a storage device, comprising: an acquisition module for acquiring monitoring data related to the failure risk of the storage device; a processing module for processing the monitoring data using a preset risk prediction model to determine multiple failure risk information of the storage device; a feature processing module for processing the input and output trajectory data of the storage device to obtain image features of the input and output trajectory data; a feature recognition module for identifying load features in the image features; and a resource allocation module for adjusting the allocation of input and output path resources of the storage device based on the load features and multiple failure risk information.
[0006] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0007] The fourth aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0008] In an embodiment of the present application, monitoring data related to storage device failure risks is acquired and processed using a preset model to determine multiple pieces of failure risk information. Image features of the storage device input and output trajectory data are processed to identify load characteristics within the image features. Based on the load characteristics and failure risk information, the allocation of storage device input and output path resources (including physical and logical resources) is adjusted to prevent failures, ensure stable operation of the storage device, and rationally allocate resources to improve resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above contents and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:
[0010] Figure 1 Schematically illustrates an application scenario diagram of a resource scheduling method, apparatus, device, and storage medium for a storage device according to an embodiment of the present application;
[0011] Figure 2 A flowchart of a resource scheduling method for a storage device according to an embodiment of the present application is schematically shown;
[0012] Figure 3 Schematically shows a flow chart for determining second fault risk information according to an embodiment of the present application;
[0013] Figure 4 Schematically shows a flow chart for determining image features according to an embodiment of the present application;
[0014] Figure 5 A flowchart for determining input and output trajectories according to an embodiment of the present application is schematically shown;
[0015] Figure 6 Schematically shows a flow chart for determining load characteristics according to an embodiment of the present application;
[0016] Figure 7A A flowchart of adjusting input and output path resources according to an embodiment of the present application is schematically shown;
[0017] Figure 7B The following schematically shows a component architecture diagram of a resource scheduling method according to an embodiment of the present application;
[0018] Figure 7C A schematic diagram of a computing architecture layered diagram of a resource scheduling method according to an embodiment of the present application is shown;
[0019] Figure 7D The following schematically shows an architectural hierarchy diagram of a resource scheduling method according to an embodiment of the present application;
[0020] Figure 8 A block diagram schematically illustrates a structure of a resource scheduling apparatus for a storage device according to an embodiment of the present application; and
[0021] Figure 9 A block diagram of an electronic device suitable for implementing a resource scheduling method for a storage device according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0022] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.
[0023] The terms used herein are only for describing specific embodiments and are not intended to limit this application. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0024] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0025] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0026] Embodiments of the present application provide a resource scheduling method, apparatus, device, and storage medium for a storage device.
[0027] Figure 1 The application scenario diagram of the resource scheduling method, apparatus, device and storage medium for storage devices according to an embodiment of the present application is schematically shown.
[0028] like Figure 1As shown, the application scenario 100 is mainly composed of a control circuit 101 and a storage device 102, and the control circuit 101 is embedded in the storage device 102, and the two are closely combined and work together.
[0029] As the core carrier of data storage, the storage device 102 is responsible for storing various types of data. Its operating status directly affects the reliability and stability of data storage. The control circuit 101 plays the role of an intelligent manager, responsible for executing the resource scheduling method for the storage device.
[0030] During operation, the control circuit 101 first obtains monitoring data related to the failure risk of the storage device 102 , processes the monitoring data using a preset risk prediction model, and accurately determines multiple failure risk information of the storage device 102 .
[0031] The control circuit 101 processes the input and output trace data of the storage device 102 , which reflects the usage of the physical path and / or logical path of the storage device during data transmission, thereby obtaining image features and identifying load features therein.
[0032] The control circuit 101 dynamically adjusts the allocation of input and output path resources of the storage device 102, including physical resources and / or logical resources, based on load characteristics and multiple fault risk information, so as to improve resource utilization and ensure efficient and stable operation of the storage device 102.
[0033] The following will be based on Figure 1 The scene described by Figures 2 to 6 、 Figure 7A to Figure 7D A resource scheduling method for a storage device according to a disclosed embodiment is described in detail.
[0034] Figure 2 The flowchart of the resource scheduling method for a storage device according to an embodiment of the present application is schematically shown.
[0035] like Figure 2 As shown, the resource scheduling for storage devices in this embodiment includes operations S210 to S250.
[0036] In operation S210 , monitoring data related to storage device failure risk is acquired.
[0037] For example, storage device failure risks include hardware failure risk, performance failure risk, and software failure risk. Monitoring data includes device status data, performance data, and log data.
[0038] For example, device status data related to storage device failure risk is obtained, including storage device temperature, Self-Monitoring, Analysis and Reporting Technology (SMART) data, etc. Performance data related to storage device failure risk is also obtained, including input / output latency, throughput, etc.
[0039] In operation S220 , the monitoring data is processed using a preset risk prediction model to determine a plurality of failure risk information of the storage device.
[0040] In an embodiment of the present application, the preset risk prediction model can be trained based on historical monitoring data, and multiple fault risk information of storage devices in future cycles can be predicted by learning the data change patterns in the historical monitoring data. The fault risk information can represent faults including physical layer failures or logical layer failures.
[0041] In operation S230 , input and output trace data of the storage device are processed to obtain image features of the input and output trace data, where the input and output trace data represent usage of physical paths and / or logical paths involved in a data transmission process of the storage device.
[0042] In the embodiments of the present application, input / output trace data is a time-series record of the storage device's use of physical and logical paths during data transmission. Image features of the input / output trace data include timing diagrams, heat maps, and the like. Physical and / or logical path usage refers to the actual state and behavior of the storage device's physical and logical resources during data transmission. This information can be used to characterize the impact of SSD hardware performance and software management efficiency on the storage device.
[0043] In operation S240 , a load feature among the image features is identified.
[0044] In the embodiment of the present application, the load feature refers to an image feature that is directly related to the current load status of the storage device. The image feature can reflect key information such as resource usage, performance pressure or task complexity in the storage medium.
[0045] In operation S250 , allocation of input / output path resources of the storage device is adjusted according to the load characteristics and the plurality of failure risk information to improve resource utilization, where the input / output path resources include physical resources and / or logical resources for data transmission.
[0046] In this embodiment of the present application, the stability and reliability of the input and output paths are determined based on load characteristics and multiple fault risk information. When the input and output path load is high and the fault risk is low, incremental physical resources are allocated accordingly, such as increasing bandwidth or adding transmission channels.
[0047] In an embodiment of the present application, monitoring data related to storage device failure risks is acquired and processed using a preset model to determine multiple pieces of failure risk information. Image features of the storage device input and output trajectory data are processed to identify load characteristics within the image features. Based on the load characteristics and failure risk information, the allocation of storage device input and output path resources (including physical and logical resources) is adjusted to prevent failures, ensure stable operation of the storage device, and rationally allocate resources to improve resource utilization.
[0048] The following combination Figure 3 , describes the process of determining multiple failure risk information of storage devices.
[0049] Figure 3 The flowchart of determining the second fault risk information according to an embodiment of the present application is schematically shown.
[0050] like Figure 3 As shown, the above operation 220 may further include operations S301 and S302.
[0051] In operation S301, device status data is analyzed using a first preset risk prediction model to determine first failure risk information, where the first failure risk information is used to evaluate the reliability status of physical components of a storage device.
[0052] In an embodiment of the present application, the first preset risk prediction model is a data-driven analysis model, which may be a machine learning model or a deep learning model. The first preset risk prediction model may determine information related to physical device failures of the storage device in the device status data by analyzing the device status data, thereby determining the first failure risk information.
[0053] In operation S302 , the resource usage data is analyzed using a second preset risk prediction model to determine second failure risk information, where the second failure risk information is used to evaluate a stability state of resource scheduling of the storage device.
[0054] In an embodiment of the present application, the first preset risk prediction model is also based on a data-driven analysis model. The first preset risk prediction model can determine information related to the load balancing of the storage device in the device status data by analyzing the resource usage data, and then determine the second fault risk information.
[0055] In this embodiment of the present application, a first predetermined risk prediction model is used to analyze device status data to obtain first failure risk information, thereby assessing physical device reliability. A second predetermined risk prediction model is used to analyze resource usage data to obtain second failure risk information, thereby assessing resource scheduling stability. Obtaining failure risk information from these two key dimensions, device status and resource usage, enables a more comprehensive and accurate assessment of the overall status of storage devices.
[0056] The following describes the process of determining the first fault risk information.
[0057] In this embodiment of the present application, a convolutional layer is used to extract device status data to obtain device status features. A normalization layer is used to normalize the device status features to obtain normalized features. A fully connected layer is used to analyze the normalized features to obtain a risk score. The risk score is compared with a preset risk threshold to determine the first fault risk information.
[0058] For example, the convolution layer uses multiple convolution kernels to slide values across preprocessed device status data, extracting key device status features, such as data change trends and abnormal fluctuations. The normalization layer processes these device status features, eliminating dimensional and scale differences between them. This transforms the device status features into normalized features with a uniform range and distribution.
[0059] The fully connected layer analyzes the normalized features through weighted connections between neurons and activation function operations, outputting a risk score that represents the degree of fault risk. This risk score is compared with a preset risk threshold to obtain a comparison result, and the first fault risk is determined based on this comparison result.
[0060] In this embodiment, the convolutional, normalization, and fully connected layers analyze device status data to generate a risk score. Comparing the risk score with a preset risk threshold allows for timely and accurate determination of failure risk, providing a reliable basis for equipment maintenance and reducing equipment failure rates and operational costs.
[0061] The following describes the process of determining the second fault risk information.
[0062] In an embodiment of the present application, resource usage data is arranged in chronological order to form a performance change sequence. The fluctuation characteristics of the performance change sequence are extracted using a feature extraction layer, and the fluctuation characteristics are processed using a prediction layer to obtain a change trend of the performance change sequence. Based on the change trend, the risk level information corresponding to the resource usage data is determined as the second fault risk information.
[0063] For example, resource usage data from storage devices is collected at preset intervals and sorted in ascending or descending order based on the timestamp to construct a performance change sequence. In the feature extraction layer, a convolutional neural network is used to extract fluctuation features from the performance change sequence. These fluctuation features include the mean, variance, and frequency domain components of the performance change sequence.
[0064] The prediction layer uses the extracted fluctuation characteristics to predict the changing trend of the performance change sequence. The gradient of the changing trend is compared with the preset gradient threshold to determine the risk level as the second fault risk information.
[0065] In the embodiments of the present application, by extracting fluctuation characteristics and predicting change trends, the stability risk of resource scheduling can be accurately assessed, and bottlenecks or abnormal situations that may occur during resource use can be discovered in advance, so as to adjust the resource allocation strategy in time and ensure efficient use of resources.
[0066] The following describes the process of determining the third fault risk information.
[0067] In an embodiment of the present application, abnormal operation information is determined in read / write data, and error log information is determined in device log data. The error log information includes error event information generated during storage device self-test or operation. Based on the abnormal operation information and the error log information, third fault risk information is generated. The third fault risk information is used to assess the immediate fault risk status of the storage device during operation.
[0068] For example, a pre-set rule engine can be used to match abnormal operation patterns corresponding to read and write data, such as unauthorized access, excessive data block transfers, or frequent instruction retries, to identify and extract abnormal operation information. Regular expressions or semantic analysis can be used to parse device log data to identify error event records within the device log data.
[0069] In the embodiment of the present application, by setting clear conditions to generate the third fault risk information, the immediate fault risk situation during the operation of the storage device can be accurately judged, thereby avoiding false alarms and missed alarms, and providing an accurate and reliable basis for fault handling.
[0070] The following describes a process for generating third fault risk information based on abnormal operation information and error record information.
[0071] In an embodiment of the present application, when the abnormal operation information is abnormal address access and the number of errors in the error record information exceeds a number threshold within a predetermined time window, third fault risk information is generated.
[0072] For example, when the abnormal operation information is determined to be an abnormal address access, error records are extracted from the device log within a preset time window and the number is counted. If the number of error records within the preset time window exceeds a preset threshold, an immediate failure risk is determined and third failure risk information is generated according to preset rules.
[0073] In an embodiment of the present application, when the abnormal operation information is a read / write data error and the device data record in the device log data exceeds the data safety threshold, third fault risk information is generated.
[0074] For example, if the abnormal operation information is determined to be a read or write data error, the device data records related to the read or write data error in the device log data are counted. When the number of device data records exceeds the data safety threshold, a fault risk is determined to exist, and third fault risk information containing the risk level and possible cause is generated according to preset rules.
[0075] In an embodiment of the present application, third fault risk information is accurately and real-timely generated based on abnormal operation information and error record information, which can help determine the immediate failure risk during the operation of the storage device, avoid false alarms and missed alarms, and provide an accurate and reliable basis for fault handling.
[0076] The following describes the process of determining the image features of input and output trajectory data.
[0077] Figure 4 The flowchart of determining image features according to an embodiment of the present application is schematically shown.
[0078] like Figure 4 As shown, the above operation S230 may further include operations S401 to S404.
[0079] In operation S401, the logical address access data is arranged in time sequence to obtain a time feature sequence.
[0080] In an embodiment of the present application, during operation of a storage device, logical address access data is collected and timestamps are added to the logical address access data. The collected logical address access data is sorted according to the order of the timestamps, and the sorted data is sequentially stored in a preset sequence structure, forming a time feature sequence containing time dimension information.
[0081] In operation S402, the logical address access data is arranged according to the address offset to obtain an address feature sequence.
[0082] In an embodiment of the present application, during operation of a storage device, logical address access data is collected and address offsets are extracted from the logical address access data. Based on the magnitude of the address offset information, the logical address access data is sorted in ascending or descending order. The sorted logical address access data is then sequentially stored in a preset sequence container to obtain an address signature sequence.
[0083] In operation S403 , a time depth feature of the protocol command queue data is obtained according to the time sequence and the queue depth of the protocol command queue data.
[0084] In an embodiment of the present application, during operation of a storage device, protocol command queue data is collected, and the timestamps and corresponding queue depth values of the protocol command queue data are recorded. The protocol command queue data is sorted in timestamp order to construct a time series, and the queue depth value corresponding to each time point is used as a feature component to obtain a time depth feature of the protocol command queue data.
[0085] In operation S404 , the time feature sequence, the address feature sequence, and the time depth feature are converted to obtain image features of the input and output trajectory data.
[0086] In an embodiment of the present application, the input and output trajectory data are converted into image features through a specific arrangement and processing method, which can provide a more intuitive and easy-to-process data form for subsequent load feature analysis, provide strong support for reasonable resource allocation, and improve the adaptability of storage devices to different load conditions.
[0087] The following describes the process of obtaining image features of input and output trajectory data in conjunction with specific embodiments.
[0088] Figure 5 The flowchart of determining input and output trajectories according to an embodiment of the present application is schematically shown.
[0089] like Figure 5 As shown, operation S404 may further include operations S501 to S503.
[0090] In operation S501 , the position information and corresponding feature value of the time depth feature are determined to obtain a feature map. The position information is determined according to the moment of the time feature sequence and the address offset of the address feature sequence. The corresponding feature value is positively correlated with the queue depth.
[0091] In the embodiment of the present application, the time information corresponding to each data point of the time feature sequence is extracted, and the address offset is extracted from the address feature sequence. The time information and the address offset are combined to obtain position information in the form of two-dimensional coordinates, which is used to represent the position in the identification feature map.
[0092] The queue depth value from the time-depth feature is obtained and converted into a corresponding feature value based on a preset positive correlation mapping rule. The mapping rule ensures that the larger the queue depth, the larger the corresponding feature value. The corresponding feature value is filled into the corresponding position using the position information as the index to generate a feature map. The feature map can intuitively reflect the spatiotemporal characteristics of the protocol command queue.
[0093] In operation S502 , a timing waveform is determined based on a relationship between characteristic values and time, where the timing waveform is used to represent a load variation characteristic of the protocol command queue.
[0094] In this embodiment, the eigenvalues corresponding to each moment in the feature graph are extracted in chronological order to form a eigenvalue sequence. A coordinate system is established with the moment as the horizontal axis and the eigenvalue as the vertical axis. The eigenvalues in the eigenvalue sequence are annotated as discrete points in the coordinate system. A curve fitting method is used to smooth the discrete points to generate a continuous curve. This generated continuous curve is used as a timing waveform to intuitively display the characteristics of the protocol command queue load changing over time.
[0095] In operation S503 , the input and output traces are determined by connecting the position information corresponding to adjacent moments in the feature graph, where the input and output traces are used to represent the access pattern characteristics of the logical address.
[0096] In this embodiment, the position information of adjacent moments is extracted from the feature graph. The position information is determined by the offset between the moments in the time feature sequence and the address feature sequence. The position information of adjacent moments is sequentially connected in chronological order to form a continuous line segment sequence. The path formed by the line segment sequence is used as the input and output trajectory.
[0097] In the embodiment of the present application, by obtaining the image features of the input and output trajectory data, rich and effective information is provided for analyzing the load characteristics, which is conducive to formulating a more scientific and reasonable resource allocation strategy.
[0098] The following describes the process of determining the load feature in the image feature in conjunction with specific embodiments.
[0099] Figure 6 The flowchart of determining load characteristics according to an embodiment of the present application is schematically shown.
[0100] like Figure 6 As shown, operation S240 may further include operations S601 to S603.
[0101] In operation S601 , input and output traces are identified, and access features corresponding to the input and output traces are obtained. The access features include continuous access features and random access features.
[0102] In an embodiment of the present application, the difference of the address offsets of adjacent accesses in the input and output traces is calculated, and the distribution of the difference is statistically analyzed. When the difference is concentrated in a smaller numerical range and shows regular fluctuations, it is determined to be a continuous access feature, indicating the existence of sequential address access behavior. When the difference distribution is scattered and there is no obvious pattern, it is determined to be a random access feature.
[0103] In operation S602 , a time series waveform is identified to obtain attribute features corresponding to the time series waveform graph, where the attribute features include periodic fluctuation features and sudden peak features.
[0104] In this embodiment, a frequency domain analysis is performed on the timing waveform, and its spectral distribution is obtained through Fourier transform. If a significant main frequency component is present, it is determined to have periodic fluctuation characteristics, reflecting regular changes in the protocol command queue load. A sliding window algorithm is used to detect instantaneous peaks in the waveform. When the peak exceeds a preset threshold and lasts for less than a set duration, it is determined to be a sudden peak characteristic, indicating a momentary increase in queue load.
[0105] In operation S603 , a load feature in the image feature is determined according to the access feature and the attribute feature, where the load feature includes at least one of a sequential access load, a random access load, and a mixed access load.
[0106] The following describes the process of determining the load feature in the image feature in conjunction with specific embodiments.
[0107] In the embodiments of the present application, by carefully identifying the input and output traces and timing waveforms in the image features, the load characteristics of the storage device can be accurately determined, providing targeted guidance for resource allocation and improving the performance and efficiency of the storage device under different load scenarios.
[0108] The following describes the process of determining the load feature in the image feature in conjunction with specific embodiments.
[0109] The above operation 603 may also include determining a combination of access characteristics and attribute characteristics; when the combination includes continuous access characteristics and periodic fluctuation characteristics, determining the load characteristics as a sequential access load; when the combination includes random access characteristics and sudden peak characteristics, determining the load characteristics as a random access load; when the combination includes multiple access characteristics and attribute characteristics, determining the load characteristics as a mixed access load.
[0110] For example, a rule base for combining access features and attribute features is pre-built, which specifies the load feature determination conditions corresponding to different combinations. The identified access features and attribute features are matched and analyzed. If both continuous access features and periodic fluctuation features are detected, the load feature is determined to be a sequential access load based on the rule base, reflecting the regular sequential access pattern of the protocol command queue load. If both random access features and sudden peak features are present, the load is determined to be a random access load, reflecting the discreteness and suddenness of the load. When multiple combinations of access features and attribute features are present, the load feature is determined to be a mixed access load based on the comprehensive determination logic of the rule base.
[0111] In the embodiment of the present application, the load feature type is determined by accurately determining the combination of access features and attribute features, which can greatly improve the accuracy and pertinence of load feature identification.
[0112] The following describes a process of adjusting allocation of input and output path resources of a storage device in conjunction with specific embodiments.
[0113] Figure 7A The flowchart of adjusting input and output path resources according to an embodiment of the present application is schematically shown.
[0114] like Figure 7A As shown, operation S250 may further include operations S701 to S703.
[0115] In operation S701 , when the load characteristic is characterized as a sequential access load and the risk level represented by the plurality of fault risk information is lower than a first level threshold, physical channel allocation and a read-ahead mechanism in input and output path resources of a storage device are adjusted.
[0116] For example, when the load characteristics are characterized as sequential access loads and the risk level represented by multiple fault risk information is lower than the first level threshold, continuous logical address blocks are allocated to multiple physical channels, and the physical channels independently process the allocated logical address blocks; based on the continuity characteristics of the sequential access load, pre-read data is dynamically calculated and cached in the cache area.
[0117] In operation S702 , when the load characteristic is a random access load and the risk level represented by the plurality of fault risk information is higher than a first level threshold and lower than a second level threshold, queue allocation and bandwidth allocation in input and output path resources of the storage device are adjusted.
[0118] For example, if the load is characterized by random access and the risk level indicated by multiple fault risk information is higher than a first threshold and lower than a second threshold, a multi-level buffer queue is established for the random access load. The multi-level buffer queue includes a real-time queue, a normal queue, and a background queue. Based on the data of input and output requests, the priority of the queues in the multi-level buffer queue is adjusted, and transmission bandwidth resources are allocated based on the load of the multi-level buffer queue.
[0119] In operation S703 , when the load characteristic is a mixed access load and the risk levels represented by the multiple fault risk information are higher than the second level threshold and lower than the third level threshold, path switching and address mapping in the input and output path resources of the storage device are adjusted.
[0120] For example, when the load characteristics are characterized as a mixed access load and the risk level represented by multiple fault risk information is higher than the second level threshold and lower than the third level threshold, the current transmission path is maintained for the sequential access load and dynamic path selection is enabled for the random access load; the address mapping relationship is reconstructed based on the risk level, and high-risk data is dispersedly mapped to different physical storage areas.
[0121] In the embodiment of the present application, targeted resource allocation adjustments are made according to different load characteristics and failure risk levels, which can give full play to the advantages of storage device resources and achieve optimal resource configuration.
[0122] The following describes the process of determining the channel allocation and bandwidth limitation strategy for input and output trajectory data in conjunction with specific embodiments.
[0123] Before operation S250 , a weight coefficient for storage resource allocation may be determined based on the risk level of the fault risk information. The weight coefficient is used to represent the channel allocation priority of each of the multiple storage areas. A channel allocation and bandwidth limitation strategy for the input and output trace data may be determined based on the weight coefficient and the resource requirement assessment value of the input and output trace data.
[0124] In an embodiment of the present application, a mapping table of fault risk levels and weight coefficients is pre-set, and the higher the risk level, the lower the corresponding weight coefficient. The weight coefficient is determined by looking up the table based on the fault risk information of each storage area, and the weight coefficient reflects the priority of channel allocation. The input and output trajectory data are evaluated for resource requirements in combination with data volume, access frequency, etc., to obtain a resource demand assessment value. The weight coefficient and the resource demand assessment value are comprehensively calculated, and the storage area with a large weight coefficient corresponds to the channel for priority allocation, and the bandwidth limit is reasonably set for each channel based on the demand assessment value.
[0125] For example, a weighted calculation is performed on the weight coefficient and the resource demand assessment value to generate a storage resource allocation score. For target storage areas with a storage resource allocation score above a first threshold, independent physical channels are allocated to the target storage areas. For target storage areas with a storage resource allocation score between the first and second thresholds, shared channels and a dynamic bandwidth adjustment mechanism are allocated to the target storage areas. For storage areas with a storage resource allocation score below the second threshold, channel multiplexing and fixed bandwidth restrictions are set for the target storage areas.
[0126] In the embodiments of the present application, reasonable allocation of storage resources is achieved, and bandwidth usage is reasonably restricted to avoid resource waste and excessive competition.
[0127] The following combination Figure 7B and specific embodiments, and introduces the component architecture of the resource scheduling method.
[0128] Figure 7B The component architecture diagram of the resource scheduling method according to an embodiment of the present application is schematically shown.
[0129] In the embodiments of this application, Figure 7B As shown, the storage device includes a central processing unit, an intelligent heterogeneous processor, a memory controller, a flash memory interface, flash memory particles, a host interface, solid state storage and a data processing engine.
[0130] The data flow within a storage device is as follows: The host interface receives host data commands and transmits them to the central processing unit (CPU), which interacts with the memory controller for temporary data storage. The CPU coordinates overall control, and the control logic manages data reading and writing of flash memory particles through the flash interface.
[0131] Intelligent heterogeneous processors retrieve data from memory for processing, and the results can be fed back or stored. A memory controller manages the memory cache, interacting with the CPU, host interface, and intelligent heterogeneous processors as needed, writing data requiring long-term storage to the flash memory chips via the flash memory interface.
[0132] Firmware storage provides control programs and parameters for the CPU. The data processing engine encrypts and compresses data before writing, and decrypts, decompresses, and corrects errors when reading, ensuring data security and reliability. All components work together to achieve efficient data flow.
[0133] The following combination Figure 7C And specific embodiments are introduced to introduce the computing architecture layering of the resource scheduling method.
[0134] Figure 7C A layered diagram of a computing architecture of a resource scheduling method according to an embodiment of the present application is schematically shown.
[0135] In the embodiments of this application, Figure 7CAs shown in the figure, the computing architecture layer of the resource scheduling method includes the host application layer, the system software layer, and the hardware abstraction layer. The host application layer deploys a deep learning framework, leveraging its powerful algorithm and model processing capabilities to provide intelligent decision-making support for resource scheduling, such as predicting storage load and optimizing resource allocation strategies. The system software layer includes the driver adaptation layer, the scheduling management layer, and the intelligent operating environment. The driver adaptation layer is used to adapt the hardware driver to the system. The scheduling management layer is responsible for coordinating resource scheduling, and the intelligent operating environment provides basic support for the operation of the entire system. The hardware abstraction layer consists of a controller, a storage interface, and an intelligent acceleration unit. The controller coordinates hardware operations, the storage interface ensures data transmission, and the intelligent acceleration unit improves data processing speed, collectively providing hardware support for the upper layer.
[0136] The following combination Figure 7D And specific embodiments are introduced to introduce the architectural level of the resource scheduling method.
[0137] Figure 7D The following schematically shows an architectural hierarchy diagram of a resource scheduling method according to an embodiment of the present application.
[0138] In the embodiments of this application, Figure 7D As shown in the figure, the architectural hierarchy of the resource scheduling method includes the application layer, the software development kit (SDK) layer, the system software layer, and the hardware abstraction layer. Within the application layer, intelligent applications are at the top of the architecture. Based on the resource scheduling capabilities of storage devices, they implement specific application functions such as intelligent data management and efficient storage services, providing users with convenient storage solutions. The SDK layer includes a model conversion tool, a promotion interface library, and a performance analysis tool. The model conversion tool facilitates the adaptation of different models within the storage system. The promotion interface library provides a standard interface for easy application layer invocation. The performance analysis tool is used to evaluate and optimize storage system performance.
[0139] In the system software layer, the intelligent operating environment provides operational support for the entire system. The heterogeneous scheduling framework rationally allocates storage resources based on information such as load characteristics and failure risks. The device driver layer directly interacts with the hardware abstraction layer to drive the normal operation of hardware devices.
[0140] The hardware of the storage device is abstracted in the hardware abstraction layer, so that the upper-layer software does not need to pay attention to hardware details, improving the compatibility and scalability of the system.
[0141] Based on the above resource scheduling method for storage devices, the present application also provides a resource scheduling device for storage devices. Figure 8 The device is described in detail.
[0142] Figure 8The structural block diagram of the resource scheduling apparatus for storage devices according to an embodiment of the present application is schematically shown.
[0143] like Figure 8 As shown, the resource scheduling apparatus 800 for storage devices of this embodiment includes an acquisition module 810 , a processing module 820 , a feature processing module 830 , a feature identification module 840 and a resource allocation module 850 .
[0144] The acquisition module 810 is used to acquire monitoring data related to the storage device failure risk. In one embodiment, the acquisition module 810 can be used to perform the operation S210 described above, which will not be repeated here.
[0145] The processing module 820 is used to process the monitoring data using a preset risk prediction model to determine multiple failure risk information of the storage device. In one embodiment, the processing module 820 can be used to perform the operation S220 described above, which will not be repeated here.
[0146] The feature processing module 830 is used to process the input and output trajectory data of the storage device to obtain image features of the input and output trajectory data. In one embodiment, the feature processing module 830 can be used to perform the operation S230 described above, which will not be repeated here.
[0147] The feature recognition module 840 is used to recognize the load feature in the image feature. In one embodiment, the feature recognition module 840 can be used to perform the operation S240 described above, which will not be repeated here.
[0148] The resource allocation module 850 is used to adjust the allocation of input and output path resources of the storage device according to the load characteristics and multiple fault risk information. In one embodiment, the resource allocation module 850 can be used to perform the operation S250 described above, which will not be repeated here.
[0149] According to embodiments of the present application, any multiple modules among the acquisition module 810, processing module 820, feature processing module 830, feature identification module 840, and resource allocation module 850 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present application, at least one of the acquisition module 810, processing module 820, feature processing module 830, feature identification module 840, and resource allocation module 850 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the acquisition module 810, the processing module 820, the feature processing module 830, the feature identification module 840 and the resource allocation module 850 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.
[0150] Figure 9 A block diagram of an electronic device suitable for implementing a resource scheduling method for a storage device according to an embodiment of the present application is schematically shown.
[0151] like Figure 9 As shown, an electronic device 900 according to an embodiment of the present application includes a processor 901, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 902 or programs loaded from a storage unit 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present application.
[0152] Various programs and data required for the operation of the electronic device 900 are stored in the RAM 903. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 performs various operations of the method flow according to the embodiment of the present application by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the programs may also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 may also perform various operations of the method flow according to the embodiment of the present application by executing the programs stored in one or more memories.
[0153] According to an embodiment of the present application, electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to bus 904. Electronic device 900 may also include one or more of the following components connected to I / O interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 908 including a hard disk; and a communication section 909 including a network interface card such as a LAN card or modem. Communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to I / O interface 905 as needed. Removable media 911, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 910 as needed, so that computer programs read from the removable media can be installed into storage section 908 as needed.
[0154] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of this application is implemented.
[0155] According to an embodiment of the present application, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer-readable storage medium may include the ROM 902 and / or RAM 903 described above and / or one or more memories other than ROM 902 and RAM 903.
[0156] The embodiments of the present application also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided in the embodiments of the present application.
[0157] The computer program executes the above functions defined in the system / device of the embodiment of the present application when the processor 901 executes the computer program. According to the embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0158] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 909, and / or installed from a removable medium 911. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0159] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from a removable medium 911. When the computer program is executed by the processor 901, the above-mentioned functions defined in the system of the embodiment of the present application are performed. According to the embodiment of the present application, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0160] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0162] It will be understood by those skilled in the art that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.
[0163] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present application, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.
Claims
1. A resource scheduling method for a storage device, characterized in that: The method comprises: Obtain monitoring data related to storage device failure risks; Processing the monitoring data using a preset risk prediction model to determine multiple failure risk information of the storage device; processing input and output trace data of the storage device to obtain image features of the input and output trace data, wherein the input and output trace data represents usage of physical paths and / or logical paths involved in a data transmission process of the storage device; identifying a load feature in the image features; According to the load characteristics and multiple pieces of failure risk information, allocation of input and output path resources of the storage device is adjusted, where the input and output path resources include physical resources and / or logical resources for data transmission.
2. The method according to claim 1, characterized in that The monitoring data includes equipment status data and resource usage data, and the preset risk prediction model includes a first preset risk prediction model and a second preset risk prediction model; The using a preset risk prediction model to process the monitoring data to determine multiple pieces of failure risk information of the storage device includes: Analyzing the device status data using the first preset risk prediction model to determine first failure risk information, where the first failure risk information is used to evaluate the reliability status of a physical component of the storage device; The resource usage data is analyzed using the second preset risk prediction model to determine second failure risk information, where the second failure risk information is used to evaluate a stability state of resource scheduling of the storage device.
3. The method according to claim 2, characterized in that The first preset risk prediction model includes a convolution layer, a normalization layer, and a fully connected layer; and analyzing the device status information using the first preset risk prediction model to determine the first fault risk information includes: Extracting the device status data using the convolutional layer to obtain device status features; Normalizing the device state features using the normalization layer to obtain normalized features; Analyzing the normalized features using the fully connected layer to obtain a risk score; The risk score is compared with a preset risk threshold to determine the first fault risk information.
4. The method according to claim 2, characterized in that The second preset risk prediction model includes a feature extraction layer and a prediction layer, and the analyzing the resource usage data using the second preset risk prediction model to determine the second fault risk information includes: Arrange the resource usage data in chronological order to form a performance change sequence; extracting the fluctuation characteristics of the performance change sequence using the feature extraction layer; Processing the fluctuation characteristics using the prediction layer to obtain a change trend of the performance change sequence; According to the change trend, risk level information corresponding to the resource usage data is determined as the second fault risk information.
5. The method according to claim 2, characterized in that The monitoring data includes read-write data and device log data, and the method further includes: Determining abnormal operation information in the read / write data; Determining error record information in the device log data, the error record information including error event information generated by a self-test or runtime of the storage device; Third failure risk information is generated according to the abnormal operation information and the error record information. The third failure risk information is used to evaluate an immediate failure risk state of the storage device during operation.
6. The method according to claim 5, characterized in that The abnormal operation information includes abnormal address access and read / write data errors; and generating third fault risk information according to the abnormal operation information and error record information includes: generating the third fault risk information when the abnormal operation information is abnormal address access and the number of errors in the error record information exceeds a number threshold within a predetermined time window; When the abnormal operation information is a read / write data error and the device data record in the device log data exceeds a data safety threshold, the third fault risk information is generated.
7. The method according to claim 1, characterized in that The input and output trace data includes logical address access data and protocol command queue data; and processing the input and output trace data of the storage device to obtain image features of the input and output trace data includes: Arranging the logical address access data in chronological order to obtain a time feature sequence; Arranging the logical address access data according to the address offset to obtain an address feature sequence; Obtaining a time depth feature of the protocol command queue data according to the time sequence and the queue depth of the protocol command queue data; The time feature sequence, the address feature sequence and the time depth feature are converted to obtain image features of the input and output trajectory data.
8. The method according to claim 7, characterized in that The image features include input and output trajectories and time series waveforms; the time feature sequence, address feature sequence and time depth feature are converted to obtain image features corresponding to the input and output trajectory data, including: Determine the position information and corresponding feature value of the time depth feature to obtain a feature map; the position information is determined according to the moment of the time feature sequence and the address offset of the address feature sequence, and the corresponding feature value is positively correlated with the queue depth; Determining the timing waveform based on a relationship between the characteristic value and time, wherein the timing waveform is used to represent a load change characteristic of the protocol command queue; The input and output traces are determined by connecting position information corresponding to adjacent moments in the feature graph, and the input and output traces are used to represent access pattern characteristics of the logical address.
9. The method according to claim 1, characterized in that The image features include input and output traces and timing waveforms, and identifying load features in the image features includes: Identifying the input and output traces, and obtaining access features corresponding to the input and output traces, wherein the access features include continuous access features and random access features; Identifying the time series waveform to obtain attribute features corresponding to the time series waveform graph, wherein the attribute features include periodic fluctuation features and sudden peak features; A load feature in the image feature is determined according to the access feature and the attribute feature, where the load feature includes at least one of a sequential access load, a random access load, and a mixed access load.
10. The method according to claim 9, characterized in that The determining, based on the access feature and the attribute feature, the load feature in the image feature includes: Determining a combination of the access characteristics and the attribute characteristics; In a case where the combination includes the continuous access feature and the periodic fluctuation feature, determining that the load feature is the sequential access load; In a case where the combination includes the random access feature and the sudden peak feature, determining that the load feature is a random access load; In a case where the combination includes a plurality of the access characteristics and attribute characteristics, the load characteristic is determined to be a mixed access load.
11. The method according to claim 1, wherein The adjusting allocation of input and output path resources of the storage device according to the load characteristics and the plurality of failure risk information includes: When the load characteristic is characterized as a sequential access load and the risk level represented by the multiple fault risk information is lower than a first level threshold, adjusting the physical channel allocation and the pre-reading mechanism in the input and output path resources of the storage device; When the load characteristic is characterized as a random access load and the risk levels represented by the multiple fault risk information are higher than the first level threshold and lower than the second level threshold, adjusting the queue allocation and bandwidth allocation in the input and output path resources of the storage device; When the load characteristic is a mixed access load and the risk levels represented by the multiple fault risk information are higher than the second level threshold and lower than the third level threshold, path switching and address mapping in the input and output path resources of the storage device are adjusted.
12. The method according to claim 1, characterized in that Before adjusting the allocation of input and output path resources of the storage device according to the load characteristics and the plurality of failure risk information, the method further includes: Determining a weight coefficient for storage resource allocation according to the risk level of the fault risk information, wherein the weight coefficient is used to represent a channel allocation priority for each of the plurality of storage areas; According to the weight coefficient and the resource requirement evaluation value of the input and output trajectory data, a channel allocation and bandwidth limitation strategy for the input and output trajectory data is determined.
13. A resource scheduling device for a storage device, characterized in that: The device comprises: An acquisition module, used to acquire monitoring data related to storage device failure risks; a processing module, configured to process the monitoring data using a preset risk prediction model to determine multiple pieces of failure risk information of the storage device; a feature processing module, configured to process the input and output trajectory data of the storage device to obtain image features of the input and output trajectory data; A feature recognition module, configured to recognize a load feature in the image feature; The resource allocation module is used to adjust the allocation of input and output path resources of the storage device according to the load characteristics and multiple failure risk information.
14. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 12.
15. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Memory management method and system, desktop computer and computer storage medium
CN118312320A
Method and device for adjusting configuration of storage system, storage medium and electronic equipment
CN119376644A
Resource information processing method and device, equipment, medium and program product
CN119415013A
Dynamic monitoring and early warning method for storage space of data center
CN120045421A
Address space mapping
US20250086107A1
Cited By
Data reconstruction method and electronic equipment
CN120929298A