Resource scheduling method and device for storage device, equipment and storage medium
By acquiring fault risk information and load characteristics of storage devices, and dynamically adjusting input and output path resources, the performance and stability issues of traditional storage device resource scheduling strategies in complex scenarios are solved, and the efficient and stable operation of the devices is achieved.
Patent Information
- Application Number
- CN202511186800.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Traditional storage devices' resource scheduling strategies struggle to cope with complex scenarios and cannot dynamically allocate resources, impacting performance and stability.
By acquiring monitoring data related to storage device failure risks, processing the monitoring data using a preset risk prediction model to determine failure risk information, processing input and output trajectory data to obtain image features, identifying load characteristics, and adjusting the allocation of input and output path resources based on load characteristics and failure risk information.
It achieves stable operation of storage devices and improves resource utilization, ensuring efficient operation of devices by preventing failures in advance.
Smart Images

Figure CN120670177B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, in particular to a resource scheduling method and device for a storage device, and a storage medium. BACKGROUND
[0002] The resource scheduling of a traditional storage device relies on a fixed strategy, which is difficult to cope with complex scenarios. When facing different loads and fault risks, dynamic resource allocation cannot be performed, thereby affecting the performance and stability of the storage device, and therefore a new resource scheduling scheme is urgently needed. SUMMARY
[0003] In view of the above problems, the present application provides a resource scheduling method and device for a storage device, which improves the utilization rate of resources.
[0004] According to a first aspect of the present application, a resource scheduling method for a storage device is provided, comprising: obtaining monitoring data related to a fault risk of the storage device; processing the monitoring data by using a preset risk prediction model to determine a plurality of fault risk information of the storage device; processing input-output trajectory data of the storage device to obtain image features of the input-output trajectory data, the input-output trajectory data representing a usage of a physical path and / or a logical path involved in a data transmission process of the storage device; identifying load features in the image features; and adjusting allocation of input-output path resources of the storage device according to the load features and the plurality of fault risk information, the input-output path resources including physical resources and / or logical resources used for data transmission.
[0005] A second aspect of the present application provides a resource scheduling device for a storage device, comprising: an obtaining module configured to obtain monitoring data related to a fault risk of the storage device; a processing module configured to process the monitoring data by using a preset risk prediction model to determine a plurality of fault risk information of the storage device; a feature processing module configured to process input-output trajectory data of the storage device to obtain image features of the input-output trajectory data; a feature identification module configured to identify load features in the image features; and a resource allocation module configured to adjust allocation of input-output path resources of the storage device according to the load features and the plurality of fault risk information.
[0006] A third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0007] A fourth aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions are executed by a processor to implement the steps of the above method.
[0008] In the embodiment of the present application, the storage device failure risk related monitoring data is acquired, and a plurality of failure risk information is determined by processing the data with a preset model. By processing the image features of the storage device input and output track data, the load features in the image features are identified. And according to the load features and the failure risk information, the storage device input and output path resource (including physical and logical resources) allocation is adjusted, the failure is prevented in advance, the stable operation of the storage device is ensured, and the resource utilization is improved by reasonable allocation of resources. BRIEF DESCRIPTION OF DRAWINGS
[0009] The above and other objects, features and advantages of the present application will become more apparent from the following description of the embodiments of the present application taken with reference to the accompanying drawings, in which:
[0010] Figure 1 An application scenario diagram of a resource scheduling method, device and equipment and storage medium for a storage device according to an embodiment of the present application is schematically shown;
[0011] Figure 2 A flowchart of a resource scheduling method for a storage device according to an embodiment of the present application is schematically shown;
[0012] Figure 3 A flowchart of determining second failure risk information according to an embodiment of the present application is schematically shown;
[0013] Figure 4 A flowchart of determining image features according to an embodiment of the present application is schematically shown;
[0014] Figure 5 A flowchart of determining input and output track according to an embodiment of the present application is schematically shown;
[0015] Figure 6 A flowchart of determining load features according to an embodiment of the present application is schematically shown;
[0016] Figure 7A A flowchart of adjusting input and output path resources according to an embodiment of the present application is schematically shown;
[0017] Figure 7B A component architecture diagram of a resource scheduling method according to an embodiment of the present application is schematically shown;
[0018] Figure 7C A computing architecture hierarchical diagram of a resource scheduling method according to an embodiment of the present application is schematically shown;
[0019] Figure 7D An architecture hierarchical diagram of a resource scheduling method according to an embodiment of the present application is schematically shown;
[0020] Figure 8 a structural block diagram of a resource scheduling apparatus for a storage device according to an embodiment of the present application is shown schematically; and
[0021] Figure 9 a block diagram of an electronic device adapted to implement a resource scheduling method for a storage device according to an embodiment of the present application is shown schematically. DETAILED DESCRIPTION
[0022] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It is to be understood, however, that the description is merely exemplary of the present application, and is not intended to limit the scope of the present application. In the following detailed description of the embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring aspects of the present application.
[0023] The terms used herein are merely used to describe specific embodiments, and are not intended to limit the present application. The terms "include" and "have" and the like used herein indicate the presence of the described features, steps, operations, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, or components.
[0024] All terms used herein, including technical and scientific terms, have the same meanings as those generally understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having meanings consistent with the context of the present description, and should not be interpreted in an idealized or overly formal manner.
[0025] In the case of using expressions similar to "at least one of A, B, and C, etc.", it is generally to be interpreted as including one or more of the same. For example, "a system having at least one of A, B, and C" should be interpreted as including a system having A alone, a system having B alone, a system having C alone, a system having A and B together, a system having A and C together, a system having B and C together, and / or a system having A, B, and C together, etc.
[0026] Embodiments of the present application provide a resource scheduling method, apparatus, device, and storage medium for a storage device.
[0027] Figure 1 An application scenario diagram of a resource scheduling method, apparatus, device, and storage medium for a storage device according to an embodiment of the present application is shown schematically.
[0028] As Figure 1As shown, the application scenario 100 mainly consists of a control circuit 101 and a storage device 102, and the control circuit 101 is embedded in the storage device 102, and the two are closely combined and work cooperatively.
[0029] The storage device 102, as a core carrier of data storage, undertakes the storage task of various data, and its running state directly affects the reliability and stability of data storage. The control circuit 101 plays the role of an intelligent manager, responsible for executing the resource scheduling method for the storage device.
[0030] In the running process, the control circuit 101 first acquires monitoring data related to the failure risk of the storage device 102, processes the monitoring data by using a preset risk prediction model, and accurately determines multiple failure risk information of the storage device 102.
[0031] The control circuit 101 processes input / output trajectory data of the storage device 102, the input / output trajectory data reflecting the use condition of the physical path and / or the logical path of the storage device in the data transmission process, and then obtains image features and identifies load features therein.
[0032] The control circuit 101 dynamically adjusts the allocation of input / output path resources, including physical resources and / or logical resources, of the storage device 102 according to the load features and the multiple failure risk information, so as to improve the utilization rate of the resources and guarantee the efficient and stable operation of the storage device 102.
[0033] The following will describe the resource scheduling method for the storage device of the disclosed embodiment based on the scenario described above. Figure 1 The resource scheduling method for the storage device of the disclosed embodiment is described in detail through Figures 2-6 , Figures 7A-7D
[0034] Figure 2 A flowchart of the resource scheduling method for the storage device according to the embodiment of the application is schematically shown.
[0035] As shown in Figure 2 , the resource scheduling method for the storage device of the embodiment includes operations S210-S250.
[0036] In operation S210, monitoring data related to the failure risk of the storage device is acquired.
[0037] For example, the failure risk of the storage device includes hardware failure risk, performance failure risk, software failure risk, etc. The monitoring data includes device state data, performance data, log data, etc.
[0038] For example, the device state data related to the storage device failure risk is acquired, including the temperature of the storage device, self-monitoring, analysis and reporting technology (SMART) data, and the like. The performance data related to the storage device failure risk is acquired, including input / output delay, throughput, and the like.
[0039] In operation S220, the monitoring data is processed by using a preset risk prediction model to determine a plurality of failure risk information of the storage device.
[0040] In the embodiments of the present application, the preset risk prediction model can be trained based on historical monitoring data, and the plurality of failure risk information of the storage device in the future period is predicted by learning the data change rule in the historical monitoring data. The failure risk information can represent a failure including a physical layer failure or a logical layer failure.
[0041] In operation S230, input / output trajectory data of the storage device is processed to obtain image features of the input / output trajectory data. The input / output trajectory data represents the usage of a physical path and / or a logical path involved in the data transmission process of the storage device.
[0042] In the embodiments of the present application, the input / output trajectory data is a time sequence record of the usage of the physical path and the logical path in the data transmission process of the storage device. The image features of the input / output trajectory data include a time sequence diagram, a heat map, and the like. The usage of the physical path and / or the logical path refers to the actual state and behavior mode of the physical resource and the logical resource called in the data transmission process of the storage device, which can represent the influence of the hardware performance and the software management efficiency of the solid state disk on the storage device.
[0043] In operation S240, a load feature in the image feature is identified.
[0044] In the embodiments of the present application, the load feature refers to the image feature directly related to the current load state of the storage device. The image feature can reflect key information such as resource occupation, performance pressure, or task complexity in the storage medium.
[0045] In operation S250, according to the load feature and the plurality of failure risk information, the allocation of input / output path resources of the storage device is adjusted to improve the utilization rate of the resources. The input / output path resources include physical resources and / or logical resources used for data transmission.
[0046] In the embodiments of the present application, according to the load feature and the plurality of failure risk information, the stability and reliability of the input / output path are determined. When the input / output path load is high and the failure risk is low, an incremental physical resource is allocated, for example, the bandwidth is improved and the transmission channel is increased.
[0047] In the embodiment of the present application, the storage device fault risk related monitoring data is acquired, and a plurality of fault risk information is determined by processing with a preset model. By processing the image features of the storage device input and output track data, the load features in the image features are identified. And according to the load features and the fault risk information, the storage device input and output path resource (including physical and logical resources) allocation is adjusted, the fault is prevented in advance, the stable operation of the storage device is ensured, and the resource utilization is improved by reasonable allocation of resources.
[0048] The following will be described in combination with Figure 3 The process of determining a plurality of fault risk information of the storage device will be introduced.
[0049] Figure 3 The flowchart for determining the second fault risk information according to the embodiment of the present application is schematically shown.
[0050] As Figure 3 shown, the above operation 220 can further include operation S301~operation S302.
[0051] In operation S301, the first preset risk prediction model is used to analyze the device state data, and the first fault risk information is determined, which is used to evaluate the reliability state of the physical device of the storage device.
[0052] In the embodiment of the present application, the first preset risk prediction model is a data-driven analysis model, which can be a machine learning model or a deep learning model. The first preset risk prediction model can determine the information related to the physical device fault of the storage device in the device state data by analyzing the device state data, and further determine the first fault risk information.
[0053] In operation S302, the second preset risk prediction model is used to analyze the resource usage data, and the second fault risk information is determined, which is used to evaluate the stability state of the resource scheduling of the storage device.
[0054] In the embodiment of the present application, the first preset risk prediction model is also a data-driven analysis model. The first preset risk prediction model can determine the information related to the load balancing of the storage device in the device state data by analyzing the resource usage data, and further determine the second fault risk information.
[0055] In the embodiments of the present application, the first failure risk information is obtained by analyzing the device state data by using a first preset risk prediction model, and the reliability of the physical device is evaluated. The second failure risk information is obtained by analyzing the resource usage data by using a second preset risk prediction model, and the stability of resource scheduling is evaluated. The failure risk information is obtained from two different key dimensions of device state and resource usage, which can more comprehensively and accurately evaluate the overall state of the storage device.
[0056] The process of determining the first failure risk information is introduced below.
[0057] In the embodiments of the present application, the device state data is extracted by using a convolution layer to obtain device state features. The device state features are normalized by using a normalization layer to obtain normalized features. The normalized features are analyzed by using a full connection layer to obtain a risk score. The risk score is compared with a preset risk threshold to determine the first failure risk information.
[0058] For example, the convolution layer extracts key device state features such as data trend and abnormal fluctuation from the preprocessed device state data by sliding and taking values on the device state data through multiple convolution kernels. The normalization layer processes the device state features to eliminate the dimensional and scale differences between the device state features. The device state features are converted into normalized features with uniform range and distribution.
[0059] The full connection layer analyzes the normalized features by weighted connection between neurons and activation function operation to output a risk score representing the degree of failure risk. The risk score is compared with a preset risk threshold to obtain a comparison result, and the first failure risk is determined according to the comparison result.
[0060] In the embodiments of the present application, the risk score is given by analyzing the device state data by using the convolution layer, the normalization layer and the full connection layer. The risk score is compared with the preset risk threshold, which can timely and accurately determine the failure risk, provide reliable basis for device maintenance, and reduce the device failure rate and operation and maintenance cost.
[0061] The process of determining the second failure risk information is introduced below.
[0062] In the embodiments of the present application, the resource usage data is arranged in time sequence to form a performance change sequence. The fluctuation features of the performance change sequence are extracted by using a feature extraction layer. The fluctuation features are processed by using a prediction layer to obtain the change trend of the performance change sequence. According to the change trend, the risk level information corresponding to the resource usage data is determined as the second failure risk information.
[0063] For example, resource usage data of the storage device is collected at preset time intervals, and the resource usage data is arranged in ascending or descending order according to the time stamp of the resource usage data to construct a performance change sequence. In the feature extraction layer, the performance change sequence is extracted by a convolutional neural network to obtain fluctuation features, including mean, variance, frequency domain components and other features of the performance change sequence.
[0064] The fluctuation features are predicted by the prediction layer to obtain the change trend of the performance change sequence. The gradient in the change trend is compared with a preset gradient threshold to determine the risk level as the second fault risk information.
[0065] In the embodiments of the present application, by extracting fluctuation features and predicting change trends, the stability risk of resource scheduling can be accurately evaluated, and possible bottlenecks or abnormal situations in the resource usage process can be found in advance, so as to adjust the resource allocation strategy in time and ensure the efficient use of resources.
[0066] The process of determining the third fault risk information is introduced below.
[0067] In the embodiments of the present application, abnormal operation information in read-write data is determined, and error record information in device log data is determined, the error record information including error event information generated by self-checking or running of the storage device. The third fault risk information is generated according to the abnormal operation information and the error record information, and the third fault risk information is used to evaluate the immediate fault risk state of the storage device in running.
[0068] For example, a preset rule engine is used to match the abnormal operation mode corresponding to the read-write data, such as unauthorized access, super-threshold data block transmission or frequent retry instruction, to identify and extract the abnormal operation information. Regular expressions or semantic analysis are used to parse the device log data to determine the error event record in the device log data.
[0069] In the embodiments of the present application, the third fault risk information is generated by setting clear conditions, which can accurately judge the immediate fault risk situation of the storage device in running, thereby avoiding false positives and false negatives, and providing accurate and reliable basis for fault handling.
[0070] The process of generating the third fault risk information according to the abnormal operation information and the error record information is introduced below.
[0071] In the embodiments of the present application, the third fault risk information is generated when the abnormal operation information is abnormal address access and the number of errors in the error record information exceeds a number threshold within a predetermined time window.
[0072] For example, when abnormal operation information is identified as abnormal address access, error records are extracted from the device log and their number is counted within a preset time window. When the number of error records within the preset time window exceeds a preset threshold, an immediate fault risk is identified, and third fault risk information is generated according to preset rules.
[0073] In this embodiment of the application, when the abnormal operation information is a read / write data error and the device data records in the device log data exceed the data security threshold, a third fault risk information is generated.
[0074] For example, when abnormal operation information is determined to be a read / write data error, the system collects relevant device data records related to read / write data errors from the device log data. When the number of device data records exceeds the data security threshold, a fault risk is identified, and third-party fault risk information, including the risk level and possible causes, is generated according to preset rules.
[0075] In this embodiment of the application, third fault risk information is generated accurately and in real time based on abnormal operation information and error record information. This can help determine the immediate fault risk during the operation of the storage device, avoid false alarms and missed alarms, and provide an accurate and reliable basis for fault handling.
[0076] The following describes the process of determining the image features of the input and output trajectory data.
[0077] Figure 4 A flowchart illustrating the determination of image features according to an embodiment of this application is shown schematically.
[0078] like Figure 4 As shown, the above operation S230 may also include operations S401 to S404.
[0079] In S401 operation, the logical address access data is arranged in chronological order to obtain a time feature sequence.
[0080] In this embodiment of the application, during the operation of the storage device, logical address access data is collected and timestamped. Based on the chronological order of the timestamps, the collected logical address access data is sorted, and the sorted data is sequentially stored in a preset sequence structure to form a time feature sequence containing time dimension information.
[0081] In S402 operation, the logical address access data is arranged according to the address offset to obtain the address feature sequence.
[0082] In the embodiment of the present application, during the running of the storage device, logical address access data is collected, and address offsets in the logical address access data are extracted. According to the size relationship of the address offset information, ascending or descending arrangement operation is performed on the logical address access data, and the sorted logical address access data is stored in a preset sequence container in order to obtain an address feature sequence.
[0083] In operation S403, according to the time sequence and the queue depth of the protocol command queue data, the time depth feature of the protocol command queue data is obtained.
[0084] In the embodiment of the present application, during the running of the storage device, protocol command queue data is collected, and the time stamp and the corresponding queue depth value of the protocol command queue data are recorded. The protocol command queue data is sorted according to the time stamp in order to construct a time sequence, and the queue depth value corresponding to each time point is taken as a feature component to obtain the time depth feature of the protocol command queue data.
[0085] In operation S404, the time feature sequence, the address feature sequence and the time depth feature are converted to obtain the image feature of the input / output trajectory data.
[0086] In the embodiment of the present application, the input / output trajectory data is converted into an image feature through a specific arrangement and processing manner, which can provide more intuitive and easy-to-process data form for subsequent load feature analysis, provide strong support for reasonable resource allocation, and improve the adaptability of the storage device to different load conditions.
[0087] The process of obtaining the image feature of the input / output trajectory data is introduced below in combination with specific embodiments.
[0088] Figure 5 A flowchart for determining the input / output trajectory according to the embodiment of the present application is schematically shown.
[0089] As shown in Figure 5 Operation S404 can also include operations S501-S503.
[0090] In operation S501, the position information and the corresponding feature value of the time depth feature are determined to obtain a feature map, the position information is determined according to the time of the time feature sequence and the address offset of the address feature sequence, and the corresponding feature value is positively correlated with the queue depth.
[0091] In the embodiment of the present application, the time information corresponding to each data point of the time feature sequence is extracted, and the address offset is extracted from the address feature sequence. The time information and the address offset are combined to obtain position information in the form of two-dimensional coordinates, and the position information is used to represent the position in the feature map.
[0092] The queue depth value is obtained from the temporal depth feature. Based on a preset positive correlation mapping rule, the queue depth value is converted into a corresponding feature value. The mapping rule ensures that the larger the queue depth, the larger the corresponding feature value. The corresponding feature value is filled into the corresponding position using the position information as an index to obtain a feature map. The feature map can intuitively reflect the spatiotemporal characteristics of the protocol command queue.
[0093] In operation S502, the timing waveform is determined based on the relationship between the characteristic value and time. The timing waveform is used to represent the load change characteristics of the protocol command queue.
[0094] In this embodiment, feature values corresponding to each time point are extracted from the feature map in chronological order to form a feature value sequence. A coordinate system is established with time point as the horizontal axis and feature values as the vertical axis. The feature values in the feature value sequence are marked as discrete points in the coordinate system. A curve fitting method is used to smooth the discrete points to generate a continuous curve. The generated continuous curve is used as a time-series waveform, which can intuitively present the characteristics of the protocol command queue load changing over time.
[0095] In operation S503, the input and output trajectories are determined by connecting the position information corresponding to adjacent time points in the feature map. The input and output trajectories are used to represent the access mode characteristics of the logical address.
[0096] In this embodiment, positional information of adjacent time points is extracted from the feature map. This positional information is determined by the time offset between the time feature sequence and the address feature sequence. The positional information of adjacent time points is then sequentially connected in chronological order to form a continuous sequence of line segments. The path formed by these line segments is used as the input and output trajectory.
[0097] In this embodiment of the application, by obtaining the image features of the input and output trajectory data, rich and effective information is provided for analyzing load characteristics, which is conducive to formulating a more scientific and reasonable resource allocation strategy.
[0098] The process of determining load features in image features is described below with reference to specific embodiments.
[0099] Figure 6 A flowchart illustrating the determination of load characteristics according to an embodiment of this application is shown schematically.
[0100] like Figure 6 As shown, operation S240 may also include operations S601 to S603.
[0101] During operation S601, the input and output trajectories are identified, and the access features corresponding to the input and output trajectories are obtained. The access features include continuous access features and random access features.
[0102] In the embodiment of the present application, the difference value of the address offset of adjacent accesses in the input / output track is calculated, the difference value distribution is counted, when the difference value is concentrated in a small value interval and presents regular fluctuation, it is determined as the continuous access feature, indicating that there is an address access behavior in order, when the difference value distribution is scattered and has no obvious regularity, it is determined as the random access feature.
[0103] In operation S602, the timing waveform is identified to obtain the attribute feature corresponding to the timing waveform graph, and the attribute feature includes the periodic fluctuation feature and the burst peak feature.
[0104] In the embodiment of the present application, the timing waveform graph is analyzed in the frequency domain, the frequency spectrum distribution is obtained through Fourier transform, if there is a significant main frequency component, it is determined as the periodic fluctuation feature, reflecting the regular change of the protocol command queue load. The sliding window algorithm is used to detect the instantaneous peak value in the waveform, when the peak value exceeds the preset threshold and the duration is shorter than the set time length, it is determined as the burst peak feature, representing the instantaneous increase of the queue load.
[0105] In operation S603, according to the access feature and the attribute feature, the load feature in the image feature is determined, and the load feature includes at least one of the sequential access load, the random access load and the mixed access load.
[0106] The process of determining the load feature in the image feature is introduced below in combination with specific embodiments.
[0107] In the embodiment of the present application, the input / output track and the timing waveform in the image feature are identified in detail, so that the load feature of the storage device can be accurately determined. The targeted guidance is provided for resource allocation, and the performance and efficiency of the storage device in different load scenarios are improved.
[0108] The process of determining the load feature in the image feature is introduced below in combination with specific embodiments.
[0109] The operation 603 can further include determining the combination mode of the access feature and the attribute feature; in the case that the combination mode includes the continuous access feature and the periodic fluctuation feature, determining that the load feature is the sequential access load; in the case that the combination mode includes the random access feature and the burst peak feature, determining that the load feature is the random access load; in the case that the combination mode includes a plurality of access features and attribute features, determining that the load feature is the mixed access load.
[0110] For example, a combination rule library of access features and attribute features is constructed in advance, in which different combinations correspond to load feature judgment conditions. The identified access features and attribute features are matched and analyzed. If continuous access features and periodic fluctuation features exist simultaneously, the load feature is determined as sequential access load according to the rule library, reflecting the regular sequential access mode of the protocol command queue load. If random access features and burst peak features exist, the load feature is determined as random access load, representing the discreteness and burstiness of the load. When multiple access features and attribute feature combinations exist, the load feature is determined as mixed access load according to the comprehensive judgment logic of the rule library.
[0111] In the embodiments of the present application, the combination of access features and attribute features is accurately determined to determine the load feature type, which can greatly improve the accuracy and pertinence of load feature identification.
[0112] The process of adjusting the allocation of input / output path resources of a storage device is introduced below in combination with specific embodiments.
[0113] Figure 7A A flowchart of adjusting input / output path resources according to an embodiment of the present application is schematically shown.
[0114] As shown in Figure 7A Operation S250 can also include operations S701-S703.
[0115] In operation S701, in the case where the load feature represents a sequential access load and the risk levels represented by the multiple fault risk information are lower than a first level threshold, the allocation of physical channels and the pre-reading mechanism in the input / output path resources of the storage device are adjusted.
[0116] For example, in the case where the load feature represents a sequential access load and the risk levels represented by the multiple fault risk information are lower than a first level threshold, continuous logical address blocks are allocated to multiple physical channels, and the physical channels independently process the allocated logical address blocks. Based on the continuity feature of the sequential access load, pre-read data is dynamically calculated and cached to a cache area.
[0117] In operation S702, in the case where the load feature represents a random access load and the risk levels represented by the multiple fault risk information are higher than a first level threshold and lower than a second level threshold, the queue allocation and bandwidth allocation in the input / output path resources of the storage device are adjusted.
[0118] For example, in the case that the load characteristic is represented as a random access load and the risk level of the plurality of failure risk information is higher than the first level threshold and lower than the second level threshold, a multi-level buffer queue is established for the random access load, the multi-level buffer queue including a real-time queue, a normal queue and a background queue. According to the data output input request, the priority of the queue in the multi-level buffer queue is adjusted, and the transmission bandwidth resource is allocated according to the load condition of the multi-level buffer queue.
[0119] In operation S703, in the case that the load characteristic is represented as a mixed access load and the risk level of the plurality of failure risk information is higher than the second level threshold and lower than the third level threshold, the path switching and address mapping in the input output path resource of the storage device are adjusted.
[0120] For example, in the case that the load characteristic is represented as a mixed access load and the risk level of the plurality of failure risk information is higher than the second level threshold and lower than the third level threshold, the current transmission path is maintained for the sequential access load, and the dynamic path selection is enabled for the random access load; the address mapping relationship is reconstructed based on the risk level, and the high-risk data is dispersedly mapped to different physical storage areas.
[0121] In the embodiments of the present application, the resource allocation is adjusted according to different load characteristics and failure risk levels, which can fully exert the advantages of the storage device resources and realize the optimal configuration of the resources.
[0122] The process of determining the channel allocation and bandwidth limitation strategy of the input output track data is introduced below in combination with specific embodiments.
[0123] Before operation S250, the weight coefficient of the storage resource allocation can also be determined according to the risk level of the failure risk information, the weight coefficient being used to represent the channel allocation priority of each of the plurality of storage areas; and the channel allocation and bandwidth limitation strategy of the input output track data is determined according to the weight coefficient and the resource demand evaluation value of the input output track data.
[0124] In the embodiments of the present application, a mapping table of the failure risk level and the weight coefficient is preset, and the higher the risk level is, the lower the weight coefficient is. The weight coefficient is determined by looking up the table according to the failure risk information of each storage area, and the weight coefficient reflects the channel allocation priority. The resource demand evaluation value is obtained by evaluating the resource demand of the input output track data in combination with the data volume, the access frequency and the like. The weight coefficient and the resource demand evaluation value are comprehensively calculated, the storage area with a large weight coefficient is preferentially allocated to the channel, and the bandwidth limitation of each channel is reasonably set according to the demand evaluation value.
[0125] For example, a storage resource allocation score is generated by weighting the weighting coefficients and the resource demand assessment value. For target storage regions with a storage resource allocation score higher than a first threshold, an independent physical channel is allocated to the target storage region. For target storage regions with a storage resource allocation score between the first and second thresholds, a shared channel and dynamic bandwidth adjustment mechanism are allocated to the target storage region. For storage regions with a storage resource allocation score lower than the second threshold, channel reuse and fixed bandwidth limits are set for the target storage region.
[0126] In this embodiment, storage resources are allocated reasonably and bandwidth usage is limited to avoid resource waste and excessive competition.
[0127] The following combination Figure 7B The component architecture of the resource scheduling method is introduced, along with specific implementation examples.
[0128] Figure 7B The schematic diagram illustrates the component architecture of a resource scheduling method according to an embodiment of this application.
[0129] In the embodiments of this application, such as Figure 7B As shown, the storage device includes a central processing unit, a smart heterogeneous processor, a memory controller, a flash memory interface, flash memory chips, a host interface, solid-state storage, and a data processing engine.
[0130] The internal data flow of the storage device is as follows: the host interface receives host data commands and transmits them to the central processing unit (CPU), which interacts with the memory controller to achieve temporary data storage. The CPU coordinates and controls the overall system, and the control logic manages data read and write operations on the flash memory chips through the flash memory interface.
[0131] The intelligent heterogeneous processor fetches data from memory for processing, and the results can be fed back or stored. The memory controller manages the memory cache, interacts with the CPU, host interface, and intelligent heterogeneous processor as needed, and writes data that needs to be stored for a long time to flash memory chips via the flash memory interface.
[0132] Firmware storage provides the CPU with control programs and parameters. The data processing engine encrypts and compresses data before writing it, and decrypts, decompresses, and corrects errors during reading, ensuring data security and reliability. All components work together to achieve efficient data flow.
[0133] The following combination Figure 7C The computational architecture of the resource scheduling method is described in detail, along with specific implementation examples.
[0134] Figure 7C The diagram illustrates a computational architecture layering diagram of a resource scheduling method according to an embodiment of this application.
[0135] In the embodiments of this application, such as Figure 7CAs shown, the computing architecture of the resource scheduling method includes a host application layer, a system software layer, and a hardware abstraction layer. The host application layer deploys a deep learning framework, uses its powerful algorithm and model processing capability to provide intelligent decision support for resource scheduling, such as predicting storage load, optimizing resource allocation strategy, etc. The system software layer includes a driver adaptation layer, a scheduling management layer, and an intelligent running environment. The driver adaptation layer is used to realize the adaptation of hardware drivers and the system. The scheduling management layer is responsible for overall resource scheduling, and the intelligent running environment provides basic support for the entire system running. The hardware abstraction layer is composed of a controller, a storage interface, and an intelligent acceleration unit. The controller coordinates hardware operations, the storage interface guarantees data transmission, and the intelligent acceleration unit improves data processing speed, which together provide hardware support for the upper layer.
[0136] The architecture hierarchy of the resource scheduling method will be introduced in the following Figure 7D and specific embodiments.
[0137] Figure 7D The architecture hierarchy diagram of the resource scheduling method according to the embodiments of the present application is schematically shown.
[0138] In the embodiments of the present application, as Figure 7D shown, the architecture hierarchy of the resource scheduling method includes an application layer, a software development kit (SDK) layer, a system software layer, and a hardware abstraction layer. Among them, the intelligent application program in the application layer is at the top of the architecture, based on the resource scheduling capability of the storage device, realizes specific application functions such as intelligent data management and efficient storage service, and provides a convenient storage solution for users. The SDK layer includes a model conversion tool, a promotion interface library, and a performance analysis tool. The model conversion tool helps the adaptation of different models in the storage system, the promotion interface library provides a standard interface for easy calling by the application layer, and the performance analysis tool is used to evaluate and optimize the performance of the storage system.
[0139] In the system software layer, the intelligent running environment provides running support for the entire system, the heterogeneous scheduling framework reasonably allocates storage resources according to load characteristics and fault risk information, and the device driver layer directly interacts with the hardware abstraction layer to drive the hardware device to work normally.
[0140] In the hardware abstraction layer, the hardware of the storage device is abstracted, so that the upper layer software does not need to pay attention to the hardware details, and the compatibility and scalability of the system are improved.
[0141] Based on the above resource scheduling method for storage devices, the present application further provides a resource scheduling device for storage devices. The device will be described in detail in the following Figure 8 .
[0142] Figure 8A structural block diagram of a resource scheduling apparatus for a storage device according to an embodiment of the present application is shown.
[0143] As shown in Figure 8 The resource scheduling apparatus 800 for a storage device of this embodiment includes an acquisition module 810, a processing module 820, a feature processing module 830, a feature recognition module 840, and a resource allocation module 850.
[0144] The acquisition module 810 is configured to acquire monitoring data related to the failure risk of the storage device. In an embodiment, the acquisition module 810 can be configured to perform operation S210 described above, and details are not repeated here.
[0145] The processing module 820 is configured to process the monitoring data using a preset risk prediction model to determine a plurality of failure risk information of the storage device. In an embodiment, the processing module 820 can be configured to perform operation S220 described above, and details are not repeated here.
[0146] The feature processing module 830 is configured to process the input / output trajectory data of the storage device to obtain image features of the input / output trajectory data. In an embodiment, the feature processing module 830 can be configured to perform operation S230 described above, and details are not repeated here.
[0147] The feature recognition module 840 is configured to recognize load features in the image features. In an embodiment, the feature recognition module 840 can be configured to perform operation S240 described above, and details are not repeated here.
[0148] The resource allocation module 850 is configured to adjust the allocation of the input / output path resources of the storage device according to the load features and the plurality of failure risk information. In an embodiment, the resource allocation module 850 can be configured to perform operation S250 described above, and details are not repeated here.
[0149] According to an embodiment of the present application, any of the modules of the obtaining module 810, the processing module 820, the feature processing module 830, the feature identifying module 840 and the resource allocating module 850 can be combined in one module, or any of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of the modules can be combined with at least part of the functions of the other modules, and implemented in one module. According to an embodiment of the present application, at least one of the obtaining module 810, the processing module 820, the feature processing module 830, the feature identifying module 840 and the resource allocating module 850 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on board, a system on package, an application specific integrated circuit (ASIC), or any other reasonable manner of integrating or packaging circuits, etc. in hardware or firmware, or implemented in any one of software, hardware and firmware or in a proper combination of any of the above. Alternatively, at least one of the obtaining module 810, the processing module 820, the feature processing module 830, the feature identifying module 840 and the resource allocating module 850 can be at least partially implemented as a computer program module which, when executed, can perform the corresponding functions.
[0150] Figure 9 A block diagram of an electronic device suitable for implementing the resource scheduling method for storage devices according to an embodiment of the present application is schematically shown.
[0151] As shown in Figure 9 The electronic device 900 according to an embodiment of the present application includes a processor 901 which can perform various appropriate actions and processes according to programs stored in a read only memory (ROM) 902 or loaded from a storage portion 908 into a random access memory (RAM) 903. The processor 901 can include, for example, a general purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special purpose microprocessor (e.g., an application specific integrated circuit (ASIC)), etc. The processor 901 can also include an on-board memory for cache use. The processor 901 can include a single processing unit or multiple processing units for executing different actions of the method processes according to embodiments of the present application.
[0152] In the RAM 903, various programs and data required for the operation of the electronic device 900 are stored. The processor 901, the ROM 902, and the RAM 903 are connected to each other via the bus 904. The processor 901 performs various operations of the method flow according to the embodiments of the present application by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the programs can also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 can also perform various operations of the method flow according to the embodiments of the present application by executing the programs stored in the one or more memories.
[0153] According to the embodiments of the present application, the electronic device 900 can further include an input / output (I / O) interface 905, which is also connected to the bus 904. The electronic device 900 can further include one or more of the following components connected to the input / output (I / O) interface 905: an input part 906 including a keyboard, a mouse, and the like; an output part 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a speaker, and the like; a storage part 908 including a hard disk, and the like; and a communication part 909 including a network interface card such as a LAN card, a modem, and the like. The communication part 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as necessary. A removable medium 911 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is mounted on the drive 910 as necessary, so that a computer program read therefrom is installed in the storage part 908 as necessary.
[0154] The present application also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, when the one or more programs are executed, the method according to the embodiments of the present application is implemented.
[0155] According to an embodiment of the present application, the computer readable storage medium can be a non-transitory computer readable storage medium, for example, can include but not limited to: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In this application, a computer readable storage medium can be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, the computer readable storage medium can include one or more of the above-described ROM 902 and / or RAM 903 and / or a memory other than the ROM 902 and the RAM 903.
[0156] Embodiments of the present application also include a computer program product, which includes a computer program containing program codes for executing the methods shown in the flowcharts. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the methods provided by the embodiments of the present application.
[0157] The above-described functions defined in the system / device / apparatus of the embodiments of the present application are performed when the computer program is executed by the processor 901. According to an embodiment of the present application, the above-described system, device, module, unit, etc. can be implemented by computer program modules.
[0158] In one embodiment, the computer program can rely on a tangible storage medium such as an optical storage device, a magnetic storage device, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of a signal on a network medium, and be downloaded and installed through the communication part 909, and / or be installed from the detachable medium 911. The program codes contained in the computer program can be transmitted by any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the foregoing.
[0159] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or be installed from the detachable medium 911. When the computer program is executed by the processor 901, the above-described functions defined in the system of the embodiments of the present application are performed. According to an embodiment of the present application, the above-described system, device, apparatus, module, unit, etc. can be implemented by computer program modules.
[0160] According to embodiments of the present application, program code for implementing the computer programs provided by embodiments of the present application can be written in any combination of one or more programming languages, and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Program code can execute entirely on a user's computing device, partly on the user's device, as a stand-alone software package, partly on a remote computing device, or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, such as through the Internet using an Internet Service Provider.
[0161] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0162] Those skilled in the art will understand that features recited in the various embodiments of the present application can be combined and / or integrated in various ways, even if such combinations or integrations are not expressly noted in the present application. In particular, features recited in the various embodiments of the present application can be combined and / or integrated in ways that are not expressly noted in the present application, without departing from the spirit and teachings of the present application. All such combinations and / or integrations are within the scope of the present application.
[0163] The embodiments of the present application have been described above. However, these embodiments are merely for the purpose of illustration, and are not intended to limit the scope of the present application. Although the embodiments are described separately above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Those skilled in the art can make various substitutions and modifications without departing from the scope of the present application, and these substitutions and modifications should all fall within the scope of the present application.
Claims
1. A resource scheduling method for a storage device, characterized by, The method comprises: obtaining monitoring data related to storage device failure risk; processing the monitoring data using a preset risk prediction model to determine a plurality of failure risk information of the storage device; processing input-output trajectory data of the storage device to obtain image features of the input-output trajectory data, the input-output trajectory data representing the use of physical and / or logical paths involved in the data transmission process of the storage device; identifying load features in the image features; in the case where the load features represent sequential access load and the risk level represented by the plurality of failure risk information is lower than a first level threshold, adjusting the physical channel allocation in the input-output path resources of the storage device and the pre-reading mechanism; in the case where the load features represent random access load and the risk level represented by the plurality of failure risk information is higher than the first level threshold and lower than a second level threshold, adjusting the queue allocation in the input-output path resources of the storage device and the bandwidth allocation; in the case where the load features represent mixed access load and the risk level represented by the plurality of failure risk information is higher than the second level threshold and lower than a third level threshold, adjusting the path switching and address mapping in the input-output path resources of the storage device.
2. The method of claim 1, wherein, The monitoring data comprises device state data and resource usage data, and the preset risk prediction model comprises a first preset risk prediction model and a second preset risk prediction model; The processing of the monitoring data using a preset risk prediction model to determine a plurality of failure risk information of the storage device comprises: analyzing the device state data using the first preset risk prediction model to determine first failure risk information, the first failure risk information being used to evaluate the reliability state of the physical devices of the storage device; analyzing the resource usage data using the second preset risk prediction model to determine second failure risk information, the second failure risk information being used to evaluate the stability state of the resource scheduling of the storage device.
3. The method of claim 2, wherein, The first preset risk prediction model comprises a convolution layer, a normalization layer and a full connection layer; the analysis of the device state data using the first preset risk prediction model to determine first failure risk information comprises: extracting the device state data using the convolution layer to obtain device state features; normalizing the device state features using the normalization layer to obtain normalized features; analyzing the normalized features using the full connection layer to obtain a risk score; comparing the risk score with a preset risk threshold to determine the first failure risk information.
4. The method of claim 2, wherein, The second preset risk prediction model comprises a feature extraction layer and a prediction layer; the analysis of the resource usage data using the second preset risk prediction model to determine second failure risk information comprises: arranging the resource usage data in chronological order to form a performance change sequence; extracting fluctuation features of the performance change sequence using the feature extraction layer; processing the fluctuation feature by using the prediction layer to obtain a change trend of the performance change sequence; determining risk level information corresponding to the resource usage data as the second fault risk information according to the change trend.
5. The method of claim 2, wherein, The monitoring data includes read-write data and device log data, and the method further includes: determining abnormal operation information in the read-write data; determining error record information in the device log data, the error record information including error event information generated by self-checking or running of the storage device; generating third fault risk information according to the abnormal operation information and the error record information, the third fault risk information being used to evaluate an instant fault risk state of the storage device in running.
6. The method of claim 5, wherein, The abnormal operation information includes abnormal address access and read-write data error; and the generating of the third fault risk information according to the abnormal operation information and the error record information includes: generating the third fault risk information in a case where the abnormal operation information is abnormal address access and the number of errors in the error record information exceeds a number threshold within a predetermined time window; generating the third fault risk information in a case where the abnormal operation information is read-write data error and device data record in the device log data exceeds a data safety threshold.
7. The method of claim 1, wherein, The input-output trajectory data includes logical address access data and protocol command queue data; and the processing of the input-output trajectory data of the storage device to obtain image features of the input-output trajectory data includes: arranging the logical address access data in time sequence to obtain a time feature sequence; arranging the logical address access data according to address offset to obtain an address feature sequence; obtaining time depth features of the protocol command queue data according to the time sequence and the queue depth of the protocol command queue data; converting the time feature sequence, the address feature sequence and the time depth features to obtain the image features of the input-output trajectory data.
8. The method of claim 7, wherein, The image features include input-output trajectory and time sequence waveform; and the converting of the time feature sequence, the address feature sequence and the time depth features to obtain the image features corresponding to the input-output trajectory data includes: determining position information and corresponding feature values of the time depth features to obtain a feature map; the position information is determined according to time of the time feature sequence and address offset of the address feature sequence, and the corresponding feature values are positively correlated with the queue depth; determining the time sequence waveform based on a change relationship of the feature values with time, the time sequence waveform being used to represent load change features of the protocol command queue; determining the input-output trajectory by connecting position information corresponding to adjacent time in the feature map, the input-output trajectory being used to represent access mode features of logical addresses.
9. The method of claim 1, wherein, The image features include input-output trajectory and time sequence waveform; and the identifying of the load features in the image features includes: identifying the input-output trajectory to obtain access features corresponding to the input-output trajectory, the access features including continuous access features and random access features; identify the attribute features corresponding to the time waveform graph from the time waveform, the attribute features including periodic fluctuation features and burst peak features; determine a load feature in the image feature according to the access feature and the attribute feature, the load feature including at least one of sequential access load, random access load and mixed access load.
10. The method of claim 9, wherein, The determining the load feature in the image feature according to the access feature and the attribute feature includes: determining a combination mode of the access feature and the attribute feature; in a case where the combination mode includes the continuous access feature and the periodic fluctuation feature, determining that the load feature is the sequential access load; in a case where the combination mode includes the random access feature and the burst peak feature, determining that the load feature is the random access load; in a case where the combination mode includes a plurality of the access feature and the attribute feature, determining that the load feature is the mixed access load.
11. The method of claim 1, wherein, After the identifying the load feature in the image feature, the method further includes: determining a weight coefficient of storage resource allocation according to a risk level of the failure risk information, the weight coefficient being used to represent respective channel allocation priorities of different storage areas; determining a channel allocation and a bandwidth limitation strategy of the input / output trajectory data according to the weight coefficient and a resource demand evaluation value of the input / output trajectory data.
12. A resource scheduling apparatus for a storage device, characterized by comprising: The apparatus includes: an acquisition module configured to acquire monitoring data related to failure risk of a storage device; a processing module configured to process the monitoring data by using a preset risk prediction model to determine a plurality of failure risk information of the storage device; a feature processing module configured to process input / output trajectory data of the storage device to obtain image features of the input / output trajectory data; a feature identification module configured to identify a load feature in the image features; a resource allocation module configured to, in a case where the load feature represents sequential access load and a risk level represented by the plurality of failure risk information is lower than a first level threshold, adjust physical channel allocation in input / output path resources of the storage device and a pre-reading mechanism; in a case where the load feature represents random access load and the risk level represented by the plurality of failure risk information is higher than the first level threshold and lower than a second level threshold, adjust queue allocation in the input / output path resources of the storage device and bandwidth allocation; in a case where the load feature represents mixed access load and the risk level represented by the plurality of failure risk information is higher than the second level threshold and lower than a third level threshold, adjust path switching in the input / output path resources of the storage device and address mapping. 13.An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, characterized in that the one or more processors execute the one or more computer programs to implement steps of the method according to any one of claims 1-11.
14. A computer readable storage medium having stored thereon a computer program or instructions, characterized in that, The computer program or instructions are executed by the processor to implement steps of the method according to any one of claims 1-11.
Citation Information
Patent Citations
Method and device for adjusting configuration of storage system, storage medium and electronic equipment
CN119376644A
Resource information processing method and device, equipment, medium and program product
CN119415013A
Dynamic monitoring and early warning method for storage space of data center
CN120045421A