Cache dynamic allocation method and device for storage hardware equipment
By performing multi-dimensional perception and dynamic decision-making on storage hardware devices, and dynamically adjusting cache allocation, the problem of insufficient multi-dimensional perception and resource contention in existing cache allocation schemes is solved, thereby improving I/O response stability and reducing hardware costs.
Patent Information
- Application Number
- CN202511439249.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-09
AI Technical Summary
In existing technologies, the cache allocation scheme of storage hardware devices relies on offline training models based on a single metric, which lacks multi-dimensional perception. This results in large latency in traffic processing and can easily cause business interference when there is competition for resource allocation. The hardware deployment cost of intelligent algorithms is high, which affects the efficiency of cache allocation.
By performing input/output pattern analysis, media latency detection, and hardware status analysis on storage hardware devices, multi-dimensional data is obtained. Combined with decision-making logic, cache allocation is dynamically adjusted, multi-level cache partitions are divided, and hot data is identified. QoS virtual channels and preemptive resource reclamation mechanisms are used for cache management.
It significantly improves I/O response stability, avoids forced degradation to pass-through mode when the write cache is full, reduces latency fluctuations and hardware deployment costs, and eliminates the risk of high-temperature downtime and cache data loss.
Smart Images

Figure CN120909529A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of cache dynamic allocation, and particularly relates to a cache dynamic allocation method and device of a storage hardware device. BACKGROUND
[0002] A RAID (Redundant Array of Independent Disks) card as a core controller of a storage system, the cache management strategy of which directly determines the upper limit of I / O (Input / Output) performance. However, the traditional static cache allocation scheme has defects such as rigid response mechanism, uncontrollable hardware risks, and deteriorating resource competition.
[0003] At present, in view of the above defects, the related technology can monitor IOPS (Input / Output Operations Per Second) or queue depth to determine whether the IOPS exceeds a threshold value, and linearly increase the write cache when the threshold value is exceeded. In addition, the related technology can also use machine learning or reinforcement learning algorithms to train the model through historical I / O patterns to predict future loads and dynamically adjust cache allocation.
[0004] However, the related technology only relies on offline training of a single indicator model, lacks multi-dimensional perception, and there is a large delay in traffic acceptance, cache, and adjustment operations. In addition, the related technology is prone to business interference when competing for resource allocation, and the hardware deployment cost of intelligent algorithms is high, which greatly affects the cache allocation efficiency, and needs to be solved urgently. SUMMARY
[0005] The present application provides a cache dynamic allocation method and device of a storage hardware device to at least solve the technical problems in the related art that overly rely on offline training of a single indicator model, lack multi-dimensional perception, and have a large delay in traffic processing. In addition, when competing for resource allocation, the related technology is prone to business interference, and the hardware cost of deploying intelligent algorithms is high, which greatly affects the cache allocation efficiency.
[0006] The application provides a cache dynamic allocation method of a storage hardware device, comprising the following steps: performing input / output mode analysis on a storage hardware device in a target server to obtain a random input / output ratio of the storage hardware device, performing medium delay detection on the storage hardware device to obtain medium detection data of the storage hardware device, and performing hardware state analysis on the storage hardware device to obtain hardware state data of the storage hardware device; determining a corresponding sequential write request proportion according to the random input / output ratio, and matching target decision logic corresponding to the storage hardware device based on the sequential write request proportion and a current temperature of the storage hardware device contained in the hardware state data; determining a read / write cache allocation ratio corresponding to the storage hardware device based on the target decision logic, in combination with the random input / output ratio, the medium detection data and the hardware state data, identifying hot data corresponding to the storage hardware device, and performing cache allocation adjustment on the storage hardware device according to the read / write cache allocation ratio and the hot data to obtain a dynamically allocated cache, and dividing the dynamically allocated cache into multi-level cache partitions of different priorities, and determining a data type of to-be-allocated cache data, so as to allocate the to-be-allocated cache data to a corresponding cache partition according to the data type and the priority of the multi-level cache partitions.
[0007] The application also provides a cache dynamic allocation device of a storage hardware device, comprising: a multi-dimensional perception module, configured to perform input / output mode analysis on a storage hardware device in a target server to obtain a random input / output ratio of the storage hardware device, perform medium delay detection on the storage hardware device to obtain medium detection data of the storage hardware device, and perform hardware state analysis on the storage hardware device to obtain hardware state data of the storage hardware device; a matching module, configured to determine a corresponding sequential write request proportion according to the random input / output ratio, and match target decision logic corresponding to the storage hardware device based on the sequential write request proportion and a current temperature of the storage hardware device contained in the hardware state data; and a cache allocation module, configured to determine a read / write cache allocation ratio corresponding to the storage hardware device based on the target decision logic, in combination with the random input / output ratio, the medium detection data and the hardware state data, identify hot data corresponding to the storage hardware device, and perform cache allocation adjustment on the storage hardware device according to the read / write cache allocation ratio and the hot data to obtain a dynamically allocated cache, and divide the dynamically allocated cache into multi-level cache partitions of different priorities, and determine a data type of to-be-allocated cache data, so as to allocate the to-be-allocated cache data to a corresponding cache partition according to the data type and the priority of the multi-level cache partitions.
[0008] The application further provides an electronic device, comprising a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the cache dynamic allocation method of any of the storage hardware devices.
[0009] The application further provides a non-volatile computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of the cache dynamic allocation method of any of the storage hardware devices.
[0010] The application further provides a computer program product, comprising a computer program, wherein the computer program is executed by a processor to implement the steps of the cache dynamic allocation method of any of the storage hardware devices.
[0011] Through the application, the input / output mode of the storage hardware device in the target server can be analyzed to obtain the random input / output ratio of the storage hardware device, the medium delay of the storage hardware device can be detected to obtain the medium detection data of the storage hardware device, and the hardware state of the storage hardware device can be analyzed to obtain the hardware state data of the storage hardware device; the corresponding sequential write request proportion is determined according to the random input / output ratio, and the target decision logic corresponding to the storage hardware device is matched based on the sequential write request proportion and the current temperature of the storage hardware device contained in the hardware state data; the read / write cache allocation ratio corresponding to the storage hardware device is determined based on the target decision logic, in combination with the random input / output ratio, the medium detection data and the hardware state data, the hot data corresponding to the storage hardware device is identified, and the cache allocation of the storage hardware device is adjusted according to the read / write cache allocation ratio and the hot data to obtain a dynamically allocated cache, the dynamically allocated cache is divided into multi-level cache partitions of different priorities, and the data type of the to-be-allocated cache data is determined, so that the to-be-allocated cache data is allocated to the corresponding cache partition according to the data type and the priority of the multi-level cache partition. Therefore, the technical problem that in related technologies, a single index offline training model is excessively relied on, multi-dimensional perception is lacked, and there is a large delay in traffic processing can be solved. In addition, in the resource allocation competition, business interference is easily caused, and the hardware cost of deploying an intelligent algorithm is high, which greatly affects the cache allocation efficiency. The technical effects of avoiding forced degradation to a transparent write mode when the write cache is full, significantly improving the I / O response stability, eliminating the high-temperature downtime and cache data loss risk, and reducing the delay fluctuation and hardware deployment cost are achieved. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort based on these drawings.
[0013] Figure 1 A flow chart of a cache dynamic allocation method of a storage hardware device according to an embodiment of the present application is provided. Figure 2 A logic architecture schematic diagram of a cache dynamic allocation method of a storage hardware device according to an embodiment of the present application is provided. Figure 3 An execution logic schematic diagram of a cache dynamic allocation method of a storage hardware device according to an embodiment of the present application is provided. Figure 4 An example diagram of a cache dynamic allocation apparatus of a storage hardware device according to an embodiment of the present application is provided.
[0014] Among them, 10 is a cache dynamic allocation apparatus of a storage hardware device, 100 is a multi-dimensional perception module, 200 is a matching module, and 300 is a cache allocation module. DETAILED DESCRIPTION
[0015] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort are within the protection scope of the present application.
[0016] It should be noted that in the description of the present application, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, not to describe a specific order or sequence.
[0017] In order to make those skilled in the art better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0018] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the cache dynamic allocation method of the storage hardware device depends, the specific application environment architecture or specific hardware architecture is described here.
[0019] Embodiments of the present application provide a cache dynamic allocation method of a storage hardware device.
[0020] As shown in the flow chart of the cache dynamic allocation method of the storage hardware device of the embodiments of the present application, the cache dynamic allocation method of the storage hardware device comprises the following steps: Figure 1 In step S101, the input and output mode of the storage hardware device in the target server is analyzed to obtain the random input and output ratio of the storage hardware device, the medium delay of the storage hardware device is detected to obtain the medium detection data of the storage hardware device, and the hardware state of the storage hardware device is analyzed to obtain the hardware state data of the storage hardware device.
[0021] Those skilled in the art should appreciate that, with the popularity of cloud computing, AI (Artificial Intelligence) training and big data analysis, storage systems are facing mixed load pressure (such as coexistence of random small file I / O and large data block sequential read and write), high concurrency access (such as thousands of virtual machines requesting in parallel), and hardware reliability challenges (such as all-flash array life, high temperature risk); storage hardware devices, such as RAID cards, as the core controller of the storage system, their cache management strategy directly determines the upper limit of I / O performance. However, the traditional cache allocation technology still has the following defects: 1. Performance bottleneck caused by lack of multi-dimensional perception: The prior art only relies on a single indicator (such as IOPS, queue depth) or an offline training model, ignores I / O mode mutations (such as AI training from sequential read to random write), medium delay differences (HDD (Hard Disk Drive) seek delay and SSD (Solid State Drive) concurrent response), and hardware state linkage (DIMM (Dual Inline Memory Modules) temperature, SSD wear level), and cannot respond to dynamic business characteristics in real time, resulting in mismatch between cache allocation and actual load, and cliff-like performance drop in burst scenarios.
[0022] 2. Reliability risk caused by response delay: The prior art cannot quickly and efficiently intercept burst traffic, is prone to cache overflow, and forces write cache to full and degrade to full write mode, resulting in cliff-like performance drop and data loss risk.
[0023] 3. Business interference caused by resource competition: When multiple virtual machines / containers share the cache, there is a lack of fine-grained QoS (Quality of Service) isolation, high-priority tasks (such as financial transaction logs) are blocked by low-priority tasks (backup jobs), and static allocation cannot dynamically adjust bandwidth quotas, thereby affecting high-priority tasks.
[0024] 4. Imbalance between hardware cost and efficiency: Intelligent algorithms require special hardware, and the deployment cost is relatively high, and the input-output ratio is low.
[0025] In view of the above defects, the embodiments of the present application can obtain multi-dimensional perception data corresponding to the storage hardware device, match the corresponding decision logic, and combine QoS virtual channels and resource isolation, and combine the pre-emptive resource recycling mechanism to perform dynamic allocation of the cache, thereby breaking through the three bottlenecks of performance, reliability and cost through software-defined cache management without hardware modification. It should be noted that the embodiments of the present application take the RAID card as an example to describe and introduce the process of dynamic allocation of the cache of the storage hardware device in detail.
[0026] Specifically, in the actual execution process, the embodiments of the present application can first obtain and fuse multi-dimensional real-time perception data of the I / O mode, medium performance (such as HDD seek delay, SSD concurrent response capability), hardware state (such as DIMM temperature, SSD wear degree) and the like corresponding to the RAID card through the multi-dimensional real-time perception layer, thereby realizing real-time capture of full-scene load characteristics and providing reliable data guidance and basis for selection of decision logic.
[0027] Optionally, in an embodiment of the present application, the input / output mode of the storage hardware device in the target server is analyzed to obtain the random input / output ratio of the storage hardware device, including: in response to a current request of the target server, determining a starting logical block address and a continuous block quantity of the current request and a starting address of a next request; calculating the sum of the starting logical block address and the continuous block quantity of the current request, and calculating the absolute value of the difference between the sum of the starting logical block address and the continuous block quantity and the starting address of the next request, and determining whether the absolute value of the difference is greater than a preset window threshold; if the absolute value of the difference is less than or equal to the window threshold, it is determined that the logical block address corresponding to the current request satisfies the preset continuity requirement, and a first target value is assigned to the logical block address index, otherwise it is determined that the logical block address does not satisfy the continuity requirement, and a second target value is assigned to the logical block address index; determining a request time point of the current request, and based on the first target value, the second target value, the request time point and a preset sliding window size, calculating the random input / output ratio.
[0028] It should be noted that in the I / O mode analysis of the server, the embodiment of the present application can first obtain the starting logical block address and the number of consecutive blocks of the current request (such as the nth request) of the server and the starting address of the next request (such as the n+1th request); secondly, the embodiment of the present application can determine whether the LBA (Logical Block Addressing, logical block addressing) is continuous according to the starting logical block address, the number of consecutive blocks and the starting address, and based on the determination result, the random input / output ratio is calculated in combination with the preset sliding window size, and the specific process is as follows: 1. Determine whether the LBA is continuous: The embodiment of the present application can construct a corresponding continuity determination expression according to the starting logical block address and the number of consecutive blocks of the nth request and the starting address of the n+1th request, as shown in the following formula: n
[0029] Among them, represents the starting logical block address of the nth request; n represents the number of consecutive blocks of the nth request; represents the starting address of the subsequent request; n is a window threshold, which can be usually set to 1 (i.e. adjacent sectors) or the size of a RAID stripe. Those skilled in the art should understand that LBA is an addressing method for hard disks, SSDs and other storage hardware devices, which can be used to locate and access data blocks on storage media; LBA assigns a unique logical address to each data block, so that the operating system and storage hardware devices can efficiently perform data read / write operations.
[0030] It can be understood that from the above continuity determination expression, the embodiment of the present application calculates the absolute value of the difference between the sum of the starting logical block address and the number of consecutive blocks and the starting address of the next request, and when the absolute value of the difference is not greater than the window threshold (i.e. the continuity requirement), the logical block address index is assigned a first target value
[0031] , otherwise the logical block address index is assigned a second target value .
[0032] 2. Calculate the random input / output ratio based on the sliding window: The embodiment of the present application can determine the request time point of the current request, and calculate the random input / output ratio based on the logical block address index , the request time point and the sliding window size, and the calculation expression of the random input / output ratio is:
[0033] wherein, is the current time point (i.e. the request time point); is the sliding window size, which can be 50-100 inputs or outputs; t W k ≤t) represents the time sequence position of the specific request in the positioning window.
[0034] Therefore, the embodiments of the present application can accurately determine the continuity of the request address, and combine the sliding window, thereby providing reliable data basis for dynamically and accurately calculating the random I / O ratio.
[0035] Optionally, in an embodiment of the present application, the medium delay detection is performed on the storage hardware device to obtain the medium detection data of the storage hardware device, including: sending a preset number of standard read commands to a first logical block of different storage hard disks in the storage hardware device within a detection period of the medium delay detection, and determining the command sending time of the standard read commands and the response time of the different storage hard disks after receiving the standard read commands; calculating the delay time between the command sending time and the response time, and calculating the statistical data corresponding to the delay time, wherein the statistical data includes the arithmetic mean value and the standard deviation corresponding to the delay time; based on the arithmetic mean value in the statistical data and a preset delay time level span threshold, calculating the performance level and the performance score corresponding to the different storage hard disks, and calculating the hard disk drive comprehensive performance score corresponding to the target server according to the performance level and the performance score; creating a plurality of input and output queues corresponding to a preset queue depth, and measuring the input and output operation performance indicators and delay information of the plurality of input and output queues under different queue depths, and determining the corresponding solid state disk comprehensive performance score according to the input and output operation performance indicators and the delay information; based on the hard disk drive comprehensive performance score and the solid state disk comprehensive performance score, determining the medium detection data.
[0036] In actual execution, the embodiment of the present application also needs to perform medium delay detection, to send a standard read command (such as a standard SCSI READ(10) command) to the first logical block (LBA 0) of different storage hard disks in the server multiple times in the detection period (i.e. test period) of the medium delay detection, and determine the command sending time of the standard read command and the response time of different storage hard disks after receiving the standard read command, and calculate the delay time between the command sending time and the response time (i.e. SCSI (Small Computer System Interface) command response time), to calculate the statistical data corresponding to the delay time, and then the embodiment of the present application can calculate the hard disk drive comprehensive performance score (HDD_Score) based on the statistical data.
[0037] It should be noted that SCSI is a communication protocol used between computers and storage hardware devices (such as hard disks, SSDs, tape drives, etc.), and the SCSI command is used to control various operations of the storage hardware device, such as reading, writing, formatting, etc.; the READ(10) command is a standard command in the SCSI protocol, which is used to read data from the storage hardware device, and the number "10" represents that the length of the command is 10 bytes. The READ(10) command can read a certain number of data blocks starting from a specified logical block address (LBA). The LBA 0 position represents the first logical block of the storage hardware device, which is usually used to store the boot record or other important information of the storage hardware device.
[0038] After that, the embodiment of the present application also needs to create a plurality of input / output queues corresponding to a queue depth of 32, and measure the input / output operation performance indicators and delay information under different queue depths, to determine the solid state drive comprehensive performance score (SSD_Score). Thus, the embodiment of the present application can determine the medium detection data corresponding to the medium delay detection according to the hard disk drive comprehensive performance score and the solid state drive comprehensive performance score. Specifically, the process of measuring the hard disk drive comprehensive performance score HDD_Score to measure the HDD seek delay is as follows: The embodiment of the present application can send a standard SCSI READ(10) command to the LBA 0 position, and use a high-precision timer (nanosecond level) to record the time difference from the command issuance to the response (i.e. the delay time of the HDD seek delay), and take statistical values for 10 times of cyclic tests, and the statistical values include the arithmetic mean and the standard deviation, and the specific calculation expression is as follows:
[0039] wherein, This represents the SCSI command response time for the i-th test (in milliseconds); HDD_Latency_Avg represents the results of 10 tests. The arithmetic mean of the values reflects the average seek performance of the hard drive. The lower the value, the faster the hard drive response. HDD_Latency_Std represents the standard deviation of 10 test results. This value measures the dispersion of latency. The larger the value, the more drastic the fluctuation in hard drive response time and the worse the performance stability. HDD_Tier represents the performance level based on the average latency. Dividing by 5 is mainly to divide the hard drive performance into 0-2 levels through quantitative segmentation, thereby simplifying performance evaluation and adapting to the needs of real-world application scenarios.
[0040] Furthermore, this application embodiment can use 5ms as a level span (i.e., latency level span threshold), thereby clearly distinguishing the performance differences of hard drives with different rotational speeds. The correspondence between hard drive performance levels and scores is shown in Table 1: Table 1
[0041] Therefore, the embodiments of this application can calculate the final hard disk drive overall performance score HDD_Score through the correspondence in Table 1.
[0042] It is understood that the embodiments of this application accurately obtain hard disk performance data through multi-dimensional detection and analysis, and combine statistical methods and preset standards to scientifically evaluate the hard disk performance level and score; in addition, the embodiments of this application can also integrate multiple performance indicators to obtain reliable media detection data, so as to help to fully understand the server storage performance.
[0043] Optionally, in one embodiment of this application, multiple input / output queues corresponding to a preset queue depth are created, and the input / output operation performance indicators and latency information of the multiple input / output queues at different queue depths are measured. A comprehensive solid-state drive (SSD) performance score is determined based on the input / output operation performance indicators and latency information. This includes: calculating the number of input / output operations per second and the average latency of the SSD at the current queue depth, and calculating the efficiency-latency ratio between the number of input / output operations per second and the average latency; calculating the maximum value corresponding to the efficiency-latency ratio, and determining the target queue depth corresponding to the maximum value; determining the peak, valley, and arithmetic mean of the number of input / output operations per second within the detection period, and calculating the difference between the peak and valley values, and calculating the performance volatility corresponding to the SSD based on the difference and the arithmetic mean; converting the performance volatility into a corresponding stability gain factor, and normalizing the target queue depth to obtain the corresponding normalization result, and multiplying the stability gain factor and the normalization result to obtain the comprehensive SSD performance score.
[0044] Further, the embodiment of the present application calculates the comprehensive performance score of the solid state disk to measure the concurrent response capability of the SSD, and the process is described as follows: Firstly, the embodiment of the present application can create an NVMe I / O queue (i.e. input / output queue) with a depth of 32 (i.e. NVMe (Non-Volatile Memory Express) queue depth), and submit parallel 4K read requests to measure the IOPS and delay information under each queue depth, and the calculation process corresponds to the mathematical expression:
[0045] wherein SSD_QD_Optimal (i.e. optimal queue depth) is the QD value corresponding to the maximum ratio of the IOPS and delay calculated under different queue depths (QD), which indicates the balance point of the SSD in throughput efficiency and delay, i.e. the operation queue depth for obtaining the highest IOPS at the unit delay cost, for example, if the ratio is the highest when QD=16, then SSD_QD_Optimal=16; is the input / output operations per second under the current queue depth to measure the throughput performance; is the average delay (unit: millisecond) under the current queue depth, which reflects the corresponding response speed; SSD_Jitter (i.e. performance fluctuation rate) indicates the proportion of the range fluctuation of IOPS to the average value, which quantifies the performance stability of the SSD and reflects the performance consistency of the SSD under continuous load; are respectively the peak value and the valley value of IOPS in the test period; is the arithmetic mean of IOPS in the test period; SSD_Score (i.e. comprehensive performance score of the solid state disk) is the normalized score calculated by combining the optimal queue depth and the fluctuation rate, and the higher the score (close to 1), the better the comprehensive performance of the SSD, which represents that the SSD can not only run efficiently at low queue depth, but also maintain high stability; is the normalization of the optimal queue depth to the 32 queue benchmark (i.e. standard high-performance SSD saturation point); 1-SSD_Jitter is the stability gain factor converted from the fluctuation rate.
[0046] Therefore, the embodiment of the present application can comprehensively evaluate the performance of the solid state disk by calculating the efficiency-delay ratio and the queue depth corresponding to the maximum value, and combining the stability factor converted from the performance fluctuation rate, so as to comprehensively reflect the running efficiency and stability of the solid state disk.
[0047] Optionally, in an embodiment of the present application, the hardware state analysis on the storage hardware device is performed to obtain the hardware state data of the storage hardware device, including: obtaining temperature change gradient information corresponding to a dual in-line memory module, and determining the current temperature of the storage hardware device through the temperature change gradient information; collecting a hard disk original life percentage and a durability factor corresponding to a solid state disk, and calculating an actual wear degree of the solid state disk according to the hard disk original life percentage and the durability factor; and determining the hardware state data based on the current temperature and the actual wear degree.
[0048] As an implementable manner, the embodiment of the present application can access a BMC (Baseboard Management Controller) sensor through an IPMI 0x2E command to read a memory slot temperature register to obtain a DIMM temperature gradient T curr, and determine the current temperature of the storage hardware device according to the DIMM temperature gradient T curr. It should be noted that the DIMM temperature of the RAID card refers to the working temperature of the DIMM memory used by the on-board cache of the RAID card. These DIMMs are not system memories on the motherboard, but cache modules (usually DDR3 / DDR4) specially inserted on the RAID card, used to cache read and write data of the RAID card, and improve I / O performance.
[0049] Then, the embodiment of the present application can collect a hard disk original life percentage and a durability factor corresponding to a solid state disk to calculate an actual wear degree of the solid state disk, so as to determine the hardware state data in combination with the current temperature and the actual wear degree of the storage hardware device.
[0050] Thus, the embodiment of the present application comprehensively and accurately reflects the hardware state by combining the memory temperature change and the actual wear degree of the solid state disk, and provides reliable data support for evaluating the server hardware health condition.
[0051] Optionally, in an embodiment of the present application, the collection of the hard disk original life percentage and the durability factor corresponding to the solid state disk to calculate the actual wear degree of the solid state disk according to the hard disk original life percentage and the durability factor includes: obtaining a theoretical data write-erase complete operation number corresponding to the solid state disk, and calculating a consumption ratio of the solid state disk according to the theoretical data write-erase complete operation number to determine the hard disk original life percentage through the consumption ratio; determining a hard disk type corresponding to the solid state disk, and calculating a durability factor of the solid state disk according to the hard disk type; and calculating a ratio between the hard disk original life percentage and the durability factor to determine the actual wear degree according to the ratio.
[0052] Specifically, the embodiment of the present application can collect the raw life percentage information of the hard disk from the SMART (Self-Monitoring Analysis and Reporting Technology) information of the NVMe solid state disk by the command nvme smart-log / dev / nvme0 | grep "percentage_used", and process the collected raw life percentage of the hard disk. The processing process is as follows:
[0053] Wherein, Wear_Level is the actual wear level, indicating the SSD life consumption percentage after the endurance factor correction, reflecting the real wear state, ranging from 0% (brand new) to 100% (exhausted); Percentage_Used represents the raw life percentage of the hard disk, which is the consumption ratio calculated by the controller based on the theoretical maximum P / E cycle number (i.e. the theoretical data write-erase complete operation number), which can be obtained from the SMART log by the above command, wherein, P / E cycle (Program / Erase Cycle) refers to the complete operation number of NAND (Not AND Flash Memory) flash memory cell from data writing to complete erasing, which is the core indicator to measure the life of solid state disk; Endurance_Factor represents the endurance factor, which quantifies the degree of life loss in actual use, which can be valued according to the type of SSD, as shown in Table 2: Table 2
[0054] Therefore, the embodiment of the present application determines the raw life percentage by combining the theoretical erase-write number, and calculates the endurance factor according to the type of hard disk, so as to determine the actual wear level by the ratio of the two, so as to accurately reflect the real wear state of the solid state disk, and provide a scientific basis for evaluating its life.
[0055] In step S102, the corresponding sequential write request proportion is determined according to the random input-output ratio, and the target decision logic corresponding to the storage hardware device is matched based on the sequential write request proportion and the current temperature of the storage hardware device contained in the hardware state data.
[0056] Further, the embodiments of the present application can calculate the sequential write request proportion by the random input-output ratio, so as to select the corresponding decision logic for the server in combination with the current temperature and the sequential write request proportion. Wherein, the sequential write means that the data is written continuously according to the physical position of the storage medium; and the random read means that the data is read from any non-continuous position of the storage medium (such as frequent head jumping of a mechanical hard disk).
[0057] Therefore, the embodiments of the present application solve the cache overflow problem caused by the response delay in the prior art through the synergistic effect of the multi-dimensional real-time perception layer and the dynamic decision engine, thereby avoiding the write cache full forced degradation to the transparent write mode (in the transparent write mode, the data is written into the cache and the underlying storage medium at the same time, and only when both of them confirm the writing completion, the write success signal is returned to the host), and significantly improving the stability of I / O response.
[0058] Optionally, in an embodiment of the present application, based on the sequential write request proportion and the current temperature of the storage hardware device contained in the hardware state data, the target decision logic corresponding to the storage hardware device is matched, including: inputting the random input-output ratio, the medium detection data and the hardware state data into the preset dynamic decision engine, so as to determine whether the current temperature is greater than the preset temperature threshold and whether the sequential write request proportion is less than the preset proportion threshold through the dynamic decision engine; if the current temperature is greater than the temperature threshold, the temperature priority decision logic is selected as the target decision logic; if the current temperature is less than or equal to the temperature threshold, and the sequential write request proportion is greater than the proportion threshold, the sequential write optimization decision logic is selected as the target decision logic; if the current temperature is less than or equal to the temperature threshold, and the sequential write request proportion is less than or equal to the proportion threshold, the elastic weight calculation decision logic is selected as the target decision logic.
[0059] In the specific implementation process, as shown in Figure 2 the embodiments of the present application can input the random input-output ratio, the medium detection data and the hardware state data into the preset dynamic decision engine, so that the dynamic decision engine selects the decision logic (i.e. decision branch) corresponding to the server according to the current temperature and the sequential write request proportion, and the specific selection process is as follows: (1) When the real-time value of the DIMM temperature sensor (i.e. the current temperature) exceeds 85℃ (i.e. the temperature threshold), the temperature priority decision logic is selected as the target decision logic; (2) When the temperature is in the safe range, i.e. the current temperature is less than or equal to 85℃, and the sequential write request proportion exceeds 60% (i.e. the proportion threshold) through the LBA continuity analysis, the sequential write optimization decision logic is selected as the target decision logic; (3) When the above two special scenarios are not met, or the current temperature is not greater than the temperature threshold and the sequential write request proportion is not greater than the proportion threshold, the elastic weight calculation decision logic is selected as the target decision logic.
[0060] Therefore, the embodiments of the present application select the corresponding decision logic according to the temperature and the proportion of sequential write requests, so as to adapt to the server state and effectively improve the rationality of server decision and the adaptability of hardware operation.
[0061] In step S103, based on the target decision logic, in combination with the random input / output ratio, the medium detection data and the hardware state data, the read / write cache allocation ratio of the storage hardware device is determined, the hot data corresponding to the storage hardware device is identified, and the cache allocation adjustment of the storage hardware device is performed according to the read / write cache allocation ratio and the hot data, so as to obtain the dynamically allocated cache, and the dynamically allocated cache is divided into multi-level cache partitions of different priorities, and the data type of the to-be-allocated cache data is determined, so as to allocate the to-be-allocated cache data to the corresponding cache partition according to the data type and the priority of the multi-level cache partition.
[0062] After that, the embodiments of the present application can construct the corresponding execution layer through the cache allocator, the migration controller and the Qos resource allocator, so as to determine the read / write cache allocation ratio and the hot data of the RAID card based on the execution layer and in combination with the above-selected decision logic, and design three-level cache partitions of different priorities as physical isolation through the QoS virtual channel, and at the same time, in combination with the preemption resource recycling mechanism, the resource allocation of the multi-level cache partition is performed, so as to solve the multi-service resource competition problem and reduce the delay fluctuation and the blocking rate.
[0063] Optionally, in an embodiment of the present application, determining the read / write cache allocation ratio and the hot data corresponding to the target redundant disk array card comprises: based on the pre-constructed multi-dimensional feature analysis model, collecting the data access frequency, access interval, data correlation and data modification frequency of the cache data in the target server, so as to generate the corresponding data heat feature set according to the data access frequency, access interval, data correlation and data modification frequency; and based on the data heat feature set and the preset dynamic threshold algorithm, identifying the hot data in the cache data.
[0064] As an implementable way, the embodiments of the present application can first construct a multi-dimensional feature analysis model fused with a time decay factor through the migration controller, which not only collects the data access frequency, interval, correlation and modification frequency, but also introduces the access timestamp weight (the recent access weight is higher), so as to generate a dynamically changing data heat feature set. Secondly, the embodiment of the present application can combine the preset dynamic threshold algorithm to update the hotness threshold value in real time through the sliding window: when the data access frequency of a certain period of time increases by more than 30%, the threshold value is automatically adjusted downward to quickly identify new hot spots; if the data access frequency is lower than the baseline for 48 consecutive hours, the threshold value is adjusted upward to avoid misjudgment. At the same time, through data correlation clustering, frequently co-accessed data is classified into the same hot spot group, improving the identification integrity, so as to identify the hot spot data in the cache data.
[0065] Therefore, the embodiment of the present application can accurately identify hot spot data by introducing time decay and dynamic threshold value and combining correlation clustering, so as to well adapt to the change of access mode and provide a scientific basis for cache allocation.
[0066] Optionally, in an embodiment of the present application, based on the target decision logic and in combination with the random input / output ratio, medium detection data and hardware state data, the read / write cache allocation ratio of the storage hardware device is determined, the hot spot data corresponding to the storage hardware device is identified, and the storage hardware device is adjusted for cache allocation according to the read / write cache allocation ratio and the hot spot data, to obtain a dynamically allocated cache, including: when the target decision logic is temperature priority decision logic, the read cache ratio is adjusted to reduce the read cache ratio to a target read cache ratio, and the hot spot data is migrated to the non-volatile memory according to the target read cache ratio, and the fan speed is adjusted to a target speed value, and the corresponding fan is controlled to rotate based on the target speed value to cool the corresponding storage hard disk; when the target decision logic is sequential write optimization decision logic, a cache reallocation operation is performed to set the read cache ratio to a target fixed ratio, and a dedicated direct memory access channel is established to perform corresponding cache allocation operations according to the target fixed ratio and the dedicated direct memory access channel; when the target decision logic is elastic weight calculation decision logic, the corresponding elastic read cache ratio is calculated based on the medium detection data, and the corresponding cache allocation operation is performed according to the elastic read cache ratio to obtain a dynamically allocated cache.
[0067] It should be noted that the embodiments of the present application can perform different decision actions according to different decision logics and in combination with the preset cache allocator and migration controller, as described in detail below: 1. When the target decision logic is temperature priority decision logic, the embodiment of the present application can adjust the cache ratio through the cache allocator to immediately reduce the read cache ratio to a target read cache ratio, such as 30%, and migrate the hot spot data to the NVMe non-volatile storage area through the migration controller, while increasing the fan speed to a target speed value, such as 80%, through the BMC interface, so as to reduce the load of the temperature sensitive area and prevent hardware failure.
[0068] 2、When the target decision logic is sequential write optimization decision logic, the embodiment of the present application can perform cache reallocation operations to set the read cache ratio to a target fixed ratio, such as 40%, and activate the direct write channel to enable direct write for data blocks larger than 128KB, establish a dedicated DMA (Direct Memory Access, direct memory access channel) channel to bypass the cache management layer, thereby reducing cache management overhead and greatly improving sequential write throughput.
[0069] 3、When the target decision logic is elastic weight calculation decision logic, the embodiment of the present application can calculate the corresponding elastic read cache ratio according to the medium detection data to perform corresponding cache allocation operations.
[0070] In summary, the embodiment of the present application can overturn the traditional passive alarm mode. When the DIMM temperature is greater than the temperature threshold of 85℃, the hot spot data is automatically migrated to the NVMe storage area and the fan speed is increased, and the JEDEC (Joint Electron Device Engineering Council, Solid State Technology Association) standard predicted 8 times failure rate is reduced to 0; when the write cache occupancy is greater than 95%, the emergency direct write channel is enabled, and the risk of high-temperature downtime and cache data loss is completely eliminated; for other scenarios, the embodiment of the present application can use the default elastic weight calculation method to scientifically calculate the appropriate read cache ratio, thereby balancing the hard disk running performance and stability.
[0071] It can be understood that the embodiment of the present application can use a scientific method to calculate the final read cache ratio based on the dynamic decision engine of the elastic weight calculation model for atypical scenarios; for high-temperature scenarios, automatically trigger hot spot data migration and cooling measures; and the embodiment of the present application supports detecting sequential write scenarios, using a large block data (>128KB) direct write channel, with a throughput of 22GB / s; the embodiment of the present application also supports limiting according to actual applications, thereby ensuring that the read cache ratio is within a reasonable range.
[0072] Optionally, in an embodiment of the present application, the fan speed is adjusted to a target speed value, and based on the target speed value, the corresponding fan rotation is controlled to cool the corresponding storage hard disk, comprising: collecting real-time temperature data and current running load information of the storage hard disk, and determining the cooling demand level of the storage hard disk according to the real-time temperature data and the current running load information; inputting the cooling demand level into a pre-constructed fan speed-heat dissipation efficiency matching model to output a target speed value corresponding to the cooling demand level; obtaining real-time running parameters of the fan, and generating a corresponding fan speed adjustment instruction based on the target speed value and the real-time running parameters; sending the fan speed adjustment instruction to a pre-set fan driving module to control the fan to rotate at the target speed value to cool the storage hard disk, while collecting temperature change data of the storage hard disk during the cooling process; based on the temperature change data, it is judged whether the temperature of the storage hard disk meets the pre-set stability requirement, if the temperature of the storage hard disk does not meet the pre-set stability requirement, the target speed value is recalculated until the temperature of the storage hard disk meets the pre-set stability requirement.
[0073] As an implementable way, in the collection of real-time temperature data and running load information of the storage hard disk, a temperature gradient sensor array is introduced, and the temperature difference of different areas of the hard disk is recorded, and a three-dimensional feature matrix is constructed combining IOPS (input / output operation times per second) and data throughput in the load information. In an embodiment of the present application, the cooling demand level division can adopt a dynamic hierarchical mechanism, specifically, when the hard disk core area temperature is more than 60℃ and the load fluctuation is more than 20%, the emergency cooling is triggered; when the temperature is 50-60℃ and the load is stable, it is moderate cooling; when the temperature is less than 50℃, it is regular heat dissipation, and the intermediate state is refined by interpolation algorithm. In addition, the fan speed-heat dissipation efficiency matching model of the embodiment of the present application can integrate a neural network prediction layer to train based on historical heat dissipation data to predict the temperature change curve under different speeds, and the output target speed value is accompanied by a buffer coefficient (such as ±5% speed fluctuation interval) to avoid frequent start and stop of the fan. During the cooling process, the embodiment of the present application can use infrared thermal imaging to record the hard disk temperature field distribution in real time, if the temperature drop is less than 1℃ / min for 3 minutes continuously and the stable threshold (such as 45℃±2℃) is not reached, the adaptive optimization algorithm is started to recalculate the speed combined with the current load trend. Therefore, the embodiment of the present application can realize precise temperature control through multi-dimensional perception and intelligent prediction, and take into account the heat dissipation efficiency and equipment loss, so as to dynamically adapt to the state change of the hard disk, thereby improving the stability and energy economy of the storage system.
[0074] Optionally, in an embodiment of the present application, when the target decision logic is the elastic weight calculation decision logic, based on the medium detection data, the corresponding elastic read cache ratio is calculated, including: determining the cache read baseline ratio corresponding to the solid state disk, the random input-output weight coefficient, the hard disk drive performance compensation coefficient and the solid state disk performance compensation coefficient; obtaining the current hard disk temperature of the solid state disk, and calculating the temperature difference between the current hard disk temperature and the preset hard disk temperature threshold, so as to calculate the temperature penalty term corresponding to the solid state disk according to the temperature difference; determining the read-write critical point corresponding to the solid state disk, and calculating the life consumption penalty term corresponding to the solid state disk according to the read-write critical point and the actual wear degree; calculating the first product result of the random input-output weight coefficient and the random input-output ratio, and calculating the second product result between the hard disk drive performance compensation coefficient and the hard disk drive comprehensive performance score, and calculating the third product result between the solid state disk performance compensation coefficient and the solid state disk comprehensive performance score; based on the first product result, the second product result, the third product result, the cache read baseline ratio, the temperature penalty term and the life consumption penalty term, the elastic read cache ratio is calculated.
[0075] In the specific implementation process, when the target decision logic is the elastic weight calculation decision logic, the elastic read cache ratio can be calculated by the following formula
[0076] , wherein, is the system default cache read baseline ratio, reflecting the ideal state in the non-interference scenario, which can generally be set to 0.5; represents the random I / O weight coefficient (i.e. the random input-output weight coefficient), which generally takes a fixed value of 0.3; represents the first product result; is the random input-output ratio, the higher the value, the more read cache is needed to cope with random access; is the HDD performance compensation coefficient (i.e. the hard disk drive performance compensation coefficient); HDD_Score represents the second product result; is the SSD performance compensation coefficient (i.e. the solid state disk performance compensation coefficient), which ranges from 0.05 to 0.15, and generally takes a fixed value of 0.1; SSD_Score represents the third product result; is the temperature penalty term, which is enabled when the temperature is greater than 70℃, the SSD high temperature causes the read-write delay to increase and the cache efficiency to decrease, when the temperature rises from 70℃ to 85℃ (the temperature variation interval is 15℃ interval), 1% of the cache access frequency needs to be reduced for every 1℃ rise, and the calculation expression of this temperature penalty term is: ; wherein, is a negative impact value of life loss (i.e., a life loss penalty term), the more the SSD cycle number consumption, the slower the cache response speed, The calculation expression of is: .
[0077] According to the calculation expression of the above It can be known that compensation exceeding 5% may cause read performance collapse, and therefore, 5% can be taken as a read-write critical point of read-write balance.
[0078] Therefore, when the HDD performance is optimal, the SSD performance is optimal, the temperature is 85℃, and the hard disk wear degree is 100%, the corresponding read cache ratio can be calculated as The minimum value = 0.5 + 0 - 0.1 - 0.1 - 0.15 - 0.05 = 0.1; when the HDD performance is worst, the SSD performance is worst, the temperature is 70℃, and the hard disk wear degree is 0%, the read cache ratio can be calculated as The maximum value = 0.5 + 0.3 -0.02 - 0 - 0 - 0 = 0.78; therefore, the read cache ratio in the embodiment of the present application is The theoretical output range of is [0.1, 0.78], so it can be forced to be constrained to a reasonable interval [0.3, 0.7] in the actual application limiting process.
[0079] Therefore, the embodiment of the present application calculates the elastic read cache ratio by comprehensively considering multi-dimensional parameters, introduces temperature and life loss penalty term, and accurately adapts to the hard disk state, so as to clearly calculate the theoretical and actual constraint interval of the elastic read cache ratio, avoid performance collapse, guarantee read-write balance, and improve the rationality of cache resource allocation and the stability of hard disk operation.
[0080] Optionally, in an embodiment of the present application, the dynamic allocation cache is divided into cache partitions of different priorities, and the data type of the cache data to be allocated is determined, so that the cache data to be allocated is allocated to the corresponding cache partition according to the data type and the priority of the multi-level cache partition, including: based on the preset cache allocation rule, the dynamic allocation cache is processed by three-level cache partitioning to divide the dynamic allocation cache into multi-level cache partitions of different priorities, wherein the multi-level cache partitions of different priorities include a dedicated area corresponding to a first priority, a normal area corresponding to a second priority, and a recycling area corresponding to a third priority, and the priority of the second priority is greater than that of the third priority, and the priority of the second priority is less than that of the first priority; the preset service level definition sensitive task is allocated in the dedicated area, and the concurrent access data is allocated in the normal area, and the compressed data or the obsolete data is allocated to the recycling area, and the remaining space size of the dedicated area is detected; it is judged whether the remaining space size is less than the preset space threshold, wherein in the case where it is detected that the remaining space size is less than the preset space threshold, a task freezing mechanism is triggered to pause the data write to the recycling area which does not meet the preset criticality requirement, and the cache data of the recycling area is released to the dedicated area.
[0081] It should be noted that the embodiments of the present application can be constructed by an execution layer of a cache allocator, a migration controller and a Qos resource allocator, and the cache allocation and data migration operations are performed in combination with the above-mentioned selected decision logic. The execution layer is described in detail as follows: 1. Cache allocator: The core function of the cache allocator is to perform real-time read / write cache ratio calculation and manage the physical address allocation of the DDR4 cache pool to perform corresponding operations according to the QoS virtual channel strategy.
[0082] It should be noted that the RAID cache QoS refers to a mechanism for prioritizing and dynamically allocating cache resources (such as read cache and write cache) in a RAID card or storage system, to ensure that critical business or high-priority tasks can obtain better cache services in resource contention, and to avoid affecting overall performance due to excessive cache occupation by low-priority tasks.
[0083] 2. Migration controller: The core function of the migration controller is to identify hot data and migrate the hot data to ensure data consistency during migration.
[0084] 3. Qos resource allocator: The Qos resource allocator can perform corresponding resource allocation operations according to the QoS virtual channel strategy (i.e., the Qos resource allocation strategy) in combination with the cache allocator and the migration controller, as described in detail below: (1) Three-level cache partitioning: 1) VIP area (i.e. exclusive area, for high priority or first priority): physically isolated contiguous cache block, dedicated to SLA (Service Level Agreement, service level agreement) sensitive tasks such as database transaction logs; 2) Normal area (for medium priority, i.e. second priority): shared cache pool, supporting multi-service concurrent access.
[0085] 3) Recycle area (for low priority, i.e. third priority): storing compressible / obsolete data (such as backup logs, etc.).
[0086] (2) Preemptive resource recycling: When the remaining space of the VIP area is <20%, trigger low-priority task freezing, i.e. suspend non-critical data writing (such as statistical reports, etc.), and release the recycle area cache to the VIP area (preferentially migrate cold data).
[0087] (3) Perform corresponding cache allocation operations according to the preset cache allocation rules, as shown in Table 3: Table 3
[0088] As can be seen from Table 3, the embodiments of the present application can control the cache allocator and the migration controller according to the minimum threshold and the maximum threshold corresponding to different cache partitions, to perform corresponding resource allocation and data migration operations.
[0089] Thus, the embodiments of the present application can avoid the occurrence of DIMM high-temperature downtime through high-temperature data migration, and can solve the problem of multi-service resource competition through the QoS three-level cache partition (i.e. VIP area / normal area / recycle area) and the preemptive resource recycling mechanism, thereby allocating exclusive cache channels (with a minimum guarantee of 30%) for high-priority tasks such as financial transaction logs, reducing delay fluctuations; secondly, the embodiments of the present application automatically freeze low-priority tasks (such as backup jobs) when the space of the VIP area is less than 20%, to reduce the blocking rate. In addition, the embodiments of the present application replace the NPU (Neural-network Processing Unit, neural network processor) / FPGA (Field Programmable Gate Array, field programmable gate array) dependence of the intelligent solution with a lightweight software algorithm, which only needs software upgrade for low-cost deployment and operation, so that the deployment cost is reduced to 1 / 3 of the intelligent solution.
[0090] The execution logic of the cache dynamic allocation method of the storage hardware device of the present application is described below in conjunction with the accompanying drawings.
[0091] Figure 3The figure shows the execution logic of the cache dynamic allocation method of the storage hardware device. Figure 3 As shown in the figure, the execution process of the cache dynamic allocation method of the storage hardware device is as follows: S301: Start the collection cycle; S302: Collect data in parallel, and simultaneously go to S303, S304 and S305; S303: Perform input / output mode analysis according to the collected data, and go to S306; S304: Perform medium delay detection according to the collected data, and go to S307; S305: Perform hardware state acquisition according to the collected data, and go to S308; S306: Calculate the random input / output ratio through input / output mode analysis, and go to S309; S307: Obtain delay data through medium delay detection, and go to S309; S308: Obtain the current temperature according to the hardware state, and calculate the actual wear degree, and go to S309; S309: Send the random input / output ratio, delay data, current temperature and actual wear degree to the dynamic decision engine; S3010: Perform decision logic selection through the dynamic decision engine; S3011: When the current temperature is not greater than the temperature threshold and the sequential write request proportion is greater than the proportion threshold, set the read cache proportion to the target read cache proportion; S3012: When the current temperature value is greater than the temperature threshold, set the read cache proportion to the target fixed proportion, and migrate the hot data; S3013: In other scene conditions, calculate the elastic read cache proportion; S3014: Update the cache allocation according to the corresponding read cache proportion; S3015: Quality of service resource allocation.
[0092] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment.
[0093] The embodiment of the application also provides a cache dynamic allocation device of a storage hardware device.
[0094] As shown in the figure, the cache dynamic allocation device of the storage hardware device 10 comprises a multi-dimensional perception module 100, a matching module 200 and a cache allocation module 300. Figure 4 As shown in the figure, the cache dynamic allocation device of the storage hardware device 10 comprises a multi-dimensional perception module 100, a matching module 200 and a cache allocation module 300. As shown in the figure, the cache dynamic allocation device of the storage hardware device 10 comprises a multi-dimensional perception module 100, a matching module 200 and a cache allocation module 300.
[0095] The multi-dimensional perception module 100 is configured to perform input / output mode analysis on the storage hardware device in the target server to obtain a random input / output ratio of the storage hardware device, perform medium delay detection on the storage hardware device to obtain medium detection data of the storage hardware device, and perform hardware state analysis on the storage hardware device to obtain hardware state data of the storage hardware device.
[0096] The matching module 200 is configured to determine a corresponding sequential write request proportion according to the random input / output ratio, and match a target decision logic corresponding to the storage hardware device based on the sequential write request proportion and a current temperature of the storage hardware device contained in the hardware state data.
[0097] The cache allocation module 300 is configured to determine a read / write cache allocation ratio corresponding to the storage hardware device based on the target decision logic and in combination with the random input / output ratio, the medium detection data, and the hardware state data, identify hot data corresponding to the storage hardware device, perform cache allocation adjustment on the storage hardware device according to the read / write cache allocation ratio and the hot data to obtain a dynamically allocated cache, divide the dynamically allocated cache into multi-level cache partitions of different priorities, determine a data type of to-be-allocated cache data, and allocate the to-be-allocated cache data to a corresponding cache partition according to the data type and the priority of the multi-level cache partitions.
[0098] Optionally, in an embodiment of the present application, the multi-dimensional perception module 100 includes a response unit, a first judgment unit, an assignment unit, and a first determination unit.
[0099] The response unit is configured to determine a starting logical block address and a number of continuous blocks of a current request and a starting address of a next request in response to a current request of the target server.
[0100] The first judgment unit is configured to calculate a sum of the starting logical block address and the number of continuous blocks of the current request, calculate an absolute value of a difference between the sum of the starting logical block address and the number of continuous blocks and the starting address of the next request, and determine whether the absolute value of the difference is greater than a preset window threshold.
[0101] The assignment unit is configured to determine that a logical block address corresponding to the current request satisfies a preset continuity requirement if the absolute value of the difference is less than or equal to the window threshold, and assign a first target value to a preset logical block address indicator, or determine that the logical block address does not satisfy the continuity requirement and assign a second target value to the logical block address indicator.
[0102] The first determination unit is configured to determine a request time point of the current request, and calculate the random input / output ratio based on the first target value, the second target value, the request time point, and a preset sliding window size.
[0103] Optionally, in an embodiment of the present application, the multi-dimensional perception module 100 further comprises a second determining unit, a first calculating unit, a second calculating unit, a creating unit and a third determining unit.
[0104] The second determining unit is configured to send a preset number of standard read commands to a first logical block of different storage hard disks in the storage hardware device within a detection period of the medium delay detection, and determine a command sending time of the standard read commands and a response time of the different storage hard disks after receiving the standard read commands.
[0105] The first calculating unit is configured to calculate a delay time between the command sending time and the response time, and calculate statistical data corresponding to the delay time, wherein the statistical data comprises an arithmetic mean value and a standard deviation corresponding to the delay time.
[0106] The second calculating unit is configured to calculate a performance grade and a performance score corresponding to the different storage hard disks based on the arithmetic mean value in the statistical data and a preset delay time grade span threshold, and calculate a hard drive comprehensive performance score corresponding to the target server according to the performance grade and the performance score.
[0107] The creating unit is configured to create a plurality of input / output queues corresponding to a preset queue depth, measure input / output operation performance indicators and delay information of the plurality of input / output queues under different queue depths, and determine a solid state drive comprehensive performance score according to the input / output operation performance indicators and the delay information.
[0108] The third determining unit is configured to determine the medium detection data based on the hard drive comprehensive performance score and the solid state drive comprehensive performance score.
[0109] Optionally, in an embodiment of the present application, the multi-dimensional perception module 100 further comprises a reading unit and a fourth determining unit.
[0110] The reading unit is configured to obtain temperature change gradient information corresponding to a preset dual in-line memory module, and determine a current temperature of the storage hardware device through the temperature change gradient information. The collecting unit is configured to collect a hard disk original life percentage and a durability factor corresponding to the solid state drive, so as to calculate an actual wear degree corresponding to the solid state drive according to the hard disk original life percentage and the durability factor.
[0111] The fourth determining unit is configured to determine the hardware state data based on the current temperature and the actual wear degree.
[0112] Optionally, in an embodiment of the present application, the creating unit comprises a first ratio calculating subunit, a maximum value calculating subunit, a volatility calculating subunit and a normalization subunit.
[0113] The first ratio calculation subunit is configured to calculate the input / output operations per second and the average latency of the solid state disk at the current queue depth, and calculate an efficiency-latency ratio between the input / output operations per second and the average latency.
[0114] The maximum value calculation subunit is configured to calculate a maximum value corresponding to the efficiency-latency ratio, and determine a target queue depth corresponding to the maximum value.
[0115] The volatility calculation subunit is configured to determine a peak value, a valley value and an arithmetic mean value of the input / output operations per second in the detection period, calculate a difference value between the peak value and the valley value, and calculate a performance volatility of the solid state disk according to the difference value and the arithmetic mean value.
[0116] The normalization subunit is configured to convert the performance volatility into a corresponding stability gain factor, normalize the target queue depth to obtain a corresponding normalization result, and multiply the stability gain factor and the normalization result to obtain a comprehensive performance score of the solid state disk.
[0117] Optionally, in an embodiment of the present application, the acquisition unit comprises a first acquisition subunit, a durability calculation subunit and a second ratio calculation subunit.
[0118] The first acquisition subunit is configured to acquire a theoretical data write-erase complete operation number corresponding to the solid state disk, and calculate a consumption ratio corresponding to the solid state disk according to the theoretical data write-erase complete operation number, so as to determine a hard disk original life percentage through the consumption ratio.
[0119] The durability calculation subunit is configured to determine a hard disk type corresponding to the solid state disk, and calculate a durability factor corresponding to the solid state disk according to the hard disk type.
[0120] The second ratio calculation subunit is configured to calculate a ratio between the hard disk original life percentage and the durability factor, so as to determine an actual wear degree according to the ratio.
[0121] Optionally, in an embodiment of the present application, the matching module 200 comprises a second judgment unit, a first selection unit, a second selection unit and a third selection unit.
[0122] The second judgment unit is configured to input the random input / output ratio, the medium detection data and the hardware state data into a preset dynamic decision engine, so as to determine whether the current temperature is greater than a preset temperature threshold value and whether the sequential write request proportion is less than a preset proportion threshold value through the dynamic decision engine.
[0123] The first selection unit is configured to select a temperature priority decision logic as a target decision logic if the current temperature is greater than the temperature threshold value.
[0124] The second selecting unit is configured to select the sequential write optimization decision logic as the target decision logic if the current temperature is less than or equal to the temperature threshold value and the sequential write request proportion is greater than the proportion threshold value.
[0125] The third selecting unit is configured to select the elastic weight calculation decision logic as the target decision logic if the current temperature is less than or equal to the temperature threshold value and the sequential write request proportion is less than or equal to the proportion threshold value.
[0126] Optionally, in an embodiment of the present application, the cache allocation module 300 comprises a control unit, a redistribution unit and a third calculation unit.
[0127] The control unit is configured to, when the target decision logic is the temperature priority decision logic, adjust the read cache proportion to reduce the read cache proportion to a target read cache proportion, migrate the hot data to the non-volatile memory according to the target read cache proportion, adjust the fan rotating speed to a target rotating speed value, and control the corresponding fan to rotate based on the target rotating speed value to perform the cooling operation on the corresponding storage hard disk.
[0128] The redistribution unit is configured to, when the target decision logic is the sequential write optimization decision logic, perform the cache redistribution operation to set the read cache proportion to a target fixed proportion, and establish a dedicated direct memory access channel to perform the corresponding cache allocation operation according to the target fixed proportion and the dedicated direct memory access channel.
[0129] The third calculation unit is configured to, when the target decision logic is the elastic weight calculation decision logic, calculate a corresponding elastic read cache proportion based on the medium detection data, and perform the corresponding cache allocation operation according to the elastic read cache proportion to obtain the dynamically allocated cache.
[0130] Optionally, in an embodiment of the present application, the third calculation unit comprises a coefficient determination subunit, a second acquisition subunit, a penalty term calculation subunit, a product calculation subunit and a cache proportion calculation subunit.
[0131] The coefficient determination subunit is configured to determine the cache reading baseline proportion, the random input-output weight coefficient, the hard disk drive performance compensation coefficient and the solid state disk performance compensation coefficient corresponding to the solid state disk.
[0132] The second acquisition subunit is configured to acquire the current hard disk temperature of the solid state disk, and calculate a temperature difference value between the current hard disk temperature and a preset hard disk temperature threshold value to calculate the temperature penalty term corresponding to the solid state disk according to the temperature difference value.
[0133] The penalty term calculation subunit is configured to determine the read-write critical point corresponding to the solid state disk, and calculate the life consumption penalty term corresponding to the solid state disk according to the read-write critical point and the actual wear degree.
[0134] a product calculation sub-unit configured to calculate a first product result of the random input-output weight coefficient and the random input-output ratio, calculate a second product result between the hard disk drive performance compensation coefficient and the hard disk drive comprehensive performance score, and calculate a third product result between the solid state disk performance compensation coefficient and the solid state disk comprehensive performance score.
[0135] a cache ratio calculation sub-unit configured to calculate the elastic read cache ratio based on the first product result, the second product result, the third product result, the cache read baseline ratio, the temperature penalty term, and the life loss penalty term.
[0136] The description of the features in the embodiments of the cache dynamic allocation apparatus of the storage hardware device can refer to the related description of the embodiments of the cache dynamic allocation method of the storage hardware device, which will not be repeated here.
[0137] The embodiments of the present application also provide an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the embodiments of the cache dynamic allocation method of the storage hardware device.
[0138] The embodiments of the present application also provide a non-volatile computer readable storage medium, which stores a computer program, and the computer program is configured to perform the steps in any of the embodiments of the cache dynamic allocation method of the storage hardware device when running.
[0139] In an example embodiment, the non-volatile computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0140] The embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in any of the embodiments of the cache dynamic allocation method of the storage hardware device.
[0141] The embodiments of the present application also provide another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in any of the embodiments of the cache dynamic allocation method of the storage hardware device.
[0142] Those skilled in the art will further realize that the mere concepts, teachings, and embodiments described herein are merely meant to provide examples of the various aspects of the present application and that various modifications, equivalents and alternatives are intended to fall within the scope of the present application. Accordingly, the appended claims are intended to embrace all such alterations, modifications, and improvements as fall within the scope of the present application. As can be seen, the application provides a novel and improved method, device and medium for dynamically allocating cache of storage hardware.
[0143] The above describes in detail the method, device, equipment and medium for dynamically allocating cache of storage hardware provided by the present application. The principles and implementation manners of the present application are described by using specific examples in the present application. The above description of the embodiments is only for helping to understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for dynamically allocating cache in a storage hardware device, characterized in that, The method comprises the following steps: performing input / output mode analysis on a storage hardware device in a target server to obtain a random input / output ratio of the storage hardware device, performing medium delay detection on the storage hardware device to obtain medium detection data of the storage hardware device, and performing hardware state analysis on the storage hardware device to obtain hardware state data of the storage hardware device; determining a corresponding sequential write request proportion according to the random input / output ratio, and matching a target decision logic corresponding to the storage hardware device based on the sequential write request proportion and a current temperature of the storage hardware device contained in the hardware state data; based on the target decision logic, in combination with the random input / output ratio, the medium detection data and the hardware state data, determining a read / write cache allocation ratio corresponding to the storage hardware device, identifying hot data corresponding to the storage hardware device, and performing cache allocation adjustment on the storage hardware device according to the read / write cache allocation ratio and the hot data to obtain a dynamically allocated cache, dividing the dynamically allocated cache into multiple cache partitions of different priorities, and determining a data type of to-be-allocated cache data to allocate the to-be-allocated cache data to a corresponding cache partition according to the data type and the priority of the multiple cache partitions.
2. The method of claim 1, wherein, The input / output mode analysis on the storage hardware device in the target server to obtain the random input / output ratio of the storage hardware device comprises: in response to a current request of the target server, determining a starting logical block address and a number of continuous blocks of the current request and a starting address of a next request; calculating a sum of the starting logical block address and the number of continuous blocks of the current request, and calculating an absolute value of a difference between the sum and the starting address of the next request, and determining whether the absolute value of the difference is greater than a preset window threshold; if the absolute value of the difference is less than or equal to the window threshold, it is determined that a logical block address corresponding to the current request satisfies a preset continuity requirement, and a first target value is assigned to a logical block address index, otherwise it is determined that the logical block address does not satisfy the continuity requirement, and a second target value is assigned to the logical block address index; determining a request time point of the current request, and calculating the random input / output ratio based on the first target value, the second target value, the request time point and a preset sliding window size.
3. The method of claim 1, wherein the step of dynamically allocating the cache of the storage hardware device comprises: The medium delay detection on the storage hardware device to obtain the medium detection data of the storage hardware device comprises: sending a preset number of standard read commands to a first logical block of different storage hard disks in the storage hardware device within a detection period of the medium delay detection, and determining a command sending time of the standard read commands and a response time of the different storage hard disks after receiving the standard read commands; Calculate the delay time between the command sending time and the response time, and calculate the statistical data corresponding to the delay time, wherein the statistical data includes the arithmetic mean value and the standard deviation corresponding to the delay time; Based on the arithmetic mean value in the statistical data and the preset delay time level span threshold, calculate the performance level and performance score corresponding to the different storage hard disks, and calculate the hard disk drive comprehensive performance score corresponding to the target server according to the performance level and the performance score; Create a plurality of input / output queues corresponding to a preset queue depth, measure the input / output operation performance indicators and delay information of the plurality of input / output queues under different queue depths, and determine the corresponding solid state disk comprehensive performance score according to the input / output operation performance indicators and the delay information; Based on the hard disk drive comprehensive performance score and the solid state disk comprehensive performance score, determine the medium detection data.
4. The method of claim 1, wherein, The hardware state analysis of the storage hardware device to obtain the hardware state data of the storage hardware device comprises: Obtain the temperature change gradient information corresponding to the preset dual inline memory, and determine the current temperature of the storage hardware device through the temperature change gradient information; Collect the hard disk original life percentage and durability factor corresponding to the solid state disk, and calculate the actual wear degree corresponding to the solid state disk according to the hard disk original life percentage and the durability factor; Determine the hardware state data based on the current temperature and the actual wear degree.
5. The method of claim 3, wherein the step of dynamically allocating the cache of the storage hardware device comprises: The creation of a plurality of input / output queues corresponding to a preset queue depth, the measurement of the input / output operation performance indicators and delay information of the plurality of input / output queues under different queue depths, and the determination of the corresponding solid state disk comprehensive performance score according to the input / output operation performance indicators and the delay information, comprising: Calculate the number of input / output operations per second and the average delay of the solid state disk under the current queue depth, and calculate the efficiency-delay ratio between the number of input / output operations per second and the average delay; Calculate the maximum value corresponding to the efficiency-delay ratio, and determine the target queue depth corresponding to the maximum value; Determine the peak value, valley value and arithmetic mean value of the number of input / output operations per second within the detection period, calculate the difference between the peak value and the valley value, and calculate the performance fluctuation rate of the solid state disk according to the difference and the arithmetic mean value; Convert the performance fluctuation rate into a corresponding stability gain factor, normalize the target queue depth to obtain a corresponding normalization result, and multiply the stability gain factor and the normalization result to obtain the solid state disk comprehensive performance score.
6. The method of claim 4, wherein, The collection of the hard disk original life percentage and the durability factor corresponding to the solid state disk to calculate the actual wear degree corresponding to the solid state disk according to the hard disk original life percentage and the durability factor, comprising: acquire a theoretical data write-erase complete operation number corresponding to the solid state disk, and calculate a consumption ratio corresponding to the solid state disk according to the theoretical data write-erase complete operation number, so as to determine the hard disk original life percentage through the consumption ratio; determine a hard disk type corresponding to the solid state disk, and calculate a durability factor corresponding to the solid state disk according to the hard disk type; calculate a ratio between the hard disk original life percentage and the durability factor, so as to determine the actual wear degree according to the ratio.
7. The method of claim 1, wherein the step of dynamically allocating the cache of the storage hardware device comprises: The matching the target decision logic corresponding to the storage hardware device based on the sequential write request proportion and the current temperature of the storage hardware device contained in the hardware state data comprises: inputting the random input-output ratio, the medium detection data and the hardware state data into a preset dynamic decision engine, so as to determine whether the current temperature is greater than a preset temperature threshold value and whether the sequential write request proportion is less than a preset proportion threshold value through the dynamic decision engine; if the current temperature is greater than the temperature threshold value, selecting a temperature priority decision logic as the target decision logic; if the current temperature is less than or equal to the temperature threshold value and the sequential write request proportion is greater than the proportion threshold value, selecting a sequential write optimization decision logic as the target decision logic; if the current temperature is less than or equal to the temperature threshold value and the sequential write request proportion is less than or equal to the proportion threshold value, selecting an elastic weight calculation decision logic as the target decision logic.
8. The method of claim 7, wherein, The determining the read / write cache allocation ratio corresponding to the storage hardware device based on the target decision logic and combining the random input-output ratio, the medium detection data and the hardware state data, and identifying the hot data corresponding to the storage hardware device, and performing cache allocation adjustment on the storage hardware device according to the read / write cache allocation ratio and the hot data, so as to obtain a dynamic allocation cache, comprises: when the target decision logic is the temperature priority decision logic, adjusting the read cache ratio to reduce the read cache ratio to a target read cache ratio, and migrating the hot data to a non-volatile memory according to the target read cache ratio, and adjusting the fan speed to a target speed value, and controlling the corresponding fan to rotate based on the target speed value to cool the corresponding storage hard disk; when the target decision logic is the sequential write optimization decision logic, performing cache reallocation operation to set the read cache ratio to a target fixed ratio, and establishing a dedicated direct memory access channel to perform corresponding cache allocation operation according to the target fixed ratio and the dedicated direct memory access channel; when the target decision logic is the elastic weight calculation decision logic, calculating a corresponding elastic read cache ratio based on the medium detection data, so as to perform corresponding cache allocation operation according to the elastic read cache ratio to obtain the dynamic allocation cache.
9. The method of claim 8, wherein, When the target decision logic is the elastic weight calculation decision logic, a corresponding elastic read cache ratio is calculated based on the medium detection data, including: determining a cache read baseline ratio, a random input-output weight coefficient, a hard disk drive performance compensation coefficient, and a solid state disk performance compensation coefficient corresponding to the solid state disk; obtaining a current hard disk temperature of the solid state disk, and calculating a temperature difference between the current hard disk temperature and a preset hard disk temperature threshold, so as to calculate a temperature penalty term corresponding to the solid state disk according to the temperature difference; determining a read-write critical point corresponding to the solid state disk, and calculating a life consumption penalty term corresponding to the solid state disk according to the read-write critical point and an actual wear degree; calculating a first product result of the random input-output weight coefficient and the random input-output ratio, calculating a second product result between the hard disk drive performance compensation coefficient and a hard disk drive comprehensive performance score, and calculating a third product result between the solid state disk performance compensation coefficient and a solid state disk comprehensive performance score; calculating the elastic read cache ratio based on the first product result, the second product result, the third product result, the cache read baseline ratio, the temperature penalty term, and the life consumption penalty term.
10. An electronic device, comprising: comprising: a memory for storing a computer program; a processor for executing the computer program to implement the steps of the cache dynamic allocation method of the storage hardware device according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method and apparatus for cache partition to allocate free pages
CN105243031A
Task processing method of artificial intelligence processor, storage medium and electronic equipment
CN120066806A
Request method and device for cache resources, equipment and medium
CN120295766A
Cited By
Self-adaptive data processing method and device for high-speed simulation data
CN121722335A