Storage system optimization method and device, electronic equipment and storage medium
By collecting and analyzing the power fingerprint characteristics of distributed storage nodes and generating dynamic energy management strategies, the problem of limited model generalization ability caused by the single feature dimension in existing technologies is solved, and more efficient energy management and data storage are achieved.
Patent Information
- Application Number
- CN202511071169.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-09-23
AI Technical Summary
Existing power fingerprint and energy management methods are relatively simple in feature dimensions, ignoring key indicators such as waveform distortion rate during equipment start-up and shutdown, resulting in limited model generalization ability.
The voltage, current and power data of distributed storage nodes are collected in real time, and the power fingerprint feature vectors including time domain, frequency domain and time-frequency domain are extracted. A dynamic energy management strategy is generated through the prediction model, and a power fingerprint labeling data storage protocol is constructed. The storage medium and the number of copies are dynamically allocated based on the energy efficiency score and waveform distortion rate.
It improves the model generalization capability, optimizes the energy management effect of the distributed storage system, and improves energy utilization efficiency and data storage accuracy.
Smart Images

Figure CN120692160A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data optimization technology, and in particular to a storage system optimization method, device, electronic device, and storage medium. Background Art
[0002] With the rapid development of cloud computing, big data, and the Internet of Things (IoT), distributed storage systems have become the core infrastructure of modern data centers, widely used in finance, communications, the internet, and other fields. Among these technologies, a collaborative system for device status awareness and energy efficiency optimization has been established through the integration of power fingerprinting, machine learning modeling, and edge computing.
[0003] However, existing power fingerprint and energy management methods are relatively simple in feature dimensions, usually relying only on basic parameters such as the mean and peak values of current or voltage, while ignoring key indicators such as waveform distortion rate and spectrum centroid during equipment start-up and shutdown, resulting in limited model generalization capabilities. Summary of the Invention
[0004] The present application provides a storage system optimization method, device, electronic device and computer-readable storage medium to at least solve the problem of limited model generalization capability in related technologies.
[0005] This application provides a storage system optimization method, including:
[0006] Collect voltage, current and power data of distributed storage nodes in real time, and extract power fingerprint feature vectors including time domain, frequency domain and time-frequency domain;
[0007] Inputting the power fingerprint feature vector into a prediction model to obtain a dynamic energy management strategy output by the prediction model;
[0008] According to the dynamic energy management strategy and the power fingerprint feature vector, a power fingerprint labeling data storage protocol is constructed, and data is divided into hot data, warm data and cold data according to access frequency;
[0009] Dynamically allocate storage media and the number of replicas based on energy efficiency scores and waveform distortion rates.
[0010] Optionally, the real-time collection of voltage, current and power data of the distributed storage nodes and the extraction of power fingerprint feature vectors including time domain, frequency domain and time-frequency domain include:
[0011] De-noising is performed on the voltage, current and power data of the distributed storage node by using a sliding window exponential moving average filter and a wavelet threshold.
[0012] Optionally, inputting the power fingerprint feature vector into a prediction model to obtain a dynamic energy management strategy output by the prediction model includes:
[0013] The prediction model predicts the load level of each node and combines with the random forest classification model to identify the node status to generate the dynamic energy management strategy; wherein, the dynamic energy management strategy includes at least one of central processing unit frequency and voltage adjustment and hard disk sleep control.
[0014] Optionally, inputting the power fingerprint feature vector into a prediction model to obtain a dynamic energy management strategy output by the prediction model includes:
[0015] The random forest classification model selects a predetermined number of features based on Gini importance; wherein the features include harmonic distortion rate, current peak value, and spectrum centroid;
[0016] Hyperparameter tuning via grid search.
[0017] Optionally, the step of constructing a power fingerprint labeled data storage protocol based on the dynamic energy management strategy and power fingerprint characteristics, dividing data into hot data, warm data, and cold data, and dynamically allocating storage media and the number of copies based on energy efficiency scores and waveform distortion rates also includes:
[0018] Adding a format tag to the data block; wherein the format tag includes at least one of power factor, harmonic distortion rate, and timestamp to guide data hierarchical storage;
[0019] The weight probability is calculated based on a preset algorithm combined with the power fingerprint weight, and the replica distribution node is selected according to the weight probability.
[0020] Optionally, the method further includes:
[0021] When it is detected that the waveform distortion rate of a node exceeds the preset threshold and lasts for a preset period of time, data migration is performed, and an operation and maintenance work order containing the faulty node name, waveform distortion rate value, and historical trend chart is generated, and pushed based on the preset push method.
[0022] Optionally, the method further includes:
[0023] A federated learning algorithm is used to collaboratively train the energy efficiency model of each node, and a differential privacy mechanism is introduced in the local gradient update process to protect the privacy of node data by Gaussian noise scrambling.
[0024] The present application also provides a storage system optimization device, comprising:
[0025] The acquisition unit is used to collect voltage, current and power data of distributed storage nodes in real time and extract power fingerprint feature vectors including time domain, frequency domain and time-frequency domain;
[0026] a calculation unit, configured to input the power fingerprint feature vector into a prediction model to obtain a dynamic energy management strategy output by the prediction model;
[0027] A construction unit, configured to construct a power fingerprint labeled data storage protocol according to the dynamic energy management strategy and the power fingerprint feature vector, and divide the data into hot data, warm data, and cold data according to access frequency;
[0028] The allocation unit is used to dynamically allocate storage media and the number of replicas based on energy efficiency scores and waveform distortion rates.
[0029] Optionally, the acquisition unit is further configured to:
[0030] De-noising is performed on the voltage, current and power data of the distributed storage node by using a sliding window exponential moving average filter and a wavelet threshold.
[0031] Optionally, the computing unit is further configured to:
[0032] The prediction model predicts the load level of each node and combines with the random forest classification model to identify the node status to generate the dynamic energy management strategy; wherein, the dynamic energy management strategy includes at least one of central processing unit frequency and voltage adjustment and hard disk sleep control.
[0033] Optionally, the computing unit is further configured to:
[0034] The random forest classification model selects a predetermined number of features based on Gini importance; wherein the features include harmonic distortion rate, current peak value, and spectrum centroid;
[0035] Hyperparameter tuning via grid search.
[0036] Optionally, the construction unit is further used to:
[0037] Adding a format tag to the data block; wherein the format tag includes at least one of power factor, harmonic distortion rate, and timestamp to guide data hierarchical storage;
[0038] The weight probability is calculated based on a preset algorithm combined with the power fingerprint weight, and the replica distribution node is selected according to the weight probability.
[0039] Optionally, the device further includes:
[0040] The execution unit is used to perform data migration when it detects that the waveform distortion rate of a node exceeds a preset threshold and lasts for a preset period of time, and generate an operation and maintenance work order containing the fault node name, waveform distortion rate value and historical trend chart, and execute push based on the preset push method.
[0041] Optionally, the device further includes:
[0042] The training unit is used to collaboratively train the energy efficiency model of each node using a federated learning algorithm, and introduce a differential privacy mechanism in the local gradient update process to protect the privacy of node data by Gaussian noise scrambling.
[0043] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned storage system optimization methods when executing the computer program.
[0044] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned storage system optimization methods are implemented.
[0045] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned storage system optimization methods when executed by a processor.
[0046] Through this application, since the voltage, current and power data of distributed storage nodes are collected in real time, and multi-dimensional power fingerprint feature vectors including time domain, frequency domain and time-frequency domain are extracted, it is no longer limited to a single basic parameter, but also incorporates key indicators such as waveform distortion rate. On this basis, a dynamic energy management strategy is obtained, and then a data storage protocol is constructed and data types are divided, and finally energy efficiency scores and waveform distortion are combined; therefore, it can solve the technical problems in the existing technology that the power fingerprint and energy management method have a single feature dimension, ignore key indicators such as waveform distortion rate during the start-up and shutdown of equipment, and lead to limited model generalization ability, so as to achieve the technical effect of improving the model generalization ability and optimizing the energy management effect of distributed storage systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0048] Figure 1 A flowchart of a storage system optimization method provided by an embodiment of the present disclosure;
[0049] Figure 2 A schematic diagram of the structure of a storage system optimization device provided in an embodiment of the present disclosure;
[0050] Figure 3 A schematic diagram of the structure of another storage system optimization device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0051] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0052] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0053] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0054] An embodiment of the present application provides a storage system optimization method, and the method is described in detail in conjunction with the execution flow of the storage system optimization method. Figure 1 A flowchart of a storage system optimization method provided by an embodiment of the present disclosure.
[0055] like Figure 1 As shown, the method comprises the following steps:
[0056] Step 101 : collect voltage, current and power data of distributed storage nodes in real time, and extract power fingerprint feature vectors including time domain, frequency domain and time-frequency domain.
[0057] In the process of collecting voltage, current, and power data from distributed storage nodes in real time and extracting power fingerprint feature vectors in the time domain, frequency domain, and time-frequency domain, it is first necessary to select an appropriate collection device to ensure high accuracy and timeliness of the data. Specifically, through the SMBus or PMBus interface of the power supply unit (PSU), the I2C protocol is used to communicate with the server, thereby reading the input and output power consumption data of the PSU in real time; current sensors are installed on the power supply lines of the motherboard and boards such as the network card and graphics card, and current consumption is detected with the help of sampling resistors. The analog signal is converted into a digital signal through an analog-to-digital converter (ADC) and then transmitted to the server's monitoring module (Baseboard Management Controller, BMC) to achieve real-time monitoring of the power consumption of each component; at the same time, IoT power monitoring equipment (such as smart meters) is deployed at each distributed storage node to collect power signals such as voltage, current, and power in real time. To capture rapidly changing power consumption fluctuations, the sampling frequency needs to be set to above 1kHz.
[0058] The collected power data is transmitted in real time to a distributed data storage platform via reliable communication protocols such as MQTT, HTTP, and CoAP. A data caching mechanism is set up during the transmission process to prevent data loss due to network fluctuations or delays. The data format is standardized to JSON, including timestamp, device ID, voltage (V), current (A), active power (W), reactive power (var), and other fields to ensure data consistency and parsability.
[0059] To improve data quality, the raw data needs to be preprocessed: a sliding window exponential moving average (EMA) filter (α = 0.2) is used to eliminate high-frequency noise. At the same time, a five-layer decomposition is performed using the Daubechies wavelet (db4). The detail coefficients are processed and reconstructed through soft thresholding to retain the valid signal, thus achieving wavelet threshold denoising. To address the problem of multi-node clock asynchrony, the Network Time Protocol (NTP) is used to calibrate the timestamp, controlling the error within ±1ms to ensure the time consistency of the data.
[0060] Step 102: Input the power fingerprint feature vector into a prediction model to obtain a dynamic energy management strategy output by the prediction model.
[0061] Inputting the power fingerprint feature vector into the prediction model to generate a dynamic energy management strategy requires first completing the prediction model's training and optimization, and then generating a specific strategy based on the model's inference results on the input feature vector. The prediction model primarily includes an LSTM (Long Short-Term Memory) time series prediction model and a random forest classification model. Both use the power fingerprint feature vector as input, learn the association between historical power consumption patterns and node status, and output prediction results to guide energy management.
[0062] During the model training phase, a training set is first constructed. A semi-automatic annotation tool, combined with the Simple Network Management Protocol (SNMP), is used to obtain storage node performance metrics (such as CPU utilization, disk IOPS, and network bandwidth usage). These metrics are then aligned with the timestamps of the power fingerprint feature vectors. Three scenario labels are generated: idle state: CPU utilization <5%, disk IOPS <100, network bandwidth usage <1%; read / write state: CPU utilization >30%, disk IOPS >1000, network bandwidth usage >50%; and fault state: WDR >5% or power dip exceeding 20%. These labels are then stored in CSV format. A sliding window (5-minute window size, 1-minute step size) is then designed to segment the time series data. A 56-dimensional power fingerprint feature vector is extracted from each window, ultimately constructing a training set containing 100,000 samples (70% training, 15% validation, and 15% test).
[0063] For the LSTM time series prediction model, its input layer receives a 56-dimensional feature sequence within the window (time step size 30, corresponding to one sample per second within a 5-minute window). The network structure consists of two LSTM layers (128 units per layer, with a tanh activation function), one dropout layer (with a dropout rate of 0.2 to prevent overfitting), and one fully connected layer (with three output neurons corresponding to low, medium, and high load levels, using softmax activation). The loss function uses categorical cross-entropy loss, and the optimizer is Adam (learning rate α = 0.001, β1 = 0.9, β2 = 0.999). The random forest classification model selects the top 20 features, such as harmonic distortion rate (WDR), current peak (I_peak), and spectrum centroid (SC), based on Gini importance. GridSearchCV is used to tune hyperparameters (number of trees 200, maximum depth 10, minimum number of leaf samples 5) to improve classification accuracy. It has been verified that the LSTM model has an accuracy of 92.3% on the test set, a recall rate of 94.5% for "fault status", and an accuracy of 89.7% for the random forest classification model, both of which can effectively achieve status prediction.
[0064] When the power fingerprint feature vector is input into the trained prediction model, the model outputs the node's load level (low / medium / high) or state category (idle / read / write / fault), which in turn generates a dynamic energy management strategy. Specifically, if the load is predicted to be "low" in the next 10 minutes, the CPU dynamic voltage and frequency scaling (DVFS) is triggered: based on the relationship between dynamic power consumption P, frequency f, and voltage V in CMOS circuit theory, the CPU frequency is reduced from 3.5GHz to 2.4GHz, and the voltage is reduced from 1.2V to 1.0V, reducing power consumption by approximately 30%. At the same time, if the load is "low" for three consecutive prediction windows and the current IOPS is less than 100, the hard disk hibernation strategy is triggered, switching the hard disk from Active mode (power consumption 10W) to Standby mode (power consumption 1W), saving approximately 8.7Wh of energy per hour for a single node. If the model predicts a harmonic anomaly on a node (e.g., WDR > 5%), the harmonic anomaly handling process is initiated. This includes screening target nodes with WDR < 2% and an Energy Efficiency Index (EEI) > 0.8, migrating data using the RSYNC incremental synchronization algorithm (bandwidth usage < 30%), and generating an operation and maintenance work order containing the faulty node ID, WDR value, and historical trend chart, which is then notified to maintenance personnel via email or SMS. Through this process, the prediction model accurately outputs dynamic energy management strategies based on the input power fingerprint feature vector, effectively optimizing the energy efficiency of the distributed storage system.
[0065] Step 103: constructing a power fingerprint labeled data storage protocol based on the dynamic energy management strategy and the power fingerprint feature vector, and dividing the data into hot data, warm data, and cold data according to access frequency;
[0066] Based on dynamic energy management strategies and power fingerprint feature vectors, a power fingerprint tagged data storage protocol is constructed. The process of classifying hot, warm, and cold data by access frequency requires integrating power characteristics with storage requirements to achieve refined data management. The core of the power fingerprint tagged data storage protocol is to attach a tag containing the power fingerprint characteristics to each data block. The tag format is based on a standardized JSON definition and includes fields such as timestamp, device ID, power factor (PF), harmonic distortion rate (WDR), and energy efficiency index (EEI). This information is derived from the time, frequency, and time-domain characteristics of the power fingerprint feature vector and is associated with the energy efficiency status of the node in the dynamic energy management strategy (such as EEI value and load level). This ensures that the tags reflect the power characteristics and energy optimization requirements of the data storage node in real time. The protocol uses a hash index mechanism to map tags to data blocks, facilitating rapid retrieval. Tags are also regularly updated as dynamic energy management strategies adjust (such as changes in node energy efficiency status), ensuring dynamic adaptation of data storage strategies to node energy status.
[0067] In terms of data division, data is divided into three levels based on access frequency and combined with the node energy efficiency information in the power fingerprint feature vector and the load prediction results in the dynamic energy management strategy: hot data refers to data that has been accessed ≥100 times in the past hour. Because it requires high-frequency access and has high response speed requirements, combined with the energy efficiency optimization goals for high-load nodes in the dynamic energy management strategy, this type of data is allocated to nodes with energy efficiency scores (EEI>0.8), and NVMe SSD (read and write latency <100μs) is used as the storage medium to match its high-frequency access requirements and utilize the energy utilization efficiency of high-efficiency nodes; warm data refers to data that has been accessed ≥10 times in the past 24 hours, with medium access frequency. Based on the energy allocation plan for medium-load nodes in the dynamic energy management strategy, it is allocated to nodes with EEI>0.6 and stored in SATA SSDs (read and write latency <1ms) balance access speed and energy consumption. Cold data, defined as data with no access history in the past seven days and with extremely low access frequency, is allocated to nodes with an EEI > 0.4 based on the dynamic energy management strategy for low-load nodes. This storage is then stored on HDDs or Ceph object storage (read and write latency <10ms) to reduce storage energy consumption. This partitioning approach not only meets data performance requirements based on access frequency, but also leverages power fingerprinting and dynamic energy management strategies to precisely match storage resources with node energy status, improving both overall storage system energy efficiency and data access efficiency.
[0068] Step 104 : Dynamically allocate storage media and the number of replicas based on the energy efficiency score and the waveform distortion rate.
[0069] The dynamic allocation of storage media and replica counts requires a quantitative assessment of the node's energy efficiency and health status, enabling precise adaptation of storage resources and ensuring reliability. The Energy Efficiency Index (EEI) is calculated using the formula "EEI = (actual throughput (IOPS) / theoretical maximum throughput) × (baseline power consumption / actual power consumption)" and reflects the node's energy efficiency under the current load (ranging from 0 to 1, with higher values indicating better energy efficiency). It also represents the proportion of harmonic components in the current waveform, indirectly reflecting the health of the node's hardware (lower values indicate more stable equipment).
[0070] In terms of storage media allocation, the system dynamically matches data storage requirements based on the node's EEI value: For hot data that requires frequent access (≥100 accesses in the past hour), nodes with an EEI>0.8 are preferred because they are more energy efficient and can meet high throughput requirements while reducing energy consumption per unit of data access. NVMe SSDs (read and write latency <100μs) are used as the storage medium to match the performance requirements of high-frequency access. For warm data that is accessed moderately frequently (≥10 accesses in the past 24 hours), nodes with an EEI>0.6 are allocated to balance energy efficiency and access speed. SATA SSDs (read and write latency <1ms) are used as the storage medium. For cold data that is accessed infrequently (no access in the past 7 days), nodes with an EEI>0.4 are allocated. Although these nodes have relatively low energy efficiency, energy consumption can be controlled through energy-saving strategies under low load (such as hard disk hibernation). HDDs or Ceph object storage (read and write latency <10ms) are used as the storage medium. This EEI-based media allocation not only ensures the performance requirements of data with different access frequencies, but also maximizes the energy advantages of high-efficiency nodes and reduces the energy consumption of the overall storage system.
[0071] In terms of adjusting the number of replicas, the system dynamically optimizes the redundancy strategy by combining EEI and WDR: 3 replicas are used by default to ensure data reliability (availability ≥ 99.99%); when the predicted storage demand is "low" and the node EEI>0.7, it means that the node energy efficiency is stable and the load pressure is small. At this time, reducing the number of replicas to 2 can save 33% of storage space. At the same time, due to the high energy efficiency of the node, the stability of data access can still be maintained; if the node WDR is detected to be>5% (exceeding the conservative threshold of the healthy node baseline value μ+3σ=4.5%), it indicates that there may be potential faults in the node hardware (such as capacitor aging). At this time, the number of replicas is temporarily increased to 4 until the faulty node returns to normal, reducing the risk of data loss by improving redundancy. In addition, for cold data, erasure coding is used instead of multiple copies to further optimize redundancy overhead. Specifically, the RS(6,3) encoding scheme is selected (data is split into 6 blocks and 3 parity blocks are generated). The storage overhead is only 1.5x (50% lower than 3 copies), and only nodes with EEI>0.5 are allowed to store parity blocks, ensuring the reliability of the verification data and low-energy maintenance.
[0072] Through the above-mentioned dynamic adjustment mechanism based on EEI and WDR, the system not only meets the data storage reliability and performance requirements, but also achieves efficient utilization of storage media and precise control of replica redundancy, further improving the energy efficiency and resource utilization of the distributed storage system.
[0073] Optionally, the real-time collection of voltage, current and power data of the distributed storage nodes and the extraction of power fingerprint feature vectors including time domain, frequency domain and time-frequency domain include:
[0074] De-noising is performed on the voltage, current and power data of the distributed storage node by using a sliding window exponential moving average filter and a wavelet threshold.
[0075] In the process of collecting voltage, current, and power data from distributed storage nodes in real time and extracting power fingerprint feature vectors in the time, frequency, and time-frequency domains, in order to eliminate high-frequency noise and interference signals in the raw power data, it is necessary to perform denoising on the data through sliding window exponential moving average (EMA) filtering and wavelet threshold denoising to ensure the accuracy and reliability of subsequent feature extraction. Among them, the sliding window EMA filter is used for raw current data (the processing logic of voltage and power data is the same) to achieve noise filtering by setting the sliding window and exponential moving average coefficient: specifically, a sliding window with a window size adapted to the sampling frequency (for example, a 1-second window corresponds to 1000 data points at a 1kHz sampling rate) is used to perform exponential moving average calculations on the raw data within the window. This method can effectively attenuate high-frequency noise (such as burr signals generated by instantaneous current fluctuations) and preserve the overall trend and steady-state characteristics of the data. Wavelet threshold denoising further refines the data after EMA filtering, selecting Daubechies wavelet (db4) as the base wavelet and decomposing the data into 5 layers: During the decomposition process, each layer will obtain approximate coefficients (reflecting the low-frequency trend of the signal) and detail coefficients (reflecting the high-frequency noise of the signal). Soft threshold processing is applied to the detail coefficients (i.e., coefficients with absolute values less than the threshold are set to zero, and coefficients with absolute values greater than the threshold are subtracted from the threshold) to filter out residual high-frequency interference. Subsequently, the processed approximate coefficients are synthesized with the detail coefficients through wavelet reconstruction, ultimately retaining the effective components in the power signal (such as steady-state power characteristics, transient start-stop waveforms, etc.), laying a data foundation for the subsequent accurate extraction of time domain, frequency domain, and time-frequency domain features.
[0076] Optionally, inputting the power fingerprint feature vector into a prediction model to obtain a dynamic energy management strategy output by the prediction model includes:
[0077] The prediction model predicts the load level of each node and combines with the random forest classification model to identify the node status to generate the dynamic energy management strategy; wherein, the dynamic energy management strategy includes at least one of central processing unit frequency and voltage adjustment and hard disk sleep control.
[0078] In the process of inputting the power fingerprint feature vector into the prediction model to obtain the dynamic energy management strategy output by the prediction model, the prediction model predicts the load level of each node through the LSTM (Long Short-Term Memory) time series prediction model, and combines the random forest (Random Forest) classification model to identify the node status, and then generates a dynamic energy management strategy including at least one of central processing unit (CPU) frequency and voltage adjustment and hard disk hibernation control. Specifically, the LSTM time series prediction model takes the power fingerprint feature vector as input. The feature vector contains 56-dimensional features extracted from the time domain, frequency domain and time-frequency domain (such as mean, standard deviation, harmonic distortion rate (WDR), spectrum centroid (SC), etc.). The model uses a network structure consisting of 2 LSTM layers (128 units per layer, activation function is tanh), 1 Dropout layer (dropout rate 0.2) and a fully connected layer (output 3 neurons, corresponding to low, medium and high load levels) to learn the feature sequence within the sliding window (window size 5 minutes, step length 1 minute) and output the load level prediction result for the next 10 minutes. This prediction is based on the learning of the association pattern between historical load and power characteristics. For example, by analyzing the low power consumption characteristics in the idle state and the power fluctuation characteristics corresponding to the high throughput in the read and write state, accurate prediction of load change trends can be achieved.
[0079] At the same time, the random forest classification model uses Gini importance to select the top 20 key features, including harmonic distortion rate (WDR), current peak (I_peak), and spectrum centroid (SC). Using an ensemble of 200 decision trees (maximum depth 10, minimum leaf size 5), it classifies node states into idle (CPU utilization <5%, disk IOPS <100, network bandwidth usage <1%), read / write (CPU utilization >30%, disk IOPS >1000, network bandwidth usage >50%), or fault (WDR >5% or power dip exceeding 20%). This classification complements the LSTM load level prediction. For example, if the LSTM predicts "low load" and the random forest identifies "idle," it confirms that the node is currently in a low energy demand state, providing a basis for subsequent energy policy adjustments.
[0080] Based on the above load level prediction and node state identification results, the dynamic energy management strategy is specifically generated as follows: For CPU frequency and voltage adjustment, when the load is predicted to be "low" in the next 10 minutes and the node state is "idle", according to the relationship between dynamic power consumption P and frequency f and voltage V in CMOS circuit theory (P∝C·V 2f, where C is the load capacitance), triggering dynamic voltage and frequency scaling (DVFS), reducing the CPU frequency from 3.5GHz to 2.4GHz and the voltage from 1.2V to 1.0V. This calculation shows a power consumption reduction of approximately 30%, meeting low-load requirements while reducing energy waste. Regarding hard drive hibernation control, if three consecutive prediction windows (each 5 minutes) show a "low" load and the random forest node status remains "idle" (current IOPS < 100), the hard drive is triggered to switch from active mode (power consumption 10W) to standby mode (power consumption 1W). Taking into account a 2-second switching delay, a single node can save approximately 8.7Wh of energy per hour. By synergizing LSTM load prediction with random forest state recognition, the system can accurately determine node energy needs and dynamically select one or both of the CPU frequency and voltage scaling and hard drive hibernation control strategies, achieving efficient energy management of distributed storage systems.
[0081] Optionally, inputting the power fingerprint feature vector into a prediction model to obtain a dynamic energy management strategy output by the prediction model includes:
[0082] The random forest classification model selects a predetermined number of features based on Gini importance; wherein the features include harmonic distortion rate, current peak value, and spectrum centroid;
[0083] Hyperparameter tuning via grid search.
[0084] When inputting the power fingerprint feature vector into the prediction model to generate a dynamic energy management strategy, the Random Forest classification model uses Gini Importance to select a predetermined number of features and uses GridSearchCV to perform hyperparameter tuning to improve the model's accuracy in identifying node states, providing a reliable basis for dynamic energy management strategies. Gini Importance measures a feature's contribution to node splitting in a random forest. A higher value indicates a greater impact on the model's classification results. Specifically, the model calculates each feature's total contribution to reducing the node's Gini impurity across all decision trees and sorts them from highest to lowest contribution value, selecting a predetermined number of key features (e.g., the top 20). These features include the Waveform Distortion Rate (WDR), current peak (I_peak), and Spectral Centroid (SC). The harmonic distortion rate, by reflecting the proportion of harmonic components in the current waveform, can effectively distinguish the health status of the device (for example, the WDR of a faulty node is usually significantly higher). The current peak captures the transient current characteristics at the moment the device starts and stops, which can reflect the node's operating mode switching (for example, a sudden change in current from idle to read / write state). The Spectral Centroid measures the concentration of frequency domain energy distribution to help identify the node's load intensity (for example, the spectrum energy distribution is more dispersed under high load). By screening these high-importance features, redundant information interference can be reduced, improving the model's computational efficiency and classification accuracy.
[0085] During the hyperparameter tuning phase, grid search is used to optimize the key hyperparameters of the random forest classification model. Grid search constructs a grid of candidate values for the hyperparameters, trains the model on all possible parameter combinations, and uses cross-validation (such as 5-fold cross-validation) to evaluate the performance of each combination, ultimately selecting the parameter combination with the best performance on the validation set. Specifically, the set hyperparameters include the number of trees (candidate values such as 100, 200, 300), the maximum depth (candidate values such as 5, 10, 15), the minimum number of leaf samples (candidate values such as 1, 5, 10), etc.: the number of trees affects the complexity and generalization ability of the model. Too few trees may lead to underfitting, while too many will increase the computational cost. The maximum depth controls the growth of the decision tree. Too deep will easily lead to overfitting, while too shallow may not be able to capture the data patterns. The minimum number of leaf samples limits the minimum sample size of the leaf node, which can prevent the model from overfitting to noisy data. Verified by grid search, the optimal hyperparameter combination for the random forest classification model is 200 trees, a maximum depth of 10, and a minimum number of leaf samples of 5. At this point, the model achieved an accuracy of 89.7% on the test set, and was able to stably identify the idle, read / write, and fault states of nodes, providing an accurate status basis for the generation of dynamic energy management strategies (such as CPU frequency and voltage adjustment and hard disk hibernation control).
[0086] Optionally, the step of constructing a power fingerprint labeled data storage protocol based on the dynamic energy management strategy and power fingerprint characteristics, dividing data into hot data, warm data, and cold data, and dynamically allocating storage media and the number of copies based on energy efficiency scores and waveform distortion rates also includes:
[0087] Adding a format tag to the data block; wherein the format tag includes at least one of power factor, harmonic distortion rate, and timestamp to guide data hierarchical storage;
[0088] The weight probability is calculated based on a preset algorithm combined with the power fingerprint weight, and the replica distribution node is selected according to the weight probability.
[0089] The process of constructing a power fingerprint-tagged data storage protocol based on dynamic energy management strategies and power fingerprint characteristics, dividing data into hot, warm, and cold data, and dynamically allocating storage media and replica counts based on energy efficiency scores (EEI) and waveform distortion rate (WDR), also includes adding format tags to data blocks, calculating weight probabilities based on a preset algorithm combined with power fingerprint weights, and selecting replica distribution nodes. The format tags attached to data blocks use a standardized JSON format and contain key information such as power factor (PF), harmonic distortion rate (WDR), and timestamp. The power factor reflects the effective utilization rate of node power (ranging from 0 to 1, with higher values indicating higher power conversion efficiency) and is an important indicator for determining node energy efficiency. The harmonic distortion rate (WDR) quantifies the health of the device by measuring the proportion of current harmonic components and is directly related to node stability. The timestamp is synchronized with the time calibrated by the NTP protocol to ensure time consistency across nodes. These tags are embedded in the metadata of the data block and quickly associated through hash index (Hash(node_id||timestamp)→storage location) to guide data tiered storage. For example, hot data is preferentially allocated to high-efficiency and healthy nodes with a power factor > 0.9 and WDR < 2%, warm data is adapted to nodes with a power factor > 0.8 and WDR < 3%, and cold data is allocated to nodes that meet basic energy efficiency and health requirements, thereby deeply coupling the tiering strategy with the power characteristics of the nodes.
[0090] The CRUSH (Controlled Replication Under Scalable Hashing) algorithm is used to calculate weighted probabilities and select replica distribution nodes based on a pre-defined algorithm combined with power fingerprint weights. This algorithm uses a hash function to map data to storage nodes, supporting the dynamic scalability of large-scale distributed systems. The power fingerprint weight (w) is calculated by combining the energy efficiency index (EEI) and the waveform distortion rate (WDR) using the formula w = α × EEI + β × (1-WDR / 100) (where α = 0.7 and β = 0.3, with weighting coefficients set based on the priority of energy efficiency and health, and the weights of all nodes are normalized to sum to 1). The higher the EEI (better energy efficiency node) and the lower the WDR (healthier device), the greater the node weight. The CRUSH algorithm calculates the probability of node selection based on this weight. Nodes with higher weighted probabilities are assigned higher priority for replica assignment, thus avoiding replicas being distributed on nodes with low energy efficiency (EEI < 0.6) or high harmonics (WDR > 5%). For example, in the default 3-copy storage strategy, the system uses the CRUSH algorithm combined with weight probability to prioritize the top 3 nodes in weight ranking to store copies, ensuring that the copy data can rely on high-efficiency nodes to reduce maintenance energy consumption and use healthy nodes to improve data reliability, achieving the optimal balance between energy efficiency and reliability in copy distribution.
[0091] Optionally, the method further includes:
[0092] When it is detected that the waveform distortion rate of a node exceeds the preset threshold and lasts for a preset period of time, data migration is performed, and an operation and maintenance work order containing the faulty node name, waveform distortion rate value, and historical trend chart is generated, and pushed based on the preset push method.
[0093] In this method, when the waveform distortion rate (WDR) of a node is detected to exceed a preset threshold and persists for a preset duration, the system triggers a series of fault handling processes, including data migration, operation and maintenance work order generation and push. The preset threshold is determined based on historical data statistics of healthy nodes: by analyzing 1000 hours of normal operation data, it is found that the WDR distribution of healthy nodes is μ = 2.1% (mean) and σ = 0.8% (standard deviation). Based on the 3σ principle, the initial threshold is set at 4.5%, and is conservatively adjusted to 5% to improve reliability, that is, the preset threshold is 5%; the preset duration is set to 5 minutes to avoid misjudgment caused by instantaneous fluctuations.
[0094] When a node's WDR is greater than 5% and persists for 5 minutes, the system first performs data migration: it selects nodes with a WDR less than 2% (to ensure the target node's hardware status is stable) and an Energy Efficiency Index (EEI) greater than 0.8 (to ensure the target node has high energy utilization efficiency) from the candidate node pool as data migration targets. The RSYNC incremental synchronization algorithm is used to implement data migration. This algorithm only transmits the difference in data, which can control bandwidth usage within 30%, avoiding significant impact on the normal operation of the system.
[0095] At the same time, the system automatically generates an operation and maintenance work order. The work order content includes the unique identifier of the faulty node (such as the node ID), the current WDR value, and a WDR historical trend chart (which visually displays the distortion rate evolution process) to help maintenance personnel quickly locate the problem. After the work order is generated, it is sent to the relevant operation and maintenance personnel via a preset push method (email or SMS), ensuring that the fault information is delivered in a timely manner so that the abnormal node can be addressed as soon as possible and the stable operation of the distributed storage system can be guaranteed.
[0096] Optionally, the method further includes:
[0097] A federated learning algorithm is used to collaboratively train the energy efficiency model of each node, and a differential privacy mechanism is introduced in the local gradient update process to protect the privacy of node data by Gaussian noise scrambling.
[0098] The optimization method of this storage system also includes the steps of collaboratively training the energy efficiency model of each node using the Federated Learning algorithm, and introducing the Differential Privacy (DP) mechanism in the local gradient update process to protect the privacy of node data by Gaussian noise scrambling. Specifically, the collaborative training of federated learning is based on the FedAvg (Federated Averaging) algorithm framework: first, the control center initializes the global energy efficiency model parameters θ_G and sends it to all distributed storage nodes; each node trains a lightweight model locally (such as the MobileNetV2 fine-tuning model, inputs a 56-dimensional feature vector, outputs EEI, and the parameter amount is <1MB to adapt to edge devices) based on the locally collected power fingerprint feature data (such as steady-state power in the time domain, harmonic components in the frequency domain, etc.) and energy efficiency labels (such as energy efficiency ratings EEI), and calculates the local gradient Update the node local model parameters; after completing local training, each node uploads the updated local model parameters to the control center. The control center performs weighted average aggregation on all local parameters according to the number of nodes K, and regularly sends the aggregated global model to each node to achieve collaborative optimization of multi-node energy efficiency models and improve the accuracy of global energy efficiency prediction.
[0099] In order to protect the privacy of the original power data of each node (to avoid information leakage caused by data transmission), a differential privacy mechanism is introduced in the local gradient update process: each node calculates the local gradient After that, add Gaussian distribution N(0,σ 2 I) (where σ = 0.1 and I is the identity matrix), i.e., the scrambled gradient, uses noise to mask the true gradient details. Simultaneously, the Moments Accountant method is used to calculate the cumulative privacy loss throughout the training process, ensuring that the performance of the model after aggregation is not significantly affected while meeting privacy protection requirements. This combination of federated learning collaborative training and differential privacy protection achieves global optimization of the cross-node energy efficiency model while avoiding the direct transmission of raw power data, meeting data security and privacy compliance requirements.
[0100] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0101] The embodiment of the present application also provides a storage system optimization device, Figure 2 A schematic diagram of the structure of a storage system optimization device provided by an embodiment of the present disclosure is shown in FIG. Figure 2 Shown, including:
[0102] The acquisition unit 21 is used to collect voltage, current and power data of distributed storage nodes in real time and extract power fingerprint feature vectors including time domain, frequency domain and time-frequency domain;
[0103] A calculation unit 22 is configured to input the power fingerprint feature vector into a prediction model to obtain a dynamic energy management strategy output by the prediction model;
[0104] A construction unit 23 is configured to construct a power fingerprint labeled data storage protocol according to the dynamic energy management strategy and the power fingerprint feature vector, and to divide the data into hot data, warm data, and cold data according to access frequency;
[0105] The allocation unit 24 is configured to dynamically allocate storage media and the number of replicas based on the energy efficiency score and the waveform distortion rate.
[0106] Furthermore, in a possible implementation of the embodiment of the present disclosure, the collection unit 21 is further configured to:
[0107] De-noising is performed on the voltage, current and power data of the distributed storage node by using sliding window exponential moving average filtering and wavelet thresholding.
[0108] Furthermore, in a possible implementation of the embodiment of the present disclosure, the calculation unit 22 is further configured to:
[0109] The prediction model predicts the load level of each node and combines with the random forest classification model to identify the node status to generate the dynamic energy management strategy; wherein, the dynamic energy management strategy includes at least one of central processing unit frequency and voltage adjustment and hard disk sleep control.
[0110] Furthermore, in a possible implementation of the embodiment of the present disclosure, the calculation unit 22 is further configured to:
[0111] The random forest classification model selects a predetermined number of features based on Gini importance; wherein the features include harmonic distortion rate, current peak value, and spectrum centroid;
[0112] Hyperparameter tuning via grid search.
[0113] Furthermore, in a possible implementation of the embodiment of the present disclosure, the construction unit 23 is further configured to:
[0114] Adding a format tag to the data block; wherein the format tag includes at least one of power factor, harmonic distortion rate, and timestamp to guide data hierarchical storage;
[0115] The weight probability is calculated based on a preset algorithm combined with the power fingerprint weight, and the replica distribution node is selected according to the weight probability.
[0116] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 3 As shown, the device also includes:
[0117] The execution unit 25 is used to perform data migration when it is detected that the waveform distortion rate of the node exceeds a preset threshold and lasts for a preset period of time, and generate an operation and maintenance work order including the fault node name, waveform distortion rate value and historical trend chart, and execute push based on the preset push method.
[0118] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 3 As shown, the device also includes:
[0119] The training unit 26 is used to collaboratively train the energy efficiency model of each node using a federated learning algorithm, and introduce a differential privacy mechanism in the local gradient update process to protect the privacy of node data by Gaussian noise scrambling.
[0120] For the description of the features in the embodiment corresponding to the storage system optimization device, please refer to the relevant description of the embodiment corresponding to the storage system optimization method, and no further details will be given here.
[0121] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned storage system optimization method embodiments.
[0122] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned storage system optimization method embodiments when running.
[0123] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0124] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned storage system optimization method embodiments are implemented.
[0125] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned storage system optimization method embodiments are implemented.
[0126] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0127] The above is a detailed introduction to the optimization, device, electronic device and storage medium of a storage system provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A storage system optimization method, characterized in that: include: Collect voltage, current and power data of distributed storage nodes in real time, and extract power fingerprint feature vectors including time domain, frequency domain and time-frequency domain; Inputting the power fingerprint feature vector into a prediction model to obtain a dynamic energy management strategy output by the prediction model; According to the dynamic energy management strategy and the power fingerprint feature vector, a power fingerprint labeling data storage protocol is constructed, and data is divided into hot data, warm data and cold data according to access frequency; Dynamically allocate storage media and the number of replicas based on energy efficiency scores and waveform distortion rates.
2. The storage system optimization method according to claim 1, characterized in that: The real-time collection of voltage, current and power data of distributed storage nodes and the extraction of power fingerprint feature vectors including time domain, frequency domain and time-frequency domain include: De-noising is performed on the voltage, current and power data of the distributed storage node by using a sliding window exponential moving average filter and a wavelet threshold.
3. The storage system optimization method according to claim 1, characterized in that: The step of inputting the power fingerprint feature vector into a prediction model to obtain a dynamic energy management strategy output by the prediction model includes: The prediction model predicts the load level of each node and combines with the random forest classification model to identify the node status to generate the dynamic energy management strategy; wherein, the dynamic energy management strategy includes at least one of central processing unit frequency and voltage adjustment and hard disk sleep control.
4. The storage system optimization method according to claim 3, characterized in that: The step of inputting the power fingerprint feature vector into a prediction model to obtain a dynamic energy management strategy output by the prediction model includes: The random forest classification model selects a predetermined number of features based on Gini importance; wherein the features include harmonic distortion rate, current peak value, and spectrum centroid; Hyperparameter tuning via grid search.
5. The storage system optimization method according to claim 1, characterized in that: The power fingerprint labeling data storage protocol is constructed based on the dynamic energy management strategy and power fingerprint characteristics, data is divided into hot data, warm data and cold data, and storage media and the number of copies are dynamically allocated based on energy efficiency scores and waveform distortion rates. The method also includes: Adding a format tag to the data block; wherein the format tag includes at least one of power factor, harmonic distortion rate, and timestamp to guide data hierarchical storage; The weight probability is calculated based on a preset algorithm combined with the power fingerprint weight, and the replica distribution node is selected according to the weight probability.
6. The storage system optimization method according to any one of claims 1 to 5, characterized in that: The method further comprises: When it is detected that the waveform distortion rate of a node exceeds the preset threshold and lasts for a preset period of time, data migration is performed, and an operation and maintenance work order containing the faulty node name, waveform distortion rate value, and historical trend chart is generated, and pushed based on the preset push method.
7. The storage system optimization method according to claim 1, characterized in that: The method further comprises: A federated learning algorithm is used to collaboratively train the energy efficiency model of each node, and a differential privacy mechanism is introduced in the local gradient update process to protect the privacy of node data by Gaussian noise scrambling.
8. A storage system optimization device, characterized in that: include: The acquisition unit is used to collect voltage, current and power data of distributed storage nodes in real time and extract power fingerprint feature vectors including time domain, frequency domain and time-frequency domain; a calculation unit, configured to input the power fingerprint feature vector into a prediction model to obtain a dynamic energy management strategy output by the prediction model; A construction unit, configured to construct a power fingerprint labeled data storage protocol according to the dynamic energy management strategy and the power fingerprint feature vector, and divide the data into hot data, warm data, and cold data according to access frequency; The allocation unit is used to dynamically allocate storage media and the number of replicas based on energy efficiency scores and waveform distortion rates.
9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the storage system optimization method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the storage system optimization method according to any one of claims 1 to 7 are implemented.