Data optimization storage method based on artificial intelligence
By combining multi-model combination prediction methods and deep reinforcement learning with SARIMA and LSTM models, the problem of imbalance in storage resource utilization in data storage systems is solved, dynamic optimization and adaptive adjustment of storage resources are realized, and the performance and efficiency of the system are improved.
Patent Information
- Application Number
- CN202510315619.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-18
AI Technical Summary
When existing data storage systems face complex and changeable data storage needs, there is a problem of imbalance in the utilization of storage resources, making it difficult to achieve dynamic optimization and adaptive adjustment.
Using an artificial intelligence-based method, through anomaly detection, time series prediction and deep reinforcement learning, combined with SARIMA and LSTM models, a multi-model combination prediction method is constructed, and a storage rate anomaly points and transmission rate residuals are combined to build a deep reinforcement learning model to generate the optimal data transmission rate control strategy.
It realizes dynamic balanced utilization of storage resources, improves data storage and access performance, reduces storage space waste, and enhances the adaptability and robustness of the system.
Smart Images

Figure CN120255804A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data storage, and particularly to an artificial intelligence-based data optimized storage method. Background Art
[0002] In the current era of rapid development of information technology, the explosive growth of data has posed unprecedented challenges to data storage systems. The efficient storage and management of massive data have become the core requirements of data centers and cloud computing platforms. However, existing data storage technologies often face the problem of unbalanced utilization of storage resources when dealing with complex and changing data storage requirements.
[0003] Traditional data storage systems usually adopt static data distribution and storage strategies, which are difficult to adapt to the dynamic changes of data volume and access patterns. In the actual data storage process, the load distribution among different storage nodes is often unbalanced, with some nodes having too high resource utilization and other nodes having idle resources. This unbalanced utilization of storage resources not only causes waste of valuable storage space but also affects the performance of data storage and access.
[0004] In addition, most existing data storage optimization technologies rely on preset rules or heuristic strategies, lacking in-depth exploration and utilization of the spatio-temporal characteristics and dynamic laws in the data storage process. These methods are difficult to cope with complex and changing storage environments and cannot achieve dynamic optimization and adaptive adjustment of storage resource utilization. Summary of the Invention
[0005] Aiming at the problem of unbalanced utilization of storage resources in the prior art, this application provides an artificial intelligence-based data optimized storage method. By means of artificial intelligence-based anomaly detection, time series prediction, and deep reinforcement learning, etc., it fully explores the spatio-temporal characteristics and dynamic laws in the data storage process and balances the utilization of storage resources.
[0006] The object of this application is achieved through the following technical solutions.
[0007] This application provides an artificial intelligence-based data optimization storage method, including: S1, collecting the storage rate data of the storage unit and the transmission rate data of the data transmission unit; S2, performing anomaly detection on the storage rate data using the density-based local outlier factor (LOF) algorithm to obtain a storage rate anomaly point sequence; S3, according to the transmission rate data and the storage rate anomaly point sequence, using a multi-model combination prediction method based on SARIMA and LSTM to obtain a predicted transmission rate sequence for the next moment; calculating the residual between the collected transmission rate data and the predicted transmission rate sequence to obtain a transmission rate residual sequence; S4, using the storage rate anomaly point sequence, the predicted transmission rate sequence, and the transmission rate residual sequence as the state space input, setting the transmission rate adjustment value as the action space output, and constructing a deep reinforcement learning model based on Proximal Policy Optimization (PPO); S5, constructing a multi-objective reward function according to the storage rate anomaly point sequence and the transmission rate residual sequence, and training the deep reinforcement learning model by maximizing the objective reward function; S6, using the trained deep reinforcement learning model to generate an optimal data transmission rate control strategy.
[0008] Among them, this application uses the SARIMA model to capture the trends, seasonality, and autocorrelation in the transmission rate. As a traditional time series model, SARIMA is good at mining linear relationships and can incorporate storage rate anomaly points as exogenous variables to make the prediction more accurate. The LSTM model is used to capture the non-linear relationships and long-term dependencies in the transmission rate. As a recurrent neural network, LSTM effectively processes time series data through a gating mechanism to mine complex patterns in the rate changes. Through model combination, the advantages of both are taken. SARIMA and LSTM model the transmission rate from different perspectives, and can form a complement after combination, reducing the limitations of a single model and improving the robustness and accuracy of the prediction.
[0009] Among them, in this application, the rate prediction and anomaly information are incorporated into the state space. Storage rate anomaly points, predicted transmission rates, and residuals, etc., reflect the current health and load trends of the system, and can provide important clues for decision-making as state representations. The transmission rate adjustment value is set as the action space. This action design directly corresponds to the control variables in the data transmission process, enabling the agent to flexibly adjust the transmission rhythm to cope with the dynamically changing storage conditions. The Actor-Critic framework is adopted for policy learning and value evaluation. The Actor is responsible for generating the action probability distribution, and the Critic is responsible for evaluating the state value. The two cooperate to efficiently explore and improve the control strategy.
[0010] Among them, traditional reinforcement learning such as REINFORCE, Actor-Critic, etc. directly updates the policy through policy gradients, but it is prone to problems such as gradient vanishing or explosion in complex environments, resulting in unstable training. While this application uses PPO, by introducing importance sampling and a surrogate objective function, it limits the policy update amplitude while ensuring monotonicity, making the training process smoother and more controllable.
[0011] Among them, in this application, the storage unit refers to the hardware device or component used to store data, such as hard disks, solid-state drives, memory, etc. It is the basic component of the data storage system and is responsible for the persistent storage and access of data. In this solution, the storage unit is the object of optimized data storage, and its storage rate data is used for anomaly detection and transmission rate prediction. The data transmission unit refers to the hardware device or component responsible for data transmission, such as network interface cards, buses, etc. It connects the storage unit and other computing devices to realize the movement and exchange of data between different nodes. In this solution, the transmission rate data of the data transmission unit is used to predict the future transmission rate and serve as the state input of the reinforcement learning model. The storage rate data refers to the speed at which the storage unit writes or reads data within a certain period of time, usually expressed in units such as MB / s or GB / s. It reflects the performance and load status of the storage device. In this solution, the storage rate data is used for anomaly detection to identify abnormal behaviors or performance bottlenecks in the storage system. The transmission rate data refers to the speed at which the data transmission unit transmits data within a certain period of time, usually expressed in units such as Mbps or Gbps. It reflects the bandwidth utilization and throughput of the data transmission channel. In this solution, the transmission rate data is used to predict the future transmission rate and serve as the state input and optimization target of the reinforcement learning model.
[0012] SARIMA is the abbreviation of "Seasonal Auto Regressive Integrated Moving Average", that is, the seasonal autoregressive moving average model. It is a classic time series prediction model that predicts the future numerical trend by modeling the characteristics of the sequence such as trends, seasonality, and random fluctuations. In this solution, the SARIMA model is used to predict the transmission rate and combined with the LSTM model to form a multi-model prediction method.
[0013] Further, in S3, based on the transmission rate data and the storage rate anomaly point sequence, a multi-model combination prediction method based on SARIMA and LSTM is adopted to obtain the predicted transmission rate sequence at the next moment, including: S31, taking the storage rate anomaly point sequence as an exogenous variable; S32, based on the transmission rate data and the exogenous variable, using the SARIMA model for single-step prediction to obtain the SARIMA predicted transmission rate sequence; S33, based on the transmission rate data, using the LSTM model for single-step prediction to obtain the LSTM predicted transmission rate sequence; S34, based on the SARIMA predicted transmission rate sequence and the LSTM predicted transmission rate sequence, through ensemble learning, obtaining the final predicted transmission rate sequence
[0014] Among them, an exogenous variable, also known as an external variable or a covariate, refers to other influencing factors introduced in time series prediction in addition to the historical data of the target variable itself. These factors are not affected by the target variable but have a certain explanatory or predictive ability for the future trend of the target variable. In this solution, the storage rate anomaly point sequence is introduced as an exogenous variable into the SARIMA model to improve the accuracy of transmission rate prediction. In this solution, the storage rate anomaly point sequence, as an exogenous variable, reflects the potential impact of the abnormal state of the storage system on the transmission rate. By incorporating it into the SARIMA model, the dynamic changes of the transmission rate can be more comprehensively modeled, improving the accuracy and reliability of the prediction.
[0015] Single-step prediction, also known as one-step prediction or rolling prediction, refers to in time series prediction, each time only predicting the value at the next time step, then feeding the true observed value back to the model, and then predicting the next step, and so on. In this solution, the SARIMA model and the LSTM model respectively adopt the single-step prediction method to predict the transmission rate value at the next moment based on the historical data of the transmission rate and the exogenous variable. This rolling prediction mechanism can make full use of real-time data, dynamically adjust the prediction results, and adapt to the real-time changes of the transmission rate. At the same time, single-step prediction also reduces the uncertainty of long-term prediction and improves the credibility of the prediction.
[0016] Among them, various abnormal situations may exist in the storage environment, such as hardware failures, load mutations, etc. A prediction model trained solely relying on historical transmission rate data may be difficult to handle these abnormal scenarios. Introducing the storage rate anomaly point as an exogenous variable can enhance the robustness of the model, enabling it to still make reliable predictions under abnormal conditions. In traditional time series prediction, usually only the historical data of the target variable itself is used for modeling. And introducing the anomaly point sequence as an exogenous variable enriches the information source of the prediction model and provides a more comprehensive decision-making reference for the model.
[0017] Further, in S32, according to the transmission rate data and exogenous variables, use the SARIMA model for single-step prediction to obtain the SARIMA predicted transmission rate sequence, including: constructing a SARIMA model based on the transmission rate data and exogenous variables; the SARIMA model includes the seasonal autoregressive order (P), seasonal differencing order (D), seasonal moving average order (Q), non-seasonal autoregressive order (p), non-seasonal differencing order (d), and non-seasonal moving average order (q), and the exogenous variables are used as the regression terms of the SARIMA model; use the conditional least squares method to estimate the parameters of the SARIMA model; the parameters include seasonal autoregressive parameters and non-seasonal autoregressive parameters, seasonal moving average parameters and non-seasonal moving average parameters, and the coefficients of exogenous variables; substitute the estimated parameters into the SARIMA model to obtain the fitted SARIMA model; according to the fitted SARIMA model and exogenous variables, predict the transmission rate at the next time step to obtain the SARIMA predicted transmission rate sequence.
[0018] Further, in S34, obtain the final predicted transmission rate sequence Through the following formula: Where represents the predicted transmission rate of the i-th model at time t, and P(M i |D) represents the posterior probability of the i-th model; Where M i represents the i-th model (SARIMA or LSTM), D represents the transmission rate data, and P(D|M i ) represents the likelihood of the data D under the model M i , and P(M i ) represents the prior probability of the model M i , and m represents the total number of models (here m = 2).
[0019] On the one hand, traditional prediction formulas usually directly take the output of a single optimal model as the final prediction, or simply average the prediction results of multiple models. This approach ignores the uncertainty in model selection and parameter estimation, which may lead to prediction bias and overfitting risks. In this application, the idea of Bayesian model averaging is introduced. By weighted combining the prediction results of different models, the uncertainty of the models is explicitly considered. The weight of each model is determined by its posterior probability, which reflects the credibility of the model given the data. This weighted average method can effectively reduce the risk of single model selection and improve the robustness of the prediction.
[0020] On the other hand, in traditional prediction formulas, if the weighted average method is adopted, fixed model weights usually need to be determined in advance. Such subjectively set weights may not be consistent with the true distribution of the data and are difficult to adapt to the dynamic changes of the data. In contrast, this formula adaptively learns the model weights through Bayesian inference. According to Bayes' theorem, the posterior probability of the model is proportional to the product of the likelihood and the prior probability. Among them, the likelihood P(D|M_i) measures the degree of fit of the data D under the model M_i, and the prior probability P(M_i) reflects the subjective preference for the model M_i. By continuously updating the likelihood and the prior probability, the model weights can be adaptively adjusted to better match the actual distribution of the data.
[0021] Further, in S2, the density-based LOF algorithm is used to perform anomaly detection on the storage rate data to obtain a sequence of storage rate anomaly points, including: S21, according to the preset minimum number of neighbors k min and the maximum number of neighbors k max , calculate the number of neighbors k(p) and the local reachability density LRD(p) of each data point p; S22, according to the local reachability density LRD(p), calculate the local outlier factor LOF(p) and the global outlier factor GOF(p) of each data point to obtain the comprehensive anomaly score OF(p); S23, according to the comprehensive anomaly score OF(p), calculate the segmentation threshold threshold through the following formula: threshold = μ OF +β×σ OF , where μ OF and σ OF are the mean and standard deviation of the comprehensive anomaly score OF(p) respectively, and β is a hyperparameter that controls the threshold, so that the threshold can adapt to different data characteristics and anomaly detection requirements; S24, mark the data points with the comprehensive anomaly score OF(p) greater than the segmentation threshold threshold as anomaly points to obtain a sequence of storage rate anomaly points.
[0022] Among them, the conventional LOF algorithm usually requires manually setting an anomaly threshold and marking the points with scores exceeding the threshold as anomalies. However, the selection of the optimal threshold depends on the data distribution and application scenarios and is difficult to determine by experience. In this application, by calculating the mean and standard deviation of the comprehensive anomaly score OF, the threshold is automatically set. The hyperparameter β controls the strictness of the threshold, enabling the threshold to adapt to different data characteristics and anomaly detection requirements.
[0023] Further, to calculate the number of neighbors k(p) and the local reachability density LRD(p) of each data point p, the following formula is used: where k min and k max are the preset minimum and maximum numbers of neighbors respectively, and λ is a hyperparameter that controls the change rate of the number of neighbors. Denotes rounding up; LRD(p) is the local reachability density of data point p, and the calculation formula for the local reachability density LRD(p) of data point p is: Among them, N k (p) represents the k-nearest neighbor set of data point p, and |N k (p)| represents the size of the set, and lrd(p,q) represents the local reachability distance of data point p relative to its k-nearest neighbor q.
[0024] Among them, the conventional LOF algorithm usually uses a fixed number of neighbors k for anomaly score calculation. However, the optimal value of k varies depending on the data distribution and anomaly type and is difficult to determine in advance. An inappropriate value of k may lead to too many false positives or false negatives. This application introduces a strategy for adaptively determining the number of neighbors k. By considering the local reachability density LRD of data points, the number of neighbors of each data point is dynamically adjusted. The smaller the LRD, the sparser the data point, and the larger its number of neighbors. This adaptive strategy can flexibly determine the number of neighbors according to the local characteristics of the data and improve the accuracy of anomaly detection.
[0025] Furthermore, calculate the local anomaly factor LOF(p) and the global anomaly factor GOF(p) of each data point to obtain the comprehensive anomaly score OF(p), using the following formula: Among them, LRD(q) represents the local reachability density of data point q, and LRD(p) represents the local reachability density of data point p; Among them, μ LOF represents the mean of the LOF sequence, and σ LOF represents the standard deviation of the LOF sequence;
[0026] OF(p) = α × LOF(p) + (1 - α) × GOF(p), where α is a balance factor that controls the weights of the local anomaly factor and the global anomaly factor.
[0027] Among them, on the one hand, the conventional LOF algorithm only considers the local anomaly degree of data points and judges anomaly by comparing the difference between its density and the density of its neighbors. However, this local perspective may ignore the global anomaly patterns and distribution information. Based on the local anomaly factor LOF, this application introduces the global anomaly factor GOF. GOF standardizes the LOF and depicts the relative position of the data point in the global LOF distribution. It comprehensively considers the anomaly information at both the local and global levels and can more comprehensively evaluate the anomaly degree of data points.
[0028] On the other hand, the anomaly score of the conventional LOF algorithm is completely based on the local outlier factor and lacks the consideration of global anomaly. This may cause some points that are locally normal but globally anomalous to be missed, or some points that are locally anomalous but globally normal to be misreported. In this application, the local outlier factor LOF and the global outlier factor GOF are weighted and fused to obtain a comprehensive anomaly score OF. The weight factor α controls the relative importance of the two outlier factors, enabling the anomaly score to take into account both local and global anomaly characteristics. This fusion strategy can more accurately identify the true anomaly points and reduce the false alarm rate.
[0029] Further, in S4, taking the storage rate anomaly point sequence, the predicted transmission rate sequence, and the transmission rate residual sequence as the state space input, and setting the transmission rate adjustment value as the action space output, a deep reinforcement learning model based on Proximal Policy Optimization (PPO) is constructed, including: S41, extracting features from the storage rate anomaly point sequence, the predicted transmission rate sequence, and the transmission rate residual sequence to obtain a multi-dimensional state vector as the state space input; setting the value range of the transmission rate adjustment value as the action space output; S42, setting the Actor-Critic network structure for policy learning and value estimation; among them, the Actor network adopts the MLP structure, taking the multi-dimensional state vector as the input and outputting the probability distribution of the transmission rate adjustment value; the Critic network adopts the MLP structure, taking the multi-dimensional state vector as the input and outputting the value estimation of the state in the state space; S43, training the Actor-Critic network structure through the Proximal Policy Optimization (PPO) algorithm to obtain a deep reinforcement learning model based on Proximal Policy Optimization (PPO).
[0030] Further, in S43, training the Actor-Critic network structure through the Proximal Policy Optimization (PPO) algorithm to obtain a deep reinforcement learning model based on Proximal Policy Optimization (PPO), including: sampling and generating multiple trajectory data in the state space according to the current policy set by the Actor network, where the trajectory data includes states, actions, and corresponding reward values; calculating the advantage function and the target value of each state-action according to the trajectory data, where the advantage function reflects the quality of the action and the target value reflects the value of the state; constructing a PPO objective function including the policy loss and the value function loss; updating the Actor-Critic network parameters by maximizing the PPO objective function to obtain a deep reinforcement learning model based on Proximal Policy Optimization (PPO).
[0031] Among them, MLP is the abbreviation of "Multi layer Perceptron", that is, a multi-layer perceptron network. It is a feedforward neural network composed of an input layer, a hidden layer, and an output layer. Information is transmitted between layers in a fully connected manner. Through non-linear activation functions and multi-layer structures, MLP can fit complex non-linear mapping relationships. In this solution, both the Actor network and the Critic network adopt the MLP structure. For the Actor network, MLP maps the multi-dimensional state vector to the probability distribution of the transmission rate adjustment value. This structure can model the complex decision-making relationship between states and actions and generate optimized strategies for different states. For the Critic network, MLP maps the multi-dimensional state vector to the value estimation of the state.
[0032] The advantage function, also known as the advantage value or advantage estimate, is an indicator used to evaluate the quality of actions in reinforcement learning. It measures the degree of advantage of taking a specific action in a certain state relative to the average policy. In this solution, the advantage function is used to evaluate the quality of different transmission rate adjustment values. For each state-action pair, the advantage estimate of the action is obtained by calculating the difference between the actual return and the state value. The advantage function guides the policy to update in a better direction, encourages actions with high advantages, and suppresses actions with low advantages. It reduces the variance of policy updates and improves the stability and sample efficiency of learning.
[0033] The target value, also known as the target function or value target, is the true label used to evaluate the state value in reinforcement learning. It represents the expected cumulative return that can be obtained by following the optimal policy in a certain state. In this solution, the target value is used to evaluate the long-term benefits in different states. For each state, the target value of the state is calculated based on the actual return and the estimated value of the next state. The target value provides a learning target for the value function, guiding it to approximate the true state value. It corrects the bias in value estimation and improves the prediction accuracy of the value function. It reflects the temporal difference characteristics of reinforcement learning, combines the immediate return with future benefits, and realizes long-term planning.
[0034] Further, in S5, a multi-objective reward function is constructed based on the storage rate anomaly point sequence and the transmission rate residual sequence, and the deep reinforcement learning model is trained by maximizing the objective reward function, including: S51, dividing the storage rate anomaly point sequence into time windows of a fixed length, calculating the number of anomaly points in each time window; setting a sub-objective reward function for minimizing the number of anomaly points; S52, dividing the transmission rate residual sequence into time windows of a fixed length, calculating the mean residual in each time window; setting a sub-objective reward function for minimizing the mean residual; S53, obtaining a multi-objective reward function through weighted combination according to the sub-objective reward function for minimizing the number of anomaly points and the sub-objective reward function for minimizing the mean residual; S54, fine-tuning the deep reinforcement learning model obtained in step S43 by maximizing the multi-objective reward function: sampling and generating multiple trajectory data in the state space according to the current policy set by the Actor network, where the trajectory data includes states, actions, and corresponding reward values, and the reward values are calculated by the multi-objective reward function; fine-tuning the Actor-Critic network parameters by maximizing the PPO objective function.
[0035] Compared with the prior art, the advantages of this application are as follows:
[0036] When calculating the local reachability density of data points, the improved LOF algorithm introduces an adaptive dynamic neighbor number selection mechanism, which can automatically adjust the number of neighbors according to the density distribution characteristics of data points, improving the accuracy and sensitivity of anomaly point detection. At the same time, by fusing the local anomaly factor and the global anomaly factor, a comprehensive anomaly score is obtained, which can more comprehensively describe the anomaly degree of data points, reducing false negatives and false positives. In addition, the improved LOF algorithm adopts a threshold-based adaptive anomaly point discrimination method. By dynamically calculating the segmentation threshold of the anomaly score, it can automatically adapt to different data distributions and anomaly ratios, improving the accuracy and robustness of anomaly point detection.
[0037] By introducing the storage rate outlier sequence as an exogenous variable, the SARIMA model can characterize the impact of storage rate outliers on the transmission rate and capture the periodic trends and exogenous interference factors in the transmission rate. Incorporating outlier information into the prediction model can better reflect the dynamic changes in the storage environment and improve the adaptability and robustness of the prediction model to abnormal events. At the same time, the LSTM model can capture the non-linear dynamic characteristics of the transmission rate and learn the long-term dependence relationships of the time series. Through the deep learning ability of the LSTM model, complex spatio-temporal patterns in the transmission rate data can be mined to achieve accurate prediction of future transmission rates. Combining the SARIMA model and the LSTM model can take into account both exogenous abnormal factors and internal dynamic characteristics to achieve multi-angle and high-precision prediction of the transmission rate. The outlier information introduced through the SARIMA model can enhance the sensitivity of the prediction model to external environmental changes, while the LSTM model can deeply explore the internal laws of the transmission rate and improve the long-term stability of the prediction.
[0038] Traditional data storage optimization methods usually rely on preset rules or heuristic strategies and are difficult to cope with complex and changing storage environments. The deep reinforcement learning model can autonomously learn and optimize the data transmission rate control strategy through continuous interaction with the storage environment, and has stronger adaptability and robustness. As an efficient deep reinforcement learning algorithm, the PPO algorithm can simultaneously learn the state value function and the policy function through the Actor-Critic network structure, balance between policy exploration and exploitation, and improve the stability and sample efficiency of policy learning. This is particularly important for data storage optimization problems because the state space and action space of the storage environment are usually complex, and the reinforcement learning model needs to be able to explore and optimize control strategies efficiently and stably. Using the storage rate outlier sequence, the predicted transmission rate sequence, and the transmission rate residual sequence as state inputs can comprehensively characterize the key dynamic features in the data storage process. The outlier sequence reflects the imbalance in the utilization of storage resources, the predicted transmission rate sequence provides the expected change trend of the future transmission rate, and the residual sequence characterizes the deviation between the actual transmission rate and the expected value. The organic integration of these state information enables the reinforcement learning model to comprehensively perceive the dynamic changes in the storage environment and capture the key influencing factors in the data storage process. Brief Description of the Drawings
[0039] This application will be further described in the form of exemplary embodiments, which will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:
[0040] Figure 1 is an exemplary flowchart of an artificial intelligence-based data optimized storage method shown in some embodiments of this application;
[0042] Figure 2 is an exemplary flowchart for generating a final predicted transmission rate sequence as shown in some embodiments of the present application;
[0043] Figure 3 is an exemplary flowchart for determining a storage rate anomaly point sequence as shown in some embodiments of the present application. Detailed implementation manners
[0044] The methods and systems provided in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0045] As Figure 1 shown, collect the storage rate data of the storage unit and the transmission rate data of the data transmission unit; perform anomaly detection on the storage rate data using the density-based local outlier factor (LOF) algorithm to obtain a storage rate anomaly point sequence; according to the transmission rate data and the storage rate anomaly point sequence, use a multi-model combination prediction method based on SARIMA and LSTM to obtain a predicted transmission rate sequence for the next moment; calculate the residual between the collected transmission rate data and the predicted transmission rate sequence to obtain a transmission rate residual sequence; use the storage rate anomaly point sequence, the predicted transmission rate sequence, and the transmission rate residual sequence as the state space input, and set the transmission rate adjustment value as the action space output to construct a deep reinforcement learning model based on proximal policy optimization (PPO); construct a multi-objective reward function according to the storage rate anomaly point sequence and the transmission rate residual sequence, and train the deep reinforcement learning model by maximizing the objective reward function; use the trained deep reinforcement learning model to generate an optimal data transmission rate control strategy.
[0046] Specifically, in S1, deploy a monitoring agent program on the storage unit and the data transmission unit to collect the storage rate and transmission rate data in real time. Design a unified data collection interface specification, including data format, collection frequency, transmission protocol, etc. The storage unit and the data transmission unit report the collected storage rate and transmission rate data to the central monitoring system according to the interface specification. Set an appropriate data collection frequency according to the requirements of data storage optimization. The collection frequency needs to balance data granularity and system overhead. A higher collection frequency can provide more fine-grained monitoring data, but it will increase the monitoring overhead of the system. For different types of storage units and data transmission units, the monitoring agent supports multiple communication protocols, such as TCP / IP, SNMP, HTTP, etc., to adapt to different system environments. Develop a data collection module on the storage unit and the data transmission unit to be responsible for collecting the local storage rate and transmission rate data. The collection module obtains the required monitoring metrics by reading system logs, calling performance counters, etc.
[0047] As Figure 2As shown, in S2, the density-based local outlier factor (LOF) algorithm is used to perform outlier detection on the storage rate data, obtaining a sequence of storage rate outlier points; in S21, according to the preset minimum number of neighbors k min and the maximum number of neighbors k max , calculate the number of neighbors k(p) and the local reachability density LRD(p) of each data point p; calculate the number of neighbors k(p) and the local reachability density LRD(p) of each data point p, using the following formula: where k min and k max are the preset minimum and maximum numbers of neighbors respectively, λ is a hyperparameter controlling the change rate of the number of neighbors, LRD(p) is the local reachability density of data point p, and the calculation formula for the local reachability density LRD(p) of data point p is: where N k (p) represents the k-nearest neighbor set of data point p, |N k (p)| represents the size of the set, and lrd(p,q) represents the local reachability distance of data point p relative to its k-nearest neighbor q; the calculation of the local reachability density LRD(p) takes into account the local reachability distance lrd(p,q) between data point p and its k-nearest neighbors, reflecting the density situation around data point p. Dynamically adjusting the number of neighbors k(p) can adaptively handle data points with different density distributions, improving the accuracy and sensitivity of outlier detection.
[0048] In S22, according to the local reachability density LRD(p), calculate the local outlier factor LOF(p) and the global outlier factor GOF(p) of each data point, obtaining the comprehensive outlier score OF(p); calculate the local outlier factor LOF(p) and the global outlier factor GOF(p) of each data point, obtaining the comprehensive outlier score OF(p), using the following formula: OF(p) = α × LOF(p) + (1 - α) × GOF(p), where LOF(p) is the local outlier factor of data point p, μ LOF and σ LOF are the mean and standard deviation of the LOF sequence respectively, GOF(p) is the global outlier factor of data point p, and α is a balance factor controlling the weights of the local outlier factor and the global outlier factor. Combining the local outlier factor and the global outlier factor can more comprehensively characterize the outlier degree of data points, reducing false negatives and false positives.
[0049] In S23, according to the comprehensive outlier score OF(p), calculate the segmentation threshold threshold through the following formula: threshold = μ OF + β × σ OF , where μ OF and σ OFare the mean and standard deviation of the comprehensive anomaly score OF(p), respectively, and β is the hyperparameter for the control threshold; the adaptive threshold can dynamically adapt to different data distributions and anomaly ratios, improving the accuracy and robustness of anomaly point detection. S24, Mark the data points with the comprehensive anomaly score OF(p) greater than the segmentation threshold threshold as anomaly points to obtain the storage rate anomaly point sequence.
[0050] As Figure 3 shown, S3, According to the transmission rate data and the storage rate anomaly point sequence, adopt a multi-model combination prediction method based on SARIMA and LSTM to obtain the predicted transmission rate sequence at the next moment, including: S31, The storage rate anomaly point sequence reflects the abnormal situation of storage resource utilization and contains the time and degree information of anomaly occurrence. Incorporating the anomaly point sequence as an exogenous variable into the prediction model can capture the impact of storage environment changes on the transmission rate.
[0051] S32, According to the transmission rate data and the exogenous variable, use the SARIMA model for single-step prediction to obtain the SARIMA predicted transmission rate sequence; including: According to the transmission rate data and the exogenous variable, construct the SARIMA model; The SARIMA model consists of a seasonal part and a non-seasonal part, which respectively characterize the seasonal and non-seasonal characteristics of the time series. The seasonal part is defined by the seasonal autoregressive order (P), the seasonal differencing order (D), and the seasonal moving average order (Q), and the non-seasonal part is defined by the non-seasonal autoregressive order (p), the non-seasonal differencing order (d), and the non-seasonal moving average order (q). By analyzing the autocorrelation function (ACF) and partial autocorrelation function (PACF) of the transmission rate data, initially determine the order range of the seasonal and non-seasonal parts. Use methods such as information criteria (such as AIC, BIC) or cross-validation to select the optimal order combination (P, D, Q, p, d, q) within the candidate order range.
[0052] Take the storage rate anomaly point sequence as an exogenous variable to reflect the impact of abnormal changes in the storage environment on the transmission rate. In the SARIMA model, the exogenous variable is introduced as a regression term and, together with the autoregressive term, differencing term, and moving average term, jointly determines the predicted value of the transmission rate. According to the determined order combination (P, D, Q, p, d, q) and the exogenous variable, construct the equation of the SARIMA model. The model equation includes the seasonal autoregressive term, seasonal differencing term, seasonal moving average term, non-seasonal autoregressive term, non-seasonal differencing term, non-seasonal moving average term, and the regression term of the exogenous variable. The model equation can be expressed as: where, y t is the transmission rate time series, x t is the exogenous variable time series, B is the lag operator, φ and Φ are autoregressive parameters, θ and Θ are moving average parameters, β is the coefficient of the exogenous variable, εt is the white noise term.
[0053] According to the structure of the SARIMA model, determine the parameters included in the model: seasonal autoregressive parameter φ, non-seasonal autoregressive parameter Φ, seasonal moving average parameter θ, non-seasonal moving average parameter Θ, and the coefficient β of the exogenous variable. Define the objective function as the sum of squares of the model prediction errors, that is, the sum of squares of the difference between the actual transmission rate and the model prediction value. The objective function can be expressed as: where: y t is the actual transmission rate. is the model prediction value, which is a function of φ, Φ, θ, Θ, and β. n is the length of the sample data.
[0054] Select a suitable numerical optimization algorithm, such as the gradient descent method, quasi-Newton method, or Levenberg-Marquardt algorithm, etc. Initialize the parameter values φ, Φ, θ, Θ, and β, which can be randomly initialized or initialized based on prior knowledge. In each iteration, calculate the gradient or Hessian matrix of the objective function J(φ, Φ, θ, Θ, β) with respect to each parameter. According to the gradient or Hessian matrix, update the parameter values to gradually reduce the value of the objective function. Set the convergence condition, such as the change in the parameter is less than the set threshold, or the change in the value of the objective function is less than the set threshold. Iteratively optimize the process until the convergence condition is reached to obtain the optimal parameter estimation values.
[0055] Substitute the optimal parameter values φ, Φ, θ, Θ, and β obtained by conditional least squares estimation into the equation of the SARIMA model. The general form of the SARIMA model is: ARIMA(p, d, q)(P, D, Q) m , where p is the non-seasonal autoregressive order, d is the non-seasonal differencing order, q is the non-seasonal moving average order, P is the seasonal autoregressive order, D is the seasonal differencing order, Q is the seasonal moving average order, and m is the seasonal period.
[0056] The exogenous variable is added to the model in the form of a regression term, which can capture the relationship between the transmission rate and other relevant factors. Represent the exogenous variable as a time series and align it with the transmission rate time series. For each exogenous variable, construct the corresponding regression term. The regression term can be a linear combination or a non-linear function of the exogenous variable, and the specific form depends on the relationship between the variables. Common forms of regression terms include: linear regression term: β1x 1,t +β2x 2,t +.......+β k x k,t ; where β1, β2,......., β k are the regression coefficients, and x 1,t , x 2,t ,......., x k,tis the value of the exogenous variable at time t. Nonlinear regression term:
[0057] f(x 1,t ,x 2,t ,.......,x k,t ; β); where f(*) is a nonlinear function, such as a polynomial function, exponential function, logarithmic function, etc., and β is the parameter of the function. The constructed regression term is added to the mean equation of the SARIMAX model. The mean equation of the SARIMAX(p, d, q)(P, D, Q)_m model can be expressed as:
[0058] where φ1, φ2,....., φ p are non-seasonal autoregressive parameters, Φ1, Φ2,......, Φ P are seasonal autoregressive parameters, θ1, θ2,......, θ q are non-seasonal moving average parameters, Θ1, Θ2,....., Θ Q are seasonal moving average parameters, B is the lag operator, and ε t is the error term. Add the regression term β1x 1,t +β2x 2,t +......+β k x k,t to the right side of the mean equation as the influence of the exogenous variable.
[0059] Use methods such as conditional least squares or maximum likelihood estimation to estimate the parameters of the SARIMAX model. During the estimation process, estimate the autoregressive parameters, moving average parameters, and regression coefficients simultaneously. After obtaining the estimated parameter values, substitute them into the SARIMAX model to obtain the fitted model. Using the fitted SARIMA model, calculate the fitted value of the transmission rate at each time step. The fitted value is the estimated value of the model for the transmission rate at the current time step given the known historical transmission rate data and exogenous variable data. By substituting the historical data into the fitted model, calculate the fitted value at each time step recursively. Obtain the storage rate anomaly point information at the latest time step as the input of the exogenous variable. Substitute the latest exogenous variable value into the fitted SARIMA model. Use the autoregressive term, moving average term, and regression term of the exogenous variable of the model to calculate the predicted value of the transmission rate at the next time step. Take the predicted transmission rate value at the next time step as the new input, combine it with the latest exogenous variable value, and recursively predict the transmission rate at subsequent time steps. Through recursive prediction, a sequence of SARIMA predicted transmission rates for multiple future time steps can be obtained. In each prediction step, use the transmission rate value predicted in the previous step and the latest exogenous variable value as the input to calculate the predicted value at the current time step. Repeat the recursive prediction process until the required number of future time steps is predicted.
[0060] This application constructs the conditional least squares objective function of the SARIMA model and uses a numerical optimization algorithm to estimate the model parameters. Substitute the estimated optimal parameter values into the SARIMA model to calculate the fitted value of the transmission rate. When performing single-step prediction, utilize the latest exogenous variable information and model parameters to recursively predict the transmission rate for multiple future time steps. This process comprehensively utilizes historical transmission rate data, exogenous variable information, and the structure of the SARIMA model to achieve dynamic prediction of the transmission rate. By continuously updating the latest data and exogenous variable information, the prediction results can be adjusted in real time to improve the accuracy and adaptability of the prediction.
[0061] S33. According to the characteristics of the transmission rate data, construct an LSTM (Long Short-Term Memory) model. The LSTM model can effectively capture the long-term dependence relationships in time series data and is suitable for the transmission rate prediction task. The model structure can include one or more LSTM layers, as well as fully connected layers and output layers. Use the training set data to train the LSTM model through the backpropagation algorithm and optimizers (such as Adam, RMS prop, etc.). During the training process, by adjusting the hyperparameters of the model (such as the number of hidden layers, the number of hidden units, the learning rate, etc.), find the optimal model structure and parameter configuration. Use the validation set data to evaluate the performance of the model and monitor the generalization ability and overfitting situation of the model. Use the trained LSTM model to perform single-step prediction on the test set data. At each time step, use the transmission rate of the previous time step as the input to predict the transmission rate of the current time step. By recursively using the predicted value as the input for the next time step, an LSTM predicted transmission rate sequence for multiple future time steps can be obtained.
[0062] S34. Respectively use the SARIMA model and the LSTM model to perform single-step prediction on the transmission rate to obtain the SARIMA predicted transmission rate sequence and the LSTM predicted transmission rate sequence. According to the principle of Bayesian Model Averaging, calculate the posterior probabilities of the SARIMA model and the LSTM model. The posterior probability represents the conditional probability of each model given the transmission rate data. The calculation formula for the posterior probability is: where M i represents the i-th model, D represents the transmission rate data, P(D|M i ) represents the likelihood of the data D under the model M i , P(M i ) represents the prior probability of the model M i , and m represents the total number of models (here m = 2).
[0063] The prior probability P(Mi ) represents the initial belief or preference for model M before observing the data. The choice of prior probability can be based on prior knowledge, expert experience, or factors such as the assumed complexity of the model. In this application, complexity penalty methods such as AIC (Akaike Information Criterion) and BIC (Bayesian Information Criterion) are used. For the SARIMA model and the LSTM model, their i can be calculated. where k i represents the number of parameters of model M i , represents the maximum likelihood value of model M i ; then the prior probability is assigned according to the complexity penalty ΔIC i = IC(M i ) - min(IC(M1), IC(M2)).
[0064] Using the prediction results of the SARIMA model and the LSTM model, the final predicted transmission rate sequence is obtained through ensemble learning. The formula for ensemble learning is: where, represents the final predicted transmission rate, represents the predicted transmission rate of the i-th model at time t, and P(M i |D) represents the posterior probability of the i-th model. By means of weighted averaging, combining the prediction results of the SARIMA model and the LSTM model, a more accurate and robust predicted transmission rate sequence is obtained.
[0065] S4. Using the storage rate anomaly point sequence, the predicted transmission rate sequence, and the transmission rate residual sequence as the state space input, and setting the transmission rate adjustment value as the action space output, a deep reinforcement learning model based on Proximal Policy Optimization (PPO) is constructed, including: S41. Feature extraction is performed on the storage rate anomaly point sequence, for example, statistical features such as the number, frequency, and amplitude of the anomaly points are calculated to obtain the anomaly point feature vector.
[0066] Extract the sequence features of storage rate anomalies, and count the total number of storage rate anomalies within a given time window. The sliding window method can be used to count the number of anomalies in different time periods. Calculate the frequency of anomalies occurring within a given time window, that is, the number of anomalies divided by the length of the time window. The frequency reflects the density of anomalies in time. Calculate the deviation degree of anomalies from the normal storage rate, which can be an absolute value or a relative value. The amplitude reflects the severity of anomalies. Combine the statistical features such as the number, frequency, and amplitude of anomalies into a feature vector. The feature vector can include the statistical features of multiple time windows to capture the time dynamic information of anomalies.
[0067] Extract the sequence features of predicted transmission rate, and calculate the arithmetic mean of the predicted transmission rate within a given time window. The mean value reflects the overall level of the predicted rate. Calculate the mean of the squares of the differences between the predicted transmission rate and the mean within a given time window. The variance reflects the fluctuation degree of the predicted rate. Use linear regression or other trend analysis methods to calculate the trend coefficient of the predicted transmission rate. The trend reflects the change direction and speed of the predicted rate. Combine the statistical features such as the mean, variance, and trend of the predicted values into a feature vector. The feature vector can include the statistical features of multiple time windows to capture the time dynamic information of the predicted rate.
[0068] Extract the sequence features of transmission rate residuals, and calculate the arithmetic mean of the transmission rate residuals within a given time window. The mean of the residuals reflects the overall level of the prediction error. Calculate the mean of the squares of the differences between the transmission rate residuals and the mean within a given time window. The variance of the residuals reflects the fluctuation degree of the prediction error. Calculate the autocorrelation coefficient of the transmission rate residuals at different time lags. The autocorrelation reflects the dependence relationship of the residuals in time. Combine the statistical features such as the mean, variance, and autocorrelation of the residuals into a feature vector. The feature vector can include the statistical features of multiple time windows to capture the time dynamic information of the residuals.
[0069] Construct a multi-dimensional state vector, and splice the anomaly feature vector, the predicted rate feature vector, and the residual feature vector in a certain order into a multi-dimensional vector. The spliced vector contains the comprehensive information of anomalies, predicted rates, and residuals. Use the spliced multi-dimensional vector as the input of the state space. The state space represents the state of the system at the current moment, including the feature information of anomalies, predicted rates, and residuals.
[0070] Set the action space and set the value range of the transmission rate adjustment value, such as [-0.1, 0.1]. The adjustment value range is determined according to specific application scenarios and requirements. Take the transmission rate adjustment value range as the output of the action space. The action space represents the actions that the agent can take, that is, the range for adjusting the transmission rate. In this application, feature extraction is performed on the storage rate anomaly point sequence, the predicted transmission rate sequence, and the transmission rate residual sequence to obtain the anomaly point feature vector, the predicted rate feature vector, and the residual feature vector. These feature vectors capture the statistical features of the anomaly points, predicted rates, and residuals, and consider the time dynamic information. Concatenate the feature vectors into a multi-dimensional state vector as the input of the state space, representing the current state of the system. At the same time, set the value range of the transmission rate adjustment value as the output of the action space, representing the actions that the agent can take. By constructing the state space and the action space, we can model the transmission rate optimization problem as a reinforcement learning problem. The agent can observe the state of the system in the state space, select appropriate actions according to the current state, that is, adjust the transmission rate, and obtain corresponding reward feedback. By continuously interacting with the environment and learning, the agent can learn the optimal transmission rate adjustment strategy to achieve the goals of minimizing anomaly points, minimizing prediction errors, and stabilizing the transmission rate.
[0071] S42. Set the Actor-Critic network structure for policy learning and value estimation; the Actor network adopts the MLP (Multi-Layer Perceptron) structure, with the multi-dimensional state vector as the input and the probability distribution of the transmission rate adjustment value as the output. The first layer: a fully connected layer with the activation function ReLU. The second layer: a fully connected layer with the activation function ReLU. The output layer: a fully connected layer with the activation function Softmax, outputting the probability distribution of the transmission rate adjustment value. The Critic network adopts the MLP structure, with the multi-dimensional state vector as the input and the value estimation of the state in the state space as the output. The first layer: a fully connected layer with the activation function ReLU. The second layer: a fully connected layer with the activation function ReLU. The output layer: a fully connected layer with the activation function being a linear function, outputting the value estimation of the state.
[0072] S43: Train with the Proximal Policy Optimization (PPO) algorithm. According to the current policy parameters of the Actor network, sample and generate multiple trajectory data in the state space. Each trajectory contains a series of states, actions, and corresponding reward values, representing the process of the agent interacting with the environment. The sampling process can be achieved by interacting with the environment or using the existing experience replay pool.
[0073] For each state-action pair in the trajectory data, the advantage function and the target value are calculated. The advantage function A(s, a) measures the quality of taking action a in state s, and is expressed as the difference between the state-action value function Q(s, a) and the state value function V(s). Q(s, a) estimates the long-term expected return of taking action a in state s. V(s) estimates the long-term expected return of state s. The advantage function A(s, a) = Q(s, a) - V(s), which reflects the quality of action a relative to the average performance. The target value is calculated by the Generalized Advantage Estimation (GAE) algorithm and is used to estimate the true value of the state. The GAE algorithm balances bias and variance by introducing a decay factor, providing a more stable estimate of the advantage function. The target value can be expressed as a weighted sum of the state value function V(s) and the advantage function A(s, a).
[0074] The PPO objective function consists of two parts: the policy loss and the value function loss. The policy loss encourages actions with a larger advantage function and penalizes actions with a smaller advantage function, guiding the policy to update in a better direction. The policy loss can be expressed as the expected value of the product of the advantage function and the ratio of policy probabilities. To prevent the policy from being updated too much, PPO introduces a clips term to limit the magnitude of the policy update. The value function loss minimizes the mean squared error between the state value function V(s) and the target value, improving the accuracy of the state value estimate.
[0075] Maximize the PPO objective function through a gradient ascent algorithm (such as Adam) to update the parameters of the Actor-Critic network. The parameter update of the Actor network aims to maximize the policy loss, causing the policy to update in the direction of actions with a larger advantage function. The parameter update of the Critic network aims to minimize the value function loss, making the state value estimate more accurate. The gradient ascent process calculates the gradient of the objective function with respect to the network parameters through the backpropagation algorithm and updates the parameters according to the gradient. Repeat steps 1 - 4 for multiple iterations to train the PPO algorithm. Each iteration samples new trajectory data, calculates the advantage function and the target value, constructs the PPO objective function, and updates the parameters of the Actor-Critic network. As the training progresses, the policy will be continuously optimized, the state value estimate will be more accurate, and the decision-making ability of the agent will gradually improve. The termination condition of the training process can be set with a certain number of iterations or determined according to the degree of policy convergence. The PPO algorithm can effectively train the Actor-Critic network to achieve policy optimization and value estimation. By introducing the advantage function and the target value, constructing a suitable objective function, and using the gradient ascent algorithm to update the network parameters, the PPO algorithm continuously improves the quality of policy and value estimation. After multiple iterations of training, the PPO algorithm can find a better policy, enabling the agent to make better decisions in the face of a complex environment.
[0076] S5. Construct a multi-objective reward function based on the storage rate anomaly point sequence and the transmission rate residual sequence, and train a deep reinforcement learning model by maximizing the objective reward function, including: S51. Divide the storage rate anomaly point sequence into time windows of a fixed length. For example, each window contains 10 time steps. For each time window, count the number of anomaly points within the window. Anomaly points can be determined according to a preset threshold or statistical method. Design a sub-objective reward function for minimizing the number of anomaly points to punish cases with a large number of anomaly points. For example: R anomaly =-k×num anomalies , where k is the penalty coefficient and num anomalies is the number of anomaly points within the time window. The more anomaly points there are, the greater the penalty and the smaller the value of R anomaly .
[0077] S52. Divide the transmission rate residual sequence into time windows of a fixed length, which is consistent with the time window in S51. For each time window, calculate the mean absolute value of the residuals within the window. The residual refers to the difference between the actual transmission rate and the predicted transmission rate. Design a sub-objective reward function for minimizing the mean of the residuals to punish cases with a large mean of the residuals. For example: R residual =-λ×mean abs,residual , where λ is the penalty coefficient and mean abs,residual is the mean of the absolute values of the residuals within the time window. The greater the mean of the residuals, the greater the penalty and the smaller the value of R residual .
[0078] S53. Combine the sub-objective reward function for minimizing the number of anomaly points and the sub-objective reward function for minimizing the mean of the residuals through weighted combination to obtain a multi-objective reward function. The multi-objective reward function can be expressed as: R = w1×R anomaly +w2×R residual , where w1 and w2 are weight coefficients. The weight coefficients w1 and w2 can be set according to the characteristics of the problem and the degree of emphasis on different sub-objectives. For example: If more emphasis is placed on reducing anomaly points, w1 can be set larger. If more emphasis is placed on reducing the residuals, w2 can be set larger. The selection of the weight coefficients can be determined through cross-validation or empirical adjustment.
[0079] S54. Fine-tune the deep reinforcement learning model obtained in step S43 by maximizing the multi-objective reward function: Use the constructed multi-objective reward function to fine-tune the deep reinforcement learning model. According to the current policy set by the Actor network, sample multiple trajectory data in the state space. Each trajectory data contains a series of states, actions, and corresponding reward values. The reward value is calculated by the multi-objective reward function, comprehensively considering the influence of the number of outliers and the mean residual. For the sampled trajectory data, calculate the advantage function and the target value. The advantage function measures the quality of each state-action pair. The target value estimates the true value of each state. Construct the PPO objective function, including the policy loss and the value function loss. The policy loss encourages actions with a larger advantage function and punishes actions with a smaller advantage function. The value function loss minimizes the mean squared error between the state value function and the target value. Maximize the PPO objective function through the gradient ascent algorithm to fine-tune the parameters of the Actor-Critic network. The parameter update of the Actor network aims to maximize the policy loss, making the policy update in the direction of actions with a larger advantage function. The parameter update of the Critic network aims to minimize the value function loss, making the state value estimation more accurate. Repeat the iterative training multiple times until the policy converges or reaches the preset number of training epochs.
[0080] This application constructs a multi-objective reward function using the storage rate outlier sequence and the transmission rate residual sequence, and uses it to fine-tune the deep reinforcement learning model. The multi-objective reward function comprehensively considers two sub-objectives: minimizing the number of outliers and minimizing the mean residual, and obtains the final reward function through weighted combination. During the model fine-tuning process, we sample trajectory data according to the policy of the Actor network and use the multi-objective reward function to calculate the reward value of each state-action pair. By maximizing the PPO objective function, we can update the parameters of the Actor-Critic network, making the policy optimize in the direction of minimizing the number of outliers and the mean residual. After multiple iterative trainings, the deep reinforcement learning model can learn a better policy, so as to effectively reduce outliers and residuals when adjusting the transmission rate.
[0081] S6. Use the trained deep reinforcement learning model to generate the optimal data transmission rate control strategy. Specifically, in S61, real-time collect the storage rate data of the storage unit and the transmission rate data of the data transmission unit: Use the data acquisition module to real-time collect the storage rate data of the storage unit and the transmission rate data of the data transmission unit, and store the collected data into the data buffer in chronological order to form a historical data sequence. Among them, the sampling frequency of the storage rate data is f s , and the sampling frequency of the transmission rate data is f t .
[0082] S62. Use the trained LOF anomaly detection model to perform anomaly detection on the real-time collected storage rate data to obtain the storage rate anomaly points at the current moment: Extract the storage rate data at the current moment from the data buffer to form a feature vector x s , and input the feature vector into the trained LOF anomaly detection model to calculate the comprehensive anomaly score OF(x s ) of the current storage rate data. Compare OF(x s ) with the preset anomaly threshold threshold. If OF(x s ) is greater than threshold, mark the current storage rate data as an anomaly point to obtain the storage rate anomaly points at the current moment.
[0083] S63. Input the historical transmission rate data and the storage rate anomaly points at the current moment into the trained SARIMA-LSTM combined prediction model to obtain the predicted transmission rate at the next moment: Extract the historical transmission rate data from the data buffer to form a time series y t , and form a feature matrix [y t , x t with the time series y s,anomaly and the storage rate anomaly points at the current moment. Input the feature matrix into the trained SARIMA-LSTM combined prediction model. After SARIMA single-step prediction, LSTM single-step prediction, and ensemble learning, obtain the predicted transmission rate y t+1 at the next moment.
[0084] S64. Calculate the residual between the actual transmission rate at the current moment and the predicted transmission rate at the next moment to obtain the transmission rate residual at the current moment: Extract the actual transmission rate data y t,real at the current moment from the data buffer, and calculate the residual e t,real between y t+1 and the predicted transmission rate y t at the next moment obtained in step S63 as the transmission rate residual at the current moment.
[0085] S65. Form a state vector with the storage rate anomaly points, predicted transmission rate, and transmission rate residual at the current moment, and input it into the trained Actor network: Form a state vector s s,anomaly = [x t+1 , y t , e t with the storage rate anomaly points x s,anomaly at the current moment, the predicted transmission rate y t+1 at the next moment, and the transmission rate residual e t at the current moment, and input the state vector into the trained Actor network.
[0086] S66, the Actor network outputs the probability distribution of the transmission rate adjustment value according to the current state vector: The Actor network receives the state vector s t After that, the probability distribution of the output transmission rate adjustment value Δy is calculated through forward propagation, and π(Δy|s t ). Specifically, the state vector s t First, it passes through the input layer, then passes through several fully connected hidden layers in sequence, and finally outputs the probability distribution π(Δy|s t ). Among them, the activation function of the hidden layer is the ReLU function, and the activation function of the output layer is the Softmax function.
[0087] S67, sample the transmission rate adjustment value from the probability distribution, use the sampled transmission rate adjustment value as the optimal control strategy, and send it to the data transmission unit: according to the probability distribution π(Δy|s t ), using the Monte Carlo sampling method, from π(Δy|s t ) to obtain the transmission rate adjustment value Δy opt , Δy opt As the optimal transmission rate control strategy under the current environmental conditions, it is sent to the data transmission unit.
[0088] S68, the data transmission unit dynamically adjusts the data transmission rate according to the received optimal transmission rate adjustment value: the data transmission unit receives the optimal transmission rate adjustment value Δy opt After that, the current transmission rate y t with Δy opt Add them together to get the adjusted transmission rate value y t,new , and according to y t,new Dynamically adjust its own data transmission rate to achieve the optimal value.
[0089] S69, repeat S61 to S68 to achieve dynamic optimization control of the data transmission rate: repeat steps S61 to S68 at a fixed time interval Δt to generate an optimal transmission rate control strategy in real time, and dynamically adjust the transmission rate of the data transmission unit according to the optimal strategy to achieve dynamic optimization control of the data transmission rate and improve data transmission efficiency.
Claims
1. A data optimization storage method based on artificial intelligence, characterized in that, Including: S1, collecting the storage rate data of the storage unit and the transmission rate data of the data transmission unit; S2, performing anomaly detection on the storage rate data using the density-based local outlier factor (LOF) algorithm to obtain a sequence of storage rate anomaly points; S3, according to the transmission rate data and the sequence of storage rate anomaly points, adopting a multi-model combination prediction method based on SARIMA and LSTM to obtain a predicted transmission rate sequence for the next moment; Calculating the residuals between the collected transmission rate data and the predicted transmission rate sequence to obtain a transmission rate residual sequence; S4, using the sequence of storage rate anomaly points, the predicted transmission rate sequence, and the transmission rate residual sequence as the state space input, and setting the transmission rate adjustment value as the action space output to construct a deep reinforcement learning model based on proximal policy optimization (PPO); S5, constructing a multi-objective reward function according to the sequence of storage rate anomaly points and the transmission rate residual sequence, and training the deep reinforcement learning model by maximizing the objective reward function; S6, using the trained deep reinforcement learning model to generate an optimal data transmission rate control strategy.
2. The data optimization storage method based on artificial intelligence according to claim 1, characterized in that: S3, adopting a multi-model combination prediction method based on SARIMA and LSTM to obtain a predicted transmission rate sequence for the next moment, including: S31, using the sequence of storage rate anomaly points as exogenous variables; S32, according to the transmission rate data and the exogenous variables, using the SARIMA model for single-step prediction to obtain a SARIMA predicted transmission rate sequence; S33, according to the transmission rate data, using the LSTM model for single-step prediction to obtain an LSTM predicted transmission rate sequence; S34. According to the SARIMA predicted transmission rate sequence and the LSTM predicted transmission rate sequence, through ensemble learning, obtain the final predicted transmission rate sequence 3. The data optimization storage method based on artificial intelligence according to claim 2, characterized in that: S32, using the SARIMA model for single-step prediction to obtain a SARIMA predicted transmission rate sequence, including: Constructing a SARIMA model according to the transmission rate data and the exogenous variables; the SARIMA model includes the seasonal autoregressive order P, the seasonal difference order D, the seasonal moving average order Q, the non-seasonal autoregressive order p, the non-seasonal difference order d, and the non-seasonal moving average order q, and the exogenous variables are used as the regression terms of the SARIMA model; Estimating the parameters of the SARIMA model using the conditional least squares method; the parameters include the seasonal autoregressive parameters and the non-seasonal autoregressive parameters, the seasonal moving average parameters and the non-seasonal moving average parameters, and the coefficients of the exogenous variables; Substituting the estimated parameters into the SARIMA model to obtain a fitted SARIMA model; According to the fitted SARIMA model and the exogenous variables, predicting the transmission rate for the next time step to obtain a SARIMA predicted transmission rate sequence.
4. The data optimization storage method based on artificial intelligence according to claim 2, characterized in that: S34, to obtain the final predicted transmission rate sequence Through the following formula: Among them, represents the predicted transmission rate of the i-th model at time t, and P(M i |D) represents the posterior probability of the i-th model; Among them, M i represents the i-th model (SARIMA or LSTM), D represents the transmission rate data, and P(D|M i ) represents the likelihood of data D under the model M i , P(M i ) represents the prior probability of the model M i , and m represents the total number of models.
5. The data optimization storage method based on artificial intelligence according to claim 2, characterized in that: S2. Use the density-based LOF algorithm to perform anomaly detection on the storage rate data, obtaining a sequence of storage rate anomaly points, including: S21, calculate the number of neighbors k(p) and the local reachability density LRD(p) of each data point p according to the preset minimum number of neighbors k min and the maximum number of neighbors k max ; S22. Calculate the local outlier factor LOF(p) and the global outlier factor GOF(p) for each data point according to the local reachability density LRD(p), obtaining the comprehensive anomaly score OF(p); S23. According to the comprehensive anomaly score OF(p), calculate the segmentation threshold threshold through the following formula: threshold = μ OF + β × σ OF Among them, μ OF and σ OF are the mean and standard deviation of the comprehensive anomaly score OF(p) respectively, and β is the hyperparameter of the control threshold; S24. Mark the data points with the comprehensive anomaly score OF(p) greater than the segmentation threshold threshold as anomaly points, obtaining a sequence of storage rate anomaly points.
6. The data optimization storage method based on artificial intelligence according to claim 5, wherein: Calculate the number of neighbors k(p) and the local reachability density LRD(p) of each data point p, using the following formula: where k min and k max are the preset minimum and maximum numbers of nearest neighbors respectively, λ is a hyperparameter controlling the change rate of the number of nearest neighbors, LRD(p) is the local reachability density of the data point p, and the calculation formula for the local reachability density LRD(p) of the data point p is: Among them, N k (p) represents the set of k-nearest neighbors of the data point p, and |N k (p)| represents the size of the set, and lrd(p, q) represents the local reachability distance of the data point p with respect to its k-nearest neighbor q.
7. The data optimization storage method based on artificial intelligence according to claim 6, wherein: Calculate the local outlier factor LOF(p) and the global outlier factor GOF(p) of each data point to obtain the comprehensive anomaly score OF(p), using the following formula: wherein, LRD(q) represents the local reachability density of data point q, and LRD(p) represents the local reachability density of data point p; Among them, μ LOF represents the mean of the LOF sequence, and σ LOF represents the standard deviation of the LOF sequence; OF(p) = α × LOF(p) + (1 - α) × GOF(p) wherein, α is a balance factor that controls the weights of the local outlier factor and the global outlier factor.
8. The data optimization storage method based on artificial intelligence according to claim 2, wherein: S4. Construct a deep reinforcement learning model based on Proximal Policy Optimization (PPO), including: S41. Extract features from the sequence of storage rate anomaly points, the predicted transmission rate sequence, and the transmission rate residual sequence to obtain a multi-dimensional state vector, which is used as the input of the state space; Set the value range of the transmission rate adjustment value as the output of the action space; S42. Set up an Actor-Critic network structure for policy learning and value estimation; wherein, the Actor network adopts an MLP structure, takes the multi-dimensional state vector as the input, and outputs the probability distribution of the transmission rate adjustment value; The Critic network adopts an MLP structure, takes the multi-dimensional state vector as the input, and outputs the value estimation of the state in the state space; S43. Train the Actor-Critic network structure through the Proximal Policy Optimization (PPO) algorithm to obtain a deep reinforcement learning model based on Proximal Policy Optimization (PPO).
9. The data optimization storage method based on artificial intelligence according to claim 8, wherein: S43. Obtain a deep reinforcement learning model based on Proximal Policy Optimization (PPO), including: According to the current policy set by the Actor network, sample and generate multiple trajectory data in the state space, and the trajectory data includes states, actions, and corresponding reward values; According to the trajectory data, calculate the advantage function and the target value of each state-action, where the advantage function reflects the quality of the action, and the target value reflects the value of the state; Construct a PPO objective function including a policy loss and a value function loss; By maximizing the PPO objective function, update the Actor-Critic network parameters to obtain a deep reinforcement learning model based on proximal policy optimization (PPO).
10. The data optimization storage method based on artificial intelligence according to claim 8, characterized in that: S5. Training the deep reinforcement learning model by maximizing the objective reward function, including: S51. Divide the storage rate anomaly point sequence into time windows of a fixed length, and calculate the number of anomaly points in each time window; Set the sub-objective reward function for minimizing the number of anomaly points; S52. Divide the transmission rate residual sequence into time windows of a fixed length, and calculate the residual mean in each time window; Set the sub-objective reward function for minimizing the residual mean; S53. According to the sub-objective reward function for minimizing the number of anomaly points and the sub-objective reward function for minimizing the residual mean, obtain a multi-objective reward function through weighted combination; S54. Fine-tune the deep reinforcement learning model obtained in step S43 by maximizing the multi-objective reward function: According to the current policy set by the Actor network, sample and generate multiple trajectory data in the state space. The trajectory data includes states, actions, and corresponding reward values, and the reward values are calculated by the multi-objective reward function; Fine-tune the Actor-Critic network parameters by maximizing the PPO objective function.
Citation Information
Patent Citations
Financial risk intelligent analysis method based on big data
CN117994026A
Solid state disk power consumption management method and solid state disk
CN118466733A
Solid state disk data transmission rate control method and solid state disk
CN118708128A
Data optimization storage method based on artificial intelligence
CN119225662A
Anomaly detection in multidimensional time series data
US20190147300A1
Cited By
New energy output data anomaly detection method and device, equipment and storage medium
CN121071322A
Multi-scale time sequence public transport passenger flow prediction method based on IC card data
CN121435195A