Data prefetching method based on multi-head attention mechanism and RNN-LSTM network

By adopting the multi-head attention mechanism and the RNN-LSTM network method in data prefetching, the accuracy and efficiency problems of traditional data prefetchers when processing complex data access modes are solved, and more efficient and accurate data prefetching is achieved, reducing data reading delay.

CN120196565AActive Publication Date: 2025-06-24QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Patent Information

Application Number
CN202510269950.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-06
Filing Date
2025-03-07
Publication Date
2025-06-24
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

Traditional hardware data prefetchers rely on simple access modes with linear and fixed steps and cannot effectively handle complex and diverse data access modes, resulting in reduced accuracy and invalid prefetch.

Method used

Using the data prefetching method based on the multi-head attention mechanism and the RNN-LSTM network, a recurrent neural network fusion model is constructed by acquiring and preprocessing the initial data set, and the multi-head attention mechanism and LSTM layer are used to capture the dependencies and patterns in the time series.

Benefits of technology

It improves the accuracy and efficiency of data prefetching, reduces the latency of data reading, effectively alleviates the storage wall problem, and improves the system's throughput and response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196565A_ABST
    Figure CN120196565A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data prefetching, and particularly provides a data prefetching method based on a multi-head attention mechanism and an RNN-LSTM network. The method comprises the following steps: acquiring an initial data set, and performing proportion division to obtain a training data set, a verification data set and a test data set; preprocessing the initial data set to obtain a preprocessed data set; constructing a recurrent neural network fusion model through a multi-head attention mechanism and an RNN-LSTM recurrent neural network; according to the method, the recurrent neural network fusion model is trained, verified and tested through the preprocessed training data set, verification data set and test data set to achieve data prefetching, and through the fusion technology, by means of the powerful basis of RNN-LSTM in processing a long-time sequence and a specific multi-head attention mechanism, the recurrent neural network fusion model is obtained. The accuracy and the high efficiency of data prefetching are improved, and the delay of data reading is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data prefetching, and particularly to a data prefetching method based on a multi-head attention mechanism and an RNN-LSTM network. Background Art

[0002] A computer system usually consists of multiple levels of memory operating at different performance levels, mainly including registers, caches, DRAM, flash memory, mechanical hard disks, etc. The purpose of the computer system storage hierarchy is to reduce the overall access time and reduce the storage cost. In the storage hierarchy system, in order to further reduce the access latency and improve the system performance, the data prefetching technology has emerged as the times require. Data prefetching is a technology that predicts the data that may be accessed in the future and loads it from the low-speed storage layer to the high-speed storage layer in advance. Its core idea is to place the data in the storage layer close to the processor in advance before the data is actually requested, so as to reduce the latency overhead caused by storage access. An effective data prefetching technology can significantly reduce the cache miss rate and the main memory access latency, and optimize the utilization rate of the system storage resources. Data prefetching is mainly implemented through hardware mechanisms or software instructions, which are applicable to automatic triggering and programming scenarios respectively. With the development of computer technology and computer hardware, the performance of the processor is also constantly improving, and the speed gap between the processor and the storage is gradually expanding, resulting in the "memory wall" problem. The processor may be idle for a long time while waiting for data to be loaded, seriously affecting the system performance. Data prefetching effectively alleviates the memory wall problem by preloading data in advance and reducing the cache miss rate.

[0003] In data-intensive applications such as high-performance computing (HPC), artificial intelligence, and big data processing, data access latency is a key bottleneck for performance optimization. The data prefetching technology accelerates task processing by shortening the transmission time of data from the storage to the processor, and significantly improves the throughput and response speed of the system. However, traditional hardware data prefetchers (such as stream prefetching and Stride prefetching) rely on simple access patterns with linear and fixed step sizes. In real-world scenario applications, the data access patterns are often more complex and diverse, which leads to a decrease in the accuracy of traditional prefetchers, and may even increase ineffective prefetching. An inefficient prefetching strategy may load unnecessary data, replace useful data, reduce the cache hit rate, and even affect the performance. Summary of the Invention

[0004] In view of this, the present invention provides a data prefetching method based on a multi-head attention mechanism and an RNN-LSTM network, so as to improve the accuracy and efficiency of data prefetching and reduce the latency of data reading.

[0005] In a first aspect, the present invention provides a data prefetching method based on a multi-head attention mechanism and an RNN-LSTM network, and the method includes:

[0006] Step 1: Obtain an initial data set, and perform proportional division to obtain a training data set, a validation data set, and a test data set; preprocess the initial data set to obtain a preprocessed data set;

[0007] Step 2: Construct a recurrent neural network fusion model through a multi-head attention mechanism and an RNN-LSTM recurrent neural network;

[0008] Step 3: Train, validate, and test the recurrent neural network fusion model through the preprocessed training data set, validation data set, and test data set respectively to achieve data prefetching.

[0009] Optionally, the obtaining the initial data set in Step 1 includes:

[0010] First, it is necessary to read the buffer area of the flash chip of the Master node to generate a time stamp. The time stamp starts from the unique address time record of the first file read from the buffer area and is generated every 5 minutes; secondly, statistics are performed every 5 minutes to record all unique pages in the current memory; record the memory offset between all unique pages; record the network occupancy rate for reading each unique page; finally, perform data integration to generate a data set form of time stamp timescamp, offset target, broadband occupancy rate dynamic_feat, and unique page mapping ID id;

[0011] The proportional division includes: collecting 10,000 pieces of data, removing the data with large offsets, and 8,000 pieces of valid data remain; and using custom code to set the training data set, validation data set, and test data set in a ratio of 6:2:2 for the entire data set;

[0012] The preprocessing of the initial data set to obtain a preprocessed data set includes:

[0013] Data normalization: Normalization helps to speed up the model training speed, improve the model convergence speed, and reduce the impact of the model initialization method on the training process; the normalization operation adopted is implemented by the MinMaxScaler function in Sklearn. MinMaxScaler scales all its numerical values to a specified range, and its formula is:

[0014]

[0015] where X is the eigenvalue of the original data; X min is the minimum value of the original data eigenvalue X in the data; X maxis the maximum value of the original data feature value X in the data; X scaled is the final result of normalization;

[0016] Perform normalization operations on the broadband occupancy rate dynamic_feat, unique page mapping ID, and offset target columns in the dataset;

[0017] Data prediction mode processing: Define a mode for data input and prediction. The mode is implemented using a custom function create_fixed_sequnences. The specific operation steps are as follows: First, input sequence_length as the input time series step of the model. Then, the model returns a prediction result, that is, a predicted time step, and returns the corresponding timestamp. Among them, the adopted sequence_length is set to 64, representing an input time series of 64 steps for the model to learn, so as to perform the prediction of a one-step time series.

[0018] Optionally, the recurrent neural network fusion model in step 2 includes:

[0019] a. Input layer: The dimension of the input layer is (sequence_length, feature_dim); where sequence_length represents the number of time steps; feature_dim represents the feature dimension of each time step. The input layer inputs the original time series data into the model for feature extraction;

[0020] b. Self-attention mechanism layer: In the attention mechanism, each input element is assigned a weight value according to its relationship with other elements. The weight value is calculated through dot product. The calculation formula of the attention mechanism is:

[0021]

[0022] where Q is the query vector query, representing the target to be focused on currently; K is the key vector key, representing the features in all inputs; V is the value vector value, that is, the actual information associated with the key vector; d k is the dimension of the key vector for scaling;

[0023] c. Residual block: The residual block is used to solve the problem of model degradation when the number of network layers increases. The introduction of the residual block is because when network degradation occurs, the effect of the shallow network is better than that of the deep network. Therefore, the residual block directly passes the calculation results of the low-level network to the high-level network, making the effect of the deep network better than that of the shallow network. The formula of the residual block is:

[0024] X l+1 = X l+A(X l +W l ));

[0025] Among them, X l is the input of the l-th layer, and A(X l +W l ) is the output of the masked multi-head self-attention mechanism module of the l-th layer;

[0026] d. Normalization layer: Layer Normalization is selected to process the data; layer normalization normalizes all neurons in a layer; assuming that the i-th input of the neurons in the l-th layer is z i (l) , the expressions for its mean and variance are respectively:

[0027]

[0028] Among them, n (l) is the number of neurons in the l-th layer;

[0029] Layer normalization is defined as:

[0030]

[0031] where γ represents the scaling parameter vector, β represents the translation parameter vector, and the dimensions of the parameter vectors γ and β are both the same as that of z (l) ;

[0032] e. RNN-LSTM recurrent neural network layer:

[0033] The RNN-LSTM recurrent neural network layer is composed of two layers of recurrent neural network RNN, three layers of long short-term memory network LSTM, and three layers of fully connected layers;

[0034] (1) The basic unit of RNN is a recurrent unit Recurrent Unit, which receives an input and a hidden state from the previous time step and outputs the hidden state of the current time step. The basic RNN structure consists of an input layer, a hidden layer, and an output layer; expanding the basic RNN structure in the time dimension, the entire network structure framework of RNN is obtained;

[0035] After the RNN network receives the input x t at time t, the value of the hidden layer is s t , and the formula is as follows:

[0036] s t = f(U × x t + W × s t-1 );

[0037] U is the weight matrix of the input x, and W is the value of the previous hidden layer s t-1 As the current input weight matrix, f is the activation function, and s t depends on x t and s t-1 ;

[0038] The output value is o t , and the calculation formula is as follows:

[0039] o t = g(V × s t );

[0040] V is the weight matrix of the output layer, and g is the activation function;

[0041] The hidden layer has two inputs. The first is the product of the weight matrix U from the input layer to the hidden layer and the input x t , and the second is the product of the value of the previous hidden layer s t-1 and W, that is, the s calculated at the previous moment t-1 needs to be cached and calculated with the input x t to jointly output the final o t ;

[0042] A two - layer RNN neural network is adopted. The first - layer RNN has 128 neurons, and the second - layer RNN has 64 neurons. The activation functions of both the first and second layers are ReLU functions. The two RNN layers are used to receive and process the output results of the multi - head attention mechanism layer, learn the relationships between multi - step time - series features, and output them to the long short - term memory network LSTM layer;

[0043] (2) The core of LSTM is a unit in which each neuron contains a forget gate, an input gate, and an output gate; each gate is a special structure used to control the flow of information and transfer the information from the previous moment to the current moment, so as to capture the time dependence in the data. The LSTM neural network includes:

[0044] Forget Gate:

[0045] The role of the forget gate is to determine what information to discard from the cell state; it is calculated by the following formula:

[0046] f t = σ(W f × [h t-1 , x t + b f );

[0047] Among them, f t represents the output of the forget gate at time t, σ is the sigmoid function, W f and bf are the weights and biases of the forget gate, h t-1 is the previous hidden state, x t is the current input;

[0048] Input Gate:

[0049] The input gate is used to update the cell state and consists of two parts: a sigmoid layer and a tanh layer;

[0050] i t = σ(W i × [h t-1 , x t + b i );

[0051]

[0052] where i t is the input of the input gate, is the candidate value vector;

[0053] Output Gate:

[0054] The output gate is used to determine the next hidden state, which contains information about the previous input and is used for prediction,

[0055] o t = σ(W o × [h t-1 , x t + b o );

[0056] h t = o t *tanh(C t );

[0057] where o t is the output of the output gate, h t is the current hidden state, C t is the current cell state;

[0058] The LSTM layer uses three layers in total. The number of neurons in the first layer of LSTM is 64, the number of neurons in the second layer of LSTM is 32, and the number of neurons in the third layer of LSTM is 32. The activation functions of the first, second, and third layers are all Tanh functions;

[0059] (3) Fully connected layer:

[0060] In the fully connected layer Dense, all input nodes and output nodes are fully connected; Three fully connected layers are used as the final fully connected layer of the RNN-LSTM recurrent neural network layer;

[0061] The number of neurons in the first fully connected layer is 32, the number of neurons in the second fully connected layer is 16, and the number of neurons in the last fully connected layer is 8. The activation function for all three fully connected layers is the ReLU function;

[0062] f. Feed-forward layer:

[0063] The feed-forward layer is used to provide more learning ability. The feed-forward network is a two-layer network. The activation function of the first layer is the ReLU function. For the vector x at each position in the input sequence, its expression is:

[0064] FFN(x) = max(0, xW1 + b1)W2 + b2.

[0065] g. Output layer:

[0066] The number of neurons in the output layer is the same as the feature dimension, and no activation function is set, directly generating the predicted value.

[0067] Optionally, step 3 includes:

[0068] (31) Initialize the model: Integrate the RMM-LSTM recurrent neural network framework and add a multi-head attention mechanism to enhance the model's parallel learning ability and response to key features;

[0069] (32) Optimizer: Select the Adam optimizer, which combines momentum and adaptive learning rate characteristics to help optimize large-scale datasets quickly and stably;

[0070] (33) Number of training epochs Epochs: Monitor the model performance through the loss of the validation dataset and stop training when the loss of the validation dataset no longer decreases or increases for several consecutive rounds; Select 200 training epochs for the parameter;

[0071] (34) Batch size Batch_Size: This parameter determines the amount of data used for each gradient update, affecting the training stability and efficiency; Select an appropriate batch size according to the hardware performance and task requirements to balance the training stability and computational efficiency; Set it to Batch_Size = 64;

[0072] (35) Evaluation metrics:

[0073] ① Accuracy

[0074] Accuracy is used to measure whether the predicted value is within the error threshold range. Its principle is as follows, providing an intuitive feedback on the comprehensive performance of the model on the training dataset, validation dataset, and test dataset;

[0075]

[0076] Among them, y i is the true value of the i-th sample, i.e., the target value, which is the actual observed value of the time series; is the predicted value of the i-th sample, i.e., the predicted output of the model; ε is the error tolerance or threshold, which sets an error range, indicating that when the gap between the predicted value and the true value is less than this threshold, the prediction is considered correct; the error tolerance is 5% or 10%; n is the total number of samples;

[0077] ② Loss rate

[0078] Select the custom binary cross-entropy loss function as the basic loss function of the recurrent neural network fusion model to judge the prediction performance of the recurrent neural network fusion model;

[0079] The formula of the binary cross-entropy loss function is:

[0080]

[0081] Among them, N is the total number of samples; y i is the true label of the i-th sample, taking values of 0 or 1; p i is the probability that the i-th sample predicted by the model belongs to class 1, and its value range is between (0, 1);

[0082] The BCE loss function adjusts the parameters of the model during the training process to make the predicted probability p i closer to the true label y i ; it encourages the model to output probability values close to the true class, thereby improving the classification accuracy;

[0083] To prevent the model from overfitting the training data, L2 regularization is introduced on the basis of the BCE loss function; L2 regularization is the weight decay, and its core idea is to punish the parameters of the model, forcing the model to maintain small weight values, thereby preventing the model from overfitting the training data;

[0084] The formula of L2 regularization is:

[0085]

[0086] Among them, λ is the hyperparameter of the regularization strength, which controls the weight of the regularization term. The larger λ is, the stronger the regularization effect; w j is the parameter of the model, i.e., the weight; n is the total number of model parameters; the L2 regularization term is the sum of the squares of all weights, and the larger the weight value, the greater the penalty;

[0087] Combine L2 regularization and the binary cross-entropy loss function BCE to obtain a new loss function, and its formula is:

[0088]

[0089] The binary cross-entropy loss function BCE calculates the prediction error of the model, while the L2 regularization term helps the model avoid overfitting by penalizing large weights;

[0090] ③R 2 value

[0091] R 2 The value measures the proportion of the variance of the target variable explained by the model, and its value range is [0,1]. The principle is shown in the following formula; R 2 reflects the goodness of fit of the model. The closer the value is to 1, the stronger the model's ability to explain the target variable;

[0092]

[0093] where y i is the true value of the i-th sample, i.e., the target value; is the predicted value of the i-th sample, i.e., the predicted output of the model; is the average value of the true values, i.e., is the residual sum of squares RSS, i.e., the sum of the squares of the differences between the true values and the predicted values; is the total sum of squares TSS, i.e., the sum of the squares of the differences between the true values and the mean of the true values;

[0094] (36) Training strategy:

[0095] To improve the multi-step prediction performance of the non-linear autoregressive neural network NARX-NN, a feedback retraining of predictions FR strategy is adopted; the feedback retraining of predictions strategy is aimed at the inconsistency between the input in the training stage and the multi-step prediction stage of the NARX-NN model to reduce the difference, thereby improving the multi-step prediction performance of the model;

[0096] The basic idea of FR is to reconstruct the training samples using the single-step prediction results with errors and retrain the model to enhance the model's robust performance against prediction errors. The specific steps are as follows:

[0097] ① First, use the initial dataset and complete the initial training using the conventional training strategy, i.e., train the network using the backpropagation algorithm. At this time, the input target variables are all actual observed values;

[0098] ② Use the trained network for single-step prediction to obtain the prediction results; single-step prediction ensures that the error of the prediction results is within a reasonable range. If multi-step prediction results are used, the reconstructed samples in the following step ③ will deviate from the dynamic characteristics of the system to be modeled due to excessive prediction errors;

[0099] ③ Replace the measured target value in the model input with its corresponding single-step prediction result to obtain a new training sample. Multiple replacements are required to reconstruct the dataset. The collective operation is as follows: Given the real data of the first 63 time steps, predict the data of the next 1 time step, replace the predicted data with the real data in the prediction time step, and complete one operation of reconstructing the dataset. The data of 64 time steps is a cycle. Repeat the operation until the data in all cycles of the entire dataset completes the replacement of the real value and the predicted value.

[0100] ④ Retrain the network using the reconstructed samples. After training, check whether the validation set loss is greater than the recorded minimum validation set loss for several consecutive rounds of training. If so, it indicates that the performance of the model on the validation set no longer improves, end the training, and obtain the multi-step prediction model. Otherwise, if the validation set loss decreases in the current round, update the minimum validation set loss, and enter steps ② and ③ to repeat single-step prediction and training the network.

[0101] In a second aspect, an embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the data prefetching method based on the multi-head attention mechanism and the RNN-LSTM network in the first aspect or any possible implementation manner of the first aspect.

[0102] In a third aspect, an embodiment of the present invention provides an electronic device, including: one or more processors; a memory; and one or more computer programs. The one or more computer programs are stored in the memory. The one or more computer programs include instructions. When the instructions are executed by the device, the device is caused to execute the data prefetching method based on the multi-head attention mechanism and the RNN-LSTM network in the first aspect or any possible implementation manner of the first aspect.

[0103] In the technical solution provided by the present invention, the method includes obtaining an initial dataset and performing proportional division to obtain a training dataset, a validation dataset, and a test dataset; preprocessing the initial dataset to obtain a preprocessed dataset; constructing a recurrent neural network fusion model through a multi-head attention mechanism and an RNN-LSTM recurrent neural network; training, validating, and testing the recurrent neural network fusion model through the preprocessed training dataset, validation dataset, and test dataset respectively to achieve data prefetching. This method improves the accuracy and efficiency of data prefetching and reduces the latency of data reading by using the powerful foundation of RNN-LSTM in processing long time series and a specific multi-head attention mechanism through a fusion technology. Description of the Drawings

[0104] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0105] Figure 1 Flowchart of the data prefetching method provided by the embodiments of the present invention;

[0106] Figure 2 Flowchart of obtaining the initial data set provided by the embodiments of the present invention;

[0107] Figure 3 Flowchart of data prediction mode processing provided by the embodiments of the present invention;

[0108] Figure 4 Schematic diagram of the recurrent neural network fusion model provided by the embodiments of the present invention;

[0109] Figure 5 Schematic diagram of the multi-head attention mechanism layer structure provided by the embodiments of the present invention;

[0110] Figure 6 Schematic diagram of the RNN-LSTM recurrent neural network provided by the embodiments of the present invention;

[0111] Figure 7 Schematic diagram of the basic RNN structure provided by the embodiments of the present invention;

[0112] Figure 8 Schematic diagram of the LSTM neural network structure provided by the embodiments of the present invention;

[0113] Figure 9 Flowchart of the training strategy provided by the embodiments of the present invention;

[0114] Figure 10 Flowchart of another training strategy provided by the embodiments of the present invention;

[0115] Figure 11 Schematic diagram of an electronic device provided by the embodiments of the present invention. Detailed implementation manners

[0116] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0117] It should be clear that the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0118] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0119] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0120] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0121] Figure 1 For the flowchart of the data prefetching method provided by the embodiments of the present invention, as Figure 1 shown, the method includes:

[0122] Step 1: Obtain an initial data set, and perform proportional division to obtain a training data set, a validation data set, and a test data set; preprocess the initial data set to obtain a preprocessed data set.

[0123] Use the buffer data in the flash chip of the Master node of the Shandong Computing Network unified storage platform as the data set; each data sequence contains a unique address, the unique page it is on, the unique offset of the upper and lower files, the time of reading the file, the network occupancy rate of reading the file, etc. However, some data has little relationship with data prefetching and may even affect the accuracy of data prefetching. Therefore, use the unique page and offset, as well as the network occupancy rate, as the data in the data set, and re-edit the time stamp.

[0124] In the embodiments of the present invention, as Figure 2 shown, the obtaining of the initial data set in the said Step 1 includes:

[0125] First, it is necessary to read the buffer of the flash chip of the Master node to generate timestamps. The timestamps start from the time record of the unique address of the first file read from the buffer and are generated every 5 minutes. Secondly, statistics are performed every 5 minutes to record all the unique pages in the current memory; record the memory offsets between all unique pages; record the network occupancy rate for reading each unique page; finally, data integration is performed to generate a dataset in the form of timestamps (timescamp), offsets (target), broadband occupancy rate (dynamic_feat), and unique page mapping IDs (id), as shown in Table 1.

[0126] Table 1 Dataset form

[0127] Timestamp (timescamp) Broadband occupancy rate (dynamic_feat) Unique page mapping ID (id) Offset (target) 2024-01-0100:00:00 0.01 0 0 2024-01-01 00:05:00 0.17 8 512 2024-01-01 00:10:00 0.13 0 512 2024-01-01 00:15:00 0.34 15 960 。

[0128] The ratio division includes: collecting 10,000 pieces of data, removing the data with large offsets, and 8,000 pieces of valid data remain; and using custom code to set the training dataset, validation dataset, and test dataset in the ratio of 6:2:2 for the entire dataset; the dataset Data.csv file, and its content format is shown in Table 2.

[0129] Table 2 Dataset Data.csv file

[0130] Timestamp (timescamp) Broadband occupancy rate (dynamic_feat) Unique page mapping ID (id) Offset (target) 2024-01-0100:00:00 0.01 0 0 2024-01-0100:05:00 0.17 8 512 2024-01-01(00:10:00 0.13 0 512 2024-01-01000:15:00 0.34 15 960 2024-01-01 00:20:00 0.17 4 704 。

[0131] Only one csv file is needed as the dataset to implement the training, validation, and testing of the model, reducing the time for data preprocessing. And such a data format can accurately reflect the changes in unique pages, page offsets, and network occupancy rates in the cache.

[0132] The preprocessing of the initial dataset to obtain the preprocessed dataset includes:

[0133] Data normalization: Normalization helps to speed up the model training speed, improve the model convergence speed, and reduce the impact of the model initialization method on the training process; the normalization operation adopted is implemented by the MinMaxScaler function in Sklearn. MinMaxScaler scales all its values to a specified range, usually [0,1], and its formula is:

[0134]

[0135] Among them, X is the eigenvalue of the original data; X min is the minimum value of the original data eigenvalue X in the data; X maxis the maximum value of the original data eigenvalue X in the data; X scaled is the final result of normalization;

[0136] Perform data normalization operations on the columns of broadband occupancy rate (dynamic_feat), unique page mapping ID (id), and offset (target) in the dataset; this standardized data is easier and faster for the neural network model to learn because it ensures that the network input features are at the same magnitude, which helps improve the efficiency and stability of gradient descent during the optimization process.

[0137] Data prediction mode processing: As Figure 3 shown, define a mode for data input and prediction, which is implemented using the custom function create_fixed_sequnences(data, time_data, sequence_length). The specific operation steps are to first input sequence_length as the input time series step length of the model, and then the model returns a prediction result, that is, a predicted time step length, and returns the corresponding timestamp; among them, the adopted sequence_length is set to 64, representing an input time series of 64 steps for the model to learn, so as to predict a time series of one step.

[0138] Step 2: Construct a recurrent neural network fusion model through the multi-head attention mechanism and the RNN-LSTM recurrent neural network.

[0139] In the embodiment of the present invention, the recurrent neural network fusion model adopts a decoder structure with a self-attention mechanism. The self-attention mechanism is adopted because it can obtain the global and local connections in one step and solves the problem that models based on recurrent neural networks or LSTM cannot perform parallel computing. The decoder structure comes from the encoder-decoder model. Predicting the next offset is not a generative task but a classification task, so the model does not require the entire encoder-decoder model. The decoder structure is selected for the model because it is considered that the offset sequence has a temporal order and each moment only has a connection with the previous sequence, so the model uses a masked decoder-only structure to mask future information.

[0140] In the embodiment of the present invention, as Figure 4 shown, the recurrent neural network fusion model in step 2 includes:

[0141] a. Input layer: The dimension of the input layer is (sequence_length, feature_dim); where sequence_length represents the number of time steps; feature_dim represents the feature dimension of each time step; the input layer inputs the original time series data into the model for feature extraction;

[0142] b. Self-attention mechanism layer: In the attention mechanism, each input element is assigned a weight value according to its relationship with other elements (e.g., the relationship between words), and the weight value is calculated by dot product. The calculation of the attention mechanism is as follows:

[0143]

[0144] where Q is the query vector, representing the target to be focused on currently; K is the key vector, representing the features in all inputs; V is the value vector, that is, the actual information associated with the key vector; d k is the dimension of the key vector for scaling; the essence of the above calculation process is to measure the correlation between the query and the key through dot product, and then weighted sum the corresponding value vector V according to these correlation weights.

[0145] The multi-head attention mechanism enhances the model's ability to capture different information by dividing the input query, key, and value into multiple heads (i.e., multiple subspaces) for parallel calculation. Its core idea is to parallelize the calculation of multiple attention mechanisms, each head focusing on different aspects of the input data, and finally aggregating this information together. As Figure 5 shown, the structure of the multi-head attention mechanism layer includes:

[0146] Ⅰ. Input: Input the preprocessed data set into the model. In the present invention, the data actually needed to be predicted is three columns (dynamic_feat, id, target), so the input shape into the model is:

[0147] Input shape = (batch_size, sequence_length, 3);

[0148] where batch_size is the number of samples put in for each training, and its value is set to 128, sequence_length is the length of the time series, which has been described in the above data processing module and the set parameter is 64, and the last parameter is the data feature dimension of each time step, corresponding to the three columns of data to be predicted in the present invention. Therefore, the size in the input model is a matrix of (128, 64, 3).

[0149] II. The multi-head attention layer performs a linear transformation on the data in the matrix internally, and this process is carried out separately in each head. As shown in Figure 5 , a total of 5 head attention mechanism layers are used. The linear transformation process will obtain three vector matrices: query Qi, key Ki, and value Vi. The linear transformation formula is as follows:

[0150] Query i = X * W i Q ;

[0151] Key i = X * W i K ;

[0152] Value i = X * W i V ;

[0153] Among them, X represents the sequence in the input dataset during the training process; W i Q is the query transformation matrix of the i-th head; W i K is the key transformation matrix of the i-th head; W i V is the value transformation matrix of the i-th head, and these weight matrices are learned parameters and are usually optimized during training.

[0154] Each head independently calculates its corresponding query, key, and value matrices and calculates their respective attentions. On this basis, a mask matrix is added so that the neural network model cannot see the information after the current moment. That is, for an offset sequence, at time step t, the output of the model should only depend on the output before this offset and cannot see the information after that, so as to ensure the generated causal relationship and increase the flexibility of the model in processing sequence data. The mask formula is as follows:

[0155] Given the sequence length T, the future information mask M is an upper triangular matrix:

[0156]

[0157] Therefore, the output head i of each head is the result of the weighted sum of the attention weights calculated based on this head and the corresponding value matrix V i .

[0158]

[0159] head i = AttentionWeight * Vi ;

[0160] Ⅲ. Put the results head of each attention layer i into the merging layer for concatenated output. The concatenation formula is as follows:

[0161] Concat(head1, head2,..., head5) = [head1, head2,..., head5];

[0162] The shape of the output after concatenation is:

[0163]

[0164] batch_size is the number of samples in each batch, and its parameter is set to 64. seq_len is the length of the time series, and its parameter is set to 64. d head is the dimension of each attention layer, and its parameter is 64. h is the number of attention heads, which is 5.

[0165] Ⅳ. To map the output after concatenation back to the model dimension d model a linear transformation is required through a weight matrix W o to map to the model dimension. The formula is as follows:

[0166] Output = Concat × W o ;

[0167] W o is a linear weight matrix used to integrate the concatenated results of multi-head attention. Its shape is (h * d head , d model ), and it is essentially a linear transformation. Through the linear transformation, this high-dimensional tensor is integrated into the total model dimension d model , further compressing the information so that the output of the attention heads can be processed by the subsequent model structure (RNN-LSTM).

[0168] c. Residual block: The residual block is used to solve the problem of model degradation that occurs when the number of network layers increases. The residual block is introduced because when network degradation occurs, the performance of the shallow network is better than that of the deep network. Therefore, the residual block directly passes the calculation results of the low-level network to the high-level network, making the performance of the deep network better than that of the shallow network. The formula for the residual block is:

[0169] X l+1 = X l + A(X l + W l );

[0170] Among them, X l is the input of the l-th layer, A(Xl +W l ) is the output of the masked multi-head self-attention mechanism module of the l-th layer;

[0171] d. Normalization layer: Layer Normalization (LN) is selected to process the data; layer normalization normalizes all neurons in a layer; assuming that the i-th input of the neurons in the l-th layer is z i (l) , and their expressions for the mean and variance are respectively:

[0172]

[0173] where n (l) is the number of neurons in the l-th layer;

[0174] Layer normalization is defined as:

[0175]

[0176] where γ represents the scaling parameter vector, β represents the translation parameter vector, and the dimensions of the parameter vectors γ and β are both the same as those of z (l) ;

[0177] e. RNN-LSTM recurrent neural network layer:

[0178] In the embodiments of the present invention, as Figure 6 shown, the RNN-LSTM recurrent neural network layer is composed of two layers of recurrent neural network RNN, three layers of long short-term memory network LSTM, and three layers of fully connected layers;

[0179] (1) The basic unit of the recurrent neural network (RNN) is a recurrent unit Recurrent Unit, which receives an input and a hidden state from the previous time step and outputs the hidden state of the current time step. As Figure 7 shown, the basic RNN structure consists of an input layer, a hidden layer, and an output layer. Figure 7 In it, x is the input vector, o is the output vector, and s represents the value of the hidden layer; U is the weight matrix from the input layer to the hidden layer, and V is the weight matrix from the hidden layer to the output layer. The value s of the hidden layer of the recurrent neural network depends not only on the current input x but also on the value s of the previous hidden layer. The weight matrix W is the weight of the value of the previous hidden layer as the input of this time; Unfolding the basic RNN structure in the time dimension gives the entire network structure framework of the RNN;

[0180] After the RNN network receives the input x t at time t, the value of the hidden layer is s t , and the formula is as follows:

[0181] s t = f(U × x t + W × s t-1 );

[0182] U is the weight matrix of the input x, and W is the value of the previous hidden layer s t-1 as the current input weight matrix, f is the activation function, and s t 's value depends on x t and s t-1 ;

[0183] The output value is o t , and the calculation formula is as follows:

[0184] o t = g(V × s t );

[0185] V is the weight matrix of the output layer, and g is the activation function;

[0186] The hidden layer has two inputs. The first is the product of the weight matrix U from the input layer to the hidden layer and the input x t ; the second is the product of the value of the previous hidden layer s t-1 and W, that is, the s calculated at the previous moment t-1 needs to be cached and calculated with the input x t to jointly output the final o t ;

[0187] A two-layer RNN neural network is adopted. The first layer of RNN has 128 neurons, and the second layer of RNN has 64 neurons; the activation functions of the first layer and the second layer are both ReLU functions; the two RNN layers are used to receive and process the output results of the multi-head attention mechanism layer, learn the relationships between multi-step time series features, and output them to the long short-term memory network LSTM layer;

[0188] In the data prefetch task, the prefetch operation usually has long-term dependencies, that is, information from multiple time steps ago is required for accurate prediction. When the RNN model processes long time series data, the problems of gradient disappearance and gradient explosion are particularly serious. With gradient disappearance, the model will have difficulty effectively learning long-distance dependencies, thus reducing the accuracy of data prefetching.

[0189] Therefore, the present invention introduces an LSTM neural network on the basis of RNN. The long short-term memory network (LSTM) is a special recurrent neural network structure, which is particularly suitable for processing sequence data and time series tasks.

[0190] (2) The core of LSTM is a unit where each neuron contains a forget gate, an input gate, and an output gate; each gate is a special structure used to control the flow of information, passing the information from previous moments to the current moment, thus capturing the temporal dependencies in the data; as Figure 8 shown, the LSTM neural network includes:

[0191] Forget Gate:

[0192] The role of the forget gate is to determine what information to discard from the cell state; it is calculated by the following formula:

[0193] f t = σ(W f × [h t-1 , x t + b f );

[0194] where f t represents the output of the forget gate at time t, σ is the sigmoid function, W f and b f are the weights and biases of the forget gate, h t-1 is the previous hidden state, and x t is the current input;

[0195] Input Gate:

[0196] The input gate is used to update the cell state and consists of two parts: a sigmoid layer and a tanh layer; the sigmoid layer determines which values will be updated; the tanh layer creates a new candidate value vector, and the new candidate value vector will be added to the state.

[0197] i t = σ(W i × [h t-1 , x t + b i );

[0198]

[0199] where i t is the input of the input gate, is the candidate value vector;

[0200] Output Gate:

[0201] The output gate is used to determine the next hidden state, and the hidden state contains information about the previous input and is used for prediction,

[0202] o t = σ(W o×[h t-1 ,x t +b o );

[0203] h t =o t *tanh(C t );

[0204] Among them, o t is the output of the output gate, h t is the current hidden state, and C t is the current cell state;

[0205] The LSTM layer adopts three layers in total. The number of neurons in the first - layer LSTM is 64, the number of neurons in the second - layer LSTM is 32, and the number of neurons in the third - layer LSTM is 32. The activation functions of the first, second, and third layers are all Tanh functions;

[0206] The working mechanism of LSTM involves several concepts in advanced mathematics:

[0207] Sigmoid function (σ): An activation function widely used in neural networks. It compresses any input in the range (-inf, inf) to a value in the interval (0, 1), which is very suitable for use in gating structures.

[0208]

[0209] Hyperbolic tangent function (tanh): This is another activation function that can map any value to between - 1 and 1. In LSTM, the tanh function helps to regulate the flow of information and maintain the stability of gradients.

[0210]

[0211] Dot - product operation: In LSTM, the dot - product operation (denoted by *) is used in the gating structure. This element - by - element multiplication operation ensures that information can flow only when the gate is open.

[0212] For data pre - fetching, since some data has a larger offset above, and both recurrent neural networks and context offset prediction are relatively complex concepts and technologies, combining them may lead to a more complex model and algorithm, increasing the difficulty of implementation and understanding. Especially when dealing with long - term dependencies, it may make training more difficult, require more computing resources and more complex optimization strategies, and also increase the overhead cost of the model.

[0213] When constructing a neural network pre - fetching model, LSTM solves two major problems:

[0214] 1. Long-term Dependence Problem: During the data offset prediction process, through the memory unit and gating mechanism, LSTM can capture long-term dependence relationships and is more suitable for processing long sequence data than traditional neural networks.

[0215] 2. Anti-gradient Vanishing and Gradient Explosion Problems: By controlling the flow of information through the gating mechanism, LSTM can effectively mitigate the effects of gradient vanishing and gradient explosion.

[0216] (3) Fully Connected Layer:

[0217] In the fully connected layer Dense, all input nodes and output nodes are fully connected; three fully connected layers are used as the final fully connected layer of the RNN-LSTM recurrent neural network layer;

[0218] The number of neurons in the first fully connected layer is 32, the number of neurons in the second fully connected layer is 16, and the number of neurons in the last fully connected layer is 8. The activation function of the three fully connected layers is the ReLU function;

[0219] f. Feedforward Layer:

[0220] The feedforward layer is used to provide more learning ability. The feedforward network is a two-layer network. The activation function of the first layer is the ReLU function. For the vector x at each position in the input sequence, its expression is:

[0221] FFN(x) = max(0, xW1 + b1)W2 + b2.

[0222] The introduction of the feedforward network can well re-learn the data output by the RNN-LSTM layer. Through the final normalization, a more accurate multi-step prediction sequence can be output.

[0223] g. Output Layer:

[0224] The number of neurons in the output layer is the same as the feature dimension, and no activation function is set, directly generating the predicted value.

[0225] Step 3: Use the preprocessed training data set, validation data set, and test data set to train, validate, and test the recurrent neural network fusion model respectively to achieve data prefetching.

[0226] In the embodiment of the present invention, the step 3 includes:

[0227] (31) Initialize the model: Integrate the RMM-LSTM recurrent neural network framework and add a multi-head attention mechanism to enhance the parallel learning ability of the model and the response to key features;

[0228] (32) Optimizer: The Adam optimizer is selected, which combines momentum and adaptive learning rate features to help optimize large-scale datasets quickly and stably;

[0229] (33) Number of training epochs: Monitor the model performance through the loss of the validation dataset and stop training when the loss of the validation dataset no longer decreases or increases for several consecutive rounds; The parameter is set to 200 training epochs;

[0230] (34) Batch size: This parameter determines the amount of data used for each gradient update and affects the training stability and efficiency; Select an appropriate batch size according to the hardware performance and task requirements to balance the training stability and computational efficiency; Set it to Batch_Size = 64;

[0231] (35) Evaluation metrics:

[0232] ① Accuracy

[0233] Accuracy is used to measure whether the predicted value is within the error threshold range. The principle is as follows, providing an intuitive feedback on the comprehensive performance of the model on the training dataset, validation dataset, and test dataset;

[0234]

[0235] Among them, y i is the true value of the i-th sample, i.e., the target value, which is the actual observed value of the time series; is the predicted value of the i-th sample, i.e., the predicted output of the model; ε is the error tolerance or threshold, setting an error range, indicating that when the gap between the predicted value and the true value is less than this threshold, the prediction is considered correct; The error tolerance is 5% or 10%; n is the total number of samples;

[0236] ② Loss rate

[0237] Select the custom binary cross-entropy loss function as the basic loss function for the recurrent neural network fusion model to judge the prediction performance of the recurrent neural network fusion model;

[0238] The formula for the binary cross-entropy loss function is:

[0239]

[0240] Among them, N is the total number of samples; y i is the true label of the i-th sample, taking values of 0 or 1; p i is the probability that the i-th sample predicted by the model belongs to class 1, with a value range between (0, 1);

[0241] The BCE loss function adjusts the model's parameters during training to make the predicted probability p i closer to the true label y i ; it encourages the model to output probability values close to the true class, thereby improving the classification accuracy;

[0242] To prevent the model from overfitting the training data, L2 regularization is introduced on the basis of the BCE loss function; L2 regularization is weight decay, and its core idea is to punish the model's parameters, forcing the model to maintain small weight values, thereby preventing the model from overfitting the training data;

[0243] The L2 regularization formula is:

[0244]

[0245] where λ is the hyperparameter of the regularization strength, controlling the weight of the regularization term. The larger λ is, the stronger the regularization effect; w j are the model's parameters, i.e., weights; n is the total number of model parameters; the L2 regularization term is the sum of the squares of all weights, and the larger the weight value, the greater the penalty;

[0246] Combining L2 regularization and the binary cross-entropy loss function BCE gives a new loss function, and its formula is:

[0247]

[0248] The binary cross-entropy loss function BCE calculates the prediction error of the model, while the L2 regularization term helps the model avoid overfitting by punishing large weights;

[0249] ③R 2 value

[0250] R 2 value measures the proportion of the variance of the target variable explained by the model, and its value range is [0, 1]. Its principle is shown in the following formula; R 2 reflects the goodness of fit of the model. The closer the value is to 1, the stronger the model's ability to explain the target variable;

[0251]

[0252] where, y i is the true value of the i-th sample, i.e., the target value; is the predicted value of the i-th sample, i.e., the predicted output of the model; is the average of the true values, i.e., is the residual sum of squares RSS, i.e., the sum of the squares of the differences between the true values and the predicted values; It is the total sum of squares TSS, that is, the sum of the squares of the differences between the true values and the mean of the true values;

[0253] (36) Training strategy:

[0254] To improve the multi-step prediction performance of the non-linear autoregressive neural network NARX-NN, a feedback retraining FR strategy for predicted values is adopted; the feedback retraining strategy for predicted values aims at the inconsistency between the input in the training stage and the multi-step prediction stage of the NARX-NN model to reduce the difference, thereby improving the multi-step prediction performance of the model;

[0255] The basic idea of FR is to reconstruct the training samples using the single-step prediction results with errors and retrain the model to enhance the model's robust performance against prediction errors. The specific steps are as follows:

[0256] ① First, use the initial dataset and complete the initial training using the conventional training strategy, that is, train the network using the backpropagation algorithm. At this time, all the target variables input are actual observed values;

[0257] ② Use the trained network for single-step prediction to obtain the prediction results; single-step prediction ensures that the error of the prediction results is within a reasonable range. If multi-step prediction results are used, the reconstructed samples in the following step ③ will deviate from the dynamic characteristics of the system to be modeled due to excessive prediction errors;

[0258] ③ Replace the measured target values in the model input with their corresponding single-step prediction results to obtain new training samples. Multiple replacements are required to reconstruct the dataset; the collective operation is as follows: Given the true data of the first 63 time steps, predict the data of the next 1 time step, replace the predicted data with the true data in the prediction time step, and complete one operation of reconstructing the dataset. The data of 64 time steps is a cycle. Repeat the operation until the data in all cycles of the entire dataset complete the replacement of true values and predicted values;

[0259] ④ Use the reconstructed samples to retrain the network again; after the training is completed, check whether the validation set loss is greater than the recorded minimum validation set loss for several consecutive rounds of training; if so, it indicates that the performance of the model on the validation set no longer improves, end the training, and obtain the multi-step prediction model; otherwise, if the validation set loss decreases in the current round, update the minimum validation set loss, and enter steps ② and ③, repeat single-step prediction and train the network.

[0260] In data-intensive applications such as high-performance computing (HPC), artificial intelligence, and big data processing, data access latency is a key bottleneck for performance optimization. Regarding traditional data prefetching strategies, there are problems such as insufficient prefetching accuracy and occupying memory space, which can lead to the data access speed not meeting expectations. To solve the above problems, the technical solution provided by the present invention not only effectively utilizes the ability of the RNN-LSTM model to effectively capture dependencies in long time series and has strong modeling capabilities for long-term trends and periodic patterns in time series data, and LSTM effectively solves the problem of gradient disappearance by introducing a gating mechanism. It also further captures the connections of features such as time correlation, user behavior, and application characteristics through the fusion with the multi-head attention mechanism. Finally, through the designed feedback value retraining strategy (FR), the multi-step prediction effect of the model is further improved, and the accuracy and efficiency of model prefetching are further enhanced. The proposed fusion technology aims to utilize the powerful foundation of RNN-LSTM in processing long time series and a specific multi-head attention mechanism to improve the accuracy and efficiency of data prefetching and reduce the latency of data reading.

[0261] In the technical solution provided by the present invention, the method includes obtaining an initial data set and performing proportional division to obtain a training data set, a validation data set, and a test data set; preprocessing the initial data set to obtain a preprocessed data set; constructing a recurrent neural network fusion model through a multi-head attention mechanism and an RNN-LSTM recurrent neural network; training, validating, and testing the recurrent neural network fusion model through the preprocessed training data set, validation data set, and test data set respectively to achieve data prefetching. This method improves the accuracy and efficiency of data prefetching and reduces the latency of data reading by using the powerful foundation of RNN-LSTM in processing long time series and a specific multi-head attention mechanism through the fusion technology.

[0262] Each step of the embodiments of the present invention can be executed by an electronic device. Among them, the electronic device includes but is not limited to mobile phones, tablet computers, portable PCs, desktop computers, etc.

[0263] The embodiments of the present invention provide a computer-readable storage medium. The computer-readable storage medium includes a stored program. Among them, when the program runs, it controls the electronic device where the computer-readable storage medium is located to execute the embodiments of the above data prefetching method based on the multi-head attention mechanism and the RNN-LSTM network.

[0264] Figure 11 It is a schematic diagram of an electronic device provided by the embodiments of the present invention, as Figure 11As shown, the electronic device 21 includes: a processor 211, a memory 212, and a computer program 213 stored in the memory 212 and executable on the processor 211. When the computer program 213 is executed by the processor 211, it implements the data prefetching method based on the multi-head attention mechanism and the RNN-LSTM network in the embodiments. To avoid repetition, details are not elaborated here.

[0265] The electronic device 21 includes, but is not limited to, a processor 211 and a memory 212. Those skilled in the art can understand that Figure 11 merely examples of the electronic device 21, which do not constitute a limitation on the electronic device 21. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0266] The so-called processor 211 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0267] The memory 212 may be an internal storage unit of the electronic device 21, such as the hard disk or memory of the electronic device 21. The memory 212 may also be an external storage device of the electronic device 21, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 21. Further, the memory 212 may also include both the internal storage unit and the external storage device of the electronic device 21. The memory 212 is used to store computer programs and other programs and data required by the network device. The memory 212 may also be used to temporarily store data that has been output or will be output.

[0268] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.

[0269] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A data pre-fetching method based on multi-head attention mechanism and RNN-LSTM network, characterized in that: The method comprises: Step 1: Obtain an initial data set and divide it into proportions to obtain a training data set, a verification data set, and a test data set; preprocess the initial data set to obtain a preprocessed data set; Step 2: Construct a recurrent neural network fusion model through the multi-head attention mechanism and RNN-LSTM recurrent neural network; Step 3: Use the preprocessed training data set, verification data set and test data set to train, verify and test the recurrent neural network fusion model respectively to achieve data pre-fetching.

2. The method according to claim 1, characterized in that The step 1 of obtaining the initial data set includes: First, you need to read the cache area of ​​the flash memory chip of the Master node and generate a timestamp. The timestamp starts from the unique address time record of the first file read from the cache area and is generated every 5 minutes. Secondly, perform statistics every 5 minutes to record all the unique pages in the current memory; record the memory offset between all unique pages; record the network occupancy rate of each unique page; finally, perform data integration to generate a data set in the form of timestamp timescamp, offset target, broadband occupancy rate dynamic_feat, and unique page mapping ID id. The ratio division includes: collecting 10,000 data, removing data with large offset, and leaving 8,000 valid data; and using custom code to set the training data set, the verification data set, and the test data set in a ratio of 6:2:2 for the entire data set; The preprocessing of the initial data set to obtain the preprocessed data set includes: Data normalization: Normalization helps speed up model training, improve model convergence, and reduce the impact of model initialization on the training process. The normalization operation used is implemented by the MinMaxScaler function in Sklearn. MinMaxScaler scales all its values ​​to a specified range. The formula is: Among them, X is the eigenvalue of the original data; X min is the minimum value of the original data eigenvalue X in the data; X max is the maximum value of the original data eigenvalue X in the data; X scaled is the final result of normalization; Normalize the data of the broadband occupancy rate dynamic_feat, unique page mapping ID, and offset target columns in the data set; Data prediction mode processing: define a data input and prediction mode, which is implemented using the custom function create_fixed_sequnences. The specific operation steps are to first input sequence_length as the input time series step of the model, and then the model returns a prediction result, that is, a prediction time step is obtained, and the corresponding timestamp is returned; among them, the sequence_length used is set to 64, which means that a time series of 64 steps is input, allowing the model to learn and thus predict a time series of one step.

3. The method according to claim 1, characterized in that: The recurrent neural network fusion model in step 2 includes: a. Input layer: The dimension of the input layer is (sequence_length, feature_dim); sequence_length represents the number of time steps; feature_dim represents the feature dimension of each time step; the input layer inputs the original time series data into the model for feature extraction; b. Self-attention mechanism layer: In the attention mechanism, each input element is assigned a weight value according to its relationship with other elements. The weight value is calculated by dot product. The calculation formula of the attention mechanism is: Among them, Q is the query vector query, which represents the current target to be focused on; K is the key vector key, which represents the features in all inputs; V is the value vector value, which is the actual information associated with the key vector; d k is the dimension of the key vector, used for scaling; c. Residual block: The residual block is used to solve the problem of model degradation when the number of network layers increases. The residual block is introduced because when network degradation occurs, the effect of the shallow network is better than that of the deep network. Therefore, the residual block directly passes the calculation results of the low-level network to the high-level network, making the effect of the deep network better than the shallow network. The formula of the residual block is: X l+1 =X l +A(X l +W l ); Among them, X l is the input of the lth layer, A(X l +W l ) is the output of the masked multi-head self-attention mechanism module of the lth layer; d. Normalization layer: Use layer normalization to process the data; layer normalization is to normalize all neurons in a layer; assuming that the i-th input of the neuron in the l-th layer is z i (l) , the expressions of its mean and variance are: Among them, n (l) is the number of neurons in the lth layer; Layer normalization is defined as: Where γ represents the scaling parameter vector, β represents the translation parameter vector, and the dimensions of the parameter vectors γ and β are the same as z (l) The dimensions are the same; e. RNN-LSTM recurrent neural network layer: The RNN-LSTM recurrent neural network layer is composed of two layers of recurrent neural network RNN, three layers of long short-term memory network LSTM and three layers of fully connected layers; (1) The basic unit of RNN is a recurrent unit, which receives an input and a hidden state from the previous time step, and outputs the hidden state of the current time step. The basic RNN structure consists of an input layer, a hidden layer, and an output layer; expand the basic RNN structure in the time dimension to obtain the entire network structure framework of the RNN; The RNN network receives the input x at time t. t After that, the value of the hidden layer is s t , the formula is as follows: s t =f(U×x t +W×s t-1 ); U is the weight matrix of input x, and W is the previous hidden layer value s t-1 As the current input weight matrix, f is the activation function, s t The value of x depends on t and t-1 ; The output value is o t , the calculation formula is as follows: o t =g(V×s t ); V is the weight matrix of the output layer, and g is the activation function; The hidden layer has two inputs. The first is the weight matrix U from the input layer to the hidden layer and the input x t The second is the product of the previous hidden layer value s t-1 The product of s and W, that is, the s calculated at the previous moment t-1 Need to cache, with input x t Calculate and output the final o together t ; A two-layer RNN neural network is used. The first layer RNN has 128 neurons and the second layer RNN has 64 neurons. The activation functions of the first and second layers are both ReLU functions. The two RNN layers are used to receive and process the output results of the multi-head attention mechanism layer, learn the relationship between multi-step time series features, and output them to the long short-term memory network LSTM layer. (2) The core of LSTM is a unit in which each neuron contains a forget gate, an input gate, and an output gate. Each gate is a special structure used to control the flow of information and pass information from the previous moment to the current moment, thereby capturing the time dependency in the data. The LSTM neural network includes: Forget Gate: The role of the forget gate is to decide what information to discard from the cell state; it is calculated by the following formula: f t =σ(W f ×[h t-1 ,x t ]+b f ); Among them, f t represents the output of the forget gate at time t, σ is the sigmoid function, W f and b f is the weight and bias of the forget gate, h t-1 is the previous hidden state, x t is the current input; Input Gate: The input gate is used to update the unit state, which consists of two parts: the sigmoid layer and the tanh layer; i t =σ(W i ×[h t-1 ,x t ]+b i ); Among them, i t is the input of the input gate, is a candidate value vector; Output Gate: The output gate is used to determine the next hidden state. The hidden state contains information about the previous input and is used for prediction. the t =σ(W o ×[h t-1 ,x t ]+b o ); h t =o t *tanh(C t ); Among them, t is the output of the output gate, h t is the current hidden state, C t is the current cell state; There are three LSTM layers in total. The number of neurons in the first LSTM layer is 64, the number of neurons in the second LSTM layer is 32, and the number of neurons in the third LSTM layer is 32. The activation functions of the first, second, and third layers are all Tanh functions. (3) Fully connected layer: In the fully connected layer Dense, all input nodes and output nodes are fully connected; three layers of fully connected layers are used as the last fully connected layer of the RNN-LSTM recurrent neural network layer; The number of neurons in the first fully connected layer is 32, the number of neurons in the second fully connected layer is 16, and the number of neurons in the last fully connected layer is 8. The activation functions of the three fully connected layers are all ReLU functions; f. Feed-forward layer: The feedforward layer is used to provide more learning capabilities. The feedforward network is a two-layer network. The activation function of the first layer is the ReLU function. For the vector x at each position in the input sequence, its expression is: FFN(x)=max(0,xW1+b1)W2+b2; g. Output layer: The number of neurons in the output layer is consistent with the feature dimension, and no activation function is set, and the predicted value is generated directly.

4. The method according to claim 1, characterized in that The step 3 comprises: (31) Initialization model: Integrate the RMM-LSTM recurrent neural network framework and add a multi-head attention mechanism to enhance the model's parallel learning ability and response to key features; (32) Optimizer: Select the Adam optimizer, which combines momentum and adaptive learning rate features to help optimize large-scale data sets quickly and stably; (33) Epochs: Monitor the model performance by the validation dataset loss, and stop training when the validation dataset loss no longer decreases or increases for several consecutive epochs. The parameter is set to 200 epochs. (34) Batch size: This parameter determines the amount of data used for each gradient update, affecting training stability and efficiency. Select an appropriate batch size based on hardware performance and task requirements to strike a balance between training stability and computational efficiency. Set it to Batch_Size = 64. (35) Evaluation indicators: ①Accuracy The accuracy is used to measure whether the predicted value is within the error threshold. The principle is as follows, providing intuitive feedback on the comprehensive performance of the model on the training dataset, validation dataset, and test dataset; Among them, y i is the true value of the i-th sample, i.e., the target value, which is the actual observed value of the time series; is the predicted value of the i-th sample, i.e., the predicted output of the model; ε is the error tolerance or threshold, which sets an error range, indicating that when the difference between the predicted value and the true value is less than the threshold, the prediction is considered correct; the error tolerance is 5% or 10%; n is the total number of samples; ② Loss rate Select the customized binary cross entropy loss function as the basic loss function of the recurrent neural network fusion model to judge the prediction performance of the recurrent neural network fusion model; The formula for the binary cross entropy loss function is: Where N is the total number of samples; y i is the true label of the i-th sample, which takes the value of 0 or 1; p i is the probability that the i-th sample predicted by the model belongs to category 1, and its value range is between (0, 1); The BCE loss function adjusts the model parameters during training so that the predicted probability p i Closer to the true label y i ; It encourages the model to output probability values ​​close to the true category, thereby improving the accuracy of classification; In order to prevent the model from overfitting the training data, L2 regularization is introduced on the basis of the BCE loss function; L2 regularization is weight decay, and its core idea is to penalize the model parameters and force the model to maintain a small weight value, thereby preventing the model from overfitting the training data; The L2 regularization formula is: Among them, λ is a hyperparameter of regularization strength, which controls the weight of the regularization term. The larger λ is, the stronger the regularization effect is. j is the parameter or weight of the model; n is the total number of model parameters; the L2 regularization term is the sum of the squares of all weights. The larger the weight value, the greater the penalty; Combining L2 regularization with the binary cross entropy loss function BCE, we get a new loss function, whose formula is: The binary cross entropy loss function BCE calculates the prediction error of the model, while the L2 regularization term helps the model avoid overfitting by penalizing larger weights; ③R 2 value R 2 The value measures the proportion of the variance of the target variable explained by the model, and its value range is [0,1]. Its principle is shown in the following formula; R 2 Reflects the goodness of fit of the model. The closer the value is to 1, the stronger the model's ability to explain the target variable is. Among them, y i is the true value of the i-th sample, i.e., the target value; is the predicted value of the i-th sample, i.e. the predicted output of the model; is the average of the true values, that is is the residual sum of squares RSS, that is, the sum of the squares of the differences between the true value and the predicted value; is the total sum of squares TSS, that is, the sum of the squares of the differences between the true value and the true value mean; (36) Training strategy: In order to improve the multi-step prediction performance of the nonlinear autoregressive neural network NARX-NN, a prediction value feedback retraining FR strategy is adopted; the prediction value feedback retraining strategy is aimed at the inconsistency of the input in the training phase and the multi-step prediction phase of the NARX-NN model to reduce the difference, thereby improving the multi-step prediction performance of the model; The basic idea of ​​FR is to reconstruct the training samples using the single-step prediction results with errors, and train the model again to enhance the model's robustness to prediction errors. The specific steps are as follows: ①First, use the initial data set and adopt the conventional training strategy to complete the initial training, that is, use the back propagation algorithm to train the network. At this time, the input target variables are all actual observation values; ②Use the trained network to perform single-step prediction to obtain the prediction result; single-step prediction ensures that the error of the prediction result is within a reasonable range. If multi-step prediction results are used, the excessive prediction error will cause the reconstructed sample in the following step ③ to deviate from the dynamic characteristics of the system to be modeled; ③Replace the measured target value in the model input with its corresponding single-step prediction result to obtain a new training sample. Multiple replacements are required to reconstruct the data set. The collective operation is: given the real data of the first 63 time steps, predict the data of the next time step, replace the real data in the prediction time step with the predicted data, and complete a reconstruction of the data set operation. The data of 64 time steps is a cycle, and the operation is repeated until the data in all cycles of the entire data set completes the replacement of the real value and the predicted value. ④ Use the reconstructed samples to train the network again; after training, check whether the validation set loss is greater than the recorded minimum validation set loss for several consecutive rounds of training; if so, it indicates that the performance of the model on the validation set is no longer improved, and the training is terminated to obtain a multi-step prediction model; conversely, if the validation set loss decreases in the current round, update the minimum validation set loss, enter steps ② and ③, and repeat the single-step prediction and network training.

5. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the data pre-fetching method based on a multi-head attention mechanism and an RNN-LSTM network as described in any one of claims 1 to 4.

6. An electronic device, characterized in that: include: one or more processors; Memory; And one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, enable the device to execute the data prefetching method based on the multi-head attention mechanism and the RNN-LSTM network as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Instructions and logic for load-indices-and-prefetch-scatters operations

    CN108369516A

  • Solid state disk data prefetching method based on attention mechanism

    CN114706798A

  • Related network flow prediction method based on multi-head attention mechanism

    CN115146732A

  • Traffic prediction method and device, electronic equipment and readable storage medium

    CN116910443A

  • Solid state disk data prefetching method based on attention mechanism

    CN119225645A

Cited By

  • Typhoon rapid enhancement prediction method based on time-space sequence and multi-modal feature fusion

    CN120633957A

  • Tailing ash value prediction system and method based on bionic neural network

    CN120635677A