ESN query-oriented decomposition reconstruction and optimization algorithm system page view prediction method and related device

The time series is decomposed through variational modal decomposition and fully integrated empirical modal decomposition algorithm, and combined with CNN-Transformer and XGBoost models for prediction, the problem of difficulty in capturing nonlinearity and long-term dependence in the ESN query service center system is solved, and the prediction accuracy is improved.

CN120216561APending Publication Date: 2025-06-27湖南工商大学
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510214963.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing time series prediction models are difficult to capture complex nonlinear patterns and long-term dependencies in ESN query service center systems, resulting in low prediction accuracy.

Method used

The time series is decomposed by variational modal decomposition algorithm and fully integrated empirical modal decomposition algorithm, high-frequency and low-frequency components are extracted, and prediction is made through the CNN-Transformer model and XGBoost model, and the model parameters are optimized in combination with the sparrow optimization algorithm.

Benefits of technology

It improves the ability to deal with nonlinear and non-stationarity problems in time series, enhances prediction accuracy, and can capture long-distance dependencies more accurately, prevent overfitting of high-precision models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216561A_ABST
    Figure CN120216561A_ABST
Patent Text Reader

Abstract

The invention provides an ESN query-oriented decomposition reconstruction and optimization algorithm system page view prediction method and a related device, and relates to the technical field of time sequence prediction. A time sequence is effectively decomposed through the combination of a variational mode decomposition algorithm and a completely integrated empirical mode decomposition algorithm, a key frequency component is extracted for a non-stationary sequence, and a prediction model is generated for a CNN-Transform model by adopting a sparrow optimization algorithm, so that the prediction precision of the model can be better improved; according to the method, the long-distance dependency relationship in the sequence is more accurately found out, and then the XGBoost model is combined for use, so that the overfitting phenomenon of a high-precision model on a simple component can be effectively prevented, the processing capability on nonlinear and non-stationary problems in time sequence data is further improved, and the prediction precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of time series prediction, and in particular to a system access volume prediction method and related devices based on a decomposition, reconstruction and optimization algorithm for ESN queries. Background Art

[0002] Time series analysis is a set of random variables obtained by observing the historical access data of the ESN query service center system and collecting them at a certain frequency. Time series prediction technology aims to mine the core laws contained in these data and make accurate estimates of future access volume based on known factors. The application scenarios of time series prediction in the ESN query service center system include access traffic analysis, system resource optimization, service response time prediction, etc.

[0003] In the field of time series forecasting, forecasting methods include traditional statistical methods (such as ARIMA models) and machine learning-based methods (such as random forests and support vector machines). Although these methods are effective in some cases, they often have difficulty capturing complex nonlinear relationships and long-term dependencies in the data.

[0004] According to the prior art, the time series data of the access volume of the ESN query service center system often exhibits complex nonlinear patterns and non-stationary characteristics, which means that the statistical characteristics of the data (such as mean and variance) will change over time. Traditional time series prediction models, such as autoregressive moving average (ARMA) and autoregressive integrated moving average (ARIMA) models, usually assume that the data is linear and stationary, which limits their prediction accuracy when dealing with nonlinear and non-stationary data. In the actual application of the ESN query service center system, the nonlinearity and non-stationarity of the data are particularly obvious, which requires the prediction model to be able to capture and handle these complexities.

[0005] According to the prior art, long-term dependencies (LTD) refer to the correlation between distant data points in a time series. In the time series data of the ESN query service center system, early access patterns may have a significant impact on the long-term future access volume. Many existing time series prediction models, especially those based on short windows, often have difficulty capturing such long-distance dependencies, resulting in the inability of the prediction model to fully utilize all the information in the time series data, thus affecting the accuracy of the prediction.

[0006] Therefore, how to achieve accurate prediction of time series has become a technical problem that needs to be solved urgently. Summary of the invention

[0007] To achieve accurate prediction of time series, this application provides a system access volume prediction method and related devices for a decomposition, reconstruction, and optimization algorithm for ESN queries.

[0008] In the first aspect, a system access volume prediction method for a decomposition, reconstruction, and optimization algorithm for ESN queries provided by this application adopts the following technical solutions:

[0009] A system access volume prediction method for a decomposition, reconstruction, and optimization algorithm for ESN queries includes:

[0010] Obtain the target time series and perform Z-score transformation on the target time series;

[0011] Use the variational mode decomposition algorithm to decompose the Z-score transformed target time series to generate a first decomposition result;

[0012] Determine the high-frequency component in the first decomposition result according to sample entropy;

[0013] Use the complete ensemble empirical mode decomposition algorithm to decompose the high-frequency component to generate a second decomposition result;

[0014] Divide the first decomposition result into complex components according to the sample entropy, and at the same time divide the second decomposition result into simple components;

[0015] Extract features from the complex components and the simple components through a preset sliding time window and determine the training set and the data set;

[0016] Construct a CNN-Transformer model according to the training set and initialize the parameters;

[0017] Use the sparrow optimization algorithm to optimize the initialized parameters and train the CNN-Transformer model to generate a prediction model;

[0018] Use the prediction model to predict the complex components to generate a first prediction result, and at the same time use the XGBoost model to predict the simple components to generate a second prediction result;

[0019] Generate a comprehensive result according to the first prediction result and the second prediction result and perform inverse Z-score transformation on the comprehensive prediction result to obtain a combined prediction result.

[0020] Optionally, the step of using the variational mode decomposition algorithm to decompose the Z-score transformed target time series to generate a first decomposition result includes:

[0021] Construct a variational problem based on the target time series after Z-score transformation, and decompose the observed signal x(t) into K modal signals, expressed as:

[0022]

[0023] Among them, u k (t) represents the k-th modal function, ω k represents the central frequency of the k-th modal function, represents the derivative with respect to time t, δ(t) represents the Dirac delta function, j represents the imaginary unit, and f(t) represents the original signal;

[0024] Equivalently transform the equality-constrained optimization problem into an unconstrained optimization problem through the augmented Lagrangian function:

[0025]

[0026] Among them, P(·) is the Lagrangian function, and λ(t) is the Lagrange multiplier operator at time t;

[0027] Solve it through the Alternating Direction Method of Multipliers (ADMM)

[0028]

[0029] Further solve:

[0030]

[0031] Among them, n is the number of iterations, ω represents the frequency domain, i is an arbitrary value obtained by setting different parameters, is the central frequency of the k-th IMF component after the n-th iteration.

[0032] Optionally, the step of determining the high-frequency component in the first decomposition result according to sample entropy includes:

[0033] Form a one-dimensional vector sequence X m (1), …, X m (N - m + 1) by serial numbers, where X m (i) = {x(i), x(i + 1), …, x(i + m - 1)}, 1 ≤ i ≤ N - m + 1, representing m consecutive x values starting from the i-th point;

[0034] Define the distance d[X m (i), X m (j)] between vector X m (i) and X m (j) as the absolute value of the maximum difference between their corresponding elements, that is:

[0035] d[Xm (i), X m (j)] = max k=0,…,m-1 (|x(i + k) - x(j + k)|)

[0036] Count X m (i) and X m the number of j (1 ≤ j ≤ N - m, j ≠ i) whose distance from X(j) is less than or equal to r, and denote it as B i , for 1 ≤ i ≤ N - m, define:

[0037]

[0038] Define B (m) (r) as:

[0039]

[0040] Increase the dimension to m + 1 and calculate the number of points whose distance between X(i) and X(j) (1 ≤ j ≤ N - m, j ≠ i) is less than or equal to r, denoted as A m+1 (i) and X m+1 (j) (1 ≤ j ≤ N - m, j ≠ i), and denote it as A i , Define it as:

[0041]

[0042] Define A m (r) as:

[0043]

[0044] B m (r) is the probability that two sequences match m points under the similarity tolerance r, while A m (r) is the probability that two sequences match m + 1 points. Sample entropy is defined as:

[0045]

[0046] When N is a finite value, estimate the high - frequency component as:

[0047]

[0048] Optionally, the step of decomposing the high - frequency component by using the complete ensemble empirical mode decomposition algorithm to generate a second decomposition result includes:

[0049] Add k - times Gaussian white noise to the original irregular sequence x(t) to construct a total of k pre - processed sequences, where k is the number of noise perturbation times, and the obtained sequence is:

[0050] x k (t) = x(t) + εk (t)

[0051] Among them, ε k (t) is the Gaussian white noise sequence added for the k-th time;

[0052] Perform EMD decomposition on all preprocessed sequences to obtain the first IMF component IMF k1 (t), and take the mean of all the first IMF components IMF k1 (t) as the first IMF component obtained by complete ensemble empirical mode decomposition, denoted as:

[0053]

[0054] The first-order residual is:

[0055] r1(t) = x(t) - IMF1(t)

[0056] Among them, K is the total number of times white noise is added; IMF k1 (t) is the first component obtained by using EMD decomposition;

[0057] Add the Gaussian white noise sequence ε k (t) to the first-order residual r1(t) to form a new sequence r 1k (t);

[0058] Perform EMD decomposition on each new sequence r 1k (t), calculate all the means to obtain the second-order IMF component, and the calculation formula is:

[0059]

[0060] The second-order residual is:

[0061] r2(t) = r1(t) - IMF2(t)

[0062] For k = 3, 4,..., K, where K is the total number of IMF components finally decomposed;

[0063] The calculation of the k-th order residual is:

[0064] r k (t) = r k-1 (t) - IMF k (t)

[0065] Add the white noise sequence to the k-th order residual to obtain the sequence r kk (t), and obtain the (k + 1)-th order IMF component as:

[0066]

[0067] When the residual sequence satisfies the preset stopping condition, the decomposition process is stopped, and the non-regular sequence x(t) is decomposed into K IMF components with different frequency characteristics and a residual component as the second decomposition result:

[0068]

[0069] Optionally, before the step of extracting features from the complex component and the simple component through a preset sliding time window and determining the training set and the data set, it includes:

[0070] Set the length of the preset sliding time window to the sum of the input sequence length and the prediction sequence length;

[0071] For a sliding time window of length d + 1, d historical data will be used as input to predict the value at the next moment, and then the window will slide forward one time step each time for a split.

[0072] Let the sequence be X = {x1, x2, …, x T}, then the first split will use x1, x2, …, x d data as input to predict x d+1 value, and a total of T - d subsequences are divided.

[0073] Optionally, the step of optimizing the initialized parameters using the sparrow optimization algorithm and training the CNN-Transformer model to generate a prediction model includes:

[0074] Initialize the population number n, the number of iterations T, and the ratio of discoverers to followers;

[0075] For the sparrow population X:

[0076]

[0077] Among them, n represents the number of the sparrow population, and d represents the dimension attached to the sparrow individual; update the position of the discoverer

[0078]

[0079] Among them, represents the position of the i-th sparrow in the d-th dimension in the t-th generation, α is a random number between [0, 1], Γ is a random number obeying the standard normal distribution, L represents a matrix of size i × d with all elements being 1, and R2 and ST represent the warning value and the safety value respectively;

[0080] Update the position of the follower

[0081]

[0082] Among them, represents the worst position of the sparrow in the d-th dimension at the t-th iteration, represents the optimal position of the sparrow in the d-th dimension at the (t + 1)-th iteration of the population. A represents a multi-dimensional matrix with elements of 1 or -1 in a row,

[0083] Update the position of the vigilant

[0084]

[0085] Among them, is the current global optimal position. β represents the step size control parameter. K is a random number between [-1, 1] representing the sparrow movement direction. δ is a very small constant to avoid the denominator being 0; f i represents the fitness value of the i-th sparrow, f g and f ω are the optimal and worst fitness values of the current sparrow population respectively.

[0086] Optionally, the step of using the prediction model to predict the complex component to generate a first prediction result and using the XGBoost model to predict the simple component to generate a second prediction result includes:

[0087] Input the complex component into the prediction model so that the CNN module uses the convolutional layer to perform local feature extraction on the time series in the complex component;

[0088] Use the pooling layer to compress the data and the number of parameters and connect the local feature representation to the Transformer module with a residual structure;

[0089] Use the self-attention mechanism to obtain the long-range dependencies in the sequence data in the Transformer module;

[0090] Reduce the sequence length and increase the data dimension through the CNN convolutional pooling layer, and then send it to the Transformer encoder layer for feature enhancement to realize the fusion of spatial features and time-domain features using multi-head attention and generate the first prediction result;

[0091] Input the simple component into the XGBoost model so that the XGBoost model outputs the second prediction result.

[0092] In a second aspect, the present application provides a system access volume prediction system for the decomposition, reconstruction and optimization algorithm for ESN queries, characterized in that the system access volume prediction system for the decomposition, reconstruction and optimization algorithm for ESN queries includes:

[0093] A data acquisition module for acquiring a target time series and performing Z - score transformation on the target time series;

[0094] A first decomposition result module for decomposing the Z - score transformed target time series using the variational mode decomposition algorithm to generate a first decomposition result;

[0095] A high - frequency component module for determining high - frequency components in the first decomposition result according to sample entropy;

[0096] A second decomposition result module for decomposing the high - frequency components using the complete ensemble empirical mode decomposition algorithm to generate a second decomposition result;

[0097] A partitioning module for partitioning the first decomposition result into complex components according to the sample entropy, and at the same time partitioning the second decomposition result into simple components;

[0098] A time window module for extracting features from the complex components and the simple components through a preset sliding time window and determining a training set and a data set;

[0099] An initialization module for constructing a CNN - Transformer model according to the training set and initializing parameters;

[0100] A prediction model module for optimizing the initialized parameters using the sparrow optimization algorithm and training the CNN - Transformer model to generate a prediction model;

[0101] A prediction result module for using the prediction model to predict the complex components to generate a first prediction result, and at the same time using the XGBoost model to predict the simple components to generate a second prediction result;

[0102] A combined prediction module for generating a comprehensive result according to the first prediction result and the second prediction result and performing inverse Z - score transformation on the comprehensive prediction result to obtain a combined prediction result.

[0103] In a third aspect, the present application provides a computer device, which includes: a memory and a processor. When the processor runs the computer instructions stored in the memory, it executes the method described above.

[0104] In a fourth aspect, the present application provides a computer - readable storage medium, including instructions. When the instructions run on a computer, the computer is caused to execute the method described above.

[0105] In summary, the present application includes the following beneficial technical effects:

[0106] This application effectively decomposes time series through the combination of variational mode decomposition algorithm and complete ensemble empirical mode decomposition algorithm. For non-stationary series, key frequency components are extracted. The sparrow optimization algorithm is used to generate a prediction model for the CNN-Transformer model, which can better improve the prediction accuracy of the model, more accurately find long-distance dependencies in the series, and then combined with the XGBoost model, it can effectively prevent the overfitting phenomenon of high-precision models for simple components, further improving the processing ability of non-linear and non-stationary problems in time series data, thus enhancing the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] Figure 1 It is a schematic structural diagram of a computer device in the hardware operating environment involved in the solution of an embodiment of this application.

[0108] Figure 2 It is a schematic flowchart of an embodiment of the system access volume prediction method for the decomposition, reconstruction and optimization algorithm for ESN queries in this application.

[0109] Figure 3 It is an algorithm flowchart of an embodiment of the system access volume prediction method for the decomposition, reconstruction and optimization algorithm for ESN queries in this application;

[0110] Figure 4 It is a structural block diagram of an embodiment of the system access volume prediction system for the decomposition, reconstruction and optimization algorithm for ESN queries in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0111] In order to make the objectives, technical solutions and advantages of this application clearer, the following further details this application through the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0112] Refer to Figure 1 , Figure 1 which is a schematic structural diagram of a computer device in the hardware operating environment involved in the solution of an embodiment of this application.

[0113] Such as Figure 1As shown in the figure, a computer device may include: a processor 1001, such as a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to implement connection communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed Random Access Memory (RAM) or a stable Non-Volatile Memory (NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0114] Those skilled in the art can understand that Figure 1 the structure shown in

[0115] does not constitute a limitation on the computer device, and may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. Figure 1 As shown in

[0116] In Figure 1 the computer device shown in

[0117] The embodiments of the present application provide a method for predicting the system access volume of a decomposition reconstruction and optimization algorithm for ESN queries. Referring to Figure 2 Figure 2 is a schematic flowchart of an embodiment of the method for predicting the system access volume of the decomposition reconstruction and optimization algorithm for ESN queries according to the present application.

[0118] ​In this embodiment, the system access volume prediction method of the decomposition, reconstruction and optimization algorithm for ESN query includes the following steps:

[0119] Step S10: Obtain the target time series and perform Z-score processing on the target time series.

[0120] It should be noted that in recent years, deep learning technologies, especially convolutional neural networks (CNNs) and Transformer models, have begun to be applied to time series prediction problems due to their excellent performance in image recognition and natural language processing. CNNs can extract local features of time series data, while Transformer models are good at capturing long-range dependencies. In addition, as an efficient ensemble learning method, XGBoost is also widely used in time series prediction, which improves the prediction accuracy by constructing multiple decision trees.

[0121] It can be understood that the technical terms in this embodiment are described as follows:

[0122] VMD: Variational Mode Decomposition (VMD) is a signal decomposition and estimation method. This method determines the frequency center and bandwidth of each component by iteratively searching for the optimal solution of the variational model, thereby being able to adaptively achieve the frequency domain dissection of the signal and the effective separation of each component.

[0123] CEEMDAN: Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN) is an advanced signal decomposition technology for processing non-linear and non-stationary signals. The CEEMDAN algorithm effectively solves the mode mixing problem in traditional Empirical Mode Decomposition (EMD) by introducing adaptive noise and multiple iterations, improving the accuracy and stability of decomposition.

[0124] SSA: Sparrow Search Algorithm (SSA) is a new swarm intelligence optimization algorithm inspired by the foraging behavior of sparrows. During the foraging process of sparrows, sparrows are divided into discoverers (explorers) and joiners (followers). Discoverers are responsible for finding food in the population and providing the foraging area and direction for the entire sparrow population, while joiners use discoverers to obtain food. The Sparrow Search Algorithm finds the optimal solution to the problem by simulating this behavior.

[0125] It can be understood that the main defects of the current existing technologies mainly include:

[0126] (1)According to the existing information, time series data often exhibits complex non-linear patterns and non-stationary characteristics, which means that the statistical characteristics of the data (such as the mean and variance) change over time. Traditional time series prediction models, such as autoregressive moving average (ARMA) and autoregressive integrated moving average (ARIMA) models, usually assume that the data is linear and stationary, which limits their prediction accuracy when dealing with non-linear and non-stationary data. In real-world applications, such as financial market analysis, weather forecasting, and industrial production process control, the non-linearity and non-stationarity of the data are particularly obvious, which requires the prediction model to be able to capture and handle these complexities.

[0127] (2)According to the prior art, long-term dependencies (LTD) refer to the correlations between distant data points in a time series. In some time series data, early events may have a significant impact on long-term future values. For example, in natural language processing, the beginning of a text may have a strong correlation with the end; in meteorology, seasonal patterns may affect the weather several months later. Many existing time series prediction models, especially those based on short windows, often have difficulty capturing such long-term dependencies, resulting in the prediction model being unable to fully utilize all the information in the time series data, thus affecting the accuracy of the prediction.

[0128] It should be noted that, as Figure 3 shown in the algorithm flow, this embodiment provides a system access volume prediction model for the decomposition and reconstruction and optimization algorithm for long-term ESN queries based on quadratic decomposition reconstruction and sparrow optimization algorithm. First, the time series data is decomposed quadratically by the VMD and CEEMDAN algorithms to extract components with multiple frequency characteristics, then the sparrow optimization algorithm is used to optimize the CNN-Transformer model, and then the CNN-Transformer and XGBoost combined model is trained using multiple frequency characteristic components, and finally the combined model is used to predict the long time series; this model has high accuracy and robustness, and can also achieve relatively ideal results for non-linear and non-stationary problems.

[0129] In specific implementation, the steps of calculating the mean and standard deviation include: calculating the mean: for a given data set, calculate the mean of all data; calculating the standard deviation: calculate the standard deviation of the data, which is the average distance of the data from the mean. The larger the standard deviation, the greater the fluctuation of the data.

[0130] Calculating the Z-score

[0131] For each data point, calculate the Z-score using the following formula:

[0132]

[0133] Among them, X is the value of a single data point, Mean is the average value of the data set, and StandardDeviation is the standard deviation of the data set.

[0134] Step S20: Use the variational mode decomposition algorithm to decompose the target time series after Z-score transformation to generate a first decomposition result.

[0135] It should be noted that the step of using the variational mode decomposition algorithm to decompose the target time series after Z-score transformation to generate a first decomposition result includes:

[0136] Construct a variational problem according to the target time series after Z-score transformation, and decompose the observed signal x(t) into K modal signals, expressed as:

[0137]

[0138] Among them, u k (t) represents the k-th modal function, ω k represents the central frequency of the k-th modal function, represents the derivative with respect to time t, δ(t) represents the Dirac δ function, j represents the imaginary unit, and f(t) represents the original signal;

[0139] Equivalently transform the equality constraint optimization problem into an unconstrained optimization problem through the augmented Lagrangian function:

[0140]

[0141] Among them, P(·) is the Lagrangian function, and λ(t) is the Lagrange multiplier operator at time t;

[0142] Solve it through the alternating direction method of multipliers ADMM

[0143]

[0144] Further solve:

[0145]

[0146] Among them, n is the number of iterations, ω represents the frequency domain, i is an arbitrary value obtained by setting different parameters, is the central frequency of the k-th IMF component after the n-th iteration.

[0147] Step S30: Determine the high-frequency component in the first decomposition result according to the sample entropy.

[0148] It should be noted that the steps to determine the high-frequency component in the first decomposition result according to sample entropy include: forming a one-dimensional vector sequence X of length m in order of serial numbers m (1),…,X m (N - m + 1), where X m (i) = {x(i), x(i + 1), …, x(i + m - 1)}, 1 ≤ i ≤ N - m + 1, representing m consecutive x values starting from the i-th point;

[0149] Define the distance d[X m (i), X m (j)] between vector X m (i) and X m (j) as the absolute value of the maximum difference between their corresponding elements. That is:

[0150] d[X m (i), X m (j)] = max k=0,…,m-1 (|x(i + k) - x(j + k)|)

[0151] Count the number of j (1 ≤ j ≤ N - m, j ≠ i) whose distance between X m (i) and X m (j) is less than or equal to r, and denote it as B i ; for 1 ≤ i ≤ N - m, define:

[0152]

[0153] Define B (m) (r) as:

[0154]

[0155] Increase the dimension to m + 1, calculate the number of elements whose distance between X m+1 (i) and X m+1 (j) (1 ≤ j ≤ N - m, j ≠ i) is less than or equal to r, and denote it as A i , Define it as:

[0156]

[0157] Define A m (r) as:

[0158]

[0159] B m (r) is the probability that two sequences match at m points under the similarity tolerance r, while A m (r) is the probability that two sequences match at m + 1 points. The sample entropy is defined as:

[0160]

[0161] When N is a finite value, the estimated high-frequency component is:

[0162]

[0163] Step S40: Use the complete ensemble empirical mode decomposition algorithm to decompose the high-frequency component to generate a second decomposition result.

[0164] It can be understood that the step of using the complete ensemble empirical mode decomposition algorithm to decompose the high-frequency component to generate a second decomposition result includes: adding Gaussian white noise to the original irregular sequence x(t) k times to construct a total of k preprocessing sequences, where k is the number of noise perturbation times, and the obtained sequences are:

[0165] x k (t) = x(t) + ε k (t)

[0166] where ε k (t) is the Gaussian white noise sequence added for the kth time; perform EMD decomposition on all preprocessing sequences to obtain the first IMF component IMF k1 (t), and take the mean of all the first IMF components IMF k1 (t) as the first IMF component obtained by the complete ensemble empirical mode decomposition, denoted as:

[0167]

[0168] The first-order residual is:

[0169] r1(t) = x(t) - IMF1(t)

[0170] where K is the total number of times white noise is added; IMF k1 (t) is the first component obtained by using EMD decomposition;

[0171] Add the Gaussian white noise sequence ε k (t) to the first-order residual r1(t) to form a new sequence r 1k (t);

[0172] Perform EMD decomposition on each new sequence r 1k (t), calculate all the means to obtain the second-order IMF component, and the calculation formula is:

[0173]

[0174] The second-order residual is:

[0175] r2(t) = r1(t) - IMF2(t)

[0176] For k = 3, 4, …, K, where K is the total number of IMF components finally decomposed;

[0177] The calculation of the k-th order residual is as follows:

[0178] r k (t) = r k-1 (t) - IMF k (t)

[0179] Add a white noise sequence to the k-th order residual to obtain the sequence r kk (t), and obtain the (k + 1)-th order IMF component as:

[0180]

[0181] When the residual sequence satisfies the preset stop condition, stop the decomposition process. The non-regular sequence x(t) is decomposed into K IMF components with different frequency characteristics and a residual component as the second decomposition result:

[0182]

[0183] Step S50: Divide the first decomposition result into complex components according to sample entropy, and at the same time divide the second decomposition result into simple components.

[0184] Step S60: Extract features from the complex components and simple components through a preset sliding time window and determine the training set and the data set.

[0185] It should be noted that before the step of extracting features from the complex components and simple components through a preset sliding time window and determining the training set and the data set, it includes: setting the length of the preset sliding time window to the sum of the input sequence length and the prediction sequence length; for a sliding time window of length d + 1, d historical data will be used as input to predict the value at the next moment, and then the window will slide forward one time step each time for a split; assuming the sequence is X = {x1, x2, …, x T}}, then the first split will use x1, x2, …, x d data as input to predict x d+1 value, and a total of T - d subsequences are divided.

[0186] Step S70: Construct a CNN-Transformer model according to the training set and initialize the parameters.

[0187] Step S80: Optimize the initialized parameters using the sparrow search algorithm and train the CNN-Transformer model to generate a prediction model.

[0188] It should be noted that the steps of optimizing the initialized parameters using the sparrow search algorithm and training the CNN-Transformer model to generate a prediction model include: initializing the population size n, the number of iterations T, and the ratio of discoverers to followers; for the sparrow population X:

[0189]

[0190] where n represents the size of the sparrow population and d represents the dimension attached to each sparrow individual; update the position of the discoverers

[0191]

[0192] where represents the position of the i-th sparrow in the d-th dimension at the t-th generation, α is a random number between [0,1], Γ is a random number following the standard normal distribution, L represents a matrix of size i×d with all elements being 1, and R2 and ST represent the warning value and the safety value respectively; update the position of the followers

[0193]

[0194] where represents the worst position of the sparrow in the d-th dimension at the t-th iteration, represents the best position of the sparrow in the d-th dimension at the (t + 1)-th iteration of the population, A represents a one-row multi-dimensional matrix with each element being 1 or -1,

[0195] Update the position of the vigilant ones

[0196]

[0197] where is the current global optimal position, β represents the step size control parameter, K is a random number between [-1,1] representing the moving direction of the sparrow, and δ is a very small constant to avoid the denominator being 0; f i represents the fitness value of the i-th sparrow, f g and f ω are the best and worst fitness values of the current sparrow population respectively.

[0198] Step S90: Use the prediction model to predict the complex components to generate the first prediction result, and at the same time use the XGBoost model to predict the simple components to generate the second prediction result.

[0199] It should be noted that the steps of using the prediction model to predict the complex component to generate the first prediction result and using the XGBoost model to predict the simple component to generate the second prediction result include: inputting the complex component into the prediction model so that the CNN module uses the convolutional layer to extract local features of the time series in the complex component; using the pooling layer to compress the data and the number of parameters and connecting the local feature representation to the Transformer module with a residual structure; using the self-attention mechanism to obtain the long-range dependencies in the sequence data in the Transformer module; reducing the sequence length and increasing the data dimension through the CNN convolutional pooling layer, and then sending it to the Transformer encoder layer for feature enhancement to realize the fusion of spatial features and time-domain features using multi-head attention and generate the first prediction result; inputting the simple component into the XGBoost model so that the XGBoost model outputs the second prediction result.

[0200] Step S100: Generate a comprehensive result based on the first prediction result and the second prediction result and perform inverse Z-score processing on the comprehensive prediction result to obtain the combined prediction result.

[0201] In this embodiment, the variational mode decomposition algorithm and the complete ensemble empirical mode decomposition algorithm are combined to effectively decompose the time series. For non-stationary sequences, the key frequency components are extracted. The sparrow optimization algorithm is used to generate a prediction model for the CNN-Transformer model, which can better improve the prediction accuracy of the model, more accurately find the long-range dependencies in the sequence, and then combined with the XGBoost model, it can effectively prevent the overfitting phenomenon of the high-precision model for simple components, further improve the processing ability of the nonlinear and non-stationary problems in the time series data, and thus improve the prediction accuracy.

[0202] In addition, the embodiment of the present application also proposes a computer-readable storage medium, on which a program for predicting the system access volume of the decomposition reconstruction and optimization algorithm for ESN queries is stored. When the program for predicting the system access volume of the decomposition reconstruction and optimization algorithm for ESN queries is executed by a processor, it realizes the steps of the method for predicting the system access volume of the decomposition reconstruction and optimization algorithm for ESN queries as described above.

[0203] Refer to Figure 4 , Figure 4 which is a structural block diagram of an embodiment of the system for predicting the system access volume of the decomposition reconstruction and optimization algorithm for ESN queries of the present application.

[0204] As Figure 4 shown, the system for predicting the system access volume of the decomposition reconstruction and optimization algorithm for ESN queries proposed in the embodiment of the present application includes:

[0205] The data acquisition module 10 is used to acquire the target time series and perform Z-score transformation on the target time series.

[0206] The first decomposition result module 20 is used to decompose the Z-score transformed target time series by using the variational mode decomposition algorithm to generate the first decomposition result.

[0207] The high-frequency component module 30 is used to determine the high-frequency component in the first decomposition result according to the sample entropy.

[0208] The second decomposition result module 40 is used to decompose the high-frequency component by using the complete ensemble empirical mode decomposition algorithm to generate the second decomposition result.

[0209] The partitioning module 50 is used to partition the first decomposition result into complex components according to the sample entropy, and at the same time partition the second decomposition result into simple components.

[0210] The time window module 60 is used to extract features from the complex components and simple components through a preset sliding time window and determine the training set and the data set.

[0211] The initialization module 70 is used to construct a CNN-Transformer model according to the training set and initialize the parameters.

[0212] The prediction model module 80 is used to optimize the initialized parameters by using the sparrow optimization algorithm and train the CNN-Transformer model to generate a prediction model.

[0213] The prediction result module 90 is used to use the prediction model to predict the complex components to generate the first prediction result, and at the same time use the XGBoost model to predict the simple components to generate the second prediction result.

[0214] The combined prediction module 100 is used to generate a comprehensive result according to the first prediction result and the second prediction result and perform inverse Z-score transformation on the comprehensive prediction result to obtain the combined prediction result.

[0215] It should be understood that the above is only for illustration and does not constitute any limitation to the technical solution of the present application. In specific applications, those skilled in the art can set according to needs, and the present application does not make any restrictions on this.

[0216] In this embodiment, the variational mode decomposition algorithm and the complete ensemble empirical mode decomposition algorithm are combined to effectively decompose the time series. For non-stationary sequences, the key frequency components are extracted. The sparrow optimization algorithm is used to generate a prediction model for the CNN-Transformer model, which can better improve the prediction accuracy of the model and more accurately find the long-distance dependence relationship in the sequence. Then, combined with the XGBoost model, it can effectively prevent the overfitting phenomenon of the high-precision model for simple components, further improving the processing ability of the nonlinear and non-stationary problems in the time series data, thereby improving the prediction accuracy.

[0217] It should be noted that the above-described work process is only illustrative and does not limit the protection scope of this application. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no limitation is made here.

[0218] In addition, for the technical details not described in detail in this embodiment, reference can be made to the method for predicting the system access volume of the decomposition, reconstruction and optimization algorithm for ESN query provided in any embodiment of this application, which will not be elaborated here.

[0219] In addition, it should be noted that in this article, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including the element.

[0220] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0221] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory (ROM) / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of this application.

[0222] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.

Claims

1. A method for predicting system access volume based on decomposition, reconstruction and optimization algorithm for ESN query, characterized in that: include: Obtaining a target time series and performing Z-score processing on the target time series; Decomposing the target time series after Z-score processing by using variational mode decomposition algorithm to generate a first decomposition result; determining a high frequency component in the first decomposition result according to sample entropy; Decomposing the high frequency component using a fully integrated empirical mode decomposition algorithm to generate a second decomposition result; Dividing the first decomposition result into complex components according to the sample entropy, and dividing the second decomposition result into simple components; Extract features from the complex component and the simple component through a preset sliding time window and determine a training set and a data set; Construct a CNN-Transformer model according to the training set and initialize parameters; Optimizing the initialized parameters using a sparrow optimization algorithm and training the CNN-Transformer model to generate a prediction model; Using the prediction model to predict the complex component to generate a first prediction result, and using the XGBoost model to predict the simple component to generate a second prediction result; A comprehensive result is generated according to the first prediction result and the second prediction result, and the comprehensive prediction result is subjected to inverse Z-score processing to obtain a combined prediction result.

2. The system access volume prediction method according to claim 1, characterized in that: The step of using the variational mode decomposition algorithm to decompose the target time series after Z-score processing to generate a first decomposition result includes: According to the variational problem constructed for the target time series after Z-score processing, the observed signal x(t) is decomposed into K modal signals, which are expressed as: Among them, u k (t) represents the kth mode function, ω k represents the center frequency of the kth mode function, represents the derivative with respect to time t, δ(t) represents the Dirac delta function, j represents the imaginary unit, and f(t) represents the original signal; The equality constraint optimization problem is equivalent to an unconstrained optimization problem by augmenting the Lagrangian function: Where P(·) is the Lagrangian function, λ(t) is the Lagrangian multiplication operator at time t; Solved by alternating direction method of multipliers ADMM Further solution: Where n is the number of iterations, ω represents the frequency domain, and i is any value obtained by setting different parameters. is the center frequency of the kth IMF component after the nth iteration.

3. The system access volume prediction method according to claim 1, characterized in that: The step of determining the high-frequency component in the first decomposition result according to the sample entropy comprises: A set of one-dimensional m vector sequences X are formed by sequence number m (1),···,X m (N-m+1), where X m (i)={x(i),x(i+1),···,x(i+m-1)},1≤i≤N-m+1, represents the m consecutive values ​​of x starting from the i-th point; Define vector X m (i) With X m (j) The distance d[X m (i),X m (j)] is the absolute value of the maximum difference between the corresponding elements of the two, that is: d[X m (i),X m (j)]=max k=0,···,m-1 (|x(i+k)-x(j+k)|) Statistics X m (i) With X m The number of j (1≤j≤Nm,j≠i) whose distance between them is less than or equal to r, and is recorded as B i , for 1≤i≤Nm, define: Definition B (m) (r) is: Increase the dimension to m+1 and calculate X m+1 (i) With X m+1 The number of distances (j)(1≤j≤Nm,j≠i) less than or equal to r, denoted as A i , Defined as: Definition A m (r) is: B m (r) is the probability that two sequences match m points under similarity tolerance r, and A m (r) is the probability that two sequences match m+1 points, and the sample entropy is defined as: When N is a finite value, the estimated high-frequency component is:

4. The system access volume prediction method according to claim 1, characterized in that: The step of decomposing the high frequency component by using a fully integrated empirical mode decomposition algorithm to generate a second decomposition result comprises: Add k Gaussian white noise to the original irregular sequence x(t) to construct a total of k preprocessed sequences, where k is the number of noise disturbances. The resulting sequence is: x k (t)=x(t)+ε k (t) Among them, ε k (t) is the Gaussian white noise sequence added for the kth time; Perform EMD decomposition on all preprocessed sequences to obtain the first IMF component IMF k1 (t), take all the first IMF components IMF k1 The mean of (t) is taken as the first IMF component obtained by fully integrated empirical mode decomposition and is expressed as: The first-order residual is: r1(t)=x(t)-IMF1(t) Where K is the total number of times white noise is added; IMF k1 (t) is the first component obtained by EMD decomposition; Add a Gaussian white noise sequence ε to the first-order residual r1(t) k (t) forms a new sequence r 1k (t); For each new sequence r 1k (t) Perform EMD decomposition and calculate all mean values ​​to obtain the second-order IMF component. The calculation formula is: The second order residual is: r2(t)=r1(t)-IMF2(t) For k = 3, 4, ..., K, where K is the total number of IMF components obtained by the final decomposition; The calculation of the k-th order residual is: r k (t)=r k-1 (t)-IMF k (t) Add a white noise sequence to the k-th order residual to obtain a sequence r kk (t), and the k+1th order IMF component is obtained as: When the residual sequence meets the preset stop condition, the decomposition process stops, and the irregular sequence x(t) is decomposed into K IMF components with different frequency characteristics and a residual component as the second decomposition result:

5. The system access volume prediction method according to claim 1, characterized in that: Before the step of extracting features from the complex component and the simple component through a preset sliding time window and determining a training set and a data set, the method includes: Set the length of the preset sliding time window to the sum of the input sequence length and the predicted sequence length; A sliding time window of length d+1 will use d historical data as input to predict the value at the next moment, and then the window will be split once each time it slides forward one time step; Let the sequence be X = {x1,x2,···,x T }, the first split will use x1,x2,···,x d Data is used as input to predict x d+1 The value of , a total of Td sub-sequences are divided.

6. The system access volume prediction method based on the decomposition, reconstruction and optimization algorithm for ESN query according to claim 1 is characterized in that: The step of optimizing the initialized parameters using the sparrow optimization algorithm and training the CNN-Transformer model to generate a prediction model includes: Initialize the population size n, the number of iterations T, and the ratio of discoverers to followers; For a sparrow population X: Among them, n represents the number of sparrow populations, and d represents the dimension attached to individual sparrows; Update finder location in, represents the position of the i-th sparrow in the t-th generation in the d-th dimension, α is a random number between [0,1], Γ is a random number that obeys the standard normal distribution, L represents a matrix of size i×d with all elements being 1, R2 and ST represent the warning value and safety value respectively; Update follower location in, represents the worst position of the sparrow in the dth dimension at the tth iteration, represents the optimal position of the sparrow in the dth dimension at the t+1th iteration of the population. A represents a multidimensional matrix with each element being 1 or -1. + =A T (AA T ) -1 Update Sentinel location in, is the current global optimal position, β represents the step size control parameter, K is a random number between [-1,1] representing the moving direction of the sparrow, and δ is a very small constant to avoid the situation where the denominator is 0; f i represents the fitness value of the i-th sparrow, f g and f ω are the optimal and worst fitness values ​​of the current sparrow population respectively.

7. The system access volume prediction method according to claim 1, characterized in that: The step of using the prediction model to predict the complex component to generate a first prediction result, and using the XGBoost model to predict the simple component to generate a second prediction result, includes: Inputting the complex component into the prediction model so that the CNN module uses a convolutional layer to extract local features of the time series in the complex component; The pooling layer is used to compress the data and the number of parameters, and the local feature representation is connected to the Transformer module using a residual structure; Using the self-attention mechanism to capture long-range dependencies in the sequence data in the Transformer module; The CNN convolutional pooling layer is used to reduce the sequence length and increase the data dimension, and then it is sent to the Transformer encoder layer for feature enhancement to achieve the fusion of spatial features and temporal features using multi-head attention and generate the first prediction result; The simple component is passed into the XGBoost model so that the XGBoost model outputs a second prediction result.

8. A system access volume prediction system based on decomposition, reconstruction and optimization algorithm for ESN query, characterized in that: The system visit volume prediction system of the decomposition, reconstruction and optimization algorithm for ESN query includes: A data acquisition module, used to acquire a target time series and perform Z-score processing on the target time series; A first decomposition result module, used to decompose the target time series after Z-score processing by using a variational mode decomposition algorithm to generate a first decomposition result; A high-frequency component module, used to determine the high-frequency component in the first decomposition result according to the sample entropy; A second decomposition result module, used for decomposing the high frequency component by using a fully integrated empirical mode decomposition algorithm to generate a second decomposition result; A division module, used for dividing the first decomposition result into complex components according to the sample entropy, and dividing the second decomposition result into simple components; A time window module, used for extracting features from the complex component and the simple component through a preset sliding time window and determining a training set and a data set; An initialization module, used to construct a CNN-Transformer model according to the training set and initialize parameters; A prediction model module, used to optimize the initialized parameters using a sparrow optimization algorithm and train the CNN-Transformer model to generate a prediction model; A prediction result module, used to use the prediction model to predict the complex component to generate a first prediction result, and use the XGBoost model to predict the simple component to generate a second prediction result; The combined prediction module is used to generate a comprehensive result according to the first prediction result and the second prediction result and to perform a de-Z-score process on the comprehensive prediction result to obtain a combined prediction result.

9. A computer device, characterized in that: The device comprises: a memory and a processor, and when the processor runs the computer instructions stored in the memory, the processor executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The method comprises instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • User generated content emotion recognition method and device, terminal and medium

    CN120492632A