Service prediction method and device based on time sequence model, equipment and storage medium
By performing Fourier transform and feature extraction on historical timing data, building a multi-layer time block model and performing feature fusion, the problem of difficulty in capturing complex timing patterns and periodic features in the prior art is solved, and higher business prediction accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510205077.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing business prediction methods are difficult to effectively capture the complex timing patterns and periodic characteristics in the data, resulting in low prediction accuracy. Especially in multi-dimensional and multi-scale business scenarios, the nonlinear relationship and long-term dependence of data make accurate prediction more difficult.
By obtaining historical timing data, preprocessing and fast Fourier transform, extracting multiple main frequency components and periodic values, recombining the training data set, building an L-layer time block model, combining feature mapping, attention mechanism and feedforward network, and establishing residual connections between adjacent time blocks, using a multi-branch convolutional network for feature extraction and weighted fusion, and training the timing model to output an analysis report.
This method can highlight the main periodic mode of data, effectively handle long sequence dependencies, realize adaptive integration of different frequency characteristics, and improve the accuracy and reliability of business prediction.
Smart Images

Figure CN120105102A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data analysis technology, and in particular to a business prediction method, device, equipment and storage medium based on a time series model. Background Art
[0002] With the advent of the big data era, enterprises are facing the challenges of analyzing and making decisions on massive amounts of business data. Traditional business forecasting methods often rely on manual experience and simple statistical models, which cannot effectively capture the complex time series patterns and periodic characteristics in the data, resulting in low prediction accuracy and difficulty in meeting the needs of modern enterprises for accurate business forecasting. Especially in multi-dimensional and multi-scale business scenarios, the nonlinear relationship and long-term dependence of data make accurate predictions more difficult. The current time series forecasting methods mainly have the following problems: First, the existing methods do not fully extract the periodic characteristics of the data, and it is difficult to effectively utilize the multi-scale periodic information in the data. The model structure is generally simple, lacking an effective fusion mechanism for multi-dimensional features, and it is difficult to fully explore the relationship between features. Traditional methods are prone to information loss and gradient vanishing problems when dealing with long sequence predictions, which affects the prediction performance of the model. These problems seriously restrict the accuracy and reliability of business forecasts, and there is an urgent need for a new forecasting method that can comprehensively consider data periodicity, feature correlation and long-term dependence. Summary of the invention
[0003] The present application provides a business prediction method, apparatus, device and storage medium based on a time series model, which are used to improve the accuracy and reliability of business prediction.
[0004] In a first aspect, an embodiment of the present application provides a service prediction method based on a time series model, the method comprising: Acquire historical time series data, and preprocess the historical time series data to obtain a training data set; Performing a fast Fourier transform on the training data set to obtain frequency domain data and frequency component intensity; Extracting a plurality of main frequency components and corresponding periodic values from the frequency domain data, and reorganizing and filling the training data set according to the periodic values to obtain a plurality of two-dimensional feature tensors; Constructing L layers of time blocks including feature mapping, attention mechanism and feedforward network according to the two-dimensional feature tensor, and establishing residual connections between adjacent time blocks to obtain a time series model to be trained; Applying a multi-branch convolutional network to each of the two-dimensional feature tensors to perform feature extraction to obtain a one-dimensional representation vector, and performing weighted fusion on the one-dimensional representation vector according to the intensity of the frequency component to obtain a fused feature; The time series model to be trained is trained by using the fusion features to obtain a trained time series model, and the trained time series model is used to output an analysis report of historical time series data.
[0005] In a second aspect, an embodiment of the present application provides a service prediction device based on a time series model, the device comprising: A data processing module is used to obtain historical time series data, pre-process the historical time series data, and obtain a training data set; A data transformation module, used for performing a fast Fourier transform on the training data set to obtain frequency domain data and frequency component intensity; A first extraction module is used to extract a plurality of main frequency components and corresponding periodic values from the frequency domain data, and to reorganize and fill the training data set according to the periodic values to obtain a plurality of two-dimensional feature tensors; A model construction module, used to construct L layers of time blocks including feature mapping, attention mechanism and feedforward network according to the two-dimensional feature tensor, and establish residual connections between adjacent time blocks to obtain a time series model to be trained; A second extraction module is used to apply a multi-branch convolutional network to each of the two-dimensional feature tensors to perform feature extraction to obtain a one-dimensional representation vector, and perform weighted fusion on the one-dimensional representation vector according to the intensity of the frequency component to obtain a fusion feature; The result output module is used to train the time series model to be trained through the fusion features to obtain a trained time series model, and the trained time series model is used to output an analysis report of historical time series data.
[0006] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the business prediction method based on the time series model as described in any one of the embodiments of the present application when executing the computer program.
[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor implements a business forecasting method based on a time series model as described in any one of the embodiments of the present application.
[0008] An embodiment of the present application provides a business prediction method based on a time series model, the method comprising: acquiring historical time series data, preprocessing the historical time series data, and obtaining a training data set; performing fast Fourier transform on the training data set to obtain frequency domain data and frequency component intensity; extracting multiple main frequency components and corresponding periodic values from the frequency domain data, and reorganizing and filling the training data set according to the periodic values to obtain multiple two-dimensional feature tensors; constructing L layers of time blocks including feature mapping, attention mechanism and feedforward network according to the two-dimensional feature tensor, and establishing residual connections between adjacent time blocks to obtain a time series model to be trained; applying a multi-branch convolutional network to each two-dimensional feature tensor for feature extraction to obtain a one-dimensional representation vector, and weightedly fusing the one-dimensional representation vector according to the frequency component intensity to obtain a fused feature; training the time series model to be trained by using the fused feature to obtain a trained time series model, and the trained time series model is used to output an analysis report of the historical time series data. Through the above method, the historical time series data is subjected to fast Fourier transform, the time domain data is converted into frequency domain data, the main frequency components and periodic values are extracted from the frequency domain data, and the training data is reorganized accordingly, which can highlight the main periodic patterns of the data. The L-layer time block structure is adopted, combined with feature mapping, attention mechanism and feedforward network, which can effectively handle long sequence dependencies, establish residual connections between adjacent time blocks, use multi-branch convolutional networks for feature extraction, and then perform weighted fusion on the one-dimensional representation vector according to the frequency component intensity, realizing the adaptive integration of different frequency features. The model and training data constructed by the above process can improve the accuracy and reliability of business forecasts. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.
[0010] Figure 1 A schematic flow chart of a business prediction method based on a time series model provided in an embodiment of the present application; Figure 2 A schematic block diagram of a business prediction device based on a time series model provided in an embodiment of the present application. DETAILED DESCRIPTION
[0011] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0012] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0013] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0014] It should be further understood that the term “and / or” used in the specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0015] See also Figure 1 , Figure 1 FIG. 1 is a schematic flow chart of a method for predicting a business transaction based on a time series model provided by an embodiment of the present application. Figure 1 As shown, the specific steps of the business prediction method based on the time series model include: S101-S106.
[0016] S101. Obtain historical time series data, preprocess the historical time series data, and obtain a training data set.
[0017] Exemplarily, historical time series data including traffic, order volume, sales volume, etc. are obtained from the enterprise database. These raw data are preprocessed comprehensively to obtain the training data set, which includes four key steps: first, use linear interpolation to handle missing values in the data, and estimate the value of the missing position by calculating the linear relationship between adjacent valid data points; second, handle outliers, set a dynamic threshold range based on data distribution, and use the moving average method to smooth data points that exceed the range; third, perform data standardization, and unify indicators of different dimensions to the same scale range by subtracting the mean and dividing by the standard deviation; fourth, slide segment the standardized data according to the preset time window size (such as 24 hours, 7 days, etc.), and each data segment contains the input feature sequence and the corresponding target prediction value.
[0018] S102, performing fast Fourier transform on the training data set to obtain frequency domain data and frequency component intensity.
[0019] Exemplarily, the data is windowed and segmented, and each data segment is weighted using a Hamming window function to reduce spectrum leakage. A discrete Fourier transform operation is performed on the weighted data segments to obtain spectral data in complex form. The amplitude spectrum is obtained by calculating the modulus of the spectral data, and the phase spectrum is obtained by calculating the argument of the complex number. These two parts together constitute a complete frequency domain representation. The amplitude spectrum is normalized and the energy distribution is calculated to obtain the intensity value of each frequency component, which reflects the importance of different periodic components in the original signal.
[0020] S103, extracting multiple main frequency components and corresponding periodic values from the frequency domain data, and reorganizing and filling the training data set according to the periodic values to obtain multiple two-dimensional feature tensors.
[0021] Exemplarily, the energy contribution of each frequency component is calculated, and K main frequency components whose energy contribution exceeds a preset threshold are selected to form a main frequency set. According to the sampling frequency and the frequency values corresponding to these main frequency components, the corresponding period value is calculated. These period values are used to reorganize the original training data set, and the one-dimensional time series data is rearranged into multiple two-dimensional matrices, where the rows of each matrix represent changes within the cycle and the columns represent changes between cycles. According to the requirements of subsequent convolution operations, these two-dimensional matrices are zero-filled to ensure that the requirements of the convolution kernel size and step size are met, and a standardized two-dimensional feature tensor is obtained.
[0022] S104. Construct L layers of time blocks including feature mapping, attention mechanism and feedforward network according to the two-dimensional feature tensor, and establish residual connections between adjacent time blocks to obtain a time series model to be trained.
[0023] Exemplarily, each time block contains four key components: the feature mapping layer maps the input features to a high-dimensional feature space; the multi-head attention layer calculates the attention weights in parallel through multiple attention heads to capture the long-distance dependencies in the sequence; the feedforward neural network layer performs nonlinear feature conversion; and the layer normalization layer maintains the stability of feature distribution. A residual connection is established between adjacent time blocks, so that the output feature of the l-th layer of time blocks is the sum of the features after the transformation of the current time block and the input features. This design effectively alleviates the gradient vanishing problem of deep networks. Through parameter configuration, the number of hidden units, the number of attention heads and other hyperparameters of each layer of time blocks are determined to form a complete time series model.
[0024] S105. Apply a multi-branch convolutional network to each two-dimensional feature tensor to perform feature extraction to obtain a one-dimensional representation vector, and perform weighted fusion on the one-dimensional representation vector according to the intensity of the frequency component to obtain a fused feature.
[0025] Exemplarily, the feature mapping results are batch normalized, and nonlinear transformations are introduced through the ReLU activation function. The spatial attention matrix and the channel attention matrix are calculated, and the comprehensive attention weight matrix is obtained through the tensor product operation to weight the importance of the features. The weighted features are subjected to maximum pooling and average pooling operations, and feature concatenation is performed on the channel dimension. According to the frequency component intensity obtained in step S102, the one-dimensional representation vector is weighted fused to obtain a fused representation that can reflect multi-scale features and periodic features.
[0026] S106. Train the time series model to be trained by fusing features to obtain a trained time series model. The trained time series model is used to output an analysis report of historical time series data.
[0027] Exemplarily, the fusion features are divided into training set, validation set and test set in a ratio of 8:1:1, and a loss function including prediction error term and model complexity penalty term is constructed. The Adam optimizer is used for model training, the initial learning rate is set to 0.001, and the EarlyStopping mechanism is used during training to prevent overfitting. When the validation set loss has not improved for 5 consecutive epochs, the learning rate decay mechanism is triggered. At the same time, the Dropout technology is used to randomly inactivate some neurons to enhance the generalization ability of the model. During the training process, the hyperparameters such as the number of time block layers, convolution kernel size, and learning rate are dynamically adjusted according to the performance indicators on the validation set. The trained model is used to analyze the historical time series data to generate a comprehensive analysis report including prediction results, trend analysis, and anomaly detection.
[0028] An embodiment of the present application provides a business prediction method based on a time series model, the method comprising: acquiring historical time series data, preprocessing the historical time series data, and obtaining a training data set; performing fast Fourier transform on the training data set to obtain frequency domain data and frequency component intensity; extracting multiple main frequency components and corresponding periodic values from the frequency domain data, and reorganizing and filling the training data set according to the periodic values to obtain multiple two-dimensional feature tensors; constructing L layers of time blocks including feature mapping, attention mechanism and feedforward network according to the two-dimensional feature tensor, and establishing residual connections between adjacent time blocks to obtain a time series model to be trained; applying a multi-branch convolutional network to each two-dimensional feature tensor for feature extraction to obtain a one-dimensional representation vector, and weightedly fusing the one-dimensional representation vector according to the frequency component intensity to obtain a fused feature; training the time series model to be trained by using the fused feature to obtain a trained time series model, and the trained time series model is used to output an analysis report of the historical time series data. Through the above method, the historical time series data is subjected to fast Fourier transform, the time domain data is converted into frequency domain data, the main frequency components and periodic values are extracted from the frequency domain data, and the training data is reorganized accordingly, which can highlight the main periodic patterns of the data. The L-layer time block structure is adopted, combined with feature mapping, attention mechanism and feedforward network, which can effectively handle long sequence dependencies, establish residual connections between adjacent time blocks, use multi-branch convolutional networks for feature extraction, and then perform weighted fusion on the one-dimensional representation vector according to the frequency component intensity, realizing the adaptive integration of different frequency features. The model and training data constructed by the above process can improve the accuracy and reliability of business forecasts.
[0029] In order to more clearly introduce the technical solution of the present application, the technical solution of the present application will be introduced through specific embodiments below. It should be noted that the specific embodiments are used to expand the technical solution of the present application, but are not intended to limit the present application.
[0030] In some embodiments, a fast Fourier transform is performed on a training data set to obtain frequency domain data and frequency component intensities, including: window segmentation and Hamming window weighting processing on the training data set to obtain a weighted data segment sequence; discrete Fourier transform operation is performed on the weighted data segment sequence to obtain spectral data represented in the complex domain; amplitude spectrum and phase spectrum are calculated based on the spectral data represented in the complex domain to obtain frequency domain data; the amplitude spectrum is normalized and the energy distribution is calculated to obtain frequency component intensities.
[0031] Exemplarily, the Hamming window is a weighted coefficient sequence in the form of a cosine function, which has a larger weight in the middle part of the data segment and gradually decreases to near zero at both ends. Through the Hamming window weighted processing, the continuous training data is divided into several shorter data segments, and the data points between adjacent segments overlap by 50%. Discrete Fourier transform (DFT) is performed on each weighted data segment to convert the time domain signal to the frequency domain. The result of DFT is a series of complex numbers, each of which corresponds to a specific frequency component, where the real part and the imaginary part represent the cosine and sine components of the frequency component respectively. These complex number spectrum data contain complete information of the original signal at each frequency point, including both amplitude information and phase information. Next, the complex number spectrum data is used to calculate the amplitude spectrum and phase spectrum. The amplitude spectrum represents the intensity of each frequency component, which is obtained by calculating the modulus of the complex number, and it reflects the relative strength of each frequency component in the signal. The phase spectrum represents the phase angle of each frequency component, which is obtained by calculating the argument of the complex number, and it reflects the phase relationship between different frequency components. These two spectra together constitute the complete frequency domain data of the signal, which can fully describe the frequency characteristics of the signal. Calculate the energy distribution of the normalized amplitude spectrum, that is, calculate the proportion of each frequency component in the total energy, which can be obtained by calculating the ratio of the square of the amplitude of each frequency point to the total sum of squares. The frequency component intensity obtained in this way reflects the distribution characteristics of the signal energy in the frequency domain, which helps to identify the main frequency components in the signal.
[0032] In some embodiments, multiple main frequency components and corresponding periodic values are extracted from the frequency domain data, and the training data set is reorganized and filled according to the periodic values to obtain multiple two-dimensional feature tensors, including: S1031-S1037.
[0033] S1031. Perform non-negative matrix decomposition on the frequency domain data to obtain a frequency component matrix and a coefficient matrix.
[0034] Exemplarily, the frequency domain data is represented as a non-negative matrix V, which is decomposed into a frequency component matrix W and a coefficient matrix H through iterative optimization, where V≈W×H, each column of the frequency component matrix W represents a basic frequency mode, and each row of the coefficient matrix H represents the weight coefficient of the corresponding mode.
[0035] S1032, calculating the energy contribution of each frequency component in the frequency component matrix, selecting a plurality of main frequency components whose energy contributions exceed a preset contribution threshold, and obtaining a main frequency set.
[0036] Exemplarily, the energy contribution of each column vector in the frequency component matrix W is calculated, and the energy contribution is measured by the product of the L2 norm of the column vector and the L2 norm of the row vector in the corresponding coefficient matrix H. All frequency components are sorted from large to small according to their energy contribution, an energy threshold (for example, 95% of the total energy) is set, and the first K frequency components whose cumulative energy contribution exceeds the threshold are selected. In this way, the most important periodic features in the signal can be retained, while filtering out noise components with smaller energy.
[0037] S1033. For each main frequency component in the main frequency set, a corresponding period value is calculated according to a reciprocal relationship between a preset sampling frequency and a frequency value corresponding to the main frequency component to obtain a period value set, where the preset sampling frequency is a frequency of a fast Fourier transform.
[0038] Exemplarily, for the selected K main frequency components, the corresponding period value T is calculated according to the signal sampling frequency fs and the frequency value f of each frequency component. The specific calculation formula is: T=fs / f, wherein fs is the sampling frequency (unit: Hz), f is the frequency value of the frequency component (unit: Hz), and the obtained T is the period value of the frequency component (unit: number of sampling points).
[0039] S1034. According to each period value in the period value set, the training data set is rearranged in segments according to the period length to obtain a one-dimensional sequence.
[0040] Exemplarily, for each calculated period value Ti, the original one-dimensional training data set is rearranged into a two-dimensional matrix. If the original sequence length is L and the period value is Ti, the number of rows and columns of the reorganized matrix is Ti and L / Ti. During the reorganization process, each Ti continuous data point is taken as a row and arranged in sequence to form a matrix, so that each column of the matrix represents a data point at the same period position.
[0041] S1035. Reorganize the one-dimensional sequence to obtain multiple initial two-dimensional feature matrices, where the number of rows of the initial two-dimensional feature matrix is equal to the period length, and the number of columns of the initial two-dimensional feature matrix is equal to the sequence length divided by the period length.
[0042] S1036, performing singular value decomposition on each initial two-dimensional feature matrix, extracting principal components and reconstructing the principal components to obtain K denoised feature matrices.
[0043] Exemplarily, singular value decomposition (SVD) is performed on each reorganized initial two-dimensional matrix to decompose the matrix into the product of three matrices U, Σ, and V. By analyzing the size of the singular values, the most significant r singular values and their corresponding singular vectors are selected to be retained, where r can be determined by setting the cumulative variance contribution rate (such as 95%). The r principal components are used to reconstruct the matrix to obtain K denoised feature matrices.
[0044] S1037. According to the convolution kernel size corresponding to the multi-branch convolutional network, the K denoised feature matrices are symmetrically filled, and filling values are added at the matrix boundaries so that the dimensions of the filled matrices meet the convolution operation requirements, and multiple two-dimensional feature tensors are obtained.
[0045] For example, if the convolution kernel size corresponding to the multi-branch convolutional network is k×k, then k / 2 rows or columns are filled at each of the four boundaries of the matrix. The filling value adopts a symmetrical filling method, that is, the value near the matrix boundary is copied to the filling area in a mirrored manner.
[0046] In some embodiments, L layers of time blocks including feature mapping, attention mechanism and feedforward network are constructed according to the two-dimensional feature tensor, and residual connections are established between adjacent time blocks to obtain a time series model to be trained, including: S1041-S1046.
[0047] S1041. Perform a linear projection transformation on the two-dimensional feature tensor to obtain a Q matrix, a K matrix, and a V matrix, where the dimension of the Q matrix is N×d, the dimension of the K matrix is N×d, and the dimension of the V matrix is N×d, where N is the input sequence length and d is the feature dimension.
[0048] S1042. Calculate the attention score matrix according to the Q matrix and the K matrix, perform softmax normalization on the attention score matrix, multiply the normalized attention weight matrix with the V matrix to obtain a weighted feature matrix, and the dimension of the weighted feature matrix is N×d.
[0049] For example, the attention score matrix is obtained by the dot product operation of the Q matrix and the K matrix. This matrix reflects the correlation between each time point in the sequence. The scores in the attention score matrix are normalized to between 0 and 1 through the softmax function to express it as a probability distribution. Then this normalized attention weight matrix is multiplied with the V matrix to obtain the weighted feature matrix.
[0050] S1043. Apply a multi-layer perceptron to perform nonlinear transformation on the weighted feature matrix, and add a normalization layer between each layer of perceptrons to obtain a time block output feature, where the dimension of the time block output feature is N×d.
[0051] For example, the weighted feature matrix obtained by the attention mechanism is input into a multi-layer perceptron for nonlinear transformation, and layer normalization is added between each layer of perceptrons to stabilize the network training process and prevent the gradient from disappearing or exploding. The role of layer normalization is similar to standardizing the data so that the output of each layer is kept within a suitable numerical range.
[0052] S1044. Construct a hierarchical structure consisting of L time blocks according to the output features of the time blocks, establish jump connections for the input and output features of adjacent time blocks, and obtain a residual network structure.
[0053] For example, L time blocks are stacked in sequence to form a deep network structure. Each time block contains the attention mechanism and nonlinear transformation components mentioned above. For example, if L=6, the data will be processed by 6 time blocks in sequence. A residual connection (also called a jump connection) is established between adjacent time blocks. The mathematical expression of this connection method can be written as: Output=F(x)+x, where x is the input feature and F(x) is the nonlinear transformation output of the current time block.
[0054] S1045. Apply a gated linear unit to each time block in the residual network structure for feature selection, and add a position code inside the time block to obtain an enhanced time block feature.
[0055] For example, the gated linear unit (GLU) divides the input features into two parts, one of which generates a gated signal (valued between 0-1) through the sigmoid function, and the other part is used as a candidate feature. After multiplying these two parts, soft selection of features can be achieved. Position encoding can use a combination of sine and cosine functions to generate a unique encoding vector for each position in the sequence. The elements of these encoding vectors will show regular changes with the change of position, allowing the model to perceive the relative position relationship between data points.
[0056] S1046. Input the enhanced time block features into a fully connected layer with learnable parameters for feature transformation to obtain a time series model to be trained.
[0057] For example, the fully connected layer is used to nonlinearly combine and map the high-dimensional features extracted previously. The fully connected layer contains multiple neurons, each of which is connected to all neurons in the previous layer to form a dense connection structure. The weights and biases of these connections are learnable parameters, which are continuously optimized during the training process through the back propagation algorithm. The design of this step needs to weigh the complexity and expressiveness of the model according to the specific task. For example, for simple trend prediction, only a simple fully connected layer may be needed, while for complex multivariate prediction, a deeper structure may be required.
[0058] In some embodiments, a hierarchical structure consisting of L time blocks is constructed based on the output features of the time blocks, and jump connections are established for the input and output features of adjacent time blocks to obtain a residual network structure, including: linearly transforming the output features of the time blocks according to preset input dimensions and output dimensions to obtain a transformed feature vector, the dimension of the transformed feature vector is N×2d, where N is the sequence length and d is the feature dimension; constructing a network skeleton with L time blocks based on the transformed feature vector, each time block includes a multi-head self-attention layer and a feedforward neural network layer, to obtain an initial hierarchical structure; performing element-level addition operations on the input features and output features of adjacent time blocks in the initial hierarchical structure, and performing layer normalization processing on the addition results to obtain a normalized feature matrix, and the normalized feature matrix The dimension is N×2d; the feature map of each time block is calculated according to the normalized feature matrix, and the feature map is processed by a nonlinear activation function to obtain the activated feature representation; an adaptive pooling operation is applied to the activated feature representation to map features of different dimensions to a unified dimensional space to obtain a pooled feature vector, and the dimension of the pooled feature vector is N×d; the weight coefficient of the residual connection is calculated according to the pooled feature vector, and the weight coefficient is weightedly combined with the original input feature to obtain a weighted residual feature; the weighted residual feature is input into a parallel structure composed of multiple sub-networks, each of which contains convolutional layers and pooling layers of different scales to obtain a multi-scale feature representation; the multi-scale feature representation is adaptively aggregated, and a dimensionality reduction transformation is performed through a fully connected layer to obtain a residual network structure.
[0059] Through a series of operations such as adjusting feature dimensions through linear transformation, processing time series information through multi-head self-attention mechanism, maintaining numerical stability through layer normalization, unifying feature dimensions through adaptive pooling, and parallel multi-scale feature extraction, the model's ability to express features of different time scales is enhanced. The residual connection is also used to effectively alleviate the gradient vanishing problem of deep networks. At the same time, the design of adaptive weights also enables the model to dynamically adjust the proportion of feature combinations according to data characteristics, thereby achieving more accurate and robust sequence feature extraction.
[0060] In some embodiments, a multi-branch convolutional network is applied to each two-dimensional feature tensor for feature extraction to obtain a one-dimensional characterization vector, and the one-dimensional characterization vector is weighted fused according to the intensity of the frequency component to obtain a fused feature, including: S1051-S1057.
[0061] S1051. Perform feature mapping on the two-dimensional feature tensor through three parallel convolution branches with a 1×1 convolution kernel, a 3×3 convolution kernel, and a 5×5 convolution kernel to obtain a feature mapping result, and perform batch normalization processing on the feature mapping result to obtain an initial feature map.
[0062] For example, a 2D feature tensor with a shape of 96×64 is input (96 represents the time step and 64 represents the feature dimension). This 2D feature tensor will be simultaneously extracted through three convolution kernels of different sizes: 1×1 convolution kernel is mainly used for feature dimension reduction and cross-channel information integration, 3×3 convolution kernel can capture local spatial dependencies, and 5×5 convolution kernel can obtain contextual information with a larger receptive field. After the convolution operation of each branch, a corresponding feature map will be obtained. The 1×1 convolution kernel obtains a 96×32 feature map result, and the 3×3 convolution kernel and the 5×5 convolution kernel respectively obtain 96×48 and 96×64 feature map results. These feature map results are batch normalized to normalize the data distribution to a normal distribution with a mean of 0 and a variance of 1, and obtain the standardized initial feature map.
[0063] S1052. Calculate the spatial attention matrix and the channel attention matrix according to the initial feature map, perform tensor product operation on the spatial attention matrix and the channel attention matrix to obtain an attention weight matrix, and weight the initial feature map according to the attention weight matrix to obtain a weighted feature map.
[0064] For example, taking the 96×64 feature mapping result as an example, the spatial attention matrix is 96×96, which represents the correlation weights between different time steps; the channel attention matrix is 64×64, which represents the dependency between different feature dimensions. The two attention matrices are tensor-producted to obtain a 96×64 attention weight matrix, where the value of each position represents the importance score of the spatiotemporal position. This attention weight matrix is then weighted with the initial feature map, which is equivalent to assigning different importance weights to each element in the feature map, highlighting the expression of important features, suppressing the influence of irrelevant features, and obtaining a weighted feature map.
[0065] S1053. Apply maximum pooling and average pooling operations to the weighted feature map, concatenate the pooling results in the channel dimension, perform dimensionality reduction through a fully connected layer, and obtain a reduced-dimensionality feature vector.
[0066] For example, the pooling operation will downsample in the time dimension. For example, if the pooling kernel size is 2 and the step size is 2, the time dimension of the feature map will be halved to 48. Maximum pooling and average pooling capture the most significant activation value and average activation intensity of the feature respectively. The two pooling results are concatenated in the channel dimension to obtain a 48×128 feature. Then, a fully connected layer is used to reduce the channel dimension from 128 to 64, and a 48×64 reduced-dimensional feature vector is obtained, which reduces the redundancy of features while retaining important information.
[0067] S1054. Construct an autoencoder network according to the reduced-dimensional feature vector, compress the reduced-dimensional feature vector into a low-dimensional latent space through an encoder, and then restore it to the original dimension through a decoder to obtain a reconstruction error.
[0068] For example, the autoencoder network contains two fully connected layers, which gradually compress the feature dimension from 64 to 32 and 16 to form a low-dimensional latent space representation. The decoder part is the opposite, and gradually restores the 16-dimensional features to 32 and 64 dimensions through two fully connected layers. The mean square error of the reconstructed features and the original reduced-dimensional feature vector is calculated to obtain the reconstruction error.
[0069] S1055. Filter the reduced-dimensional feature vector according to the reconstruction error to obtain a one-dimensional representation vector.
[0070] For example, according to the reconstruction error calculated in the previous step, the importance of each feature dimension in the 48×64 dimensionality reduction feature vector is evaluated. An adaptive threshold is set, such as taking the mean of the reconstruction error, and the features with reconstruction errors greater than the threshold are regarded as noise features. By resetting the weights of these noise features to 0, the key features with small reconstruction errors are retained, and automatic feature screening is achieved. The dimension of the one-dimensional representation vector obtained after screening is still 48×64, but some unimportant features have been filtered out, which improves the signal-to-noise ratio of the feature.
[0071] S1056. Perform softmax normalization processing on the frequency component intensity to obtain weight coefficients, and perform linear combination on the one-dimensional representation vectors according to the weight coefficients to obtain initial fusion features.
[0072] S1057, performing time series modeling on the initial fusion features through the residual connected gated recurrent unit, introducing an adaptive threshold mechanism in the hidden layer state of the gated recurrent unit for feature selection, and obtaining fusion features.
[0073] Exemplarily, the obtained initial fusion features are input into a gated recurrent unit, which includes three gating mechanisms: input gate, forget gate, and output gate. In the process of processing the sequence of initial fusion features, an adaptive threshold is introduced to control the feature selection strength at different time steps. The output fusion features retain multi-cycle information and have good time series modeling capabilities.
[0074] In some embodiments, the time series model to be trained is trained by fusing features to obtain a trained time series model, including: S1061-S1066.
[0075] S1061. Divide the fused features into a training sample set, a validation sample set, and a test sample set in a ratio of 8:1:1, and group the training sample set into batches of 32 samples to obtain training batches.
[0076] S1062. Construct a loss function based on the training batch and the validation sample set. The loss function includes a mean square error term, an L1 regularization term, and an L2 regularization term. The weight of the mean square error term is 0.6, the weight of the L1 regularization term is 0.2, and the weight of the L2 regularization term is 0.2, and the optimized objective function is obtained.
[0077] S1063. Apply the Adam optimization algorithm to update the parameters of the optimization objective function. The learning rate of the Adam optimization algorithm is 0.001, the beta1 parameter is 0.9, the beta2 parameter is 0.999, and the epsilon parameter is 1e-8 to obtain the parameter gradient.
[0078] S1064. Update the weights of the time series model according to the parameter gradient. During the updating process, randomly inactivate the neurons of the hidden layer with a probability of 0.5 to obtain the model parameters of the current round.
[0079] For example, in the actual training process, the model uses the Adam optimization algorithm to update the weight parameters of the neural network according to the parameter gradient. For example, for a time series model with 3 hidden layers and 128 neurons in each layer, the total amount of parameter gradients is about 500,000. The Adam optimizer will adaptively adjust the update step size of each parameter according to the set hyperparameters (learning rate 0.001, beta1 is 0.9, beta2 is 0.999, epsilon is 1e-8). At each parameter update, the 128 neurons in the hidden layer will be randomly inactivated with a probability of 0.5, that is, the output of about 64 neurons will be temporarily set to 0. This dropout mechanism can effectively prevent the model from overfitting.
[0080] S1065. Calculate the validation loss value for the validation sample set. When the validation loss value does not decrease for 5 consecutive rounds, the learning rate is decayed by 0.1 times. If it does not decrease for 10 consecutive rounds, terminate the training and obtain the candidate model.
[0081] For example, during the validation loss calculation phase, the model will use the validation sample set to evaluate the generalization performance of the current model after each round of training. Assuming that the validation set contains 1,000 samples, the average loss value of these samples will be calculated each time. The system maintains a queue of length 5 to record the validation losses of the most recent rounds. When it is found that the validation loss has not decreased for 5 consecutive rounds, the learning rate will be reduced from 0.001 to 0.0001, and then to 0.00001, and more refined parameter adjustments will be achieved by gradually reducing the learning rate. If the validation loss has not improved for 10 consecutive rounds, it means that the model has reached a convergence state. At this time, training will be stopped and the current model will be used as a candidate model.
[0082] S1066. Apply an ensemble learning method to the candidate model, and perform weighted averaging of the five model parameters with the lowest validation loss values during the training process according to the validation set performance scores to obtain a trained time series model.
[0083] For example, in the final ensemble learning phase, the model will select the five model checkpoints with the lowest validation loss from the entire training process. For example, these five checkpoints may come from the training states of the 100th, 150th, 200th, 250th, and 300th rounds. For each checkpoint model, its performance score (such as accuracy, F1 score, etc.) will be calculated on the validation set. Assuming that the performance scores of the five models are 0.92, 0.93, 0.91, 0.94, and 0.90, respectively, the system will calculate the normalized weights based on these scores, and then perform a weighted average on the parameters of the five models to obtain a trained time series model.
[0084] See also Figure 2 , Figure 2 1 is a schematic block diagram of a time series model-based business prediction device provided by an embodiment of the present application, wherein the time series model-based business prediction device 200 is used to execute the time series model-based business prediction method described above. The time series model-based business prediction device 200 can be configured in a server.
[0085] Among them, the server can be an independent server or a server cluster, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN) and big data and artificial intelligence platforms.
[0086] like Figure 2 As shown, the business prediction device 200 based on the time series model includes: a data processing module 201, a data transformation module 202, a first extraction module 203, a model construction module 204, a second extraction module 205 and a result output module 206.
[0087] The data processing module 201 is used to obtain historical time series data, pre-process the historical time series data, and obtain a training data set.
[0088] The data transformation module 202 is used to perform fast Fourier transformation on the training data set to obtain frequency domain data and frequency component intensity.
[0089] The first extraction module 203 is used to extract multiple main frequency components and corresponding periodic values from the frequency domain data, and reorganize and fill the training data set according to the periodic values to obtain multiple two-dimensional feature tensors.
[0090] The model construction module 204 is used to construct L layers of time blocks including feature mapping, attention mechanism and feedforward network according to the two-dimensional feature tensor, and establish residual connections between adjacent time blocks to obtain the time series model to be trained.
[0091] The second extraction module 205 is used to apply a multi-branch convolutional network to each two-dimensional feature tensor to perform feature extraction to obtain a one-dimensional characterization vector, and perform weighted fusion on the one-dimensional characterization vector according to the intensity of the frequency component to obtain a fused feature.
[0092] The result output module 206 is used to train the time series model to be trained by fusing features to obtain a trained time series model, and the trained time series model is used to output an analysis report of historical time series data.
[0093] The embodiment of the present application provides an electronic device, the electronic device comprising a memory and a processor. The memory is used to store a computer program. The processor is used to execute the computer program and implement any of the time series model-based business prediction methods in the embodiments of the present application when executing the computer program.
[0094] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor implements a business prediction method based on a time series model as described in any one of the embodiments of the present application.
[0095] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A business prediction method based on a time series model, characterized in that: The method comprises: Acquire historical time series data, and preprocess the historical time series data to obtain a training data set; Performing a fast Fourier transform on the training data set to obtain frequency domain data and frequency component intensity; Extracting a plurality of main frequency components and corresponding periodic values from the frequency domain data, and reorganizing and filling the training data set according to the periodic values to obtain a plurality of two-dimensional feature tensors; Constructing L layers of time blocks including feature mapping, attention mechanism and feedforward network according to the two-dimensional feature tensor, and establishing residual connections between adjacent time blocks to obtain a time series model to be trained; Applying a multi-branch convolutional network to each of the two-dimensional feature tensors to perform feature extraction to obtain a one-dimensional representation vector, and performing weighted fusion on the one-dimensional representation vector according to the intensity of the frequency component to obtain a fused feature; The time series model to be trained is trained by using the fusion features to obtain a trained time series model, and the trained time series model is used to output an analysis report of historical time series data.
2. The business prediction method based on the time series model according to claim 1, characterized in that: The performing fast Fourier transform on the training data set to obtain frequency domain data and frequency component intensity includes: Performing window segmentation and Hamming window weighting processing on the training data set to obtain a weighted data segment sequence; Performing a discrete Fourier transform operation on the weighted data segment sequence to obtain spectrum data represented in the complex domain; Calculate the amplitude spectrum and the phase spectrum according to the frequency spectrum data represented in the complex domain to obtain the frequency domain data; The amplitude spectrum is normalized and the energy distribution is calculated to obtain the frequency component intensity.
3. The business prediction method based on the time series model according to claim 1, characterized in that: The extracting of a plurality of main frequency components and corresponding periodic values from the frequency domain data, and reorganizing and filling the training data set according to the periodic values to obtain a plurality of two-dimensional feature tensors includes: Performing non-negative matrix decomposition on the frequency domain data to obtain a frequency component matrix and a coefficient matrix; Calculating the energy contribution of each frequency component in the frequency component matrix, selecting a plurality of main frequency components whose energy contribution exceeds a preset contribution threshold, and obtaining a main frequency set; For each main frequency component in the main frequency set, a corresponding period value is calculated according to a reciprocal relationship between a preset sampling frequency and a frequency value corresponding to the main frequency component to obtain a period value set, wherein the preset sampling frequency is the frequency of the fast Fourier transform; According to each period value in the period value set, the training data set is rearranged in sections according to the period length to obtain a one-dimensional sequence; Reorganizing the one-dimensional sequence to obtain a plurality of initial two-dimensional feature matrices, wherein the number of rows of the initial two-dimensional feature matrix is equal to the period length, and the number of columns of the initial two-dimensional feature matrix is equal to the sequence length divided by the period length; Performing singular value decomposition on each of the initial two-dimensional feature matrices, extracting principal components and reconstructing the principal components to obtain K denoised feature matrices; According to the convolution kernel size corresponding to the multi-branch convolutional network, the K denoised feature matrices are symmetrically filled, and filling values are added to the matrix boundaries so that the dimensions of the filled matrices meet the convolution operation requirements, thereby obtaining the multiple two-dimensional feature tensors.
4. The business prediction method based on time series model according to claim 1, characterized in that: The step of constructing L layers of time blocks including feature mapping, attention mechanism and feedforward network according to the two-dimensional feature tensor and establishing residual connections between adjacent time blocks to obtain a time series model to be trained includes: Performing a linear projection transformation on the two-dimensional feature tensor to obtain a Q matrix, a K matrix, and a V matrix, wherein the dimension of the Q matrix is N×d, the dimension of the K matrix is N×d, and the dimension of the V matrix is N×d, where N is the input sequence length and d is the feature dimension; Calculate an attention score matrix according to the Q matrix and the K matrix, perform softmax normalization on the attention score matrix, multiply the normalized attention weight matrix by the V matrix to obtain a weighted feature matrix, where the dimension of the weighted feature matrix is N×d; Applying a multi-layer perceptron to perform nonlinear transformation on the weighted feature matrix, and adding a normalization layer between each layer of perceptrons to obtain a time block output feature, wherein the dimension of the time block output feature is N×d; Constructing a hierarchical structure consisting of L time blocks according to the output features of the time blocks, establishing skip connections between the input and output features of adjacent time blocks, and obtaining a residual network structure; Applying a gated linear unit to each of the time blocks in the residual network structure to perform feature selection, and adding a position code inside the time block to obtain enhanced time block features; The enhanced time block features are input into a fully connected layer with learnable parameters for feature transformation to obtain the time series model to be trained.
5. The business prediction method based on time series model according to claim 4, characterized in that: The step of constructing a hierarchical structure consisting of L time blocks according to the output features of the time blocks, establishing skip connections between the input and output features of adjacent time blocks, and obtaining a residual network structure includes: Performing a linear transformation on the time block output feature according to a preset input dimension and output dimension to obtain a transformed feature vector, wherein the dimension of the transformed feature vector is N×2d, where N is the sequence length and d is the feature dimension; Constructing a network skeleton having L time blocks according to the transformed feature vector, each of the time blocks comprising a multi-head self-attention layer and a feedforward neural network layer, to obtain an initial hierarchical structure; Performing element-wise addition operation on input features and output features of adjacent time blocks in the initial hierarchical structure, and performing layer normalization processing on the addition result to obtain a normalized feature matrix, wherein the dimension of the normalized feature matrix is N×2d; Calculating a feature map of each time block according to the normalized feature matrix, and performing nonlinear activation function processing on the feature map to obtain an activated feature representation; Applying an adaptive pooling operation to the activated feature representation to map features of different dimensions to a unified dimensional space to obtain a pooled feature vector, wherein the dimension of the pooled feature vector is N×d; Calculate the weight coefficient of the residual connection according to the pooled feature vector, and perform weighted combination of the weight coefficient and the original input feature to obtain a weighted residual feature; Inputting the weighted residual features into a parallel structure composed of multiple sub-networks, each of which includes convolutional layers and pooling layers of different scales, to obtain multi-scale feature representation; Adaptive feature aggregation is performed on the multi-scale feature representation, and a dimensionality reduction transformation is performed through a fully connected layer to obtain the residual network structure.
6. The business prediction method based on time series model according to claim 1, characterized in that: The step of applying a multi-branch convolutional network to each of the two-dimensional feature tensors to extract features to obtain a one-dimensional characterization vector, and performing weighted fusion on the one-dimensional characterization vector according to the intensity of the frequency component to obtain a fusion feature comprises: Performing feature mapping on the two-dimensional feature tensor through three parallel convolution branches having a 1×1 convolution kernel, a 3×3 convolution kernel, and a 5×5 convolution kernel to obtain a feature mapping result, and performing batch normalization processing on the feature mapping result to obtain an initial feature map; Calculating a spatial attention matrix and a channel attention matrix according to the initial feature map, performing a tensor product operation on the spatial attention matrix and the channel attention matrix to obtain an attention weight matrix, and weighting the initial feature map according to the attention weight matrix to obtain a weighted feature map; Applying maximum pooling and average pooling operations to the weighted feature map, concatenating the pooling results in the channel dimension, and performing dimensionality reduction through a fully connected layer to obtain a reduced-dimensionality feature vector; An autoencoder network is constructed according to the reduced-dimensional feature vector, the reduced-dimensional feature vector is compressed into a low-dimensional latent space by an encoder, and then restored to the original dimension by a decoder to obtain a reconstruction error; The reduced-dimensional feature vector is screened according to the reconstruction error to obtain a one-dimensional representation vector; Performing softmax normalization processing on the intensity of the frequency component to obtain a weight coefficient, and performing linear combination on the one-dimensional representation vector according to the weight coefficient to obtain an initial fusion feature; The initial fusion feature is subjected to time series modeling through a residual-connected gated recurrent unit, and an adaptive threshold mechanism is introduced into the hidden layer state of the gated recurrent unit to perform feature selection to obtain the fusion feature.
7. The business prediction method based on time series model according to claim 1, characterized in that: The step of training the time series model to be trained by using the fusion features to obtain a trained time series model includes: The fusion features are divided into a training sample set, a verification sample set and a test sample set in a ratio of 8:1:1, and the training sample set is grouped into batches of 32 samples to obtain training batches; Constructing a loss function according to the training batch and the validation sample set, the loss function includes a mean square error term, an L1 regularization term, and an L2 regularization term, wherein the weight of the mean square error term is 0.6, the weight of the L1 regularization term is 0.2, and the weight of the L2 regularization term is 0.2, to obtain an optimized objective function; Applying the Adam optimization algorithm to the optimization objective function to update the parameters, the learning rate of the Adam optimization algorithm is 0.001, the beta1 parameter is 0.9, the beta2 parameter is 0.999, and the epsilon parameter is 1e-8, to obtain the parameter gradient; The weight of the time series model is updated according to the parameter gradient, and during the updating process, the neurons of the hidden layer are randomly deactivated with a probability of 0.5 to obtain the model parameters of the current round; Calculate the validation loss value for the validation sample set. When the validation loss value does not decrease for 5 consecutive rounds, perform a 0.1-fold decay process on the learning rate. If the validation loss value does not decrease for 10 consecutive rounds, terminate the training and obtain a candidate model. An ensemble learning method is applied to the candidate model, and the five model parameters with the lowest verification loss values during the training process are weighted averaged according to the verification set performance scores to obtain the trained time series model.
8. A business prediction device based on a time series model, characterized in that: The service prediction device based on the time series model comprises: A data processing module is used to obtain historical time series data, pre-process the historical time series data, and obtain a training data set; A data transformation module, used for performing a fast Fourier transform on the training data set to obtain frequency domain data and frequency component intensity; A first extraction module is used to extract a plurality of main frequency components and corresponding periodic values from the frequency domain data, and to reorganize and fill the training data set according to the periodic values to obtain a plurality of two-dimensional feature tensors; A model construction module, used to construct L layers of time blocks including feature mapping, attention mechanism and feedforward network according to the two-dimensional feature tensor, and establish residual connections between adjacent time blocks to obtain a time series model to be trained; A second extraction module is used to apply a multi-branch convolutional network to each of the two-dimensional feature tensors to perform feature extraction to obtain a one-dimensional representation vector, and perform weighted fusion on the one-dimensional representation vector according to the intensity of the frequency component to obtain a fusion feature; The result output module is used to train the time series model to be trained through the fusion features to obtain a trained time series model, and the trained time series model is used to output an analysis report of historical time series data.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the business prediction method based on the time series model as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the business prediction method based on a time series model as described in any one of claims 1 to 7.
Citation Information
Cited By
Fourier series time sequence prediction method based on KoImogorov-ArnoId theory
CN120448822A
Agricultural meteorological disaster time sequence prediction system and method based on multi-modal data fusion
CN120559760A
Porosity prediction method based on time-frequency characteristics
CN120630342A
Multi-feature fusion-based multi-element ocean observation data prediction method
CN120653940A
Welding quality diagnosis method and platform considering multi-sensing time sequence characteristics
CN120724291A