Diffusion model-based time series data prediction method and device, equipment and medium
By aligning and splicing multimodal data, combining the self-attention mechanism and pre-trained diffusion model, the problem of insufficient analysis of single modal data is solved, and more accurate time series data prediction is achieved.
Patent Information
- Application Number
- CN202510714263.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-19
AI Technical Summary
Existing diffusion models in the analysis of time series data in the financial and medical fields rely only on single-modal data, making it difficult to accurately capture complex time series characteristics and multi-factor interactions, resulting in inaccurate prediction results.
By obtaining target time series data and multimodal evaluation data, aligning the time dimension and then splicing the features, the self-attention mechanism is used to generate a self-attention feature matrix, and the pre-trained diffusion model is used to perform diffusion simulation to capture the time series dependency and multimodal feature association.
It improves the accuracy and reliability of time series data prediction, generates richer feature representations, and outputs more accurate target prediction data.
Smart Images

Figure CN120670754A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis technology, and in particular to a time series data prediction method, device, equipment and medium based on a diffusion model. Background Art
[0002] Time series data is widespread and valuable in numerous fields, such as financial market forecasting, meteorological data processing, and industrial equipment monitoring. Diffusion models, as an emerging analytical tool, can effectively mine complex patterns and trends in time series data, providing strong support for accurate modeling and forecasting.
[0003] In data analysis and prediction tasks in the financial and medical fields, data often exhibits characteristics such as multi-source heterogeneity, high dimensionality, strong noise, and complex time-series dependencies. Existing diffusion models rely solely on a specific type of data for analysis and prediction, such as using only historical price data in financial market forecasts or relying solely on medical records in medical diagnoses. This single-modality approach to data utilization is insufficient for exploring the long-term dependencies and dynamic evolution patterns of time-series data. It also struggles to accurately capture the complex time-series characteristics and multi-factor interactions of the data, resulting in inaccurate predictions. Summary of the Invention
[0004] The present invention provides a time series data prediction method, apparatus, computer equipment and medium based on a diffusion model to solve the technical problem that the existing diffusion model has insufficient analysis capability for time series data, resulting in inaccurate prediction results.
[0005] In a first aspect, a time series data prediction method based on a diffusion model is provided, comprising:
[0006] Obtain target time series data and multimodal evaluation data;
[0007] Aligning the target time series data with the multimodal evaluation data in time dimension, performing feature concatenation on the multimodal data corresponding to each time step, and obtaining a fused feature vector for each time step;
[0008] Based on the self-attention mechanism, the fusion feature vector of each time step is weighted summed to obtain the self-attention feature matrix;
[0009] Based on the pre-trained diffusion model, the self-attention feature matrix is subjected to diffusion simulation to output target prediction data.
[0010] In a second aspect, a time series data prediction device based on a diffusion model is provided, comprising:
[0011] Data acquisition module, used to obtain target time series data and multimodal evaluation data;
[0012] A feature fusion module is used to align the target time series data with the multimodal evaluation data in the time dimension, perform feature splicing on the multimodal data corresponding to each time step, and obtain a fused feature vector for each time step;
[0013] The self-attention matrix acquisition module is used to perform weighted summation of the fused feature vectors of each time step based on the self-attention mechanism to obtain the self-attention feature matrix;
[0014] The data prediction module is used to perform diffusion simulation on the self-attention feature matrix based on a pre-trained diffusion model and output target prediction data.
[0015] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned method for predicting time series data based on a diffusion model are implemented.
[0016] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned diffusion model-based time series data prediction method are implemented.
[0017] In the scheme implemented by the time series data prediction method, device, computer equipment and storage medium based on the above-mentioned diffusion model, by aligning the time dimension of the target time series data with the multimodal evaluation data, it is possible to ensure that the different modal data have consistency and correspondence in time. The multimodal data corresponding to each time step is feature concatenated to form a fused feature vector for each time step, which can integrate multimodal information and enrich the data features of each time step. The self-attention feature matrix is obtained by weighted summation of the fused feature vector based on the self-attention mechanism, which can capture the temporal dependencies in the feature sequence and the mutual correlations between multimodal features, and generate a richer feature representation. The diffusion simulation of the self-attention feature matrix using the pre-trained diffusion model can simulate the data generation process to capture the complex distribution and internal laws of the data, thereby outputting more accurate target prediction data and improving the accuracy of the prediction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0019] Figure 11 is a schematic diagram of an application environment of a time series data prediction method based on a diffusion model in an embodiment of the present invention;
[0020] Figure 2 A schematic flow chart of a first embodiment of a time series data prediction method based on a diffusion model provided by an embodiment of the present invention;
[0021] Figure 3 A schematic flow chart of a second embodiment of a time series data prediction method based on a diffusion model provided by an embodiment of the present invention;
[0022] Figure 4 1 is a structural diagram of a time series data prediction device based on a diffusion model in one embodiment of the present invention;
[0023] Figure 5 is a structural diagram of a computer device in one embodiment of the present invention;
[0024] Figure 6 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] The time series data prediction method based on the diffusion model provided by the embodiment of the present invention can be applied in the following fields: Figure 1 In an application environment, the client communicates with the server through a network. The server obtains target time series data and multimodal evaluation data from the client; aligns the target time series data with the multimodal evaluation data in time dimension, performs feature concatenation on the multimodal data corresponding to each time step, and obtains a fused feature vector for each time step; based on the self-attention mechanism, performs weighted summation on the fused feature vector for each time step to obtain a self-attention feature matrix; based on a pre-trained diffusion model, performs diffusion simulation on the self-attention feature matrix and outputs target prediction data. The server sends the target prediction data to the client for viewing by the user of the client.
[0027] In the present invention, in order to solve the technical problem that the existing diffusion model is insufficient in analyzing time series data in scenarios such as financial markets and medical diagnosis, resulting in inaccurate prediction results, the target time series data can be aligned with the multimodal evaluation data in the time dimension to ensure that the different modal data have consistency and correspondence in time. The multimodal data corresponding to each time step is spliced with features to form a fused feature vector for each time step, which can integrate the multimodal information and enrich the data features of each time step. The self-attention feature matrix is obtained by weighted summation of the fused feature vector based on the self-attention mechanism, which can capture the temporal dependencies in the feature sequence and the mutual correlation between the multimodal features, and generate a richer feature representation. The diffusion simulation of the self-attention feature matrix using the pre-trained diffusion model can simulate the data generation process to capture the complex distribution and internal laws of the data, thereby outputting more accurate target prediction data, effectively solving the technical problems of the existing diffusion model's insufficient analysis ability for time series data and inaccurate prediction results, and improving the accuracy and reliability of the prediction results.
[0028] The client can be, but is not limited to, various personal computers, laptops, smartphones, tablet computers, and portable wearable devices. The server can be implemented as an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0029] See also Figure 2 As shown, Figure 2 The flowchart of the first embodiment of the time series data prediction method based on the diffusion model provided by the embodiment of the present invention includes the following steps:
[0030] S101: Acquire target time series data and multimodal evaluation data;
[0031] In one embodiment, target time series data and multimodal evaluation data are collected, wherein the multimodal evaluation data is multimodal data such as text, images, and voice that is related to the target time series data or has a significant impact on changes in the target time series data.
[0032] For example, in the financial sector, for example, stock price time series data, such as historical stock price data including opening price, closing price, highest price, lowest price, and trading volume, can include textual information such as news articles, analyst reports, and trading volume data related to the stock market. In the medical sector, for example, for patient medical record time series data, such as chronologically recorded symptoms, diagnosis results, and treatment processes, multimodal evaluation data can include textual data such as patient medical record descriptions and medical reports, or image data such as patient medical images.
[0033] In one embodiment, after collecting the target time series data and multimodal evaluation data, data preprocessing can be performed on each modality based on the data type, including cleaning, noise removal, and outlier removal. For example, noise can be smoothed using methods such as sliding average for time series data, word segmentation and stop word removal can be performed on text data, and cropping and normalization can be performed on image data.
[0034] For example, in financial data, more advanced statistical methods or machine learning-based anomaly detection algorithms are used to identify and address abnormal stock price fluctuations. More sophisticated feature engineering is performed on multimodal data to extract more representative and discriminative features. For example, when processing text data, natural language processing techniques such as BERT, word embedding, and Transformer pre-trained language models are used to convert text data into feature representations in vector form; deep learning models such as convolutional neural networks (CNNs) are used to extract key features of images.
[0035] After completing the data preprocessing of the target time series data and multimodal evaluation data, normalization methods such as minimum-maximum normalization and Z-Score normalization can be used to normalize the data of different modalities to the same scale for subsequent fusion processing.
[0036] S102: Aligning the target time series data with the multimodal evaluation data in time dimension, performing feature concatenation on the multimodal data corresponding to each time step, and obtaining a fused feature vector for each time step;
[0037] In one embodiment, the preprocessed multimodal data are aligned in the time dimension to ensure that the modal data corresponding to each time step have a one-to-one correspondence.
[0038] At each time step, the feature vectors of the corresponding multimodal data are concatenated to form a fused feature vector. For example, if the feature dimension of the time series data is d t , the feature dimension of text data is d w , the feature dimension of the image data is d i , then the dimension of the concatenated fusion feature vector is d t +d w +d i .
[0039] The modal data are aligned in the time dimension, and then the feature vectors of different modalities are spliced at each time step to form a fused feature vector as the input of the model.
[0040] S103: Based on the self-attention mechanism, perform weighted summation on the fused feature vectors of each time step to obtain a self-attention feature matrix;
[0041] In one embodiment, a self-attention layer may be constructed before the input layer of the pre-trained diffusion model, and the output of the self-attention layer may be used as the input of the pre-trained diffusion model.
[0042] Specifically, the structure of the self-attention layer is designed, including the initialization of the query, key, and value matrices. The dimensions of these matrices should match the dimensions of the fused feature vector. For example, if the fused feature vector dimension is D, then the query matrix W Q , bond matrix W K Sum matrix W V The dimension can be set to D×D k , where D k is the feature dimension in the attention mechanism.
[0043] The fused feature vector of each time step is transformed linearly through the query matrix, key matrix and value matrix respectively to obtain the query vector Q, key vector K and value vector V.
[0044] Calculate the dot product between the query vector and the key vector to obtain the attention score matrix A. The self-attention score matrix can be expressed as:
[0045] A=Q·K T
[0046] Among them, each element in the attention score matrix represents the degree of correlation between different time steps.
[0047] Scale the attention score matrix, usually by dividing by the square root of the key vector dimension Compute with stable gradients:
[0048]
[0049] Apply the Softmax function to the scaled attention score matrix to convert it into the attention weight matrix W att Each element in the attention weight matrix represents the importance weight of the corresponding value vector in feature fusion, and the sum of all weights is 1.
[0050] Use the attention weight matrix to perform weighted summation on the value vector to obtain the weighted feature representation matrix H att :
[0051] H att =W att ·V
[0052] In this way, the feature vector of each time step integrates the feature information of other time steps and highlights the features that are more important to the prediction target.
[0053] When processing time series data, the attention mechanism enables the model to automatically focus on the time steps and feature dimensions that are more important to the prediction target, while reducing the interference of irrelevant information. For example, in financial market forecasting, the model can learn which economic indicators or historical price trends are more critical to future stock price predictions at different stages of the economic cycle, thereby allocating more attention resources and improving forecast accuracy.
[0054] S104: Based on the pre-trained diffusion model, perform diffusion simulation on the self-attention feature matrix and output target prediction data.
[0055] The weighted feature representation matrix H att Input into the pre-trained diffusion model. The diffusion model uses its internal neural network structure (such as convolutional layer, fully connected layer or Transformer layer, etc.) to train H att Perform feature extraction and map it into a potential feature space to obtain the initial feature representation.
[0056] For example, taking the pre-trained Transformer diffusion model as an example, the fused feature vector will be divided into feature sequences of several time steps. The feature vector of each time step is input into the Transformer encoder. After being processed by the multi-head self-attention mechanism and the feedforward neural network, the temporal dependencies in the feature sequence and the mutual correlations between multimodal features are captured to generate a richer feature representation.
[0057] Using the core idea of the diffusion model, the diffusion process of the weighted feature representation is simulated. In the trained diffusion model, the feature representation is gradually diffused in the feature space according to the learned diffusion law.
[0058] Specifically, noise can be added or other diffusion operations can be performed over multiple time steps, while the pre-trained diffusion model can recover the original features or target prediction features from the diffused feature states. For example, when predicting future time series data, starting from the initial fused feature vector, the diffused feature states are gradually generated according to the diffusion steps learned in the pre-trained diffusion model. The pre-trained diffusion model then predicts the target data for the future time steps based on these diffused feature states through the learned inverse diffusion process.
[0059] At each stage of the diffusion process, the features obtained by diffusion are combined with the features processed by the self-attention mechanism (H att ) or other related features. Through feature fusion mechanisms (such as weighted summation, feature concatenation followed by a fully connected layer, etc.), the feature information at different stages is fully integrated to obtain a comprehensive feature representation.
[0060] The integrated feature representation is input into the prediction layer (such as the fully connected layer, regression layer or classification layer, etc.), and the target prediction data is output according to the specific prediction task (such as predicting numerical stock prices, categorized disease types, etc.).
[0061] When making predictions, the pre-trained diffusion model not only inputs time series data, but also other relevant multimodal data, allowing the model to make predictions based on comprehensive fusion information, allowing different types of data to influence and diffuse each other in a common state space, thereby more comprehensively capturing the comprehensive characteristics of the data and improving the model's understanding and prediction capabilities of complex phenomena.
[0062] In one embodiment, the target prediction data is the prediction result obtained by the pre-trained diffusion model based on the comprehensive evaluation of the target time series data and multimodal evaluation data. It is the data value of a certain time step in the future, or the data change trend within a certain period of time in the future. When outputting the target prediction data, a visual data display option can be provided for the client user to select, and the target prediction data can be visualized to the user. The user can display the target prediction data in the form of an intuitive visual chart according to their own needs and usage scenarios, so as to have a clearer understanding of the prediction results.
[0063] Furthermore, based on the data type of the target prediction data, a visual display mode of the target prediction data is determined; based on the visual display mode, the target prediction data is formatted to achieve visual display of the target prediction data.
[0064] In one embodiment, data visualization refers to a technical means of displaying data in the form of intuitive charts, graphs, etc. Visualization can more clearly display the characteristics, trends and patterns of data, making it easier for users to understand and analyze.
[0065] Exemplarily, data visualization forms may include line charts, bar charts, pie charts, scatter plots, and the like.
[0066] For example, when the target prediction data is numerical, if it is a continuous numerical type, it can be visualized through a line graph or a scatter plot.
[0067] Line charts are suitable for displaying trends in data over time or other continuous variables. For example, to predict daily temperature changes over the next week, a line chart can clearly show rising or falling trends. Scatter charts are used to show relationships between data points when there are multiple influencing factors or when observing the distribution of data. For example, when predicting the relationship between product sales and advertising investment, a scatter chart can be used to show the correlation between the two.
[0068] For example, when converting to a line chart, the forecast data can be organized into a table with two main columns: one column for the value of the time or continuous variable (such as date, timestamp, etc.), and the other column for the corresponding forecast value. When converting to a scatter plot, the data is similarly organized into two columns: one column for the independent variable (such as advertising investment amount) and the other column for the dependent variable (such as product sales volume). In a scatter plot, each data point represents an observation or forecast sample, and these points are plotted to show the relationship between the variables.
[0069] If the target prediction data is a discrete numerical type, it can be visualized using a bar chart or pie chart.
[0070] Among them, bar charts can be used to visually compare the size of data across different categories or time points. For example, to display the projected sales volume of different products over the next month, a bar chart can clearly show the sales differences between products. Pie charts are suitable for showing the proportion of each component to the whole. For example, to predict the contribution of different regions to the total sales volume of a product in a certain quarter, a pie chart can visually show the proportion of each region.
[0071] For example, when converting a bar chart, create a table with one column containing the name of the category or time point (such as product name, month, etc.) and the other column containing the corresponding predicted value. In the visualization tool, map the category name to the horizontal axis and the predicted value to the vertical axis to generate a bar chart. When converting a pie chart, calculate the proportion of the predicted value of each category to the total, and draw the corresponding sector area in the pie chart based on these proportions. The angle of each sector represents the proportion of the category.
[0072] If the target forecast data is continuous, time-series data, a line chart can be used to show the data's changing trends over time. For example, to forecast a company's quarterly revenue for the coming year, a line chart can show both seasonal fluctuations and long-term trends. Alternatively, an area chart can be used, adding color to the line chart to more intuitively represent the data's cumulative or percentage distribution. For example, to forecast the distribution of daily website visits (e.g., visits from different marketing channels) for the coming week, an area chart can show how visits from each channel change over time and their proportion of total visits.
[0073] In this way, by choosing a suitable visualization method based on the data type of the target forecast data and converting the data to the appropriate format, we can achieve an intuitive and effective visualization of the forecast results. This helps users better understand and analyze the forecast data, providing strong support for decision-making.
[0074] It can be seen that in the above scheme, by aligning the target time series data with the multimodal evaluation data in the time dimension, the temporal consistency and correspondence of the different modal data can be ensured. The multimodal data corresponding to each time step is concatenated to form a fused feature vector for each time step. This can integrate multimodal information, enrich the data features of each time step, make up for the lack of information in a single modal data, and improve the representation ability of the features. Based on the self-attention mechanism, the weighted summation of the fused feature vectors is used to obtain the self-attention feature matrix. This can capture the temporal dependencies in the feature sequence and the interrelationships between multimodal features, highlight important feature information, weaken redundant feature interference, and generate a richer feature representation. Using a pre-trained diffusion model to perform diffusion simulation on the self-attention feature matrix can simulate the data generation process to capture the complex distribution and inherent laws of the data, thereby outputting more accurate target prediction data and improving the accuracy of the prediction results.
[0075] See also Figure 3 As shown, Figure 3 A flowchart of a second embodiment of a time series data prediction method based on a diffusion model provided by an embodiment of the present invention.
[0076] like Figure 3 As shown, based on the above Figure 2 In the embodiment shown, before step S104, the following steps are further included:
[0077] S201: Acquire historical time series data and historical multimodal evaluation data;
[0078] Historical time series data refers to a sequence of data collected in the past with a chronological order, such as historical stock price data, patient medical history time series records, etc.
[0079] Historical multimodal evaluation data refers to other types of data associated with historical time series data. These data come from different modalities (such as text, images, genes, etc.) and can provide supplementary information for the analysis and prediction of time series data.
[0080] For example, in the financial sector, historical stock price data (such as daily opening and closing prices, highest and lowest prices, and trading volume over the past few years) can be collected as historical time series data. At the same time, relevant news text data, analyst reports, macroeconomic indicators, etc. can be collected as historical multimodal evaluation data.
[0081] For example, in the medical field, we can collect patients' time-series medical records (such as changes in patients' symptoms, diagnosis results, treatment processes, etc.) and collect patients' medical imaging data (such as X-rays, CT scans), genetic data, laboratory test results, etc. as historical multimodal evaluation data.
[0082] It is understandable that the embodiments of the present application can also be applied to data analysis and prediction application scenarios in other fields, such as weather forecasting, power forecasting, etc. The financial field and medical field provided in the embodiments of the present application are only examples of specific application scenarios and are not intended to limit the application fields and application scenarios.
[0083] S202: Preprocess the historical time series data and the historical multimodal evaluation data based on a data processing algorithm to obtain model training data;
[0084] Data processing algorithms are used to perform preprocessing operations on data, such as cleaning, normalization, and feature extraction, in order to improve the quality and usability of data.
[0085] For example, for cleaning historical time series data and historical multimodal evaluation data in the financial sector, preprocessing operations can be performed based on the data type. For example, stock price data can be cleaned to remove obvious errors or abnormal price fluctuations (such as extreme values caused by data entry errors). For news text data, preprocessing operations such as word segmentation and stop word removal can be performed to extract the main semantic information of the text.
[0086] In medical application scenarios, medical record time series data can be checked to correct possible recording errors, while medical imaging data can be denoised to improve image quality.
[0087] After data cleaning, time series data and various assessment data are normalized to the same scale. For example, stock price data and trading volume data are converted to the [0, 1] range using the min-max normalization method. Medical data such as medical records and genetic data are also processed using appropriate normalization methods.
[0088] After data normalization, historical data from different modalities are aligned along the time dimension. For example, in the financial sector, this ensures a one-to-one correspondence between stock price data, news release times, and the recording times of macroeconomic indicators. In the medical field, this aligns the recording times of patient symptoms, imaging examinations, and genetic testing, all on the same timeline.
[0089] At each time step, the feature vectors of the aligned multimodal data are concatenated to form a fused feature vector, which serves as model training data.
[0090] Specifically, the historical time series data is recorded as T = {t1, t2, ..., t n}, where t iRepresents the data value at the i-th time point. Remove outliers and noise from the data. By setting a reasonable threshold, identify and correct or delete data points that significantly deviate from the normal range. For example, for a temperature time series data, if a value differs from the previous and next values by more than 10°C, it may be considered an outlier. Map the data to a specific interval, such as [0,1] or [-1,1]. For example, the minimum-maximum normalization method can be expressed as:
[0091]
[0092] Among them, t i is a value in the data set T, min(T) is the minimum value in the data set T, max(T) is the maximum value in the data set T, is the normalized value. The time series data T after cleaning and normalization norm .
[0093] S203: constructing an initial diffusion model;
[0094] In one embodiment, a diffusion mode of data points is defined; model parameters of an initial diffusion model are set to construct a model network structure of the initial diffusion model; and the diffusion mode of the data points is integrated with the model network structure to construct the initial diffusion model.
[0095] First, define the diffusion method of data points. The diffusion process of data points from time step k to k+1 can be set as:
[0096]
[0097] Among them, x k represents the data point at time step k, x k+1 represents the data point at time step k+1, ∈ k is a random variable that follows a standard normal distribution and is used to introduce randomness to simulate the diffusion characteristics of the data. Δt is the time step that controls the speed of diffusion.
[0098] Secondly, setting the model parameters of the initial diffusion model can include determining the hyperparameters of the diffusion model, such as the diffusion coefficient (which controls the diffusion rate) and the number of time steps (which determines the total duration of the diffusion process). The choice of these parameters affects the model's ability to capture data features and its computational complexity.
[0099] Next, a diffusion model structure based on a neural network is constructed. For example, a recurrent neural network (RNN) or its variant, the long short-term memory (LSTM) network, can be used to process time series information. Integrating the diffusion process into the network structure enables the network to learn the diffusion pattern of the data. In the LSTM unit, the diffused state can be used as part of the input and processed together with the original data.
[0100] Finally, the defined data point diffusion method is integrated with the constructed model network structure to form a complete initial diffusion model. For example, in the LSTM model, the input data at each time step is first diffused according to the diffusion formula to obtain the diffused data points, which are then input into the LSTM unit for processing, allowing the model to simultaneously learn the temporal characteristics and diffusion characteristics of the data.
[0101] S204: Training the initial diffusion model based on the model training data to obtain the pre-trained diffusion model.
[0102] In one embodiment, the model training data includes a training set and a validation set.
[0103] Specifically, the preprocessed time series data T norm Divide into training set T train and validation set T val For example, the data can be divided into 70% as a training set and 30% as a validation set.
[0104] Furthermore, the time series data in the training set are sequentially input into the initial diffusion model; based on the initial diffusion model, the data point state of the time series data at each time step is calculated, and the predicted data value of the target time step is output; the difference between the real data value corresponding to the target time point in the validation set and the predicted data value of the target time step is compared to obtain a model evaluation index; based on the model evaluation index, the initial diffusion model is optimized until the model loss value is less than a preset threshold, thereby obtaining the pre-trained diffusion model.
[0105] The weight parameters in the diffusion model are randomly initialized. The time series data in the training set are sequentially input into the initial diffusion model. For each sample, the input is a sequence consisting of fused feature vectors from multiple time steps.
[0106] According to the diffusion method defined in the initial diffusion model, the input time series data is diffused to simulate the diffusion state changes of the data points at the time step. The initial diffusion model extracts features and calculates the state of the diffused data through its network structure (such as LSTM). At each time step, the initial diffusion model calculates the state of the data point and passes it to the next time step. For example, for the input initial data point x0, x1, x2, ..., x are calculated through the diffusion formula. n .
[0107] Based on the processing and learning of the time series data by the initial diffusion model, a prediction result for the target time step is output. This prediction result can be a prediction of the data value at a certain point in the future or a reconstruction of the entire diffusion process, depending on the design goal of the model. For example, in the financial field, it can predict the stock price at a certain point in the future; in the medical field, it can predict the value of a physiological indicator or disease status of a patient at a certain point in the future.
[0108] Compare the actual data values corresponding to the target time point in the validation set with the predicted data values output by the initial diffusion model to calculate model evaluation metrics. Model loss can be calculated using loss functions, such as mean squared error (MSE), root mean squared error (RMSE), mean absolute error (MAE), and cross entropy loss.
[0109] For example, the mean square error (MSE) is calculated as:
[0110]
[0111] Where N is the number of samples, is the predicted value, y i is the true value.
[0112] Based on the model evaluation metrics, the initial diffusion model is optimized. Optimization methods include adjusting model parameters (such as the learning rate and diffusion coefficient), using different optimization algorithms (such as stochastic gradient descent (SGD) and the Adam optimization algorithm), and adjusting the model structure (such as increasing or decreasing the number of network layers or changing the number of neurons in a network layer).
[0113] For example, the gradient of the model parameters is calculated by back propagation algorithm according to the gradient of the loss function. An optimization algorithm such as stochastic gradient descent (SGD) or its variants Adagrad, Adadelta, etc. is used to update the model parameters according to the calculated gradient to reduce the loss value. For example, in SGD, the parameter update formula is
[0114]
[0115] Among them, θ t is the parameter of the current time step, η is the learning rate, is the loss function with respect to the parameter θ t gradient.
[0116] On the validation set T valEvaluate the model's performance on the dataset, monitoring the loss and other evaluation metrics (such as accuracy and root mean square error). If the model overfits (performs well on the training set but degrades on the validation set), optimize it by employing regularization methods (such as L1 or L2 regularization), increasing the amount of data, or adjusting the model structure. If the model underfits (performs poorly on both the training and validation sets), increase model complexity and adjust hyperparameters.
[0117] Through multiple iterations of training and optimization, until the model loss value is less than a preset threshold, a pre-trained diffusion model is obtained. The parameters of the trained pre-trained diffusion model are optimized to better capture the diffusion patterns and characteristics in time series data.
[0118] Through the above-mentioned detailed pre-training diffusion model training process, this embodiment can make full use of historical time series data and multimodal evaluation data, mine the complex patterns and dynamic characteristics in the data, and thus improve the accuracy and reliability of time series data prediction.
[0119] In one embodiment, after the pre-trained diffusion model is obtained, the pre-trained diffusion model may be further tested and optimized.
[0120] Furthermore, historical time series data and historical multimodal evaluation data within a preset period are collected as a test set; based on the pre-trained diffusion model, diffusion analysis is performed on the test set to obtain a test set prediction result; based on the denormalization operation, the prediction result is transformed to obtain a denormalized prediction result; the denormalized prediction result is compared with the corresponding real data in the test set to obtain a prediction error; based on the prediction error, the pre-trained diffusion model is retrained to optimize the pre-trained diffusion model.
[0121] In one embodiment, during the application of the pre-trained diffusion model, historical time series data and historical multimodal evaluation data of a preset period can be collected as a test set, and data preprocessing can be performed on the data in the test set. text As the input of the model, each input data point of the pre-trained diffusion model is a normalized feature vector.
[0122] The pretrained diffusion model processes input data based on learned diffusion patterns and parameters. Within the model, data points undergo diffusion transformations according to the defined diffusion process. The pretrained diffusion model captures the temporal dependencies and characteristic patterns of the data through its neural network architecture (such as LSTM or Transformer). The pretrained diffusion model outputs predictions for data at future time points. These predictions are in a normalized space and require further transformation to obtain predicted values in the original data space.
[0123] Specifically, the opposite operation of normalization is used to transform the prediction results Convert back to the original data space. The formula is as follows:
[0124]
[0125] Among them, max(T) and min(T) are the maximum and minimum values obtained from the model training data T, ensuring that the denormalization operation is consistent with the target time series data.
[0126] The prediction results after denormalization The true value of the test data Compare and calculate the evaluation indicators to quantify the prediction accuracy of the model. Common evaluation indicators include root mean square error (RMSE), mean absolute error (MAE), etc. For example, the formula of RMSE is:
[0127]
[0128] Where N is the number of test samples, is the predicted value of the i-th sample, is the true value of the i-th sample.
[0129] Furthermore, based on a visualization tool, a comparison graph of the denormalized prediction result and the real data is drawn; and based on the comparison graph, the prediction error is determined.
[0130] In one embodiment, a data visualization tool (such as Matplotlib, Seaborn, etc.) can be used to draw a comparison chart of the predicted results and the true values. The horizontal axis represents the time step or sample number, and the vertical axis represents the data value.
[0131] Comparison charts can visually demonstrate the model's prediction performance and show how close the predicted curve is to the true curve. Through visual analysis, you can quickly determine the model's ability to capture data trends and fluctuations.
[0132] Analyze the distribution of forecast errors. For example, calculate statistical indicators such as the mean and standard deviation of the errors to understand the concentration and fluctuation range of the errors. Plot a time series graph of the forecast errors to observe the trends of the errors over different time periods. If you find that the errors are larger in certain time periods, you can further analyze the data characteristics of these time periods.
[0133] For example, in the financial sector, unexpected market news, policy changes, or the release of economic indicators may cause unusual fluctuations in financial data such as stock prices. In these cases, you can check whether the model has adequately accounted for the impact of these external factors and whether further optimization is needed to better address such situations.
[0134] In the medical field, significant changes in physiological indicators and other data may occur due to special turning points in a patient's condition, such as the occurrence of complications or adjustments to treatment plans. In these cases, it is possible to analyze the model's shortcomings in handling complex disease changes and the interaction of multiple factors, and consider whether it is necessary to introduce more relevant medical data or improve the model structure to improve prediction accuracy.
[0135] Based on the analysis results, identify the model's shortcomings and potential areas for improvement. For example, if the model performs poorly in predicting long-term data trends, consider increasing the model's depth or introducing a more complex time series modeling structure. If the model is insensitive to data from certain modalities, optimize the data fusion strategy or adjust the weights of each modality. If errors are primarily concentrated in certain scenarios or data types, perform targeted data augmentation or model fine-tuning to address these situations.
[0136] Through the detailed testing and result analysis process described above, we can comprehensively evaluate the time series data prediction performance of the pre-trained diffusion model and provide a basis for further improvement of the pre-trained diffusion model. This will help to continuously improve the model's prediction accuracy and reliability in practical applications, and better meet the data analysis and decision support needs of fields such as finance and healthcare.
[0137] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0138] In one embodiment, a time series data prediction device based on a diffusion model is provided. The time series data prediction device based on a diffusion model corresponds one-to-one to the time series data prediction method based on a diffusion model in the above embodiment. Figure 4 As shown, the time series data prediction device based on the diffusion model includes: a data acquisition module 301, a feature fusion module 302, a self-attention matrix acquisition module 303 and a data prediction module 304. The functional modules are described in detail as follows:
[0139] Data acquisition module 301, used to acquire target time series data and multimodal evaluation data;
[0140] A feature fusion module 302 is configured to align the target time series data with the multimodal evaluation data in a time dimension, perform feature concatenation on the multimodal data corresponding to each time step, and obtain a fused feature vector for each time step;
[0141] A self-attention matrix acquisition module 303 is used to perform weighted summation on the fused feature vector of each time step based on the self-attention mechanism to obtain a self-attention feature matrix;
[0142] The data prediction module 304 is used to perform diffusion simulation on the self-attention feature matrix based on the pre-trained diffusion model and output target prediction data.
[0143] In one embodiment, the time series data prediction device based on the diffusion model further includes a model construction module, including:
[0144] A historical data acquisition unit, used to acquire historical time series data and historical multimodal evaluation data;
[0145] A data preprocessing unit, configured to perform data preprocessing on the historical time series data and the historical multimodal evaluation data based on a data processing algorithm to obtain model training data;
[0146] A model building unit, used to build an initial diffusion model;
[0147] A model training unit is used to train the initial diffusion model based on the model training data to obtain the pre-trained diffusion model.
[0148] In one embodiment, the model building unit includes:
[0149] The diffusion mode definition subunit is used to define the diffusion mode of data points;
[0150] A model parameter setting subunit is used to set the model parameters of the initial diffusion model and construct the model network structure of the initial diffusion model;
[0151] The initial diffusion model construction subunit is used to fuse the diffusion mode of the data points with the model network structure to construct the initial diffusion model.
[0152] In one embodiment, the model training data includes a training set and a validation set; the model training unit includes:
[0153] A data input subunit, configured to sequentially input the time series data in the training set into the initial diffusion model;
[0154] A data prediction subunit, configured to calculate the data point state of the time series data at each time step based on the initial diffusion model, and output a predicted data value at a target time step;
[0155] a difference comparison subunit, configured to compare the difference between the actual data value corresponding to the target time point in the validation set and the predicted data value of the target time step to obtain a model evaluation index;
[0156] The model optimization subunit is used to optimize the initial diffusion model based on the model evaluation index until the model loss value is less than a preset threshold, thereby obtaining the pre-trained diffusion model.
[0157] In one embodiment, the time series data prediction device based on the diffusion model further includes a model optimization module, including:
[0158] A test set acquisition unit is used to collect historical time series data and historical multimodal evaluation data within a preset period as a test set;
[0159] A test set analysis unit, configured to perform diffusion analysis on the test set based on the pre-trained diffusion model to obtain a test set prediction result;
[0160] an inverse normalization operation unit, configured to transform the prediction result based on an inverse normalization operation to obtain an inverse normalized prediction result;
[0161] a prediction error determination unit, configured to compare the denormalized prediction result with corresponding real data in the test set to obtain a prediction error;
[0162] A model optimization unit is used to retrain the pre-trained diffusion model based on the prediction error to optimize the pre-trained diffusion model.
[0163] In one embodiment, the prediction error determination unit includes:
[0164] A comparison graph drawing subunit, used for drawing a comparison graph between the denormalized prediction result and the real data based on a visualization tool;
[0165] The prediction error determination subunit is configured to determine the prediction error based on the comparison graph.
[0166] In one embodiment, the time series data prediction device based on the diffusion model further includes a data visualization module, including:
[0167] a display mode determining unit, configured to determine a visual display mode of the target prediction data based on a data type of the target prediction data;
[0168] A visualization display unit is used to convert the format of the target prediction data based on the visualization display mode to achieve visualization display of the target prediction data.
[0169] The present invention provides a time series data prediction device based on a diffusion model, which can ensure the consistency and correspondence of different modal data in time by aligning the time dimension of the target time series data with the multimodal evaluation data. The multimodal data corresponding to each time step is feature concatenated to form a fused feature vector for each time step, which can integrate the multimodal information and enrich the data features of each time step. The self-attention feature matrix is obtained by weighted summation of the fused feature vectors based on the self-attention mechanism, which can capture the temporal dependencies in the feature sequence and the mutual correlations between multimodal features, and generate richer feature representations. The diffusion simulation of the self-attention feature matrix using a pre-trained diffusion model can simulate the data generation process to capture the complex distribution and internal laws of the data, thereby outputting more accurate target prediction data and improving the accuracy of the prediction results.
[0170] For the specific definition of the time series data prediction device based on the diffusion model, please refer to the definition of the time series data prediction method based on the diffusion model above, which will not be repeated here. The various modules in the above-mentioned time series data prediction device based on the diffusion model can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0171] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a time series data prediction method based on a diffusion model.
[0172] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a time series data prediction method based on a diffusion model.
[0173] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0174] Obtain target time series data and multimodal evaluation data;
[0175] Aligning the target time series data with the multimodal evaluation data in time dimension, performing feature concatenation on the multimodal data corresponding to each time step, and obtaining a fused feature vector for each time step;
[0176] Based on the self-attention mechanism, the fusion feature vector of each time step is weighted summed to obtain the self-attention feature matrix;
[0177] Based on the pre-trained diffusion model, the self-attention feature matrix is subjected to diffusion simulation to output target prediction data.
[0178] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0179] Obtain target time series data and multimodal evaluation data;
[0180] Aligning the target time series data with the multimodal evaluation data in time dimension, performing feature concatenation on the multimodal data corresponding to each time step, and obtaining a fused feature vector for each time step;
[0181] Based on the self-attention mechanism, the fusion feature vector of each time step is weighted summed to obtain the self-attention feature matrix;
[0182] Based on the pre-trained diffusion model, the self-attention feature matrix is subjected to diffusion simulation to output target prediction data.
[0183] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0184] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0185] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0186] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A time series data prediction method based on a diffusion model, characterized in that: The method comprises: Obtain target time series data and multimodal evaluation data; Aligning the target time series data with the multimodal evaluation data in time dimension, performing feature concatenation on the multimodal data corresponding to each time step, and obtaining a fused feature vector for each time step; Based on the self-attention mechanism, the fusion feature vector of each time step is weighted summed to obtain the self-attention feature matrix; Based on the pre-trained diffusion model, the self-attention feature matrix is subjected to diffusion simulation to output target prediction data.
2. The time series data prediction method based on the diffusion model according to claim 1, characterized in that: Before acquiring the target time series data and the multimodal evaluation data, the method further includes: Obtain historical time series data and historical multimodal evaluation data; Based on a data processing algorithm, preprocess the historical time series data and the historical multimodal evaluation data to obtain model training data; Construct an initial diffusion model; The initial diffusion model is trained based on the model training data to obtain the pre-trained diffusion model.
3. The time series data prediction method based on the diffusion model according to claim 2, characterized in that: The constructing of the initial diffusion model includes: Define how data points are spread; Setting model parameters of the initial diffusion model and constructing a model network structure of the initial diffusion model; The diffusion mode of the data points is integrated with the model network structure to construct the initial diffusion model.
4. The method for predicting time series data based on a diffusion model according to claim 2, wherein: The model training data includes a training set and a validation set; The step of training the initial diffusion model based on the model training data to obtain the pre-trained diffusion model includes: inputting the time series data in the training set into the initial diffusion model in sequence; Based on the initial diffusion model, calculating the data point state of the time series data at each time step, and outputting the predicted data value of the target time step; Comparing the difference between the actual data value corresponding to the target time point in the validation set and the predicted data value of the target time step to obtain a model evaluation index; Based on the model evaluation index, the initial diffusion model is optimized until the model loss value is less than a preset threshold, thereby obtaining the pre-trained diffusion model.
5. The time series data prediction method based on the diffusion model according to claim 1 is characterized in that: After performing diffusion simulation on the self-attention feature matrix based on the pre-trained diffusion model and outputting target prediction data, the method further includes: Collect historical time series data and historical multimodal evaluation data within a preset period as a test set; Based on the pre-trained diffusion model, performing diffusion analysis on the test set to obtain a test set prediction result; Based on the denormalization operation, transforming the prediction result to obtain a denormalized prediction result; Comparing the denormalized prediction result with the corresponding real data in the test set to obtain a prediction error; Based on the prediction error, the pre-trained diffusion model is retrained to optimize the pre-trained diffusion model.
6. The time series data prediction method based on the diffusion model according to claim 5 is characterized in that: Comparing the denormalized prediction result with the corresponding real data in the test set to obtain a prediction error includes: Based on a visualization tool, a comparison graph of the denormalized prediction result and the real data is drawn; Based on the comparison graph, the prediction error is determined.
7. The time series data prediction method based on the diffusion model according to claim 1, characterized in that: After performing diffusion simulation on the self-attention feature matrix based on the pre-trained diffusion model and outputting target prediction data, the method further includes: Determining a visual display mode of the target prediction data based on a data type of the target prediction data; Based on the visual display mode, the target prediction data is formatted to achieve visual display of the target prediction data.
8. A time series data prediction device based on a diffusion model, characterized in that: The time series data prediction device based on the diffusion model includes: Data acquisition module, used to obtain target time series data and multimodal evaluation data; A feature fusion module is used to align the target time series data with the multimodal evaluation data in the time dimension, perform feature splicing on the multimodal data corresponding to each time step, and obtain a fused feature vector for each time step; The self-attention matrix acquisition module is used to perform weighted summation of the fused feature vectors of each time step based on the self-attention mechanism to obtain the self-attention feature matrix; The data prediction module is used to perform diffusion simulation on the self-attention feature matrix based on a pre-trained diffusion model and output target prediction data.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the time series data prediction method based on the diffusion model as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the time series data prediction method based on the diffusion model as claimed in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Streaming speech recognition method and device and electronic equipment
CN116741160A
Liver cancer longitudinal recurrence prediction and treatment effect evaluation system based on multi-modal fusion
CN120878240A
Multi-modal time series data generation method and device for model training, equipment and medium
CN121882122A
Method and apparatus for generating multi-modal time series data for model training, and medium
CN121882122B