Establishment method of time sequence prediction model for industrial multi-modal data
Through multi-dimensional spatiotemporal mapping and improved SeqGAN model, multi-modal data integration and sparseness problems in industrial timing prediction are solved, and a prediction model for multi-modal industrial timing data is built, realizing accurate data prediction and intelligent decision support.
Patent Information
- Application Number
- CN202510282876.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-04
AI Technical Summary
Existing industrial timing prediction models cannot effectively integrate multimodal data from different sources and characteristics, and lead to inaccurate predictions in the absence of data sparsity and non-randomness, affecting production decisions and operational efficiency.
Using a heterogeneous data fusion perception method with multi-dimensional spatiotemporal mapping and an improved SeqGAN model, a prediction model for multimodal industrial timing data is constructed through data cleaning, enhancement, fusion and prediction model training, and data interpolation and prediction are used for AI big models.
Accurate prediction of multimodal industrial timing data is achieved, the generalization ability and data integrity of the model are improved, and intelligent decision-making and optimization of industrial processes are supported.
Smart Images

Figure CN120256815A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial time series prediction, and in particular to a method for establishing a time series prediction model for industrial multi-modal data. Background Art
[0002] In the industrial Internet of things, industrial data may include data from different sensors, different time resolutions, or different types (such as text, images, sounds). A large amount of such data often contains important information that can help workers understand and predict key indicators in industrial processes, thereby achieving accurate prediction and intelligent decision-making.
[0003] Industrial time series prediction refers to the process of using a large amount of historical industrial data collected from different sensors in the industrial field, with different time resolutions and covering various types such as text, images, and sounds. Through specific algorithms and models, the laws and information contained therein are mined, and the change trends of key indicators (such as production volume, equipment status, etc.) in the industrial process over a period of time in the future are inferred and estimated, so as to provide key basis for accurately predicting the industrial process trend and making intelligent decisions. This method is particularly important in Industry 4.0 and intelligent manufacturing, and can help enterprises optimize production processes, reduce downtime due to failures, and improve production efficiency. Existing industrial time series prediction models include statistical models such as AR, MA, ARMA, ARIMA, etc., machine learning models, deep learning models, etc. However, the current prediction models have the following problems:
[0004] First, due to the widely different structures of industrial time series data from different acquisition sources, existing industrial time series prediction models cannot effectively integrate data from different sources with different characteristics and time scales;
[0005] Second, in the complex data environment of the industrial Internet of things, the integrity of time series data is crucial for ensuring the accuracy of analysis and prediction. However, problems of information loss inevitably occur during the data acquisition process, especially for time series data collected from sensors and devices, and the loss of these data often exhibits non-random characteristics. Traditional interpolation techniques have problems with data interpolation estimation bias because they cannot align heterogeneous data. This bias may lead to inaccurate results in subsequent analysis and prediction, and further affect the entire production decision-making and operation efficiency. Especially in the industrial environment, the requirement for data accuracy is extremely high, and any small deviation may be amplified, causing serious consequences;
[0006] In summary, existing industrial time series prediction models face problems of data diversity, data sparsity, and insufficient model generalization ability. Therefore, there is currently a lack of a time series prediction method for industrial multi-modal data. Summary of the Invention
[0007] In view of the above-mentioned shortcomings of the prior art, the present invention proposes a prediction model for multi-modal industrial time series data, aiming to solve the problems of data diversity and sparsity faced by industrial time series prediction models, and to improve the generalization ability of the models.
[0008] The present invention adopts a method for establishing a time series prediction model for industrial multi-modal data to construct the above-mentioned prediction model, which includes the following steps:
[0009] Step 1: Obtain multi-source heterogeneous industrial time series data, and clean the obtained data;
[0010] Step 2: Perform data augmentation on the cleaned data to generate an augmented time series data set;
[0011] Step 3: Use a heterogeneous data fusion perception method based on multi-dimensional spatio-temporal mapping to fuse the multi-source heterogeneous time series data in the obtained augmented time series data set;
[0012] Step 4: Use the fused data to build a prediction model for multi-modal industrial time series data based on an AI large model, and perform model training;
[0013] Step 5: Predict the change trend of key indicators in the industrial process in the future for a period of time according to the trained prediction model for multi-modal industrial time series data.
[0014] Further, the specific operation steps of Step 2 include:
[0015] Step 2.1: Construct an improved SeqGAN model;
[0016] Step 2.2: Input the cleaned industrial time series data into the trained improved SeqGAN model to interpolate the missing sample data of the industrial time series data, and obtain enhanced multi-type industrial time series data, thereby constructing an augmented time series data set.
[0017] Further, the improved SeqGAN model in Step 2.1 is based on the original SeqGAN model architecture, adds a time-series dependent gated recurrent unit GRU to the generator G, introduces a self-attention mechanism based on time series features in the discriminator D, and calculates the optimization policy gradient through a reward function based on a time window. Finally, by introducing noise or perturbation to simulate the uncertainty in industrial data, an enhanced adversarial training strategy is used for model training.
[0018] Further, the reward function based on the time window is:
[0019] R window (t)=α1Rsim (t) + α2R smooth (t) + α3R periodic (t) - α4R abnormal (t)
[0020] where R window (t) represents the reward function based on the time window; α1, α2, α3, α4 represent weight parameters used to adjust the influence of each reward; R sim (t) represents the similarity reward function, which is used to measure the similarity between the generated data and the real data within the time window; R smooth (t) represents the smoothness reward function, which is used to measure the smoothness of the generated data and prevent drastic changes in the data within the time window; R periodic (t) represents the periodic consistency reward function; R abnormal (t) represents the anomaly detection reward function.
[0021] Furthermore, the heterogeneous data fusion perception method for multi-dimensional spatio-temporal mapping in step 3 includes the following steps:
[0022] Step 3.1: Register the industrial time series data obtained from different data sources in the enhanced time series dataset;
[0023] Step 3.2: Use the heterogeneous graph attention network to fuse the industrial time series data of each registered data source.
[0024] Furthermore, the specific steps of step 3.1 include:
[0025] Step 3.1.1: Extract the key features in their respective data sources from the time series data of different data sources, and use them as the registration control points for each data source;
[0026] Step 3.1.2: Preprocess the time series data in each data source;
[0027] Step 3.1.3: Set the target window centered on the registration control point, and obtain the feature points with the same name as the registration control point according to the target window;
[0028] Step 3.1.4: Based on the feature points with the same name, use the Delaunay triangulation algorithm to construct a dense triangular mesh;
[0029] Step 3.1.5: Use the differential operator to describe the deformation in the triangular mesh, and perform deformation compensation on each triangular mesh area to achieve the registration of multi-source heterogeneous data.
[0030] Furthermore, the specific steps of step 3.2 include:
[0031] Step 3.2.1: Use the registered industrial time series data as input and input it into the heterogeneous graph attention network;
[0032] Step 3.2.2: Map the industrial time series data into embedding vectors according to the following formula Realize data fusion:
[0033]
[0034] where, v i ∈V, v j ∈V, is the parameter learning matrix; w s x ∈R 1×d and w n x ∈R 1×d are two learnable parameters, is the correlation coefficient between node v i and its neighbor node v j , The specific form of is:
[0035]
[0036] where, σ is the activation function implemented by LeakyReLU, || represents the concatenation operation, a ∈ R 1×d is the trainable attention parameter.
[0037] Furthermore, the specific operation steps of Step 4 include:
[0038] Step 4.1: Obtain the industrial time series data after data augmentation and fusion, perform normalization preprocessing on it, and divide the industrial time series data into multiple small pieces, each small piece being a patch;
[0039] Step 4.2: Perform patch reprogramming based on the multi-head cross-attention mechanism, convert the behavior of the time series into natural language, and output the translated patch;
[0040] Step 4.3: Send the translated patch into the multi-head attention mechanism, and then perform linear projection to align the dimension of the reprogrammed patch with the dimension of the backbone of the AI large model;
[0041] Step 4.4: Use the translated patch as the prompt prefix of the input data and input it into the frozen AI large model;
[0042] Step 4.5: The AI large model uses the intrinsic knowledge to deeply analyze the time series data and generate prediction values;
[0043] Step 4.5: Map the predicted values output by the AI large model to the numerical space of the time series through a linear layer, so as to generate specific time series prediction values for future development.
[0044] Further, the specific steps of the patch reprogramming in Step 4.2 include:
[0045] Step 4.2.1: Obtain n patches of each industry time series;
[0046] Step 4.2.2: Generate corresponding query, key, and value vectors for each patch;
[0047] Step 4.2.3: Input the query vector of the time series patch and the key and value vectors of the natural language description into the multi-head cross-attention mechanism;
[0048] Step 4.2.4: In the multi-head cross-attention mechanism, calculate the attention weights of each head through the following formula:
[0049]
[0050] where Q, K, and V represent the query, key, and value vectors respectively, and d k is the dimension of the key vector;
[0051] Step 4.2.5: Concatenate the outputs of all heads and, through linear transformation, align the dimension with the dimension of the large model backbone:
[0052] MultiHead(Q, K, V) = Concat(head1,..., head h )W O
[0053] where W i Q , W i K , W i V and W O represent learnable parameter matrices; Concat represents the concatenation operation; head h represents the h-th head;
[0054] Step 4.2.6: Output the reprogrammed n patches.
[0055] Therefore, the present invention adopts the above method for establishing a time series prediction model for industrial multi-modal data, and has the following beneficial effects:
[0056] First, the present invention reconstructs the feature space of heterogeneous industrial time series data through a heterogeneous data fusion perception method of multi-dimensional spatiotemporal mapping, converts the heterogeneous data fusion problem into a modeling problem of feature kernel function and attention discriminant network, and can directly process the complex original distribution without variational lower bound, which greatly avoids the distribution deviation problem of training samples;
[0057] Second, the present invention proposes a semantically-aware time series data enhancement algorithm, which first uses the decision gradient method to learn the distribution of the original trajectory data and updates the status in real time; then trains samples that obey the true original trajectory data distribution as much as possible, and interpolates the missing sample data of the original trajectory; finally, uses statistical methods to estimate the variance of the interpolation statistic for the interpolated data. Next, a deep neural network architecture is used to extract low-level semantic information from multimodal fusion data sets such as equipment work log data and sensor data, and share a set of network parameters. High-level semantic information is then extracted using a high-level network, thereby solving the semantic gap problem. This allows the algorithm to fully perceive the semantic information in multimodal fusion data and mine the semantic features and laws therein to achieve data enhancement;
[0058] Third, the present invention constructs a time series prediction model for industrial multimodal data based on the AI big language model. The AI big model has reliable pattern recognition and reasoning capabilities when processing complex tag sequences. It can process and predict complex time series data in industrial environments by extracting and learning rich semantic features and rules in the data set, and provide accurate prediction results to support industrial decision-making and optimization.
[0059] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a technical framework diagram of the present invention.
[0061] Figure 2 Schematic diagram of semantic-aware embedding in heterogeneous graph attention networks.
[0062] Figure 3 It is a time series prediction method based on AI big model.
[0063] Figure 4 This is the schematic diagram of multi-head attention. DETAILED DESCRIPTION
[0064] In the description of the present invention, it should also be noted that unless otherwise clearly specified and defined, these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0065] The present invention proposes a method for establishing a time series prediction model for industrial multi-modal data, as Figure 1 shown, which includes the following steps:
[0066] Step 1: Data collection;
[0067] Obtain multi-source heterogeneous industrial time series data from data generated by sensors, device operation log environment data, and other industrial Internet of Things devices, and perform data cleaning on it;
[0068] The collected data includes four heterogeneous data sets: factory power consumption data set, weather data set, worker work data set, and raw material data set. These four data sets come from different acquisition devices and systems, have inconsistent timestamps, and different data types;
[0069] When performing data cleaning, first delete the outliers in the data, and then delete the records with missing values exceeding 1 / 2, so as to output the obtained multiple types of original time series data as cleaned multi-type time series data;
[0070] Step 2: Data augmentation;
[0071] First, construct an improved SeqGAN model, and then use the improved SeqGAN model to perform data augmentation on the cleaned data to generate an augmented time series data set;
[0072] Step 3: Data fusion
[0073] Based on the industrial time series data in the augmented time series data set, use the heterogeneous data fusion perception method of multi-dimensional spatio-temporal mapping to perform fusion perception on the obtained multi-source heterogeneous time series data;
[0074] Step 4, use the fused time series data to build a prediction model for multi-modal industrial time series data based on the AI large model;
[0075] Step 5, predict the change trend of key indicators in the industrial process in the future based on the prediction model for multi-modal industrial time series data, so as to guide industrial decision-making and optimization, such as realizing intelligent allocation of industrial resources, intelligent monitoring of industrial equipment, and intelligent scheduling of industrial production.
[0076] The key methods in the above steps are introduced as follows:
[0077] 1. Semantic-aware Temporal Data Augmentation Algorithm
[0078] Missing information is a common problem in industrial datasets. Filling in the missing information or generating a new dataset based on the original dataset can provide sufficient data guarantee for training an accurate prediction model subsequently. For this purpose, the present invention first constructs an improved SeqGAN model, and then uses the SeqGAN model to interpolate the missing sample data of the original industrial time series data, and outputs enhanced multi-type time series data. The specific steps are as follows:
[0079] (1) Data Integrity Construction Based on Generative Adversarial Networks
[0080] Generative Adversarial Networks (GAN) are often used for data integrity construction of a large amount of original industrial time series data collected in reality. GAN consists of a generator G and a discriminator D, and its objective function model is:
[0081]
[0082] Among them, Pdata(x) represents the sampling of x and the real data distribution, P z (z) represents the sampling of z and the prior distribution, and E(·) represents calculating the expected value.
[0083] When the generation target is discrete data, since the discrete output of the generator is difficult to transfer the gradient update from the discriminator D to the generator G, and the discriminant model can only evaluate a complete sequence and it is difficult to evaluate a partial generated sequence. To solve these problems, in the prior art, SeqGAN regards the generator model as a reinforcement learning process, uses the discriminator D to calculate the sequence score, and feeds the score back to the generator G in training. To solve the problem that the discrete output data cannot backpropagate the gradient to the generation model, SeqGAN regards the generation model as a randomly parameterized strategy. In policy gradient, the Monte Carlo method is used to search for approximate state-action values. By directly training the generation model using policy gradient, the discrimination difficulty of discrete data in traditional GAN is avoided.
[0084] Since industrial time-series data usually has complex temporal characteristics and strong dependencies, traditional SeqGAN may not be able to capture these characteristics when generating such data. Therefore, the present invention further improves the traditional SeqGAN model. By introducing temporal modeling and optimizing the policy gradient, it can improve the performance of the generator in constructing time-series data, ensure that the generated data is more coherent in time series, and thus enhance the integrity and accuracy of the generated data. The improvements of the improved SeqGAN model compared with the traditional SeqGAN model include:
[0085] First, in view of the characteristics of industrial time-series data (such as periodicity, temporality, and multidimensionality), a gated recurrent unit (GRU) with temporal dependencies is added to the generator G to enhance the model's ability to model time-series data, making it more suitable for processing industrial time-series data with periodicity, temporality, and multidimensionality;
[0086] Second, a self-attention mechanism based on temporal features is introduced for the discriminator D to help it better identify the patterns and anomalies in time-series data and improve the discriminative ability of the discriminator; at the same time, the calculation of the policy gradient is optimized, and a reward function based on a time window is introduced, enabling the model to consider the integrity and coherence of time-series data when generating data and improving the quality of the generated data;
[0087] Among them, the reward function based on a time window is a strategy optimized for time-series data generation models, aiming to ensure the integrity, coherence, and regularity of the generated data in the time series. In industrial time-series data, the data usually has strong time dependencies, periodicity, and multidimensionality. Therefore, the generated data needs to maintain a consistent trend and periodic pattern within a local time window, avoiding unreasonable jumps or fluctuations.
[0088] R window (t) = α1R sim (t) + α2R smooth (t) + α3R periodic (t) - α4R abnormal (t)
[0089] Among them, R window (t) is the reward function based on a time window; α1, α2, α3, α4 are weight parameters used to adjust the influence of each reward; R sim (t) is a similarity reward function used to measure the similarity between the generated data and the real data within a time window, usually calculated using the mean square error (MSE) or the mean absolute error (MAE):
[0090]
[0091] Among them, and respectively represent the values of the generated data and the real data at the \(i\)-th moment, and \(\Delta t\) represents the size of the time window;
[0092] R smooth (t) is the smoothness reward function, which is used to measure the smoothness of the generated data and prevent drastic changes in the data within the time window:
[0093]
[0094] R periodic (t) is the periodic consistency reward function, which is optimized by calculating the periodic consistency between the generated data and the real data:
[0095]
[0096] where \(P\) represents the period of the data. The higher the periodic consistency, the higher the reward;
[0097] R abnormal (t) is the anomaly detection reward function. By introducing an anomaly detection mechanism, it punishes the outliers in the generated data to avoid generating unreasonable data points:
[0098]
[0099] where \(I(\cdot)\) is the indicator function. When the generated data is determined to be an anomaly, it returns 1, otherwise it returns 0. Outliers will result in negative rewards, thus suppressing the model from generating abnormal data.
[0100] Finally, by introducing noise or perturbations to simulate the uncertainty in industrial data and adopting an enhanced adversarial training strategy, the diversity and robustness of the generated data are further improved.
[0101] One iteration of training the discriminator \(D\) requires the following steps: Extract \(m\) samples from the noise data \(z\) input, denoted as \(\{z 1 ,z 2 ,\cdots,z m \}, and take out \(m\) real sample data \(\{x 1 ,x 2 ,\cdots,x m \}\) from the real database \(x\). Take the partial derivatives of all parameters in the discriminator \(D\) using the objective function \(L(\theta)\) to obtain the gradient of each parameter. Since here we are training the discriminator \(D\), that is, \(D(x i )\) is as large as possible and \(D(G(z t ))\) is as small as possible, that is, the objective function is maximized, so an ascending gradient should be used here. Use the backpropagation algorithm and the batch gradient descent algorithm to optimize the discriminator \(D\), and its optimization function is:
[0102]
[0103] Among them, m is the number of real sample data;
[0104] Taking the real original time series data x and the noise data z as the input of the generator G, the generator G continuously learns the distribution of the original time series data (i.e., the time series data initially obtained through sensors), and uses the discriminator D to determine whether the input data comes from the original data or from the generator G. The generator G updates the state by adopting policy gradient and Monte Carlo search, and updates according to the expected score received from the discriminator D. Finally, the generator G trains samples that obey the distribution Pdata(x) of the real original time series data as much as possible to interpolate the missing sample data of the original industrial time series data.
[0105] Performing data registration and fusion on the enhanced multi-type industrial time series data according to the above method, and outputting the embedding vectors corresponding to the data nodes in the multi-type industrial time series after registration and fusion.
[0106] 2. Heterogeneous data fusion perception method for multi-dimensional spatio-temporal mapping
[0107] In order to achieve the fusion of multi-source heterogeneous time series data, the present invention analyzes the data from different time scales, first registers the heterogeneous data, and then fuses the heterogeneous data, so as to realize the integration of time series data from different sources with different characteristics, and finally form a data set rich in semantic information. The method includes the following steps:
[0108] (1) Registration of heterogeneous data
[0109] In order to solve the fusion problem of multi-dimensional heterogeneous features and multi-scale coupling correlation, it is necessary to first complete the registration of heterogeneous data. The registration of multi-source heterogeneous data refers to converting data with inconsistent spatial dimensions to the same dimension. And the fusion of multi-source heterogeneous data refers to extracting important features from data in multiple fields, enhancing data features and reducing redundant data at the same time.
[0110] The registration of multi-source heterogeneous data includes:
[0111] The first step is feature extraction, that is, extracting the key features from each data source in the industrial Internet of Things, and using the key features as the registration control points for each field (data source).
[0112] The second step is to preprocess the time series data of each field. The preprocessing includes data translation, flipping and scaling processing. Through preprocessing, the differences in the plane position, orientation and scale of different data can be reduced, which is convenient for subsequent data matching.
[0113] Step 3: Set the target window centered on the registration control points. The size of the target window can be set to the time window size corresponding to the time series data with the minimum sampling frequency, and then obtain the feature points with the same name as the registration control points according to the target window.
[0114] Specifically, first, based on prior knowledge, establish a search window for any one of the data sources. The search window can be set to twice the sampling frequency window of the data source. Then calculate the gray matrix of each sub-window with the same size as the target window in the search window, and then compare it with the gray matrix of the target window. Select the center of the sub-window with the most similar gray matrix to the target window gray matrix among all the calculated gray matrices as the feature points with the same name as the registration control points.
[0115] Step 4: Registration correction. Based on the feature points with the same name, use the Delaunay triangulation algorithm to construct a dense triangular mesh. When constructing the triangular mesh, the time series data needs to be extended to two-dimensional or three-dimensional space. The extracted feature points can be expressed as two-dimensional coordinates (t i , y i ), where t i is time and y i is the value. The coordinates (positions in space) of these feature points with the same name become the nodes of the triangular mesh.
[0116] Since the geometric deformation between data sources is reflected by the relative positions of the feature points with the same name, differential operators can be used to describe these deformations, such as gradient, curvature, distortion, etc. These deformations can be achieved by solving a local differential equation, and then the relative changes of each triangle can be found. By comparing the local deformation values (such as displacement, rotation, scaling, etc.) between the target data source and the reference data source, the geometric adjustment of each triangular mesh element will be accumulated into the overall registration result. By compensating for the deformation of each triangular mesh area, the geometric distortion between different data sources can be effectively reduced, and finally accurate multi-source heterogeneous data registration can be achieved.
[0117] Taking the registration of power consumption data as an example, register the power consumption data collected by different acquisition devices for the same observation target. First, register the same type of data. First, find the key features (such as maximum and minimum values) in each dataset, and then align the timestamps of all data sources to a unified time granularity (by hour). It is necessary to use an interpolation formula to interpolate and fill the coarse-grained data into fine-grained data. The interpolation formula is:
[0118]
[0119] where x′ t represents the interpolated value, are the integers before and after the time point.
[0120] Next, perform normalization preprocessing operations on the time series data in each field. The normalization formula is:
[0121]
[0122] where μ is the mean and σ is the standard deviation.
[0123] Again, set the target window and extract the corresponding feature points of the control points; then use moving average or wavelet transform to remove noise. In this invention, moving average is adopted, and the moving average formula is:
[0124]
[0125] where w represents the window size;
[0126] Again, align the feature points of each data source into a unified time window. The target window size is set to the time step of the data source with the minimum sampling frequency. Search for the point closest to the feature points of other data sources in the target window, and use the similarity metric formula to judge the proximity of two points:
[0127]
[0128] where ti and xi are the time and value of the data point Pi;
[0129] Again, map the time and value of all feature points to a two-dimensional space, construct a point set P = {(t i , x i )}, and then generate a Delaunay triangulation network containing all points in the constructed point set P, and record the edge and vertex information of each triangle.
[0130] The specific construction method of this Delaunay triangulation network is as follows:
[0131] 1) Obtain multi-source heterogeneous feature points and represent them as a two-dimensional coordinate point set:
[0132] P = {(t1, x1), (t2, x2),..., (t n , x n )}
[0133] where t i represents the time dimension and x i represents the eigenvalue;
[0134] 2) Randomly select three vertices of the initial triangle;
[0135] 3) Adjust the edge connection method (edge flipping) to ensure that the new triangulation network satisfies the maximum minimum angle property, that is:
[0136] If P i , P j , P k is the vertex of a triangle, and the point then the part needs to satisfy:
[0137] The point P l is not inside the circumcircle of △P i P j P k .
[0138] 4) Output the Delaunay triangulation containing all points in P.
[0139] Based on the Delaunay triangulation, describe the deformation in the triangulation according to the area of each triangle. The calculation method for the area of each triangle is as follows:
[0140] Represent the vertices of each triangle T as (t1, x1), (t2, x2), (t3, x3), then the triangle area A:
[0141]
[0142] Adjust the triangulation nodes according to the relative positions of the feature points to reduce error accumulation, and perform deformation compensation on each triangular grid area, so as to achieve multi-source heterogeneous data registration. The specific method is as follows:
[0143] 1) Calculate the geometric deviation △P between each feature point i :
[0144]
[0145] Among them, P i =(t i , x i ) represents the coordinates of the original feature point; t i represents the time dimension, x i represents the numerical dimension, represents the coordinates of the reference point, usually from the estimation of the homologous point of P i in other data sources.
[0146] 2) Then distribute the error △P of each feature point i to all triangular elements related to it. The distribution weight is based on the triangle area, and the influence degree of the feature point on the triangle is determined by the size of the area. The calculation of the error value is as follows:
[0147]
[0148] Among them, Ai represents the area of the triangle related to the feature point Pi Adjacent triangle area; ∑ j∈triangles A j Denotes the sum of the areas of all triangles adjacent to P i ; ΔT i Denotes the error value assigned to the triangular element.
[0149] Finally, the positions of all nodes are adjusted using the least squares method to minimize the registration error:
[0150]
[0151] After the registration of data of the same type is completed, registration is also required between different types of data. Since the power consumption data is the main data for this prediction and its collection frequency is once per hour. Therefore, it is necessary to set the data node frequency in other data sets to be the same as that of the power consumption data, that is, there is a sampling data every hour, and the specific value of each sampling data is set using average value interpolation, and finally multi-source heterogeneous data registration is achieved.
[0152] (2) Heterogeneous data fusion
[0153] After the registration of heterogeneous data is completed, the features of multi-source heterogeneous data are tiled in a suitable space using a heterogeneous information network, and it is transformed into the expression of a similarity matrix, so as to solve the fusion problem of data expression and fully explore the multi-level correlations existing between industrial multi-modal data;
[0154] Specifically, the present invention uses the multi-layer mapping strategy in the existing heterogeneous graph attention network to achieve vector mapping and semantic perception embedding, Figure 2 which is the embedding strategy for the semantic perception process. Denote the number of mapping layers of the heterogeneous graph attention network as the variable S, and from Figure 2 it can be seen that at the x-th mapping layer, is the embedding vector of node v j , and the mapping formula of is:
[0155]
[0156] where, v i ∈V, v j ∈V, is the parameter learning matrix;
[0157] Assume that N i is the set of neighbor nodes directly connected to node v i in the heterogeneous network G. Based on the mapping vector i of node v at the x-th layer, the embedding vector value of node v i at the x + 1-th layer can be obtained:
[0158]
[0159] where w sx ∈ R 1×d and w nx ∈ R 1×d are two learnable parameters, is the correlation coefficient between node v i and its neighbor node v j and its specific form is: The specific form of
[0160]
[0161] where σ is an activation function implemented by LeakyReLU, || represents the concatenation operation, and a ∈ R 1×d is a trainable attention parameter.
[0162] Finally, the embedding vector value h of each element in the registered time series can be obtained. This vector value overcomes the heterogeneity between multimodal data and retains the correlation between heterogeneous data.
[0163] 3. Prediction Model for Multimodal Industrial Time Series Data
[0164] The original design intention of large AI models is to understand and generate natural language used by humans, rather than directly processing time series data with continuous numerical characteristics. During their pre-training stage, they mainly learn language structures and semantic information, and are not specifically optimized for the statistical characteristics of time series. Therefore, they lack an in-depth understanding of common patterns in industrial time series data, such as trends, periodicity, etc. Moreover, large AI models cannot directly process long sequence data by recursively using the results of their own predictions for multi-step prediction. Therefore, large AI models cannot be directly applied to industrial time series data processing.
[0165] In order to utilize the capabilities of large AI models in feature extraction and knowledge transfer learning and apply them to industrial time series data processing, there are two ways as follows: One is to train a brand-new large AI model specifically for time series data; the other is to perform time series prediction based on existing open-source large AI models. Considering the computing power and time required for the training process of large AI models, the present invention uses an open-source pre-trained large AI model (LLaMA 2 - 7B) to achieve time series prediction. By aligning industrial continuous data with natural language discrete data, for example, converting time series into text description forms, large AI models can be used for specific time series and spatio-temporal tasks. By processing numerical sequence data across text modalities with large AI models, rich semantic features and patterns in the dataset are extracted and learned, and powerful reasoning and prediction capabilities are exerted in time series.
[0166] Specifically, the present invention reprograms the input time series data using a text prototype, represents the semantic information of the time series data using natural language representations, and then aligns the two different data modalities, so that the large language model can understand the information behind another data modality without any modification. Based on this, as Figure 3 shown, the method for predicting industrial time series data based on an AI large model includes the following steps:
[0167] (1) Input transformation
[0168] The time series data after data augmentation is imported into the system and undergoes necessary preprocessing, such as normalization, to ensure data consistency and stability. Then, the time series data is sliced into multiple small pieces, each small piece is called a "patch", so that the continuous time series data can be processed by the AI large model in a similar way to text processing. This transformation not only makes the data format suitable for the input requirements of the AI large model, but also helps to capture local patterns in the time series.
[0169] (2) Patch reprogramming
[0170] Align all the patches sliced from the time series with the same pre-trained text prototype (the BERT language model in the prior art), and encode the behavior of the time series as a natural language description, so that it can be naturally integrated into the language space of the AI large model.
[0171] Specifically, by learning the mapping relationship between the time series data and the natural language description, the AI large model can utilize its pre-trained knowledge in natural language understanding to process the time series data. Patch reprogramming is based on the multi-head cross-attention mechanism, which specifically includes:
[0172] Step 1: Obtain n patches of each time series;
[0173] Step 2: Generate corresponding query (Query), key (Key), and value (Value) vectors for each patch;
[0174] Step 3: Input the query vectors of the time series patches and the key and value vectors of the natural language description into the multi-head cross-attention mechanism;
[0175] Step 4: In the multi-head cross-attention mechanism, calculate the attention weights of each head, and the formula is as follows:
[0176]
[0177] where Q, K, and V represent the query, key, and value vectors respectively, and d k is the dimension of the key vector;
[0178] Step 5: Concatenate the outputs of all heads and align the dimensions with those of the large model backbone through a linear transformation. The formula is as follows:
[0179] MultiHead(Q, K, V) = Concat(head1,..., head h )W O (7)
[0180] where head i = Attention(QW i Q , KW i K , VW i V ), and W i Q , W i K , W i V and W O are learnable parameter matrices;
[0181] Step 6: Output the reprogrammed n patches as the input of the model.
[0182] Through the above steps, the model can effectively map time series data with natural language descriptions by using the multi-head cross-attention mechanism, so as to process time series data in the pre-trained language model.
[0183] Patch reprogramming allows the model to dynamically select and combine different text prototypes to best represent the input time series data. This step is the most important step and the key to achieving cross-modal data alignment. In this way, we effectively encode the behavior of the time series as natural language input. The reprogrammed patches translated in step 6 above are sent to the multi-head cross-attention mechanism (Equation 6), and then linearly projected to align the dimensions of the reprogrammed patches with those of the large model backbone, where the principle of the multi-head attention mechanism is as Figure 4 shown.
[0184] The multi-head attention model is designed in the Query-Key-Value (QKV) mode. First, map the input sequence X to three different vector spaces to obtain vectors Q, K, and V respectively, where Q = W q X, K = W k X, V = W v X. Then, calculate the weight matrix M using the Scaled Dot-Production:
[0185]
[0186] Then, the output vector H is obtained by using the weight matrix M and the vector V:
[0187]
[0188] Finally, the output vectors h of all heads are aggregated by the following formula:
[0189] head i = Attention(Qw i Q , Kw i k , Vw i v )
[0190] MultiHead(Q, K, V) = concat(head1, …, head h )w O (10)
[0191] where w i Q , w i k , w i v , w O are projection matrix parameters.
[0192] (3) Prompt Prefix
[0193] Based on the literature "PromptCast: A New Prompt-based Learning Paradigm for TimeSeries Forecasting", the design idea of the "prompt prefix" of the present invention is to use the time series data that is reprogrammed from patches and converted into natural language prompts as the prefix of the input data, providing additional context information and task instructions for the model. These prompts not only enhance the understanding of time series data by the AI large model, but also guide the AI large model on how to reason according to the given tasks. For example, the prompts can include statistical information about the time series, trend descriptions, or specific prediction instructions, thereby helping the AI large model to more accurately capture the characteristics of the time series and the prediction target.
[0194] (4) Model Inference
[0195] At this stage, the reprogrammed and enhanced input data is fed into the frozen large AI model. Utilizing the powerful pattern recognition and reasoning capabilities of the large AI model while avoiding adjusting the model weights, this reduces the training cost and maintains the stability of the model. In this way, the large AI model can utilize its inherent knowledge to conduct in-depth analysis of time series data and generate predicted values.
[0196] (5) Output transformation
[0197] Convert the output of the large AI model into the final time series prediction.
[0198] Map the output of the model from the language space of the large AI model back to the numerical space of the time series. Through a simple linear layer, project and transform the output of the model to generate specific numerical predictions for future development.
[0199] The formula for the linear transformation performed by the linear layer is:
[0200] Given an input vector x and a weight matrix W, the output y of the linear transformation can be expressed by the following formula:
[0201] y = Wx + b (11)
[0202] Where, is the input vector, with dimension n; is the weight matrix, with dimension m×n, where m is the output dimension and n is the input dimension; is the bias vector, with dimension m; is the output vector, with dimension m;
[0203] Embodiment
[0204] To verify the feasibility of the present invention, this embodiment conducts experiments on industrial electricity consumption datasets and industrial water consumption datasets. Randomly select 80% of the experimental data as the training set, and the remaining 20% as the test set; set the output size of the embedding layer to 10; the number of model training times is 1000; the initial value of the learning rate is 0.001; the batch size for model training is 64; the model uses the Adam algorithm as the optimizer; at the same time, an early stopping strategy is set. When the loss of the test set is equal to the loss of the training set, the model training reaches the optimum, that is, stop training to prevent the model from overfitting. Three commonly used evaluation indicators are selected to evaluate the experimental results, namely: Recall, F1-score, and Mean Average Precision value.
[0205] Table 1 shows the performance comparison results of the algorithm proposed in the present invention and other baseline methods (including: LSTM, GRU, STGCN, ASTGCN, GMAN, LAMA2, ST-LLM) on the industrial electricity consumption dataset and the industrial water consumption dataset. The results of all algorithms are the average of ten test results. Compared with the baseline algorithms, the algorithm of the present invention has the optimal values in terms of hit rate, F1, and MRR metrics. The experimental results demonstrate the effectiveness of the algorithm of the present invention.
[0206] Table 1 Comparison Results between the Present Invention and Baseline Algorithms
[0207]
[0208] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for establishing a time series prediction model for industrial multi-modal data, characterized in that, It includes the following steps: Step 1: Obtain multi-source heterogeneous industrial time series data and clean the obtained data; Step 2: Perform data augmentation on the cleaned data to generate an augmented time series data set; Step 3: Use a heterogeneous data fusion perception method based on multi-dimensional spatio-temporal mapping to fuse the multi-source heterogeneous time series data in the obtained augmented time series data set; Step 4: Use the fused data to construct a prediction model for multi-modal industrial time series data based on an AI large model and perform model training; Step 5: Predict the change trend of key indicators in the industrial process in the future for a period of time according to the trained prediction model for multi-modal industrial time series data.
2. The method for establishing a time series prediction model for industrial multi-modal data according to claim 1, characterized in that, The specific operation steps of Step 2 include: Step 2.1: Construct an improved SeqGAN model; Step 2.2: Input the cleaned industrial time series data into the trained improved SeqGAN model to interpolate the missing sample data of the industrial time series data, and obtain enhanced multi-type industrial time series data, thereby constructing an augmented time series data set.
3. The method for establishing a time series prediction model for industrial multi-modal data according to claim 2, wherein The improved SeqGAN model in Step 2.1 is based on the original SeqGAN model architecture. A gated recurrent unit GRU with time series dependence is added to the generator G, and a self-attention mechanism based on time series features is introduced into the discriminator D. The optimization policy gradient is calculated through a reward function based on a time window. Finally, by introducing noise or perturbations to simulate the uncertainty in industrial data, an enhanced adversarial training strategy is used for model training.
4. The method for establishing a time series prediction model for industrial multi-modal data according to claim 3, wherein, The reward function based on a time window is: R window R(t) = α1R sim (t) + α2R smooth (t) + α3R periodic (t) - α4R abnormal (t) where, R window (t) represents the reward function based on the time window; α1, α2, α3, α4 represent the weight parameters used to adjust the influence of each reward; R sim (t) represents the similarity reward function, which is used to measure the similarity between the generated data and the real data within the time window; R smooth (t) represents the smoothness reward function, which is used to measure the smoothness of the generated data and prevent drastic changes in the data within the time window; R periodic (t) represents the periodic consistency reward function; R abnormal (t) represents the anomaly detection reward function.
5. The method for establishing a time series prediction model for industrial multi-modal data according to claim 1, characterized in that, The heterogeneous data fusion perception method based on multi-dimensional spatio-temporal mapping in Step 3 includes the following steps: Step 3.1: Register the industrial time series data obtained from different data sources in the augmented time series data set; Step 3.2: Use a heterogeneous graph attention network to fuse the industrial time series data of each registered data source.
6. The method for establishing a time series prediction model for industrial multi-modal data according to claim 5, wherein, The specific steps of Step 3.1 include: Step 3.1.1: Extract the key features in their respective data sources from the time series data of different data sources and use them as the registration control points for each data source; Step 3.1.2: Preprocess the time series data in each data source; Step 3.1.3: Set a target window centered on the registration control point, and obtain feature points with the same name as the registration control point according to the target window; Step 3.1.4: Based on the feature points with the same name, use the Delaunay triangulation algorithm to construct a dense triangular network; Step 3.1.5: Use a differential operator to describe the deformation in the triangular network and perform deformation compensation on each triangular mesh area, thereby realizing the registration of multi-source heterogeneous data.
7. The method for establishing a time series prediction model for industrial multi-modal data according to claim 5, characterized in that The specific steps of Step 3.2 include: Step 3.2.1: Input the registered industrial time series data as input into the heterogeneous graph attention network; Step 3.2.2: Map the industrial time series data into embedding vectors according to the following formula Achieve data fusion: where, v i ∈V, v j ∈V, is the parameter learning matrix; w s x ∈R 1×d and w n x ∈R 1×d are two learnable parameters, is the correlation coefficient between node v i and its neighbor node v j , and its specific form is: where σ is the activation function implemented by LeakyReLU, || represents the concatenation operation, and a ∈ R 1×d is the trainable attention parameter.
8. The method for establishing a time series prediction model for industrial multi-modal data according to claim 1, wherein The specific operation steps of Step 4 include: Step 4.1: Obtain the industrial time series data after data augmentation and fusion, perform normalization preprocessing on it, and divide the industrial time series data into multiple small pieces, with each small piece being a patch; Step 4.2: Perform patch reprogramming based on the multi-head cross-attention mechanism, convert the behavior of the time series into natural language, and output the translated patches; Step 4.3: Send the translated patches into the multi-head attention mechanism, and then perform linear projection to align the dimension of the reprogrammed patches with the dimension of the backbone of the AI large model; Step 4.4: Use the translated patches as the prompt prefix of the input data and input them into the frozen AI large model; Step 4.5: The AI large model utilizes the internal knowledge to deeply analyze the time series data and generate predicted values; Step 4.5: Map the predicted values output by the AI large model to the numerical space of the time series through a linear layer, thereby generating specific time series predicted values for future development.
9. The method for establishing a time series prediction model for industrial multi-modal data according to claim 8, wherein, The specific steps of the patch reprogramming in Step 4.2 include: Step 4.2.1: Obtain n patches of each industry time series; Step 4.2.2: Generate corresponding query, key, and value vectors for each patch; Step 4.2.3: Input the query vector of the time series patch and the key and value vectors of the natural language description into the multi-head cross-attention mechanism; Step 4.2.4: In the multi-head cross-attention mechanism, calculate the attention weights of each head through the following formula: where Q, K, and V represent the query, key, and value vectors respectively, and dk is the dimension of the key vector; Step 4.2.5: Concatenate the outputs of all heads and perform linear transformation to align the dimension with the dimension of the backbone of the large model: MultiHead(Q, K, V) = Concat(head1,..., head h )W O Among them, W i Q , W i K , W i V and W O represent learnable parameter matrices; Concat represents the concatenation operation; head h represents the h-th head; Step 4.2.6: Output the reprogrammed n patches.
Citation Information
Cited By
Smart city environment health and safety risk prediction system and method based on multi-mode perception
CN121860426A
A Smart City Environmental Health and Safety Risk Prediction System and Method Based on Multimodal Perception
CN121860426B
Cross-working-condition multivariable time sequence anomaly detection method based on stage perception migration diffusion
CN121980476A
Cross-condition multivariate time series anomaly detection method based on phase-aware migration diffusion
CN121980476B
Pot overflow early warning method and computer program product
CN122290021A