Time sequence prediction method and system based on dual tasks and random data disruption

By introducing dual tasks and random data disruption mechanisms in time series prediction, the problem of distribution offset in meteorological time series data is solved, and the accuracy and reliability of precipitation prediction are significantly improved.

CN120045874APending Publication Date: 2025-05-27XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510104566.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the distribution offset problem in meteorological timing data, resulting in insufficient accuracy in precipitation prediction.

Method used

A time series prediction method based on dual tasks and random data disruption is adopted. Through the data random disruption attention module and dual task module, a time series prediction model is built to solve the distribution offset problem and improve prediction accuracy.

Benefits of technology

It effectively alleviates the distribution offset problem in meteorological timing data, improves the accuracy and reliability of precipitation prediction, especially in short-term precipitation prediction in the northern Xinjiang region.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045874A_ABST
    Figure CN120045874A_ABST
Patent Text Reader

Abstract

The invention discloses a time series prediction method and system based on dual tasks and random data disruption, and the method comprises the steps: obtaining a historical meteorological data set of a to-be-predicted region, and carrying out the text assignment processing, and obtaining meteorological time series data; preprocessing the meteorological time sequence data, and calculating the preprocessed meteorological time sequence data and the rainfall correlation coefficient to obtain rainfall correlation meteorological factor characteristics; analyzing a distribution offset problem existing in the preprocessed meteorological time series data, and constructing a time series prediction model based on dual tasks and a random data disruption mechanism; and according to the weight vector of the rainfall-related meteorological factor feature, performing prediction by using the time series prediction model to obtain a time series prediction result corresponding to the target prediction time. According to the method, the distribution offset problem in the meteorological time series data can be effectively solved, and the accuracy of rainfall prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data mining, and particularly relates to a time series prediction method and system based on dual tasks and random data shuffling. Background Art

[0002] Precipitation prediction is an important research field in meteorology. It involves analyzing and predicting the movement and transformation processes of moisture in the atmosphere to predict the precipitation situation in a certain area in the future for a period of time. Precipitation prediction is of great significance for aspects such as social and economic development, agricultural production, water resource management, and disaster prevention and mitigation. In the context of frequent extreme events due to global warming, Xinjiang, located in arid and semi-arid regions, is particularly affected, with an increased risk of freezing damage, drought, snow disaster, rainstorm flood, and snowmelt flood. The annual average precipitation in the northern Xinjiang region shows an increasing trend year by year, and the precipitation tendency rate is about 12 mm / 10 years. Since the ecology of Xinjiang is relatively fragile and the infiltration capacity of the underlying surface is poor, continuous extreme precipitation often causes natural disasters such as debris flow and flood. Therefore, the research on precipitation prediction in the Xinjiang region is particularly important. Accurate precipitation prediction is based on the effective analysis of various meteorological data. By comprehensively considering the changes of various meteorological elements such as temperature, humidity, wind direction, wind speed, and air pressure, more accurate precipitation prediction results can be obtained. These meteorological data can be obtained through devices such as meteorological observation stations, satellite remote sensing, and meteorological radars. The meteorological data is processed and analyzed to establish a meteorological model for precipitation prediction.

[0003] Meteorological data usually has different time resolutions, such as hourly, daily, monthly, etc. A smaller time resolution (hourly level) can provide more detailed prediction information and can better capture short-term meteorological changes. By analyzing precipitation and other meteorological data in a specific area or a wider area, the seasonal and interannual changes hidden in the data can be mined, so as to analyze the precipitation changes in different seasons and years, providing a useful reference for precipitation forecasting. By observing multiple variables of meteorological data, for complex meteorological data, deep learning methods can capture long-term trends such as seasonal and interannual changes in the data, which helps the model learn the complex relationships of the meteorological system, providing a reliable data basis for carrying out precipitation prediction. Therefore, it is of great significance to propose a precipitation prediction method on this basis. Summary of the Invention

[0004] To achieve the above object, the present invention provides a time series prediction method and system based on dual tasks and random data shuffling. Among them, a time series prediction method based on dual tasks and random data shuffling includes:

[0005] Obtain the historical meteorological data set of the area to be predicted, and perform text assignment processing on the historical meteorological data set to obtain meteorological time series data;

[0006] Perform preprocessing operations on the meteorological time series data, including denoising, cleaning, and missing value handling, to obtain preprocessed meteorological time series data;

[0007] Analyze the preprocessed meteorological time series data to obtain the correlation coefficient of rainfall;

[0008] Calculate the correlation coefficient between the preprocessed meteorological time series data and the rainfall to obtain precipitation-related meteorological factor characteristics;

[0009] Analyze the distribution shift problem in the preprocessed meteorological time series data and construct a time series prediction model based on a dual-task and random data shuffling mechanism;

[0010] According to the weight vector of the precipitation-related meteorological factor characteristics, use the time series prediction model to make a prediction and obtain the time series prediction result corresponding to the target prediction time.

[0011] Preferably, the process of performing text assignment processing on the historical meteorological data set includes:

[0012] Convert the text data in the historical meteorological data set of the area to be predicted into time series data;

[0013] First, create a mapping dictionary, and map the unique values in each text column to numerical values through the mapping dictionary; assign new values to each unique value according to the length of the dictionary so that each unique value has a unique numerical representation;

[0014] Then, use the apply method combined with the lambda function to replace the values in a specific column with the corresponding numerical representations through the mapping dictionary.

[0015] Preferably, the process of mapping through the mapping dictionary includes:

[0016] Directly create a dictionary according to the precipitation RRR characteristics, map "no precipitation" to -1; map "sign of precipitation" to 0; if it is not "no precipitation" or "sign of precipitation", keep it unchanged.

[0017] Preferably, the correlation coefficient of rainfall is the Spearman correlation coefficient, and the correlation between each meteorological factor and precipitation RRR is evaluated through the Spearman correlation coefficient.

[0018] Preferably, the formula expression of the Spearman correlation coefficient is:

[0019]

[0020] where d iis the rank difference between two variables, n is the number of observations, and its value range is [-1, 1], where

[0021] -1 indicates a perfect negative correlation, 1 indicates a perfect positive correlation, and 0 indicates no correlation.

[0022] Preferably, the correlation threshold of the Spearman correlation coefficient is 0.1;

[0023] The precipitation-related meteorological factor characteristics are meteorological factors whose absolute value of the correlation coefficient with the precipitation amount RRR is greater than or equal to 0.1.

[0024] Preferably, the process of constructing a time series prediction model based on a dual-task and random data shuffling mechanism includes:

[0025] Introduce a data random shuffling attention module and a dual-task module. Randomly shuffle and rearrange the original time series data through the data random shuffling attention module, and use it together with the original data as the training features of the model;

[0026] Through the dual-task mechanism of the dual-task module, while predicting the future sequence, use the future sequence to restore the past sequence for bidirectional prediction and reconstruction.

[0027] Preferably, the working process of the data random shuffling attention module includes:

[0028] Embed the input time series X n to obtain the corresponding feature vector representation:

[0029]

[0030] At the same time, embed the randomly shuffled time series to obtain the shuffled feature vector:

[0031]

[0032] Perform linear projection on the embedded vectors to generate the query vector Q n of the original feature, the key vector K n the value vector V n and the query vector of the shuffled feature the key vector

[0033]

[0034] Calculate the feature fusion result through the attention mechanism, and the formula expression is:

[0035]

[0036] where, Represents the time series after random shuffling, The vector representations of the original distribution feature and the shuffled distribution feature are respectively denoted as Q n , K n , V n They respectively represent the query, key, and true value of the original distribution feature, which are the query and key of the shuffled distribution feature.

[0037] Preferably, the working process of the dual-task module includes:

[0038] Define X 1:t+T and Y 1:t+T which respectively represent the original time series and the target value during the entire time period. For the forward prediction task, the sequence input to the model is defined as X' 1:t = X 1:t , and the output result is The true value is Y' t+1:t+T = Y t+1:t+T ; then the loss function of the forward prediction task is defined as:

[0039]

[0040] For the backward prediction task, the input sequence is defined as X″ 1:t = X t+T:1+T:-1 , where the -1 in the subscript indicates that the index decreases within the continuous interval; the predicted output result is The true value is Y″ t+1:t+T = Y T:1:-1 , then the loss function of the backward prediction task is expressed as:

[0041]

[0042] Combining the forward task and the backward task, the total loss function of the bidirectional task is defined as:

[0043]

[0044] where λ is a hyperparameter for balancing the losses of the forward prediction task and the backward prediction task.

[0045] The present invention also provides a time series prediction system based on dual tasks and random data shuffling, including:

[0046] A data acquisition module, configured to acquire the historical meteorological data set of the area to be predicted, and perform text assignment processing on the historical meteorological data set to obtain meteorological time series data;

[0047] A data preprocessing module, connected to the data acquisition module, is configured to perform preprocessing operations of denoising, cleaning, and missing value processing on the meteorological time series data to obtain preprocessed meteorological time series data;

[0048] A feature selection module, connected to the data preprocessing module, is configured to analyze the preprocessed meteorological time series data to obtain the correlation coefficient of rainfall; calculate the correlation coefficient between the preprocessed meteorological time series data and rainfall to obtain precipitation-related meteorological factor features;

[0049] A time series prediction module, connected to the feature selection module, is configured to analyze the distribution shift problem existing in the preprocessed meteorological time series data, and construct a time series prediction model based on a dual-task and random data shuffling mechanism; according to the weight vector of the precipitation-related meteorological factor features, use the time series prediction model to perform prediction to obtain a time series prediction result corresponding to the target prediction time.

[0050] Compared with the prior art, the present invention has the following advantages and technical effects:

[0051] The present invention performs preprocessing operations of denoising, cleaning, and missing value processing on the meteorological time series data by converting meteorological text data into time series data, calculates the correlation coefficient between the preprocessed meteorological time series data and rainfall through a feature selection method to obtain precipitation-related meteorological factor features; constructs a time series prediction model in combination with a dual-task and random data shuffling mechanism for precipitation prediction. The present invention mainly aims at the distribution shift problem existing in the existing precipitation data, and adopts a time series method based on deep learning for short-term precipitation prediction in the northern Xinjiang region, which can effectively solve the distribution shift problem in meteorological time series data and improve the accuracy of precipitation prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0053] Figure 1 is a schematic flowchart of the method according to an embodiment of the present invention;

[0054] Figure 2 is an analysis diagram of data non-linearity and monotonicity, and non-parametric property analysis according to an embodiment of the present invention;

[0055] Figure 3 is a distribution shift diagram according to an embodiment of the present invention;

[0056] Figure 4 is a Urumqi hyperparameter experiment diagram according to an embodiment of the present invention;

[0057] Figure 5 Experimental diagram of the superparameters in Tacheng for the embodiments of the present invention;

[0058] Figure 6 Experimental diagram of the superparameters in Altay for the embodiments of the present invention;

[0059] Figure 7 Experimental diagram of the superparameters in Karamay for the embodiments of the present invention;

[0060] Figure 8 Schematic diagram related to the time series prediction model based on the dual-task and random data shuffling mechanism for the embodiments of the present invention. Detailed implementation manners

[0061] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will describe this application in detail with reference to the drawings and in combination with the embodiments.

[0062] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0063] Embodiment 1

[0064] As Figure 1-8 shown, in this embodiment, a time series prediction method based on dual-task and random data shuffling is provided, including:

[0065] Obtain the historical meteorological data set of the area to be predicted, and perform text assignment processing on the historical meteorological data set to obtain meteorological time series data;

[0066] Perform preprocessing operations of denoising, cleaning, and missing value processing on the meteorological time series data to obtain the preprocessed meteorological time series data;

[0067] Analyze the preprocessed meteorological time series data, and correspondingly obtain the correlation coefficient of rainfall;

[0068] Calculate the preprocessed meteorological time series data and the correlation coefficient of rainfall to obtain precipitation-related meteorological factor features;

[0069] Analyze the distribution shift problem existing in the preprocessed meteorological time series data, and construct a time series prediction model based on the dual-task and random data shuffling mechanism;

[0070] According to the weight vector of the precipitation-related meteorological factor features, use the time series prediction model to perform prediction to obtain the time series prediction result corresponding to the target prediction time.

[0071] For further optimized solutions, the method of this embodiment specifically includes the following steps:

[0072] S1. Obtain the historical meteorological data set of the northern Xinjiang region through a meteorological website, and perform preliminary text assignment processing on it to convert it into time series data that can be recognized by a time series model. Save the meteorological data sets of 4 regions as the original data set, namely the meteorological data set D of Urumqi City u , the meteorological data set D of Karamay City k , the meteorological data set D of Tacheng Prefecture t and the meteorological data set D of Altay Prefecture a ;

[0073] S2. Preprocess the original data sets D u , D k , D t and D a obtained in step 1, including denoising, cleaning, and missing value processing, to obtain the preprocessed data sets D' u , D' k , D' t and D' a ;

[0074] S3. Analyze the preprocessed data sets D' u , D' k , D' t and D' a obtained in step 2, and select appropriate correlation coefficients.

[0075] S4. Use the correlation coefficients selected in step 3 to perform feature selection of precipitation-related meteorological factors.

[0076] S5. Analyze the data in the preprocessed data set, analyze the distribution shift problem existing in the data, and construct a time series prediction model, that is, a time series prediction model combining a dual task and a random data shuffling mechanism.

[0077] S6. Perform hyperparameter experiments on the constructed model and select the hyperparameters with the best performance.

[0078] S7. Use the optimal hyperparameters selected in step 6 to perform comparative experiments and ablation experiments to evaluate the performance of the model.

[0079] Furthermore, the preliminary text assignment processing of the data set described in step 1 specifically includes:

[0080] Text data to time series data: First, create a mapping dictionary to map the unique values in each text column to numerical values. Assign new values to each unique value according to the length of the dictionary to ensure that each unique value has a unique numerical representation. Then, use the apply method combined with a lambda function to replace the values in a specific column with their corresponding numerical representations, mapping according to the previously created dictionary.

[0081] Specifically, according to the characteristics of precipitation RRR, directly create a dictionary, map "no precipitation" to -1, map "precipitation indication" to 0, and keep the value unchanged if it is not "no precipitation" or "precipitation indication".

[0082] Furthermore, the preprocessing of the dataset described in step 2 includes:

[0083] Data cleaning: Due to reasons such as sensor failures, network communication problems, data storage and transmission in meteorological stations, there are missing values in the acquired original meteorological data. After statistics, it is found that there are a large number of missing values in the early years, and in the relatively recent years, the phenomenon of data missing is significantly reduced, and there is no data or a large amount of missing data in the feature columns of ff10, N, Cl, Nh, Cm, Ch, E, Tg, sss, tR, E. Therefore, in this embodiment, the data from July 2014 to March 2023 is selected as the experimental data, and the columns with no data or a large number of missing values are deleted. There are still some missing values in the selected data. Since the time interval is short (recorded every three hours), and the weather changes are relatively smooth at a fine-grained time scale, the previous data item of the missing value is used to fill it.

[0084] Furthermore, the feature selection of precipitation-related meteorological factors includes:

[0085]

[0086] where d i is the rank difference between two variables, n is the number of observations, and the value range is [-1, 1], where

[0087] -1 indicates a perfect negative correlation, 1 indicates a perfect positive correlation, and 0 indicates no correlation.

[0088] More specifically, calculate the correlation coefficients between each meteorological feature according to the Spearman correlation coefficient. The correlation coefficients between each meteorological feature and precipitation RRR are shown in Table 1.

[0089] Table 1

[0090]

[0091] Feature selection using the correlation coefficient: To screen out the meteorological factors that have the most influence on precipitation prediction, in this embodiment, the correlation threshold is set to 0.1. The basis for choosing 0.1 is that through multiple experiments, it is proved that this value can not only ensure the screening of factors with a certain degree of correlation with precipitation, but also avoid introducing too many factors with low correlation, thereby increasing the computational complexity and reducing the prediction accuracy. Finally, the meteorological factors with the absolute value of the correlation coefficient greater than or equal to 0.1 with precipitation RRR are selected as the input features of the model to improve the generalization ability and prediction accuracy of the model.

[0092] As can be seen from Table 1, the meteorological characteristics with relatively high correlation with rainfall at the Urumqi station are: U, ff3, WW, W1, W2, Tn, Tx, H, VV, Td; the meteorological characteristics with relatively high correlation with rainfall at the Karamay station are: T, Po, P, U, ff3, WW, W1, W2, Tn, Tx, H, Td; the meteorological characteristics with relatively high correlation with rainfall at the Altay station are: U, ff3, WW, W1, W2, H, VV, Td; the meteorological characteristics with relatively high correlation with rainfall at the Tacheng station are: Td, H, W2, W1, WW, ff3, DD, U, Pa. If the meteorological characteristics with relatively high precipitation correlation common to all stations are taken as input features, the meteorological characteristics with the highest precipitation correlation at each station may be lost, thus affecting the prediction accuracy. Therefore, different meteorological characteristics are selected as the input of the model according to the conditions of each station.

[0093] Furthermore, for the dataset D' u , D' k , D' t and D' a are analyzed, the existing problems are proposed, and the process of constructing a time series prediction model, that is, a time series prediction model based on a dual-task and random data shuffling mechanism, includes:

[0094] Distribution Shift is a common problem in time series prediction. Distribution Shift refers to the phenomenon that the distributions of training data and test data are inconsistent. This inconsistency may be caused by reasons such as sampling method differences, environmental changes, or data imbalance, resulting in a decline in the performance of the model on the test set because the features it learned during training may not be applicable to the test data. By analyzing the collected meteorological time series data, it is found that it also has the problem of distribution shift.

[0095] Such as Figure 2-3As shown, due to the distribution shift in meteorological data, that is, the distribution characteristics of the prediction window data are inconsistent with those of the historical data in the lookback window, and the data is sparse, resulting in limited training samples, it is impossible to learn the distribution characteristics of the prediction window well. To alleviate the above problems, this embodiment proposes a training strategy combining random shuffling and dual tasks. Specifically, first, the original time series data is randomly shuffled, and then it is used as the training feature of the model together with the original data. This random shuffling method can enable the model to learn more distribution characteristics and alleviate the problem of distribution shift. At the same time, a dual-task mechanism is designed, which requires the model to use the future sequence to reconstruct the past sequence while predicting the future sequence. Through this two-way prediction and reconstruction method, the model can not only understand the internal laws and dependence structures of time series more profoundly, but also capture potential distribution shift signals during the reconstruction process, further improving the adaptability of the model to complex time series. In summary, the proposed method provides an effective solution to alleviate the distribution shift problem through the collaborative optimization of data random shuffling and dual tasks.

[0096] More specifically, to solve the above problems, this embodiment proposes the PFDformer model, as Figure 8 shown. This model alleviates the distribution shift problem and improves the prediction accuracy of the time series prediction model by introducing the data random shuffling attention and dual-task module. Among them, the random shuffling attention mechanism randomly rearranges the original time series data and uses it as the input feature of the model together with the original data, helping the model capture richer distribution characteristics, thus alleviating the distribution shift problem and also alleviating the problem of limited training samples caused by data sparsity. The dual-task learning strategy requires the model to use the future sequence to reconstruct the past sequence while predicting the future sequence. Through the two-way prediction and reconstruction process, the model can understand the internal laws and dependence structures of time series more profoundly and capture potential distribution shift signals during the reconstruction process. Specifically, if the model cannot accurately reconstruct the past sequence through the future sequence, it indicates that there may be a deviation in its modeling of the time dependence relationship, and this deviation usually stems from the distribution shift problem of the data. At this time, by optimizing the objective function of the dual task, the model parameters can be dynamically adjusted, thereby enhancing its adaptability and robustness to distribution shift. The random shuffling attention mechanism further helps the model extract feature information from more distributions, effectively improving the representation ability for complex time series data. Experiments show that compared with traditional models that only perform one-way prediction, PFDformer shows significant advantages in dealing with the distribution shift problem, improving the accuracy and reliability of time series prediction, and providing more scientific and reliable technical support for short-term precipitation prediction in the northern Xinjiang region.

[0097] Among them, the specific implementation process of the data random shuffling attention module is as follows:

[0098] Embed the input time series X n to obtain the corresponding feature vector representation:

[0099]

[0100] At the same time, embed the randomly shuffled time series to obtain the shuffled feature vector:

[0101]

[0102] Perform a linear projection on the embedding vectors to generate the query vector Q n , key vector K n , value vector V n of the original features and the query vector key vector

[0103]

[0104] Calculate the feature fusion result through the attention mechanism, and the formula expression is:

[0105]

[0106] where represents the randomly shuffled time series, represent the vector representations of the original distribution feature and the shuffled distribution feature respectively, and Q n , K n , V n represent the query, key, and true value of the original distribution feature respectively, are the query and key of the shuffled distribution feature.

[0107] The specific implementation steps of the dual-task module are as follows:

[0108] Define X 1:t+T and Y 1:t+T to represent the original time series and the target value within the entire time period respectively. For the forward prediction task, the sequence input to the model is defined as X' 1:t = X 1:t , and the output result is The true value is Y' t+1:t+T = Y t+1:t+T ; then the loss function of the forward prediction task is defined as:

[0109]

[0110] For the backward prediction task, the input sequence is defined as X″ 1:t = X t+T:1+T:-1, where -1 in the subscript indicates that the index decreases within a continuous interval; the predicted output result is The true value is Y″ t+1:t+T = Y T:1:-1 , then the loss function of the inverse prediction task is expressed as:

[0111]

[0112] Combining the forward task and the inverse task, the total loss function of the bidirectional task is defined as:

[0113]

[0114] where λ is a hyperparameter that balances the losses of the forward prediction task and the inverse prediction task.

[0115] Furthermore, hyperparameter experiments are carried out on the constructed model, and the optimal hyperparameters are selected for the next experiment.

[0116] λ is a hyperparameter that balances the losses of the forward prediction task and the inverse prediction task. In this embodiment, λ = 0.05, 0.1, 0.3, 0.5, 0.8, and 1 are taken for hyperparameter experiments, as shown in Table 2. And we visualize the time results, and the results are as Figure 4-Figure 7 shown. Thus, we conclude that on the meteorological datasets of Urumqi, Altay, and Tacheng, when λ = 0.3, the model has the best prediction effect, and on the Karamay dataset, when λ = 0.5, the model has the best prediction effect.

[0117] Table 2

[0118]

[0119] First, collect and organize the meteorological data of the northern Xinjiang region, including Urumqi City, Karamay City, Tacheng Prefecture, and Altay Prefecture. By accessing public meteorological websites such as rp5.ru, download the meteorological datasets of these regions. The downloaded data includes basic meteorological parameters such as temperature, humidity, and wind speed. Conduct preliminary text parsing processing on the downloaded data, such as formatting date and time tags, to ensure that the data is correctly read and processed in the form of a time series. These processed datasets are respectively named the Urumqi City Meteorological Dataset D u , the Karamay City Meteorological Dataset D k , the Tacheng Prefecture Meteorological Dataset D t and the Altay Prefecture Meteorological Dataset D aSecondly, preprocessing operations are performed on the above original dataset to improve data quality and prepare for subsequent analysis. This includes noise removal, data cleaning to remove outliers, and filling in missing values, etc. This step is crucial because high-quality data input will directly affect the prediction accuracy of the model. The preprocessed datasets are respectively labeled as D' u , D' k , D' t and D' a . Next, feature selection is performed on the preprocessed dataset, especially selecting meteorological factors with a relatively high correlation coefficient with precipitation. In this step, the Spearman correlation coefficient is used to calculate the meteorological factors that have the greatest impact on predicting precipitation, thereby simplifying the complexity of the model while maintaining prediction accuracy. Finally, a time series prediction model is constructed, which incorporates a dual-task and data random shuffling module to predict future meteorological conditions. By learning different distribution characteristics in the time series, this model can more effectively predict meteorological changes in the future for a period of time. Through this method, the problem of distribution shift in time series data is effectively solved, and short-term precipitation prediction in the northern Xinjiang region is carried out more accurately.

[0120] More specifically, in this embodiment, experiments are carried out on the meteorological datasets in 4 regions of northern Xinjiang and the Hami region, and seven models such as iTransformer, PatchTST, Nonstationary_Transformer, FEDformer, Autoformer, Informer, and Transformer are selected as benchmark models for comparative experiments. Experiments are carried out with λ = 0.3 on the Urumqi, Tacheng, and Altay datasets, and experiments are carried out with λ = 0.5 on the Karamay dataset. The experimental results are shown in Table 3. The research results show that PFDformer has achieved good improvements on all 4 datasets.

[0121] Table 3 Comparative experimental results of λ = 0.3, Urumqi, Tacheng, and Altay datasets

[0122]

[0123] Comparative experimental results of λ = 0.5, Karamay dataset

[0124]

[0125] The performance of PFDformer is better than that of 7 baseline models on 5 meteorological datasets. Among them, on the Urumqi dataset, the MSE is improved by 23.70% - 46.23%, with an average improvement of 14.07%, and the MAE is improved by 14.97% - 48.30%, with an average improvement of 14.49%; on the Tacheng dataset, the MSE is improved by 3.79% - 35.04%, with an average improvement of 17.16%, and the MAE is improved by 3.90% - 43.23%, with an average improvement of 18.31%; on the Altay dataset, the MSE is improved by 3.86% - 53.01%, with an average improvement of 23.45%, and the MAE is improved by 15.23% - 40.72%, with an average improvement of 20.76%; on the Karamay dataset, the MSE is improved by 3.78% - 35.71%, with an average improvement of 13.61%, and the MAE is improved by 0.97% - 30.72%, with an average improvement of 14.44%. Based on this, we can draw the conclusion that the PFDformer model has achieved significant improvement in the prediction performance on meteorological datasets in regions such as Urumqi, Karamay, Tacheng, and Altay. Compared with the baseline models, PFDformer shows better results in the MSE and MAE indicators in each region.

[0126] To study the different impacts of the two methods on solving the distribution shift problem, we designed ablation experiments, removing the dual-task module (w / o Dual), the data random shuffling module (w / o Random), and removing both methods simultaneously (w / o Dual+Random). The experiments were conducted on the Urumqi and Karamay datasets, and the results are shown in Table 4. It can be seen that the performance of the model decreases regardless of which method is removed. In addition, the PFDformer model achieved the best results at most prediction lengths, which further verifies the effectiveness and superiority of the proposed method in solving the distribution shift problem.

[0127] Table 4

[0128] On the Urumqi and Karamay datasets, the results of the ablation experiments, where λ = 0.3 on the Urumqi dataset and λ = 0.5 on the Karamay dataset

[0129]

[0130] Embodiment 2

[0131] Based on the same inventive concept, this embodiment also provides a time series prediction system based on dual-task and random data shuffling, including:

[0132] A data acquisition module, configured to acquire the historical meteorological dataset of the area to be predicted, and perform text assignment processing on the historical meteorological dataset to obtain meteorological time series data;

[0133] A data preprocessing module, connected to the data acquisition module, for performing preprocessing operations of denoising, cleaning, and missing value processing on meteorological time series data to obtain preprocessed meteorological time series data;

[0134] A feature selection module, connected to the data preprocessing module, for analyzing the preprocessed meteorological time series data to obtain the correlation coefficient of rainfall; calculating the preprocessed meteorological time series data with the correlation coefficient of rainfall to obtain precipitation-related meteorological factor features;

[0135] A time series prediction module, connected to the feature selection module, for analyzing the distribution shift problem existing in the preprocessed meteorological time series data, constructing a time series prediction model based on a dual-task and random data shuffling mechanism; using the time series prediction model for prediction according to the weight vector of precipitation-related meteorological factor features to obtain a time series prediction result corresponding to the target prediction time.

[0136] A time series prediction system based on dual-task and random data shuffling provided in this embodiment has all the advantages of the time series prediction method based on dual-task and random data shuffling provided in Embodiment 1.

[0137] Embodiment 3

[0138] This embodiment also discloses a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method described in Embodiment 1.

[0139] Embodiment 4

[0140] This embodiment also discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the method described in Embodiment 1.

[0141] Embodiment 5

[0142] This embodiment also discloses a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the method described in Embodiment 1.

[0143] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A time series prediction method based on dual tasks and random data disruption, characterized in that: include: Acquire a historical meteorological data set of the area to be predicted, and perform text assignment processing on the historical meteorological data set to obtain meteorological time series data; Performing preprocessing operations of denoising, cleaning and missing value processing on the meteorological time series data to obtain preprocessed meteorological time series data; Analyzing the pre-processed meteorological time series data to obtain a corresponding correlation coefficient of rainfall; Calculating the correlation coefficient between the preprocessed meteorological time series data and rainfall to obtain precipitation-related meteorological factor characteristics; Analyze the distribution shift problem in preprocessed meteorological time series data, and build a time series prediction model based on dual tasks and random data disruption mechanism; According to the weight vector of the precipitation-related meteorological factor characteristics, the time series prediction model is used to perform prediction to obtain a time series prediction result corresponding to the target prediction time.

2. The method according to claim 1, characterized in that The process of performing text assignment processing on the historical meteorological data set includes: Convert the text data of the historical meteorological dataset of the area to be predicted into time series data; First, a mapping dictionary is created, by which unique values ​​in each text column are mapped to numeric values; a new value is assigned to each unique value according to the length of the dictionary, so that each unique value has a unique numeric representation; Then, the apply method is used in conjunction with the lambda function to replace the values ​​in the specific column with the corresponding numerical representations, and the mapping is performed through the mapping dictionary.

3. The method according to claim 2, characterized in that The process of mapping through the mapping dictionary includes: Create a dictionary directly based on the precipitation RRR feature, map "no precipitation" to -1, map "sign of precipitation" to 0; if it is not "no precipitation" or "sign of precipitation", it remains unchanged.

4. The method according to claim 1, characterized in that The correlation coefficient of the rainfall is the Spearman correlation coefficient, and the correlation between each meteorological factor and the precipitation RRR is evaluated by the Spearman correlation coefficient.

5. The method according to claim 4, characterized in that The formula expression of the Spearman correlation coefficient is: Among them, d i is the rank difference between the two variables, n is the number of observations, and the value range is [-1, 1], where -1 indicates a perfect negative correlation, 1 indicates a perfect positive correlation, and 0 indicates no correlation.

6. The method according to claim 4, characterized in that The correlation threshold of the Spearman correlation coefficient is 0.1; The precipitation-related meteorological factor characteristic is a meteorological factor whose absolute value of the correlation coefficient with the precipitation RRR is greater than or equal to 0.

1.

7. The method according to claim 1, characterized in that The process of building a time series prediction model based on dual-task and random data scrambling mechanism includes: The data random shuffle attention module and the dual-task module are introduced. The original time series data are randomly shuffled and rearranged through the data random shuffle attention module, and used together with the original data as the training features of the model; Through the dual-task mechanism of the dual-task module, while predicting the future sequence, the past sequence is restored using the future sequence, thereby performing bidirectional prediction and reconstruction.

8. The method according to claim 7, characterized in that The working process of the data random shuffle attention module includes: For the input time series X n Embed and get the corresponding feature vector representation: At the same time, the randomly disrupted time series Embed and get the disrupted feature vector: Perform linear projection on the embedded vector to generate the query vector Q of the original feature n , key vector K n , value vector V n and the query vector with shuffled features Key Vector The feature fusion result is calculated through the attention mechanism, and the formula is: in, represents the randomly disrupted time series, The vector representations of the original distribution features and the shuffled distribution features, Q n ,K n ,V n denote the query, key and true value of the original distribution features respectively, The query and key for the shuffled distribution features.

9. The method according to claim 7, characterized in that: The working process of the dual-task module includes: Define X 1:t+T and Y 1:t+T Represent the original time series and target value in the entire time period respectively. For the forward prediction task, the sequence of the input model is defined as X' 1:t =X 1:t , the output is The true value is Y' t+1:t+T =Y t+1:t+T ; Then the loss function of the forward prediction task is defined as: For the backward prediction task, the input sequence is defined as X″ 1:t =X t+T:1+T:-1 , where -1 in the subscript indicates that the index decreases in the continuous interval; the output prediction result is The true value is Y″ t+1:t+T =Y T:1:-1 , then the loss function of the reverse prediction task is expressed as: Combining the forward task and the reverse task, the total loss function of the bidirectional task is Defined as: Among them, λ is a hyperparameter that balances the loss of the forward prediction task and the backward prediction task.

10. A time series prediction system based on dual tasks and random data disruption, characterized in that: include: A data acquisition module is used to acquire a historical meteorological data set of a region to be predicted, and to perform text assignment processing on the historical meteorological data set to obtain meteorological time series data; A data preprocessing module, connected to the data acquisition module, is used to perform preprocessing operations of denoising, cleaning and missing value processing on the meteorological time series data to obtain preprocessed meteorological time series data; A feature selection module is connected to the data preprocessing module and is used to analyze the preprocessed meteorological time series data to obtain a corresponding correlation coefficient of rainfall; the correlation coefficient between the preprocessed meteorological time series data and rainfall is calculated to obtain precipitation-related meteorological factor characteristics; The time series prediction module is connected to the feature selection module and is used to analyze the distribution deviation problem existing in the preprocessed meteorological time series data, and to construct a time series prediction model based on a dual-task and random data scrambling mechanism; according to the weight vector of the characteristics of the precipitation-related meteorological factors, the time series prediction model is used to perform predictions to obtain the time series prediction results corresponding to the target prediction time.