Carbon asset price prediction method under multi-source fusion

Through the multi-source fusion method, a multi-step prediction multivariate heterogeneous model is designed, which solves the problem of failing to effectively consider multiple complex factors in the existing technology, improves the accuracy and stability of the market price prediction of carbon assets, and enhances the interpretability of the model.

CN119941277AInactive Publication Date: 2025-05-06UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510070960.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing carbon asset price prediction model fails to effectively comprehensively consider a variety of complex factors affecting the carbon asset market, resulting in instability and inaccuracy of the prediction results.

Method used

Using the multi-source fusion method, multi-step prediction multivariate heterogeneous models are designed by inputting time series data of static variables, historical observations and future known values, and components such as gating mechanism, gated residual network, variable selection network and timing fusion decoder are used for feature extraction and prediction.

Benefits of technology

It improves the accuracy and stability of the market price prediction of carbon assets, can capture factors that affect carbon asset prices from multiple dimensions, and provides feature importance scores through variable selection networks, enhancing the interpretability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941277A_ABST
    Figure CN119941277A_ABST
Patent Text Reader

Abstract

The invention discloses a carbon asset price prediction method under multi-source fusion. The method comprises multi-feature support: multivariate heterogeneous variables including exchange information, carbon asset prices, environmental indexes, macroeconomic indicators, social media information, policy change records, holidays and festivals and non-transaction time can be processed; interpretability: the influence of each feature is measured by using a variable selection network (VSN), which time sequence variables are most important in a historical time interval are revealed, and a visual feature importance score is provided for each feature variable; according to the method, prediction intervals under different percentiles can be output, so that more information is provided for risk management and decision making; advanced machine learning technology fusion: a plurality of advanced machine learning technologies are comprehensively applied, different feature layers are utilized to extract space-time difference features of data, and the combination improves the data feature extraction capability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of financial analysis, and in particular to a carbon asset price prediction method under multi-source fusion. Background Art

[0002] As an important mechanism to respond to global climate change, the carbon market has a high demand for accurate carbon asset price forecasts. The current carbon asset market price forecasts mainly include traditional economic models, machine models, and economic-machine learning hybrid models. Traditional economic models include supply and demand analysis and historical price trends. Although they have good explanatory power, the accuracy of the model is low; machine learning models mainly analyze historical price trends, belong to black box models, and are prone to overfitting problems. The hybrid model of economic-machine learning models can improve the accuracy of predictions through data-driven, and can also provide better explanations for the results. However, these models often fail to comprehensively consider the various complex factors that affect the carbon asset market, such as meteorological changes, policy adjustments, industry dynamics, etc., resulting in instability and inaccuracy in the prediction results. In addition, the existing carbon asset price prediction models mainly rely on historical price data, and insufficient consideration of key factors that have a greater impact on the carbon asset market, such as the environment, policies, and media sentiment.

[0003] Therefore, a carbon asset price prediction method based on multi-source fusion is proposed to solve this problem. Summary of the invention

[0004] To achieve the above object, the present invention provides the following technical solution: a carbon asset price prediction method under multi-source fusion, comprising the following steps: Step 1: Input data: The input variables of the model are: static variables that do not change with time during the study period, time series variables of past observations that are known in the past but unknown in the future, and a priori known time series variables that are known in both the past and the future; Step 2: Data preprocessing: Convert text variables such as policies, news events, and social media information into numerical sequences. This function can be achieved by scoring the sentiment tendency of text information. Each output value represents the sentiment score of the corresponding word's impact on the carbon asset market price. Encode categorical variables (such as carbon asset trading market information) using unique codes; Process outliers and missing values ​​for other numerical data, and then standardize all numerical data; convert all variable data into time series based on timestamps, where each timestamp corresponds to a specific date and time point, accurate to the day; All data were converted into daily data through resampling method; All time series are divided into three parts: training set for learning, validation set for hyperparameter tuning, and holdout test set for performance evaluation; Step 3: Model structure and principle analysis: In time series tasks, there are mainly two types of data sources: static variables and time-varying variables. Time-varying variables can be divided into time-varying variables observed in the past and time-varying variables with known future values. Designing network structures for different data sources helps improve the prediction effect of the model. Establishment of multivariate heterogeneous model for multi-step forecasting: Assume that there is tradable products, the future price of each product is related to a set of static variables , each time step The input time-varying variables under and the target variable related; Time-varying variables are divided into two categories, namely ,in represents a time-varying variable with historical observations but unknown future values. Indicates that past and future values ​​are known time-varying variables; Quantile prediction is used for the prediction of target variables affected by multi-source heterogeneous variables, such as inputting the 10th, 50th, and 90th quantiles at each time step; Each quantile prediction model is as follows ; in, Indicated in Predict the future at all times The qth quantile value of the carbon asset i price in the next step, A carbon asset price prediction model; , that is, output simultaneously , , , The predicted price of carbon assets at the moment; Indicates the past Carbon assets The price, Indicates the past Impact on carbon assets at every moment The observed input of the price, Indicates the past Expiration Impact of carbon assets The price of the future known input, Indicates that the time span of the research object is the past Period, with 1 day as 1 period, Indicates impact on carbon assets Static variable of price; Model architecture and principle analysis The input data is divided into three categories: one is static variable s, and the other is historical value , one is the future known value ,MTFT can construct feature representations for three different types of data sets using different components; The main components and functions of MTFT include gating mechanism, gated residual network (GRN), variable selection network (VSN) and variable selection module; MTFT used different VSNs to process static variables, historical observations, and future inputs; Static Covariate Encoder (SCE): This generates four different context vectors by integrating information from static variable metadata using a separate GRN encoder , the input of SCE is the result of static variables after VSN, where Gave it to VSN, Give LSTM an initialization state. gave the static enrichment layer (SEL); Temporal Fusion Decoder (TFD): learns short-term and long-term temporal relationships between variables from time-varying variable inputs with historical observations and known future values; Sequence-to-sequence layers are used for local processing, while long-term dependencies are captured using a masked interpretable multi-head attention module; TFD consists of three modules: static enrichment layer, time series self-attention layer, and position-based front network layer; Static enrichment layer (SEL): input static covariates to improve time series features; Time Series Self-Attention Layer (TSL): learns long-term dependencies of time series data and makes the model interpretable; Position-based pre-layer (PFL): applies additional non-linear processing to the output of the TSL layer; Step 4: The entire model processing process is: Static variables are feature engineered by the variable selection network (VSN module) and the static covariate encoder (SCE), and time-varying variables are feature engineered by VSN and then output features; The VSN method embeds the categorical variables as a whole and performs linear transformation on the continuous variables. The transformed variables are dimension vector, the converted variable is , for historical input, the flattened result is (in is the number of variables), and these transformed variables are used for variable selection. The variable selection formula is: ,in is the weight of variable selection; Weight by Calculate and obtain; is the feature after nonlinear processing; Nonlinear processing features are achieved through Get it by Provided by SCE; The historical observation features Input LSTM encoder, future known value features Input LSTM decoder, both generate a set of consistent time series features ,in as a position index; Before entering the temporal fusion decoder (TFD), it is processed by the standard layer-normalized gating layer. is the standard layer-normalized function; In order to ignore unnecessary features in a given data set, TFD adds a component GLU based on a linear gating module. The reason why the GLU component can suppress the nonlinear contribution of the data set is that the output of GLU may be close to 0, so the initial input can also completely skip the gating layer and directly enter the next module. The standardization-normalization gating layer is the preparation work before the feature enters TFD. The specific formula is: ; After the feature data enters the TFD module, the first step is to enhance the data features. This step is completed through the static enrichment layer (SEL), which uses static covariates to improve the time series features. Specifically, it uses the GRN and SCE provided The processing method is: ; Enter the time series self-attention layer (TSL), which applies an interpretable multi-head attention mechanism at each prediction time step. In addition to preserving the causal information flow through "masking", the self-attention layer also allows the MTFT to pick up long-range dependencies. After the self-attention layer, an additional gating layer is used to filter information and improve training effects; This module is equivalent to an interpretable multi-head attention layer plus a standard layer-normalized gating layer. Its main function is to learn the long-term correlation of time series data and mask the interpretable multi-head attention mechanism learning function as shown in the formula As shown, the formula is the gating layer function; As mentioned earlier, the TSL layer preserves the causal information flow by “masking” to make the model interpretable; in ,The final output of the masked interpretable multi-head attention layer in the TSL layer is similar to that of the single attention layer. The key difference is that the former generates attention in a different way; Entering the position-based pre-network layer (PFL), the PFL layer uses an additional nonlinear function to process the output of the time series self-attention layer. The processing function is as follows: As shown, the features are filtered again using the gated residual network (GRN); The model takes into account that some data may not require additional complex processing, and the effect of directly using a simple model is better. Therefore, a gating layer that can skip the TFD module is set up to provide a direct path from sequence to sequence. The learning function of the gating layer is as follows: As shown; Enter the fully connected layer, output the predicted value at each time step as the point prediction, add percentiles to turn the point prediction into an interval prediction, and the fully connected layer uses the processing function as shown in the formula to make the output of the model the predicted carbon asset market price range; in is the linear coefficient for a particular quantile q, , and finally output the carbon asset price forecast intervals at different quantiles and time steps at the same time (taking the 10th, 50th and 90th quantiles as examples); Model loss function: The model is trained by minimizing the joint quantile loss, and all quantile outputs are summed. The model loss function is as follows: ; ; in, refers to the domain of training data containing M samples, W represents the weight parameter of MTFT, is the set of output quantiles, ; In order to avoid errors caused by inconsistent prediction dimensions at different prediction points, the normalized quantile loss in the entire prediction range is evaluated, that is, regularization is performed: ; in, is the domain of the test sample.

[0005] Preferably, in step 1, static variables that do not change over time are studied: such as carbon asset trading market mechanisms and macro policies; time series variables of past observations that are known in the past but unknown in the future: such as historical carbon asset prices, meteorological data, economic indicators, policy change records, industrial activities, environmental variables, and social media information; a priori known time series variables that are known in the past and in the future: such as non-trading time intervals and holidays.

[0006] Preferably, in step 2, the accuracy is as follows: Represents a variable of a certain year, month and day The value of .

[0007] Preferably, in step 1, the working temperature for cleaning with a detergent is 23° C.-25° C., and the working humidity is RH53-55RH.

[0008] Preferably, in step 3, a gating mechanism is used: the mechanism allows features to enter only the modules that need to be used, so that the model can automatically adjust the complexity and network depth to adapt to different scenarios and data structures; Gated Residual Networks (GRNs): Simplify models by giving them the flexibility to apply nonlinear processing only when necessary; Variable Selection Network (VSN): selects relevant input variables at each time step, greatly optimizing modeling performance by leveraging learning capabilities only on the most significant features; The variable selection module improves models by removing noisy inputs that can negatively impact performance, and also provides insight into which variables are most important for your prediction problem.

[0009] Preferably, in step 3, different VSNs process static variables, historical observations and future inputs, and parameters of the VSN modules of the three are not shared.

[0010] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention has multi-feature support, explainability, prediction interval output and advanced machine learning technology integration; Multi-feature support: It can process multiple heterogeneous variables including exchange information, carbon asset prices, environmental indices, macroeconomic indicators, social media information, policy change records, holidays and non-trading time. Compared with the existing carbon asset price prediction model, this invention can capture the factors affecting carbon asset prices from multiple dimensions; Interpretability: Compared with most black-box machine learning models, this paper uses variable selection networks (VSNs) to measure the impact of each feature, revealing which time series variables are most important in the historical time interval, and provides an intuitive feature importance score for each feature variable, which not only enhances the transparency and reliability of the model, but also makes the model interpretable; Prediction interval output: Different from the existing carbon asset price prediction models that can only provide point predictions, the present invention can output prediction intervals at different percentiles, which provides more information for risk management and decision-making; Fusion of advanced machine learning technologies: A variety of advanced machine learning technologies are used in combination to extract the spatiotemporal differences of data using different feature layers. This combination improves the model's data feature extraction capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1A schematic diagram of heterogeneous data sources used in multi-step prediction of the present invention; Figure 2 The MTFT model architecture and principle diagram of the present invention. DETAILED DESCRIPTION

[0012] The present invention will be described in more detail below by way of examples, which are merely illustrative and do not in any way limit the scope of the present invention.

[0013] Embodiment 1: The present invention provides a technical solution: a carbon asset price prediction method under multi-source fusion, comprising the following steps: Step 1: Input data: The input variables of the model are: static variables that do not change with time during the study time, time series variables of past observations that are known in the past but unknown in the future, and a priori known time series variables that are known in the past and in the future; Specifically, static variables that do not change over time during the study period: such as carbon asset trading market mechanisms and macro policies; time series variables of past observations that are known in the past but unknown in the future: such as historical carbon asset prices, meteorological data, economic indicators, policy change records, industrial activities, environmental variables, and social media information; a priori known time series variables that are known in the past and in the future: such as non-trading time intervals and holidays; Step 2: Data preprocessing: Convert text variables such as policies, news events, and social media information into numerical sequences. This function can be achieved by scoring the sentiment tendency of text information. Each output value represents the sentiment score of the corresponding word's impact on the carbon asset market price. Encode categorical variables (such as carbon asset trading market information) using unique codes; Process outliers and missing values ​​for other numerical data, and then standardize all numerical data; convert all variable data into time series based on timestamps, where each timestamp corresponds to a specific date and time point, accurate to the day; The accuracy is as high as Variable representing a certain year, month and day The value of All data were converted into daily data through resampling method; All time series are divided into three parts: training set for learning, validation set for hyperparameter tuning, and holdout test set for performance evaluation; Step 3: Model structure and principle analysis: In time series tasks, there are mainly two types of data sources: static variables and time-varying variables. Time-varying variables can be divided into time-varying variables observed in the past and time-varying variables with known future values. Designing network structures for different data sources helps improve the prediction effect of the model. Establishment of multivariate heterogeneous model for multi-step forecasting: Assume that there is tradable products, each with a future price and a set of static variables , each time step The input time-varying variables under and the target variable related; Time-varying variables are divided into two categories, namely ,in represents a time-varying variable with historical observations but unknown future values. Indicates that past and future values ​​are known time-varying variables; Quantile prediction is used for the prediction of target variables affected by multi-source heterogeneous variables, such as inputting the 10th, 50th, and 90th quantiles at each time step; Each quantile prediction model is as follows ; in, Indicated in Predict the future at all times The qth quantile value of the carbon asset i price in the next step, A carbon asset price prediction model; , that is, output simultaneously , , , The predicted price of carbon assets at the moment; Indicates the past Carbon assets The price, Indicates the past Impact on carbon assets at every moment The observed input of the price, Indicates the past Expiration Impact of carbon assets The price of the future known input, Indicates that the time span of the research object is the past Period, with 1 day as 1 period, Indicates impact on carbon assets Static variable of price; Model heterogeneous data source analysis logic such as Figure 1 shown.

[0014] Model architecture and principle analysis The input data is divided into three categories: one is static variable s, and the other is historical value , one is the future known value ,MTFT can construct feature representations for three different types of data sets using different components; Figure 2 The architecture and principle of MTFT are shown in detail. The main components and functions of MTFT include gating mechanism, gated residual network (GRN), variable selection network (VSN) and variable selection module; Specifically, the gating mechanism allows features to enter only the modules that need to be used, so that the model can automatically adjust the complexity and network depth to adapt to different scenarios and data structures; correspond Figure 2 The "Gating Layer" module in provides a direct path from sequence to sequence; Gated Residual Networks (GRNs): Simplify models by giving them the flexibility to apply nonlinear processing only when necessary; Variable Selection Network (VSN): selects relevant input variables at each time step, greatly optimizing modeling performance by leveraging learning capabilities only on the most significant features; The variable selection module improves models by removing noisy inputs that can negatively impact performance. It also provides insight into which variables are most important for a prediction problem. MTFT used different VSNs to process static variables, historical observations, and future inputs; Different VSNs process static variables, historical observations, and future inputs, and the parameters of the three VSN modules are not shared; Static Covariate Encoder (SCE): This generates four different context vectors by integrating information from static variable metadata using a separate GRN encoder , the input of SCE is the result of static variables after VSN, where Gave it to VSN, Give LSTM an initialization state. gave the static enrichment layer (SEL); Temporal Fusion Decoder (TFD): learns short-term and long-term temporal relationships between variables from time-varying variable inputs with historical observations and known future values; Sequence-to-sequence layers are used for local processing, while long-term dependencies are captured using a masked interpretable multi-head attention module; TFD consists of three modules: static enrichment layer, time series self-attention layer, and position-based front network layer; Static enrichment layer (SEL): input static covariates to improve time series features; Time Series Self-Attention Layer (TSL): learns long-term dependencies of time series data and makes the model interpretable; Position-based pre-layer (PFL): applies additional non-linear processing to the output of the TSL layer; Step 4: The entire model processing process is (see Figure 2 ): Static variables are passed through the variable selection network (i.e. Figure 2 The VSN module) and the static covariate encoder (SCE) are used for feature engineering, and the time-varying variables are output after feature engineering by VSN; The VSN method embeds the categorical variables as a whole and performs linear transformation on the continuous variables. The transformed variables are dimension vector, the converted variable is , for historical input, the flattened result is (in is the number of variables), and these transformed variables are used for variable selection. The variable selection formula is: ,in is the weight of variable selection; Weight by Calculate and obtain; is the feature after nonlinear processing; Nonlinear processing features are achieved through Get it by Provided by SCE; The historical observation features Input LSTM encoder, future known value features Input LSTM decoder, both generate a set of consistent time series features ,in as a position index; Before entering the temporal fusion decoder (TFD), it is processed by the standard layer-normalized gating layer. is the standard layer-normalized function; In order to ignore unnecessary features in a given data set, TFD adds a component GLU based on a linear gating module. The reason why the GLU component can suppress the nonlinear contribution of the data set is that the output of GLU may be close to 0, so the initial input can also completely skip the gating layer and directly enter the next module. The standardization-normalization gating layer is the preparation work before the feature enters TFD. The specific formula is: ; After the feature data enters the TFD module, the first step is to enhance the data features. This step is completed through the static enrichment layer (SEL), which uses static covariates to improve the time series features. Specifically, it uses the GRN and SCE provided The processing method is: ; Enter the time series self-attention layer (TSL), which applies an interpretable multi-head attention mechanism at each prediction time step. In addition to preserving the causal information flow through "masking", the self-attention layer also allows the MTFT to pick up long-range dependencies. After the self-attention layer, an additional gating layer is used to filter information and improve training effects; This module is equivalent to an interpretable multi-head attention layer plus a standard layer-normalized gating layer. Its main function is to learn the long-term correlation of time series data and mask the interpretable multi-head attention mechanism learning function as shown in the formula As shown, the formula is the gating layer function; As mentioned earlier, the TSL layer preserves the causal information flow by “masking” to make the model interpretable; in ,The final output of the masked interpretable multi-head attention layer in the TSL layer is similar to that of the single attention layer. The key difference is that the former generates attention in a different way; Entering the position-based pre-network layer (PFL), the PFL layer uses an additional nonlinear function to process the output of the time series self-attention layer. The processing function is as follows: As shown in , the features are filtered again using the gated residual network (GRN). Figure 2 It can be seen that the model takes into account that some data may not require additional complex processing, and the effect of directly using a simple model is better. Therefore, a gating layer that can skip the TFD module is set to provide a direct path from sequence to sequence. The learning function of the gating layer is as follows: As shown; Enter the fully connected layer, output the predicted value at each time step as the point prediction, add percentiles to turn the point prediction into an interval prediction, and the fully connected layer uses the processing function as shown in the formula to make the output of the model the predicted carbon asset market price range; in is the linear coefficient for a particular quantile q, , and finally output the carbon asset price forecast intervals at different quantiles and time steps at the same time (taking the 10th, 50th and 90th quantiles as examples); Model loss function: The model is trained by minimizing the joint quantile loss, and all quantile outputs are summed. The model loss function is as follows: ; ; in, refers to the domain of training data containing M samples, W represents the weight parameter of MTFT, is the set of output quantiles, ; In order to avoid errors caused by inconsistent prediction dimensions at different prediction points, the normalized quantile loss in the entire prediction range is evaluated, that is, regularization is performed: ; in, is the domain of the test sample; In summary, the present invention integrates multi-source heterogeneous data, comprehensively considers carbon asset trading market information, trading data, macroeconomic indicators, environmental indexes, policy change records, social media information and future known holiday information, etc., and utilizes the powerful data processing, feature extraction and time series modeling capabilities of the multi-source heterogeneous time series fusion model (MTFT), so that the present invention has a higher accuracy rate in predicting carbon asset market prices than existing models. At the same time, the variable selection network in the MTFT model is used for feature interpretation to enhance the interpretability of the model and make up for the black box disadvantage of most machine learning models.

[0015] Embodiment 2: The present invention also provides an example implementation process of the method: First, data collection: determine the time span from December 1, 2019 to December 1, 2024, and predict the carbon asset price in the next three months at the time point of December 1, 2024. Obtain relevant information about China's local carbon emission trading markets from the official websites of local carbon emission exchanges, including but not limited to the city where the exchange is located, exchange trading hours, information on various carbon asset products, and daily data on carbon asset prices for the past five years; obtain monthly data on various economic indicators for the time period from the National Bureau of Statistics, including but not limited to GDP indicators, urbanization rate, and population growth rate; obtain policy change records and important news events for the time period from the official websites of local governments and national official websites; obtain monthly data on industrial activities for the time period from the official websites of various industry associations and government statistical departments; obtain various environmental variables for the time period from the National Environmental Monitoring Center, including but not limited to climate variables such as temperature, precipitation, and humidity, and atmospheric components such as greenhouse gases and air pollutants; use web crawler technology to capture media information data for the time period from major social media platforms; obtain holiday information from December 1, 2019 to March 1, 2025 from the national holiday schedule announcement; Then, the data is preprocessed by scoring the sentiment tendency of text data such as policies, news events, and social media information, and converting them into numerical sequences; Each value represents the sentiment score of the corresponding word’s impact on the carbon asset market price, which specifically includes the following steps: Data acquisition: Using web crawler technology to capture news data related to the carbon market from news websites, financial websites and other platforms; Data Screening: Screen the acquired data and only retain the data related to carbon trading. Specifically, the data can be selected according to the common keywords of carbon trading (common keywords: carbon emissions, carbon price, carbon trading, carbon credit, carbon offset, carbon tax, carbon market, emission rights, cap and trade, carbon neutrality, carbon capture and storage, voluntary emission reduction, renewable energy certificate, and clean development mechanism); Data Cleaning: Use regular expressions to remove irrelevant content such as HTML tags, URL links, special symbols, and spaces in the text; Word Segmentation: Use the jieba library in python to split continuous text into individual words; Stop Word Removal: By defining a stop word list, retrieve and delete common words that contribute little to the meaning of the text, such as "of", "is", "in", etc.; Text Vectorization: Evaluate the importance of each word in the thesaurus. Specifically, it is achieved by calculating the product of the term frequency (TF) of a word in a document and the inverse document frequency (IDF) in the entire document set. Use the Word2Vec model to convert words into fixed-size vectors, which capture the semantic information of the words; Call the Deep Learning Model: Use the SnowNLP tool to convert the text vector into a corresponding sentiment score. The sentiment score ranges from [0, 1], where 0 represents very negative and 1 represents very positive; Use unique encoding for categorical variables, such as encoding the carbon asset exchange market information; First, use the IQR method to remove outliers from all numerical data, then use the cubic spline function to impute missing values, and finally use the Z-score method to process all numerical data to make the data comparable; Convert all data sets into time series based on timestamps, where each timestamp corresponds to a specific date, accurate to the day. Specifically, the timestamp refers to the number of days between each time point and December 1, 2019. For example, represents the variable value on December 2, 2019 ; Convert all data into daily data through resampling methods; Data Set Division: Divide all time series into three parts: a training set for learning, a validation set for hyperparameter tuning, and a hold-out test set for performance evaluation.

[0016] Subsequently, Model Construction: Use MTFT to process multivariate heterogeneous data, including static variable s - carbon trading market exchange information, historical observations ——Carbon asset prices, economic indicators, industrial activities, environmental variables, important news events, policy change records and social media sentiment from December 1, 2019 to December 1, 2024, with a priori known future variables —— Non-trading days and holidays from December 2, 2024 to March 1, 2025, namely ; Construct a multi-step prediction model, such as formula As shown, the prediction results of the 10th, 50th and 90th quantiles are output. Then, through the GRN, VSN, SCE, TED and other components of the MTFT model, the feature representation and time relationship in the dataset are learned.

[0017] After model training and validation: First, the MTFT model is trained using the training set data and optimized by minimizing the joint quantile loss function; then the model hyperparameters are adjusted using the validation set data to improve model performance. Hyperparameter optimization is performed by random search, iterated 60 times, and the complete search range of all hyperparameters is set as shown in Table 1; finally, the prediction accuracy and stability of the model are evaluated on the test set to determine the optimal model parameters. In order to maintain the interpretability of the model, a masked interpretable multi-head attention layer is used. Similar code can be found on GitHub3 and can be modified according to the actual data field.

[0018] Table 1 Hyperparameter search range

[0019] After that, the result is output and applied: the model with learned variable features and optimal parameters is fitted on the test set and the model is evaluated, and different step lengths are output at the same time. The daily price forecast range of carbon assets under the 10th, 50th and 90th percentiles - that is, The forecast results provide risk management tools for carbon asset market participants. The forecast results are used in asset trading strategies to improve the scientificity and effectiveness of trading decisions.

[0020] Finally, the model is explained: By analyzing the formula The variable selection weights in quantify the importance of variables; Specifically, the variable selection weight for each variable in the test set is The sum is taken and the 10th, 50th and 90th percentiles of each sample are recorded. The importance of the variables is measured by comparing and analyzing the selection weights of each variable. The degree of influence of different variables on carbon asset prices can be directly obtained.

[0021] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A carbon asset price prediction method based on multi-source fusion, characterized by: The following steps are involved: Step 1: Input data: The input variables of the model are: static variables that do not change with time during the study period, time series variables of past observations that are known in the past but unknown in the future, and a priori known time series variables that are known in both the past and the future; Step 2: Data preprocessing: Convert text variables such as policies, news events, and social media information into numerical sequences. This function can be achieved by scoring the sentiment tendency of text information. Each output value represents the sentiment score of the corresponding word's impact on the carbon asset market price. Encode categorical variables (such as carbon asset trading market information) using unique codes; Process outliers and missing values ​​for other numerical data, and then standardize all numerical data; convert all variable data into time series based on timestamps, where each timestamp corresponds to a specific date and time point, accurate to the day; All data were converted into daily data through resampling method; All time series are divided into three parts: training set for learning, validation set for hyperparameter tuning, and holdout test set for performance evaluation; Step 3: Model structure and principle analysis: In time series tasks, there are mainly two types of data sources: static variables and time-varying variables. Time-varying variables can be divided into time-varying variables observed in the past and time-varying variables with known future values. Designing network structures for different data sources helps improve the prediction effect of the model. Establishment of multivariate heterogeneous model for multi-step forecasting: Assume that there is tradable products, each with a future price and a set of static variables , each time step The input time-varying variables under and the target variable related; Time-varying variables are divided into two categories, namely ,in represents a time-varying variable with historical observations but unknown future values. Indicates that past and future values ​​are known time-varying variables; Quantile prediction is used for the prediction of target variables affected by multi-source heterogeneous variables, such as inputting the 10th, 50th, and 90th quantiles at each time step; Each quantile prediction model is as follows ; in, Indicated in Predict the future at all times The qth quantile value of the carbon asset i price in the next step, A carbon asset price prediction model; , that is, output simultaneously , , , The predicted price of carbon assets at the moment; Indicates the past Carbon assets The price, Indicates the past Impact on carbon assets at every moment The observed input of the price, Indicates the past Expiration Impact of carbon assets The price of the future known input, Indicates that the time span of the research object is the past Period, with 1 day as 1 period, Indicates impact on carbon assets Static variable of price; Model architecture and principle analysis The input data is divided into three categories: one is static variable s, and the other is historical value , one is the future known value ,MTFT can construct feature representations for three different types of data sets using different components; The main components and functions of MTFT include gating mechanism, gated residual network (GRN), variable selection network (VSN) and variable selection module; MTFT used different VSNs to process static variables, historical observations, and future inputs; Static Covariate Encoder (SCE): This generates four different context vectors by integrating information from static variable metadata using a separate GRN encoder , the input of SCE is the result of static variables after VSN, where Gave it to VSN, Give LSTM an initialization state. gave the static enrichment layer (SEL); Temporal Fusion Decoder (TFD): learns short-term and long-term temporal relationships between variables from time-varying variable inputs with historical observations and known future values; Sequence-to-sequence layers are used for local processing, while long-term dependencies are captured using a masked interpretable multi-head attention module; TFD consists of three modules: static enrichment layer, time series self-attention layer, and position-based front network layer; Static enrichment layer (SEL): input static covariates to improve time series features; Time Series Self-Attention Layer (TSL): learns long-term dependencies of time series data and makes the model interpretable; Position-based pre-layer (PFL): applies additional non-linear processing to the output of the TSL layer; Step 4: The entire model processing process is: Static variables are feature engineered by the variable selection network (VSN module) and the static covariate encoder (SCE), and time-varying variables are feature engineered by VSN and then output features; The VSN method embeds the categorical variables as a whole and performs linear transformation on the continuous variables. The transformed variables are dimension vector, the converted variable is , for historical input, the flattened result is (in is the number of variables), and these transformed variables are used for variable selection. The variable selection formula is: ,in is the weight of variable selection; Weight by Calculate and obtain; is the feature after nonlinear processing; Nonlinear processing features are achieved through Get it by Provided by SCE; The historical observation features Input LSTM encoder, future known value features Input LSTM decoder, both generate a set of consistent time series features ,in as a position index; Before entering the temporal fusion decoder (TFD), it is processed by the standard layer-normalized gating layer. is the standard layer-normalized function; In order to ignore unnecessary features in a given data set, TFD adds a component GLU based on a linear gating module. The reason why the GLU component can suppress the nonlinear contribution of the data set is that the output of GLU may be close to 0, so the initial input can also completely skip the gating layer and directly enter the next module. The standardization-normalization gating layer is the preparation work before the feature enters TFD. The specific formula is: ; After the feature data enters the TFD module, the first step is to enhance the data features. This step is completed through the static enrichment layer (SEL), which uses static covariates to improve the time series features. Specifically, it uses the GRN and SCE provided The processing method is: ; Enter the time series self-attention layer (TSL), which applies an interpretable multi-head attention mechanism at each prediction time step. In addition to preserving the causal information flow through "masking", the self-attention layer also allows the MTFT to pick up long-range dependencies. After the self-attention layer, an additional gating layer is used to filter information and improve training effects; This module is equivalent to an interpretable multi-head attention layer plus a standard layer-normalized gating layer. Its main function is to learn the long-term correlation of time series data and mask the interpretable multi-head attention mechanism learning function as shown in the formula As shown, the formula is the gating layer function; As mentioned earlier, the TSL layer preserves the causal information flow by "masking" to make the model interpretable; in ,The final output of the masked interpretable multi-head attention layer in the TSL layer is similar to that of the single attention layer. The key difference is that the former generates attention in a different way; Entering the position-based pre-network layer (PFL), the PFL layer uses an additional nonlinear function to process the output of the time series self-attention layer. The processing function is as follows: As shown, the features are filtered again using the gated residual network (GRN); The model takes into account that some data may not require additional complex processing, and the effect of directly using a simple model is better. Therefore, a gating layer that can skip the TFD module is set up to provide a direct path from sequence to sequence. The learning function of the gating layer is as follows: As shown; Enter the fully connected layer, output the predicted value at each time step as the point prediction, add percentiles to turn the point prediction into an interval prediction, and the fully connected layer uses the processing function as shown in the formula to make the output of the model the predicted carbon asset market price range; in is the linear coefficient for a particular quantile q, , and finally output the carbon asset price forecast intervals at different quantiles and time steps at the same time (taking the 10th, 50th and 90th quantiles as examples); Model loss function: The model is trained by minimizing the joint quantile loss, and all quantile outputs are summed. The model loss function is as follows: ; ; in, refers to the domain of training data containing M samples, W represents the weight parameter of MTFT, is the set of output quantiles, ; In order to avoid errors caused by inconsistent prediction dimensions at different prediction points, the normalized quantile loss in the entire prediction range is evaluated, that is, regularization is performed: ; in, is the domain of the test sample.

2. The method for predicting carbon asset prices under multi-source fusion according to claim 1 is characterized by: In step 1, static variables that do not change over time are studied, such as carbon asset trading market mechanisms and macro policies; time series variables of past observations that are known in the past but unknown in the future, such as historical carbon asset prices, meteorological data, economic indicators, policy change records, industrial activities, environmental variables, and social media information; time series variables that are known a priori in both the past and the future, such as non-trading time periods and holidays.

3. The method for predicting carbon asset prices under multi-source fusion according to claim 1 is characterized in that: In step 2, the exact date is Represents a variable of a certain year, month and day The value of .

4. The method for predicting carbon asset prices under multi-source fusion according to claim 1 is characterized in that: In step 3, the gating mechanism: This mechanism allows features to enter only the modules that need to be used, so that the model can automatically adjust the complexity and network depth to adapt to different scenarios and data structures; Gated Residual Networks (GRNs): Simplify models by giving them the flexibility to apply nonlinear processing only when necessary; Variable Selection Network (VSN): selects relevant input variables at each time step, greatly optimizing modeling performance by leveraging learning capabilities only on the most significant features; The variable selection module improves models by removing noisy inputs that can negatively impact performance, and also provides insight into which variables are most important for your prediction problem.

5. The method for predicting carbon asset prices under multi-source fusion according to claim 1 is characterized by: In step 3, different VSNs process static variables, historical observations, and future inputs, and the parameters of the three VSN modules are not shared.