BERT language-based carbon emission prediction method
By combining the BERT language model and VMD-EMD decomposition method, using the LSTM model to process carbon emission data, the problem of inaccurate carbon emission prediction in the existing technology is solved, and a higher accuracy and stable prediction effect is achieved.
Patent Information
- Application Number
- CN202510644968.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing carbon emission forecasting methods fail to fully understand the text information, cannot accurately reflect the trend of carbon emission changes, and the model prediction is difficult.
The BERT language model is used to process exogenous features, combined with the VMD-EMD method to decompose carbon emission sequences, use the LSTM model to predict, and integrate factors such as meteorological and urban development.
It improves the accuracy and stability of carbon emission forecasts, reduces the difficulty of prediction, and captures the implicit relationships in carbon emission data.
Smart Images

Figure CN120450152A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of carbon emission prediction, and in particular to a carbon emission prediction method based on BERT language. Background Art
[0002] Observational data show that global average temperatures have shown a clear upward trend over the past century. This temperature rise is closely related to greenhouse gas emissions, particularly carbon dioxide (CO2), caused by human activities. To address climate change, it is necessary to accurately understand carbon emissions and predict their future trends so that effective emission reduction measures can be implemented. With the development of big data technology, massive amounts of historical data on carbon emissions in various regions have accumulated, providing a foundation for data-driven forecasting methods. However, extracting effective features from this massive data and building efficient and accurate forecasting models are current research priorities.
[0003] The patent application with publication number CN 119129832A discloses a carbon emission prediction method based on an LSTM network model, including: decomposing the original carbon emission sequence into a series of VMF modal components and residuals with limited bandwidth; introducing white noise weighting and multiple reconstructions to obtain IMF components; adding the Luong attention mechanism to the LSTM long short-term memory network model; inputting the VMF modal components and IMF components into the improved LSTM long short-term memory network model, superimposing and reconstructing each prediction result to obtain an attention output vector.
[0004] During the prediction process, this existing technology cannot fully understand text information and capture the implicit relationships in carbon emission data, which is not conducive to more accurately reflecting the trend of carbon emission changes.
[0005] Therefore, a new technical solution is needed to solve the above technical problems. Summary of the Invention
[0006] The purpose of the present invention is to provide a carbon emission prediction method based on BERT language, which uses the BERT language model to process exogenous features, thereby improving the accuracy of the carbon emission prediction model. At the same time, the VMD-EMD method is used to decompose the carbon emission sequence, and the complex time series signal is decomposed into multiple simple modal components, which reduces the prediction difficulty of the model and improves the stability and accuracy of the prediction. This combined method not only improves the prediction performance of the model, but also provides a new idea and technical means for carbon emission prediction.
[0007] The technical solution adopted in the present invention is:
[0008] A BERT-based carbon emission prediction method includes the following steps:
[0009] Step 1: Decompose the original carbon emission sequence using the variational mode decomposition algorithm to obtain a series of modal components and a residual component;
[0010] Step 2: Use empirical mode decomposition to decompose the residual component to obtain a series of modal components and a residual component;
[0011] Step 3: Organize the factors affecting carbon emissions, including meteorological indicators, urban development, national management policies, vegetation coverage, industrial structure, and the popularization of carbon emission knowledge among the population. Use the BERT algorithm to process the exogenous feature text data to obtain exogenous features.
[0012] Step 4: Combine the multiple subsequences obtained after the VMD-EMD algorithm with exogenous feature information and perform normalization to eliminate the impact of different dimensions. Then, establish an LSTM model to predict carbon emissions.
[0013] Step 5: Superimpose the prediction results of all component models to obtain the final carbon emission prediction.
[0014] By employing these prediction methods, the BERT model is able to fully understand textual information and capture implicit relationships within carbon emissions data, thereby more accurately reflecting trends in carbon emissions. Furthermore, the VMD-EMD method is used to decompose the carbon emissions sequence, breaking down the complex time series signal into multiple simple modal components. This reduces the model's prediction difficulty and improves both stability and accuracy. This combined approach not only enhances the model's predictive performance but also provides a new approach and technical approach for carbon emissions forecasting.
[0015] Preferably, in step 1, the variational mode decomposition (VMD) method is used to decompose the original carbon emission sequence f into several mode functions {u k (t)}, k=1, 2,…, K.
[0016] Preferably, the operating steps of step 1 are:
[0017] Step 1.1: For each mode u k (t) Perform Hilbert transform to obtain the one-sided spectrum:
[0018]
[0019] Step 1.2: By mixing the center frequency ω k The estimated index term Shift the spectrum of a mode to the corresponding baseband:
[0020]
[0021] Step 1.3: Estimate the bandwidth of each mode by calculating the L2 norm of the above signal, and obtain the constrained variational problem:
[0022]
[0023] Where: {u k}={u1,u2,...,u k} are the K modal components after decomposition; {ω k}={ω1,ω2,...,ω k} is the center frequency corresponding to each mode; is the sum of all modes.
[0024] Step 1.4: Make the constrained variational problem unconstrained by introducing the quadratic penalty factor α and the Lagrange multiplication operator λ(t):
[0025]
[0026] Step 1.5: Update by alternation and λ n+1 To solve the above variational problem, the update methods for modal components and center frequencies are:
[0027]
[0028] in, is the Wiener filter of the current residual; The real part of the inverse Fourier transform is {u k (t)}; is the center of gravity of the current modal power spectrum; are the Fourier transform values of f(ω), u(ω), and λ(ω) respectively.
[0029] By adopting the above prediction method, during the prediction process, the algorithm uses VMD to decompose the original carbon emission sequence to obtain subsequences AFS1, AFS2, ..., AFS m and residual terms, making it easier to obtain a model with higher prediction accuracy.
[0030] Preferably, in step 2, EMD is used to decompose the VMD residual term with strong non-stationarity. Assuming that the VMD residual term is s(t), the EMD decomposition steps are as follows:
[0031] Step 2.1: Determine all extreme points of s(t). Connect all the maximum and minimum points with a curve to form the upper and lower envelopes of VMD-EMD-LSTM, with all the signal data in the middle. The average of the upper and lower envelopes is called m1, and the difference between s(t) and m1 is called h1:
[0032] s(t)-m1=h1 (6)
[0033] Step 2.2: Replace s(t) with h1 and repeat step 1.2.1k times until h1k satisfies the IMF assumptions, obtaining the first IMF component c1:
[0034] c1=h1k (7)
[0035] Step 2.3: Separate c1 from s(t) and record it as r1:
[0036] r1=s(t)-c(1) (8)
[0037] Step 2.4: r1 is considered as the new s(t), and steps 1.2.1 to 1.2.3 are repeated until the termination condition is met, and n IMF components c1, c2, ..., c n and the residual component r n . Thus, EMD decomposes the residual term s(t) into:
[0038]
[0039] By adopting the above prediction method and using EMD to decompose the VMD residual term with strong non-stationarity, the difficulty of carbon emission prediction can be further reduced.
[0040] Preferably, step 3: using the BERT algorithm to process the text data to obtain exogenous features includes the following steps:
[0041] Step 3.1: Organize textual data on factors influencing carbon emissions, including: meteorological indicators, urban development, national management policies, vegetation coverage, industrial structure, and the level of knowledge about carbon emissions among the population;
[0042] Step 3.1: Preprocess the text data to remove punctuation, special symbols, and spaces, and then feed the dataset into the BERT model for training.
[0043] The formula of the attention mechanism is expressed as:
[0044]
[0045] Where, d k Indicates the dimension of k, QK T Indicates that similarity calculation is achieved through matrix multiplication.
[0046] By adopting the above prediction method, the BERT language model is used to process exogenous features. BERT is a pre-trained language model that has achieved remarkable results in the field of natural language processing and can be widely used in various NLP tasks. The BERT model can fully understand text information and capture the implicit relationships in carbon emission data, thereby more accurately reflecting the trend of carbon emission changes.
[0047] Preferably, step 4: LSTM is an improved model of RNN, which adds a forget gate f on the basis of RNN t , input gate i t AND output gate O t , which alleviates the gradient vanishing problem. At each moment, LSTM accepts the current moment input x through the gate t , the hidden state h at the previous moment t-1 and the old memory state C t-1 , get the summary status and the new memory state C t And update the output h t , the internal operation process of the LSTM unit is as follows:
[0048] (1) Obtain f through the forget gate t :
[0049] f t =σ(ω f ·[h t-1 , x t ]+b f ) (11)
[0050] (2) Input gate to get i t :
[0051] i t =σ(ω i ·[h t-1 , x t ]+b i ) (12)
[0052] (3) Output gate gets O t :
[0053] O t =σ(ω O ·[h t-1 , x t ]+b O ) (13)
[0054] (4) Calculate summary status
[0055]
[0056] (5) Calculate the new memory state C t :
[0057]
[0058] (6) Update output h t :
[0059] h t =O t tanh(C t ) (16)
[0060] Where σ represents the Sigmoid activation function; tanh represents the hyperbolic tangent function; ω f 、ω i 、ω O 、ω C represents the weight matrix; b f 、b i 、b O 、b C represents the bias term.
[0061] By adopting the above prediction method, RNN is proposed to process time series data. The output is jointly determined by the current input and the memory state of the historical moment. It has certain advantages in processing time series data. The improved RNN model LSTM is used to alleviate the gradient vanishing problem and optimize the prediction results.
[0062] Preferably, in step 5, the prediction results of all component models are synthesized, reconstructed, and superimposed to obtain the final predicted carbon emissions.
[0063] Compared with the prior art, the present invention has the following beneficial effects:
[0064] 1. This paper uses the BERT language model to process exogenous features, improving the accuracy of the carbon emission prediction model. The BERT model can fully understand textual information and capture the implicit relationships in carbon emission data, thereby more accurately reflecting the trend of carbon emission changes. At the same time, the VMD-EMD method is used to decompose the carbon emission sequence, breaking down the complex time series signal into multiple simple modal components, reducing the model's prediction difficulty and improving the stability and accuracy of the prediction.
[0065] 2. In the prediction process of the present invention, the algorithm uses VMD to decompose the original carbon emission sequence to obtain subsequences AFS1, AFS2, ..., AFS m and residual terms, which makes it easier to obtain a model with higher prediction accuracy, and using EMD to decompose the VMD residual terms with strong non-stationarity can further reduce the difficulty of carbon emission prediction.
[0066] 3. The prediction method of the present invention uses an improved RNN model LSTM. RNN is proposed to process time series data. The output is determined by the current input and the memory state of the historical moment. It has certain advantages in processing time series data. The use of LSTM alleviates the gradient vanishing problem and optimizes the prediction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 The internal structure of the BERT model of the present invention;
[0068] Figure 2 This is the internal structure diagram of the LSTM of the present invention;
[0069] Figure 3 It is a complete technical flow chart of the present invention. DETAILED DESCRIPTION
[0070] like Figure 1-3 As shown, a BERT-based carbon emission prediction method includes the following steps:
[0071] Step 1: Use the variational mode decomposition (VMD) algorithm to decompose the original carbon emission sequence to obtain a series of modal components and a residual component;
[0072] Step 2: Use Empirical Mode Decomposition (EMD) to decompose the residual component to obtain a series of modal components and a residual component;
[0073] Step 3: Organize the factors affecting carbon emissions, including meteorological indicators, urban development, national management policies, vegetation coverage, industrial structure, and the popularization of carbon emission knowledge among the population. Use the BERT algorithm to process the exogenous feature text data to obtain exogenous features.
[0074] Step 4: Combine the multiple subsequences obtained after the VMD-EMD algorithm with exogenous feature information and perform normalization to eliminate the impact of different dimensions. Then, establish an LSTM model to predict carbon emissions.
[0075] Step 5: Superimpose the prediction results of all component models to obtain the final carbon emission prediction.
[0076] Regarding step 1, the variational mode decomposition (VMD) method is used to decompose the original carbon emission sequence f into several mode functions {u k (t)}, k=1, 2, ..., K. In the VMD-EMD algorithm, VMD is used to decompose the original carbon emission sequence to obtain subsequences AFS1, AFS2, ..., AFS mAnd residual terms. Then use EMD to decompose the residual terms and get the subsequences IMF1, AFS2, ..., AFS n .
[0077] Step 1.1: For each mode u k (t) Perform Hilbert transform to obtain the one-sided spectrum:
[0078]
[0079] Step 1.2: By mixing the center frequency ω k The estimated index term Shift the spectrum of a mode to the corresponding baseband:
[0080]
[0081] Step 1.3: Estimate the bandwidth of each mode by calculating the L2 norm of the above signal, and obtain the constrained variational problem:
[0082]
[0083] Where: {u k}={u1,u2,...,u k} are the K modal components after decomposition; {ω k}={ω1,ω2,...,ω k} is the center frequency corresponding to each mode; is the sum of all modes.
[0084] Step 1.4: Make the constrained variational problem unconstrained by introducing the quadratic penalty factor α and the Lagrange multiplication operator λ(t):
[0085]
[0086] Step 1.5: Update by alternation and λ n+1 To solve the above variational problem, the update methods for modal components and center frequencies are:
[0087]
[0088] in, is the Wiener filter of the current residual; The real part of the inverse Fourier transform is {u k (t)}; is the center of gravity of the current modal power spectrum; are the Fourier transform values of f(ω), u(ω), and λ(ω) respectively.
[0089] Regarding step 2, EMD is used to decompose the VMD residual term with strong non-stationarity, aiming to further reduce the difficulty of carbon emission prediction. Assuming that the VMD residual term is s(t), the EMD decomposition steps are as follows:
[0090] Step 2.1: Determine all extreme points of s(t). Connect all the maximum and minimum points with a curve to form the upper and lower envelopes of VMD-EMD-LSTM, with all the signal data in the middle. The average of the upper and lower envelopes is called m1, and the difference between s(t) and m1 is called h1:
[0091] s(t)-m1=h1 (6)
[0092] Step 2.2: Replace s(t) with h1 and repeat step 1.2.1k times until h1k satisfies the IMF assumptions, obtaining the first IMF component c1:
[0093] c1=h1k (7)
[0094] Step 2.3: Separate c1 from s(t) and record it as r1:
[0095] r1=s(t)-c(1) (8)
[0096] Step 2.4: r1 is considered as the new s(t), and steps 1.2.1 to 1.2.3 are repeated until the termination condition is met, and n IMF components c1, c2, ..., c n and the residual component r n . Thus, EMD decomposes the residual term s(t) into:
[0097]
[0098] Regarding step 3, the BERT algorithm is used to process the text data to obtain exogenous features. BERT is a pre-trained language model that has achieved remarkable results in the field of natural language processing (NLP) and has been widely used in various NLP tasks, such as Figure 1 The internal structure of the BERT model shown includes the following processes:
[0099] Step 3.1: Organize textual data on factors influencing carbon emissions, including: meteorological indicators, urban development, national management policies, vegetation coverage, industrial structure, and the level of knowledge about carbon emissions among the population;
[0100] Step 3.1: Preprocess the text data to remove punctuation, special symbols, and spaces, and then feed the dataset into the BERT model for training.
[0101] The formula of the attention mechanism is expressed as:
[0102]
[0103] Where, d k Indicates the dimension of k, QK T Indicates that similarity calculation is achieved through matrix multiplication.
[0104] Regarding step 4: RNN was proposed to process time series data. The output is determined by the current input and the memory state of the past. It has certain advantages in processing time series data. However, RNN has the problem of vanishing gradient due to its structural defects. LSTM is an improved model of RNN. It adds a forget gate f on the basis of RNN. t , input gate i t AND output gate O t , which alleviates the gradient vanishing problem, such as Figure 2 The internal structure diagram of LSTM is shown.
[0105] At each moment, LSTM accepts the current moment input x through the gate t , the hidden state h at the previous moment t-1 and the old memory state C t-1 , get the summary status and the new memory state C t And update the output h t The internal operation process of the LSTM unit is as follows:
[0106] (1) Obtain f through the forget gate t :
[0107] f t =σ(ω f ·[h t-1 , x t ]+b f ) (11)
[0108] (2) Input gate to get i t :
[0109] i t =σ(ω i ·[h t-1 , x t ]+b i ) (12)
[0110] (3) Output gate gets O t :
[0111] O t =σ(ω O ·[h t-1 , x t ]+b O ) (13)
[0112] (4) Calculate summary status
[0113]
[0114] (5) Calculate the new memory state C t :
[0115]
[0116] (6) Update output h t :
[0117] h t =O t tanh(C t ) (16)
[0118] In the formula, σ represents the Sigmoid activation function; tanh represents the hyperbolic tangent function; ω f 、ω i 、ω O 、ω C represents the weight matrix; b f 、b i 、b O 、b C represents the bias term.
[0119] Regarding step 5, Figure 3 As shown in the figure, it is a complete technical flow chart, which synthesizes, reconstructs and superimposes the prediction results of all component models to obtain the final predicted carbon emissions.
[0120] In the specific prediction process of the present invention, in step 1, text information such as carbon emission data, meteorological indicators, urban development, national management policies, vegetation coverage, industrial structure, and carbon emission knowledge popularization of a certain region in a certain year is collected. The last day of the data in this year is used as the test set, and the remaining 364 days are used as the training set. In order to eliminate the impact of differences between multiple dimensions on prediction accuracy, the data is first normalized:
[0121]
[0122] In the formula, x′ is the normalized result, x is the actual result, and x min and x max are the minimum and maximum values respectively.
[0123] The present invention selects Mean Square Error (MSE), Mean Absolute Error (MAE) and Symmetrical Mean Absolute Percentage Error (SMAPE) to evaluate the prediction model.
[0124]
[0125] Where y represents the real electricity price, Represents the predicted electricity price, and n represents the number of samples.
[0126] In step 2, the original carbon emission data has strong volatility. In order to obtain a model with higher prediction accuracy, the original carbon emission data is first decomposed by VMD-EMD to obtain the decomposed carbon emission sequence.
[0127] In step 3, to verify the predictive effectiveness of the VMD-EMD-BERT-LSTM algorithm, we compared it with the LSTM model and the VMD-EMD-LSTM model. The three models were used to predict carbon emissions for the selected region on the last day of the year. The prediction results are shown in Table 1 below:
[0128] Table 1 Statistics of carbon emission prediction model indicators
[0129]
[0130] The prediction results of VMD-EMD-BERT-LSTM are closer to the real data, which shows that the prediction performance of VMD-EMD-BERT-LSTM is better than that of LSTM and VMD-EMD-LSTM models. In the present invention, the decomposition of the carbon emission sequence by VMD-EMD improves the accuracy of the model, and the processing of carbon emission influencing factors by the BERT algorithm further improves the accuracy of the prediction. Combining the two with LSTM realizes the optimization of the prediction model.
[0131] The embodiments described above are merely descriptions of preferred implementations of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by those skilled in the art should be included in the scope of protection determined by the claims of the present invention.
Claims
1. A BERT-based carbon emission prediction method, characterized by: The following steps are involved: Step 1: Decompose the original carbon emission sequence using the variational mode decomposition algorithm to obtain a series of modal components and a residual component; Step 2: Use empirical mode decomposition to decompose the residual component to obtain a series of modal components and a residual component; Step 3: Organize the factors affecting carbon emissions, including meteorological indicators, urban development, national management policies, vegetation coverage, industrial structure, and the popularization of carbon emission knowledge among the population. Use the BERT algorithm to process the exogenous feature text data to obtain exogenous features. Step 4: Combine the multiple subsequences obtained after the VMD-EMD algorithm with exogenous feature information and perform normalization to eliminate the impact of different dimensions. Then, establish an LSTM model to predict carbon emissions. Step 5: Superimpose the prediction results of all component models to obtain the final carbon emission prediction.
2. The BERT-based carbon emission prediction method according to claim 1, characterized in that: In step 1, the variational mode decomposition (VMD) method is used to decompose the original carbon emission sequence f into several mode functions {u k (t)}, k=1, 2,…, K.
3. The BERT-based carbon emission prediction method according to claim 2, characterized in that: The operation steps of step 1 are: Step 1.1: For each mode u k (t) Perform Hilbert transform to obtain the one-sided spectrum: Step 1.2: By mixing the center frequency ω k The estimated index term Shift the spectrum of a mode to the corresponding baseband: Step 1.3: Estimate the bandwidth of each mode by calculating the L2 norm of the above signal, and obtain the constrained variational problem: Where: {u k }={u1,u2,...,u k } are the K modal components after decomposition; {ω k }={ω1,ω2,...,ω k } is the center frequency corresponding to each mode; is the sum of all modes. Step 1.4: Make the constrained variational problem unconstrained by introducing the quadratic penalty factor α and the Lagrange multiplication operator λ(t): Step 1.5: Update by alternation and λ n+1 To solve the above variational problem, the update methods for modal components and center frequencies are: in, is the Wiener filter of the current residual; The real part of the inverse Fourier transform is {u k (t)}; is the center of gravity of the current modal power spectrum; are the Fourier transform values of f(ω), u(ω), and λ(ω) respectively.
4. The BERT-based carbon emission prediction method according to claim 1, characterized in that: In step 2, EMD is used to decompose the VMD residual term with strong non-stationarity. Assuming that the VMD residual term is s(t), the EMD decomposition steps are as follows: Step 2.1: Determine all extreme points of s(t). Connect all the maximum and minimum points with a curve to form the upper and lower envelopes of VMD-EMD-LSTM, with all the signal data in the middle. The average of the upper and lower envelopes is called m1, and the difference between s(t) and m1 is called h1: s(t)-m1=h1 (6) Step 2.2: Replace s(t) with h1 and repeat step 1.2.1k times until h1k satisfies the IMF assumptions, obtaining the first IMF component c1: c1=h1k (7) Step 2.3: Separate c1 from s(t) and record it as r1: r1=s(t)-c(1) (8) Step 2.4: r1 is considered as the new s(t), and steps 1.2.1 to 1.2.3 are repeated until the termination condition is met, and n IMF components c1, c2, ..., c n and the residual component r n . Thus, EMD decomposes the residual term s(t) into:
5. The BERT-based carbon emission prediction method according to claim 1, characterized in that: Step 3: Using the BERT algorithm to process text data to obtain exogenous features includes the following steps: Step 3.1: Organize textual data on factors influencing carbon emissions, including: meteorological indicators, urban development, national management policies, vegetation coverage, industrial structure, and the level of knowledge about carbon emissions among the population; Step 3.1: Preprocess the text data to remove punctuation, special symbols, and spaces, and then feed the dataset into the BERT model for training. The formula of the attention mechanism is expressed as: Where, d k Indicates the dimension of k, QK T Indicates that similarity calculation is achieved through matrix multiplication.
6. The BERT-based carbon emission prediction method according to claim 1, characterized in that: Step 4: LSTM is an improved model of RNN, which adds a forget gate f on the basis of RNN t , input gate i t AND output gate O t , which alleviates the gradient vanishing problem. At each moment, LSTM accepts the current moment input x through the gate t , the hidden state h at the previous moment t-1 and the old memory state C t-1 , get the summary state C~ t and the new memory state C t And update the output h t , the internal operation process of the LSTM unit is as follows: (1) Obtain f through the forget gate t : f t =σ(ω f ·[h t-1 ,x t ]+b f ) (11) (2) Input gate to get i t : I t =σ(ω i ·[h t-1 ,x t ]+b i ) (12) (3) Output gate gets O t : The t =σ(ω O ·[h t-1 ,x t ]+b O ) (13) (4) Calculate summary status (5) Calculate the new memory state C t : (6) Update output h t : h t =O t ·tanh(C t ) (16) In the formula, σ represents the Sigmoid activation function; tanh represents the hyperbolic tangent function; ω f 、ω i 、ω O 、ω C represents the weight matrix; b f 、b i 、b O 、b C represents the bias term.
7. The BERT-based carbon emission prediction method according to claim 1, characterized in that: In step 5, the prediction results of all component models are synthesized, reconstructed, and superimposed to obtain the final predicted carbon emissions.
Citation Information
Patent Citations
Carbon emission prediction method based on LSTM network model
CN119129832A