A method for predicting carbon emissions of transportation vehicles operating in high-altitude areas
By using the LSTM network model to train and correlate the environmental and vehicle data in high-altitude areas, calculate the carbon emission correlation coefficient and conduct secondary training, the accuracy of carbon emission prediction in high-altitude areas is solved, and more accurate carbon emission prediction and emission reduction measures are achieved.
Patent Information
- Application Number
- CN202510495371.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The existing carbon emission forecasting model cannot fully consider the differences in carbon emission characteristics of different types of vehicles under different operating conditions in high-altitude areas, resulting in a large deviation from the actual carbon emission conditions, making it difficult to meet the demand for accurate carbon emission forecasts in the transportation industry in high-altitude areas.
The LSTM network model is adopted to collect and preprocess environmental data, operational transportation vehicle data and operation data in high altitude areas, conduct training and correlation analysis, calculate the comprehensive correlation coefficients of carbon emissions, concentrations and ranges, and use them as additional features for secondary training, optimize model parameters to improve prediction accuracy.
It improves the accuracy of carbon emission forecasts in high-altitude areas, can more accurately capture the impact of different factors on carbon emissions, provide targeted emission reduction measures, and support energy conservation and emission reduction in transportation industries in high-altitude areas.
Smart Images

Figure CN120031211B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of carbon emission prediction, and particularly to a method for predicting carbon emissions of transportation vehicles operating in high-altitude areas. Background Art
[0002] In high-altitude areas, the fuel performance and carbon emissions of transportation vehicles are affected by special environmental conditions. Environmental conditions in this area, such as altitude, temperature, air pressure, oxygen content, and other factors, have a significant impact on the performance and carbon emissions of transportation vehicles. For example, as the altitude increases, the air gradually thins, which leads to a decrease in the intake air volume of the vehicle engine, thereby affecting the combustion efficiency, causing the vehicle performance to decline, and at the same time, the carbon emissions also change; the change in temperature will affect the fluidity of the fuel and the viscosity of the vehicle lubricating oil, further interfering with the normal operating state of the vehicle, and ultimately affecting the carbon emission situation; the fluctuation of air pressure cannot be ignored either, as it will affect the power output and energy consumption of the vehicle, thereby indirectly affecting the carbon emission level.
[0003] In addition, the types of transportation vehicles operating in high-altitude areas are rich and diverse, covering passenger cars, trucks, etc. Different types of vehicles have significant differences in carbon emission characteristics due to differences in their design purposes, engine powers, load capacities, etc. For example, passenger cars are mainly used for personnel transportation, and their driving conditions are relatively stable, while trucks have more complex and variable operating conditions due to the diversity of loads and transportation routes, which makes there be obvious differences in their carbon emission characteristics.
[0004] However, when the existing carbon emission prediction models are applied to high-altitude areas, there are obvious limitations. These models usually conduct unified carbon emission predictions for all transportation vehicles operating in high-altitude areas, and fail to fully consider the differences in carbon emission characteristics of different types of vehicles under different operating conditions. This "one-size-fits-all" approach cannot accurately reflect the actual situation, resulting in a large deviation between the prediction results and the actual carbon emission situation, and it is difficult to meet the demand for accurate carbon emission prediction in the transportation industry in high-altitude areas.
[0005] Therefore, there is an urgent need for a method for predicting carbon emissions of transportation vehicles operating in high-altitude areas to improve the accuracy of carbon emission prediction in plateau areas and provide strong support for energy conservation, emission reduction, and sustainable development of the transportation industry in high-altitude areas. Summary of the Invention
[0006] In order to solve the above technical problems, the present invention provides a method for predicting carbon emissions of transportation vehicles operating in high-altitude areas, which can improve the accuracy of carbon emission prediction.
[0007] The present invention provides a method for predicting carbon emissions of transportation vehicles operating in high-altitude areas, including the following steps:
[0008] Step 1: Collect multiple sets of historical data and perform preprocessing;
[0009] The historical data includes: environmental data at the current moment, transportation vehicle operation data, operation data of transportation vehicles, carbon emissions, carbon emission concentration, and carbon emission range after T time;
[0010] Step 2: Use the multiple sets of preprocessed historical data to train the LSTM network model to obtain a trained LSTM network model;
[0011] Step 3: Analyze the environmental data at the current moment, transportation vehicle operation data, operation data of transportation vehicles, carbon emissions, carbon emission concentration, and carbon emission range after T time to obtain the comprehensive correlation coefficient of carbon emissions, the comprehensive correlation coefficient of carbon emission concentration, and the comprehensive correlation coefficient of carbon emission range;
[0012] Step 4: Use the comprehensive correlation coefficient of carbon emissions, the comprehensive correlation coefficient of carbon emission concentration, and the comprehensive correlation coefficient of carbon emission range as additional features and input them into the trained LSTM network model for secondary training to optimize the model parameters until the preset maximum number of iterations is reached or the loss value of the loss function reaches the minimum, obtaining a secondarily trained LSTM network model;
[0013] Step 5: Collect real-time environmental data, transportation vehicle operation data, and operation data of transportation vehicles, input them into the secondarily trained LSTM network model, and predict the carbon emissions, carbon emission concentration, and carbon emission range of transportation vehicles after T time.
[0014] Further, the preprocessing includes data cleaning, feature encoding, data standardization, data normalization, and time alignment.
[0015] Further, Step 2 specifically includes:
[0016] Step 21: Construct the LSTM model architecture and initialize it;
[0017] Step 22: Divide the multiple sets of preprocessed historical data into a training set, a validation set, and a test set according to a preset ratio;
[0018] Step 23: Input the training set data into the initialized LSTM network model, use the backpropagation algorithm, and update the model parameters by minimizing the loss function to obtain a trained LSTM model.
[0019] Further, the loss function is the mean square error loss function.
[0020] Further, step 3 specifically includes:
[0021] Step 31: Group the preprocessed historical data according to the type of operating transportation vehicle, operating scenario, and altitude.
[0022] Step 32: Conduct a correlation analysis on each group of classified data and create interaction features within each group.
[0023] Step 33: Calculate the correlation coefficients between each group of grouped data and the interaction features within the group and the carbon emissions, carbon emission concentration, and carbon emission range after T time.
[0024] Step 34: Analyze the data and interaction features within each group to obtain the comprehensive correlation coefficients of carbon emissions, carbon emission concentration, and carbon emission range within the group.
[0025] Step 35: According to the proportion of the number of data samples in each group to the total number of samples, assign corresponding weights to the comprehensive correlation coefficients of each group, and perform weighted summation on the comprehensive correlation coefficients of carbon emissions, carbon emission concentration, and carbon emission range of all groups respectively to obtain the comprehensive correlation coefficients of carbon emissions, carbon emission concentration, and carbon emission range.
[0026] Further, step 32, conducting a correlation analysis on each group of classified data and creating interaction features within each group, specifically includes:
[0027] Step 321: Calculate the correlation coefficients of the environmental data, operating transportation vehicle data, and operating data of the operating transportation vehicle at the current moment within each group of classified data.
[0028] Step 322: Preset a correlation coefficient threshold, and create interaction features for the data with a correlation coefficient greater than or equal to the correlation coefficient threshold.
[0029] Further, step 34 specifically includes:
[0030] Step 341: Analyze the importance of the data and interaction features within each group to obtain the importance scores of each data and interaction feature within the group.
[0031] Step 342: Assign weights to the data and interaction features within each group according to the importance scores.
[0032] Step 343: Obtain the comprehensive carbon emission correlation coefficient, carbon emission concentration correlation coefficient, and carbon emission range correlation coefficient within the group based on the data within each group, the weights of the intra-group interaction characteristics, and the correlation coefficients.
[0033] Further, in Step 314, using environmental data, operation traffic vehicle data, operation traffic vehicle operation data, and intra-group interaction characteristics as input features, and the carbon emissions, carbon emission concentration, and carbon emission range after T time as output targets respectively, calculate the average impurity reduction amount of each feature, and determine the importance score of each feature according to the average impurity reduction amount.
[0034] The embodiments of the present invention have the following technical effects:
[0035] The present invention adopts the LSTM (Long Short-Term Memory Network) model, which is an excellent algorithm for processing time series data in the field of deep learning. By collecting environmental data, operation traffic vehicle data, and operation traffic vehicle operation data in high-altitude areas and using these data as input features, it can accurately capture the influence law of different factors on carbon emissions over time. This process not only considers static factors such as operation vehicle data, but also fully incorporates dynamic factors such as weather condition changes and real-time operation status, effectively improving the accuracy of carbon emission prediction in the complex and changeable high-altitude environment. Further, by calculating the correlation between each type of data (environmental data, vehicle data, operation data) and the output results (carbon emissions, concentration, range), the obtained comprehensive carbon emission correlation coefficient, carbon emission concentration correlation coefficient, and carbon emission range correlation coefficient are fed back to the model as additional features for secondary training. Greatly enhancing the prediction ability of the model, the model can further adjust its own weights and structure according to the importance of the indicators reflected by the correlation coefficients to better capture the carbon emission laws of different vehicles. At the same time, it also provides information on which factors have a significant impact on carbon emissions, helping to formulate more targeted and effective emission reduction measures. Description of the Drawings
[0036] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 It is a flowchart of a method for predicting carbon emissions of operation transportation vehicles in high-altitude areas provided by an embodiment of the present invention;
[0038] Figure 2 It is the structural diagram of a carbon emission prediction system for transportation vehicles operating in high-altitude areas provided by an embodiment of the present invention. Specific implementation manners
[0039] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be described clearly and completely below. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without any creative efforts shall fall within the scope protected by the present invention.
[0040] Embodiment 1
[0041] Figure 1 It is the flowchart of a carbon emission prediction method for transportation vehicles operating in high-altitude areas provided by an embodiment of the present invention. Refer to Figure 1 , which specifically includes:
[0042] Step 1: Collect multiple groups of historical data and perform preprocessing;
[0043] The historical data includes: environmental data at the current moment, transportation vehicle data for operation, operation data of the transportation vehicle for operation, carbon emissions, carbon emission concentration, and carbon emission range after T time;
[0044] In some embodiments, the environmental data at the current moment includes but is not limited to: altitude, temperature, air pressure, humidity, oxygen content; the transportation vehicle data for operation at the current moment includes but is not limited to: type, model, rated power, driving mileage, fuel consumption; the operation data of the transportation vehicle for operation at the current moment includes but is not limited to: EGR rate, boost pressure, engine speed, fuel flow rate.
[0045] The higher the altitude, the thinner the air and the relatively lower the oxygen content, which will cause incomplete fuel combustion. For example, for an airplane flying at high altitude or a car driving in a plateau area, due to insufficient oxygen, the combustion process deteriorates, resulting in an increase in carbon emissions. Moreover, the low atmospheric pressure in high-altitude areas may affect the intake air volume and combustion efficiency of the engine, and is also closely related to the carbon emission concentration and range. Extreme temperature conditions will affect the performance of the engine. At low temperatures, the atomization effect of the fuel becomes poor, the combustion is incomplete, and the carbon emissions increase; at high temperatures, the engine may experience overheating problems, affecting its working efficiency and thus affecting carbon emissions. At the same time, temperature also affects the convection and diffusion ability of the atmosphere, and has an impact on the carbon emission range.
[0046] Different types of transportation vehicles (such as cars, trains, airplanes, etc.) and specific models have different power systems, design concepts, and usage purposes. For example, compared with small cars, large trucks usually have larger engines and higher power requirements, and their carbon emissions will also be higher; the carbon emission performance of new energy-saving models of transportation vehicles may be better than that of old models. Fuel consumption directly reflects the amount of energy used by transportation vehicles. The more fuel consumed, the greater the theoretical carbon emissions, and there is a positive correlation between it and carbon emissions. Moreover, a large fuel consumption may lead to an increase in the carbon emission concentration and an expansion of the emission range in a local area.
[0047] An appropriate EGR rate (exhaust gas recirculation rate) can reduce the combustion temperature and reduce the generation of nitrogen oxides, but an excessively high EGR rate may lead to unstable combustion, increase the emissions of hydrocarbons and carbon monoxide, thereby affecting the carbon emissions and carbon emission concentration. Too high or too low engine speed may affect the combustion efficiency. At high speeds, the engine may not have enough time to fully burn the fuel, resulting in an increase in carbon emissions; at low speeds, if the load is large, incomplete combustion will also occur.
[0048] The above data cover multiple key factors affecting the carbon emissions of transportation vehicles. By analyzing these data, the model can learn the carbon emission laws under different environmental conditions, different transportation vehicles, and different operating states. In this way, the model can predict future carbon emissions, concentrations, and ranges based on real-time data, thus providing data support for traffic management, energy conservation and emission reduction, and environmental protection.
[0049] In some embodiments, the preprocessing includes:
[0050] Data cleaning: To handle missing values, methods such as deleting records containing missing values, filling in missing values (such as using the mean, median, or predicted values by a specific algorithm) can be adopted; to remove outliers, identify and process those data points that significantly deviate from the normal range, which can be completed through statistical methods or machine learning algorithms.
[0051] Feature encoding: For categorical variables (such as vehicle types), they may need to be converted into numerical forms for easy processing by the model. Commonly used encoding methods include one-hot encoding, label encoding, etc.
[0052] Data standardization: Perform standardization processing on the data to make the data of different features have the same scale and avoid affecting the model training effect due to large differences in feature scales. Commonly used methods include Z-Score standardization.
[0053] Data normalization: Use the maximum-minimum normalization method to normalize the index values to obtain the normalized values;
[0054] The normalization formula is as follows:
[0055] ;
[0056] where x new is the normalized value,
[0057] X is the original data value of the index value,
[0058] x min is the minimum value in this index,
[0059] x max is the maximum value in this index.
[0060] Time alignment: Ensure that all data is consistent in the time dimension. For data with different sampling frequencies, interpolation or downsampling methods are used for processing so that the data can correspond to the same time point. For example, for the operation data collected at a high frequency and the environmental data collected at a low frequency, they can be synchronized in time through methods such as linear interpolation.
[0061] Step 2: Use the preprocessed multiple groups of historical data to train the LSTM network model to obtain a trained LSTM network model;
[0062] In some embodiments, step 2 specifically includes:
[0063] Step 21: Construct the LSTM model architecture and initialize it;
[0064] Exemplarily, for determining the input dimension, after the data collected in step 1 is preprocessed, the dimension of the input layer of the LSTM model should be determined according to the number of its features. The environmental data has 5 features including altitude, temperature, air pressure, humidity, and oxygen content; the operation data of transportation vehicles includes 5 features such as type, model, rated power, driving mileage, and fuel consumption; the operation data of transportation vehicles covers 4 features including EGR rate, boost pressure, engine speed, and fuel flow. Then the feature dimension of the input layer is 5 + 5 + 4 = 14. When constructing the LSTM model, the dimension of the input layer needs to be set to 14. For determining the output dimension: The target variables are carbon emissions, carbon emission concentration, and carbon emission range after T time, so the number of neurons in the output layer should be set to 3.
[0065] Initialize the weight matrix and bias vector in the model. Commonly used initialization methods include random initialization (such as randomly generating initial values using uniform distribution or normal distribution), Xavier initialization, Kaiming initialization, etc.
[0066] Step 22: Divide the preprocessed multiple groups of historical data into a training set, a validation set, and a test set according to a preset ratio;
[0067] Divide the preprocessed multiple groups of historical data into input features and target variables. The input features are the environmental data (altitude, temperature, air pressure, humidity, oxygen content) at the current moment, the operation data of transportation vehicles (type, model, rated power, mileage, fuel consumption), and the operation data of transportation vehicles (EGR rate, boost pressure, engine speed, fuel flow). The target variables are the carbon emissions, carbon emission concentration, and carbon emission range after T time.
[0068] A common division ratio is 70% for the training set, 15% for the validation set, and 15% for the test set, but it can be adjusted according to the data volume and actual needs. For example, if the data volume is small, the proportion of the training set can be appropriately increased to improve the training effect of the model; if more accurate evaluation of the model performance is required, the proportions of the validation set and the test set can be appropriately increased.
[0069] Step 23: Input the training set data into the initialized LSTM network model, and use the backpropagation algorithm to update the model parameters by minimizing the loss function to obtain a trained LSTM model.
[0070] Specifically, input the data in the training set into the initialized LSTM model. Since LSTM is a recurrent neural network, it can effectively capture the long-term dependencies in time series data, so it is very suitable for dealing with this kind of carbon emission prediction problem that changes over time.
[0071] Use the backpropagation algorithm. Backpropagation starts from the output layer, calculates the gradients layer by layer, and reversely transmits the error information to each parameter of the network. Update the model parameters by minimizing the loss function. By continuously adjusting the parameters, the predicted values of the model are getting closer and closer to the actual values. Optionally, the loss function selected is the mean squared error loss function (MSE), which calculates the average of the squared differences between the predicted values and the actual values of the model.
[0072] The formula for the mean squared error loss function (MSE) is: ; where y i is the actual value of the i-th sample, is the predicted value of the i-th sample, and n is the number of samples.
[0073] Step 3: Analyze the environmental data, operation data of transportation vehicles, operation data of transportation vehicles, carbon emissions, carbon emission concentration, and carbon emission range after T time at the current moment to obtain the comprehensive correlation coefficient of carbon emissions, the comprehensive correlation coefficient of carbon emission concentration, and the comprehensive correlation coefficient of carbon emission range;
[0074] In some embodiments, step 3 specifically includes:
[0075] Step 31: Group the preprocessed historical data according to the type of operating transportation vehicle, the operating scenario, and the altitude;
[0076] Specifically, for the type of operating transportation vehicle, its specific category such as car, truck, etc. is determined through the basic information of the vehicle, such as the vehicle driving license, the vehicle registration file, etc. The operating scenario is judged by combining the vehicle operation trajectory data, the geographic information data, and the road classification information provided by the traffic management department. For example, using GPS positioning data combined with map data, when the vehicle is within the urban area and the road speed limit, traffic flow, etc. conform to the characteristics of urban roads, it is determined as an urban road operation scenario; if the vehicle is in the mountainous area where the road has a large slope and many curves, etc., it is determined as a mountain highway operation scenario. The geographic longitude and latitude information is used to divide into multiple high-altitude regions.
[0077] Through the groupby function, grouping is performed according to the three dimensions of vehicle type, operating scenario, and geographic area to obtain multiple grouped data subsets. Each data subset represents the data of a specific type of transportation vehicle within a specific altitude in a specific operating area, ensuring that the data within each subset has similar conditions.
[0078] Exemplarily, assuming that the operating scenarios are divided into urban roads and mountain roads, and the geographic area is divided into four altitude intervals below 2000m, [2000, 3000m), [3000, 4000m), and above 4000m according to the altitude, the grouping results are shown in Table 1 below:
[0079] Table 1 Statistical table of grouping results
[0080] Due to differences in power systems, load capacities, and technical levels among different types of vehicles, their carbon emissions are inherently different. In terms of operating scenarios, factors such as driving speed, start-stop frequency, slope road conditions, operating time, and environmental temperature affect the engine working state and fuel consumption, thereby changing carbon emissions. At different altitudes, the change in air oxygen content due to altitude affects combustion, the climate conditions increase additional energy consumption, and the geographical environment and traffic planning influence the driving efficiency, all of which result in different carbon emissions. After grouping the data, in-depth analysis can be carried out for each specific group to explore the relationship between environmental data, vehicle data, operation data, and carbon emission indicators under specific conditions.
[0081] Step 32: Perform a correlation analysis on each group of classified data and create interaction features within each group;
[0082] Step 321: Calculate the correlation coefficient of the environmental data, operation traffic vehicle data, and operation traffic vehicle running data at the current moment within each group of classified data.
[0083] The sub-parameters of environmental data include altitude, temperature, air pressure, humidity, and oxygen content; the sub-parameters of operation traffic vehicle data include type, model, rated power, driving mileage, and fuel consumption; the sub-parameters of operation traffic vehicle running data include EGR rate, boost pressure, engine speed, and fuel flow.
[0084] The correlation calculation formula is:
[0085] ;
[0086] where R xy represents the correlation coefficient between sub-parameter x and sub-parameter y; r xy represents the Pearson correlation coefficient between sub-parameter x and sub-parameter y; MI(x; y) represents the mutual information content between sub-parameter x and sub-parameter y; H(x) represents the entropy of sub-parameter x; H(y) represents the entropy of sub-parameter y.
[0087] Exemplarily, taking Group 5 (car, urban road, [2000, 3000m)) as an example, assuming there are 100 data records in this group, including the environmental data (altitude, temperature, air pressure, humidity, oxygen content), operation traffic vehicle data (type, model, rated power, driving mileage, fuel consumption), and operation traffic vehicle running data (EGR rate, boost pressure, engine speed, fuel flow) at the current moment.
[0088] Select the temperature (sub-parameter x) in the environmental data and the driving speed (sub-parameter y) in the operation traffic vehicle running data to calculate the correlation coefficient. First, calculate the Pearson correlation coefficient, which measures the linear correlation degree between two variables. Then calculate the mutual information content, which represents the information content that one variable contains about the other variable. Next, calculate the entropy of sub-parameter x and the entropy of sub-parameter y. Entropy represents the uncertainty of a variable. Finally, obtain the correlation coefficient between temperature and driving speed according to the correlation calculation formula.
[0089] Similarly, perform the same calculation for other sub-parameter pairs within this group (such as temperature and vehicle load, driving speed and vehicle load, etc.) to obtain the correlation coefficients between all sub-parameter pairs.
[0090] Step 322: Preset a correlation coefficient threshold, and create interaction features for the data with a correlation coefficient greater than or equal to the correlation coefficient threshold.
[0091] The formula for the interaction feature is: z = log(x + 1) × e (y / max(y)) ; z represents the interaction feature.
[0092] For example, the preset correlation coefficient threshold is 0.6. In group 1, after calculation in step 321, it is found that the correlation coefficient between temperature and driving speed is 0.7, which is greater than the threshold of 0.6, so an interaction feature is created for these two variables.
[0093] Let temperature be x and driving speed be y. Calculate the interaction feature z using the interaction feature formula. For other variable pairs within the group whose correlation coefficients are greater than or equal to the threshold, create corresponding interaction features using this formula.
[0094] Step 33: Calculate the correlation coefficient between each group of data and the interaction characteristics within the group and the carbon emissions, carbon emission concentration, and carbon emission range after time T;
[0095] In this embodiment, for each group of data, various types of characteristic data (environmental data, operational transportation vehicle data, operational transportation vehicle operation data, and interaction characteristics) and three target variables (carbon emissions after time T, carbon emission concentration, and carbon emission range) are clearly defined.
[0096] Taking group 5 (car, urban road, [2000, 3000m)) as an example, assume that there are n data records in the group. The feature data includes altitude H, temperature M, air pressure P, humidity S, and oxygen content O in the environmental data; type A, model W, rated power Q, mileage B, and fuel consumption C in the operational vehicle data; EGR rate E, boost pressure F, engine speed D, and fuel flow K in the operational vehicle data; and the interactive feature z created in step 32 (assuming it is an interactive feature of temperature and driving speed). The target variables are carbon emissions C after time T. e , carbon emission concentration C c and carbon emissions scope C r .
[0097] First, calculate the temperature T and carbon emissions C e Pearson correlation coefficient r T,Ce ; Then calculate the temperature T and carbon emissions C e Mutual information, and then calculate the temperature T and carbon emissions C respectively e The entropy of the final calculation is based on the correlation formula to calculate the temperature T and carbon emissions C. e The correlation coefficient R T,Ce Repeat this step for all feature variables (including original features and interaction features) in each group and the three target variables (carbon emissions, carbon emission concentration, and carbon emission range) to calculate all correlation coefficients.
[0098] Step 34: Analyze the data and intra-group interaction features within each group to obtain the comprehensive correlation coefficient of carbon emissions, the comprehensive correlation coefficient of carbon emission concentration, and the comprehensive correlation coefficient of carbon emission range within the group.
[0099] Step 341: Analyze the importance of the data and intra-group interaction features within each group to obtain the importance scores of each data and interaction feature within the group.
[0100] In this example, using environmental data, operating transportation vehicle data, operating transportation vehicle operation data, and intra-group interaction features as input features, and the carbon emissions, carbon emission concentration, and carbon emission range after time T as output targets respectively, calculate the average impurity reduction amount of each feature. The higher the score, the greater the impact of the feature on the output target. The average impurity reduction amount measures the average degree of impurity reduction of a certain feature during node splitting in all decision trees of the random forest. The larger the reduction amount, the more the feature can reduce the impurity when splitting the samples, which means the greater the impact of the feature on the output target and the higher the importance. Then, normalize the average impurity reduction amounts of each feature so that the sum of the importance scores of all features is 1. This can more intuitively compare the relative importance of each feature in the whole.
[0101] Step 342: Assign weights to the data and intra-group interaction features within each group according to the importance scores.
[0102] ;
[0103] Among them, S i represents the importance score of the i-th data; m represents the total number of data; W i represents the weight of the i-th data.
[0104] Taking group 1 as an example, assume that after the analysis in step 341, the importance score of temperature for carbon emissions is 0.3, the importance score of driving speed for carbon emissions is 0.2, and the importance score of interaction feature z for carbon emissions is 0.5. Normalize these importance scores so that their sum is 1. After normalization, the weight of temperature is 0.3 / (0.3 + 0.2 + 0.5) = 0.3, the weight of driving speed is 0.2 / (0.3 + 0.2 + 0.5) = 0.2, and the weight of interaction feature z is 0.5 / (0.3 + 0.2 + 0.5) = 0.5.
[0105] For carbon emission concentration and carbon emission range, also use the same method to assign weights to each feature according to the corresponding importance scores.
[0106] Step 343: Obtain the comprehensive correlation coefficient of carbon emissions, the comprehensive correlation coefficient of carbon emission concentration, and the comprehensive correlation coefficient of carbon emission range within each group based on the data within each group and the weights and correlation coefficients of the intra-group interaction characteristics.
[0107] The calculation formula for the comprehensive correlation coefficient of carbon emissions within each group is:
[0108] ;
[0109] W i represents the weight of the i-th data; ρ i,Ce represents the correlation coefficient between the i-th data and carbon emissions. Similarly, calculate the comprehensive correlation coefficient R Cc of carbon emission concentration and the comprehensive correlation coefficient R Cr of carbon emission range within this group.
[0110] Step 35: Assign corresponding weights to the comprehensive correlation coefficients of each group according to the proportion of the number of data samples in each group to the total number of samples, and perform weighted summation on the comprehensive correlation coefficients of carbon emissions, carbon emission concentration, and carbon emission range of all groups respectively to obtain the comprehensive correlation coefficient of carbon emissions, the comprehensive correlation coefficient of carbon emission concentration, and the comprehensive correlation coefficient of carbon emission range.
[0111] The number of data samples in different groups may be different. Groups with a larger number of samples should have a greater influence in the overall analysis. Therefore, assign weights to the comprehensive correlation coefficients of each group according to the proportion of the number of data samples in each group to the total number of samples, and then perform weighted summation on the comprehensive correlation coefficients of all groups to obtain the final comprehensive correlation coefficient of carbon emissions, the comprehensive correlation coefficient of carbon emission concentration, and the comprehensive correlation coefficient of carbon emission range.
[0112] Exemplarily, there are a total of k groups, the number of data samples in the j-th group is n j , and the total number of samples is ; the comprehensive correlation coefficient of carbon emissions in the j-th group is R Ce,j , the comprehensive correlation coefficient of carbon emission concentration is R Cc,j , the comprehensive correlation coefficient of carbon emission range is R Cr,j , then the comprehensive correlation coefficient of carbon emissions is: ; the comprehensive correlation coefficient of carbon emission concentration is: ; the comprehensive correlation coefficient of carbon emission concentration is: .
[0113] The comprehensive correlation coefficient of carbon emissions is determined through correlation analysis between the data of all groups (including environmental data, operation data of transportation vehicles, and their operation data) and the carbon emissions after time T. First, interaction features are created within each group, and the correlation coefficients between these features and carbon emissions are calculated. Then, based on the importance scores of each feature (calculated by methods such as average impurity reduction), weights are assigned to the data within each group and the interaction features within the group. Finally, the comprehensive correlation coefficient of carbon emissions for all groups is obtained by weighted summation according to the proportion of the data sample size of each group in the total sample size; the comprehensive correlation coefficient of carbon emission concentration is determined through correlation analysis between the data of all groups and the carbon emission concentration after time T, and the whole process is the same as that of the comprehensive correlation coefficient of carbon emissions, but the focus is on carbon emission concentration rather than carbon emissions; the comprehensive correlation coefficient of carbon emission range is determined through correlation analysis between the data of all groups and the carbon emission range after time T. Similarly, this process is similar to the previous two, but it focuses on the carbon emission range. These three comprehensive correlation coefficients all consider the influence of various factors, including but not limited to the characteristics of different vehicle models at different altitudes, reflecting how these factors of environmental data, operation data of transportation vehicles, and their operation data jointly act on different aspects of carbon emissions (quantity, concentration, range), providing a broader and more comprehensive perspective to understand the overall effect of various factors affecting carbon emissions, and in the subsequent steps, this information is fed back as additional features to the LSTM model for secondary training, which can more accurately capture the complex relationships between various factors affecting carbon emissions.
[0114] Step 4: Take the comprehensive correlation coefficient of carbon emissions, the comprehensive correlation coefficient of carbon emission concentration, and the comprehensive correlation coefficient of carbon emission range as additional features and input them into the trained LSTM network model for secondary training to optimize the model parameters until the preset maximum number of iterations is reached or the loss value of the loss function reaches the minimum, and obtain the LSTM network model after secondary training.
[0115] By introducing these additional features, the model can obtain more information for learning, thereby improving the accuracy of predicting future carbon emission situations. Considering the influence of different types of transportation vehicles, different operation scenarios, and different altitudes, the model becomes more general and flexible.
[0116] Step 5: Collect real-time environmental data, operation data of transportation vehicles, and operation data of transportation vehicles, and input them into the LSTM network model after secondary training to predict the carbon emissions, carbon emission concentration, and carbon emission range of transportation vehicles after time T.
[0117] By grouping the preprocessed historical data according to the type of operating transportation vehicle, operating scenario, and altitude, the data within each grouped data subset has similar conditions. This enables subsequent analysis to focus on the data characteristics under specific environments and conditions, avoiding interference from different types, scenarios, and altitude factors, greatly improving the homogeneity of the data, and helping to more accurately explore the internal laws and relationships within the data. Conducting correlation analysis within each group and creating interaction features can deeply study the internal connections among the environmental data, operating transportation vehicle data, and operating data of the transportation vehicle at the current moment. The creation of interaction features can capture the complex interactions between variables, providing rich data for further understanding the data and establishing more accurate models. By analyzing the correlations between the data and interaction features within different groups and the carbon emissions, concentration, and range after time T, the comprehensive correlation coefficients of carbon emissions, the comprehensive correlation coefficient of carbon emission concentration, and the comprehensive correlation coefficient of carbon emission range are obtained, and this information is fed back as additional features to the LSTM model for secondary training, which can more precisely capture the complex relationships among various factors affecting carbon emissions. This method helps to improve the accuracy of carbon emission prediction.
[0118] Embodiment 2
[0119] Figure 2 is a structural diagram of a carbon emission prediction system for operating transportation vehicles in high-altitude areas provided by an embodiment of the present invention. Refer to Figure 2 The present invention also provides a carbon emission prediction system for operating transportation vehicles in high-altitude areas, which is used for the carbon emission prediction method for operating transportation vehicles in high-altitude areas described above, and includes the following modules:
[0120] Data collection module: used to collect multiple groups of historical data and perform preprocessing; the historical data includes: environmental data at the current moment, operating transportation vehicle data, operating data of the transportation vehicle, and carbon emissions, carbon emission concentration, and carbon emission range after time T;
[0121] Training module: used to train the LSTM network model with multiple groups of preprocessed historical data to obtain a trained LSTM network model;
[0122] Correlation calculation module: used to analyze the environmental data at the current moment, operating transportation vehicle data, operating data of the transportation vehicle, carbon emissions, carbon emission concentration, and carbon emission range after time T to obtain the comprehensive correlation coefficient of carbon emissions, the comprehensive correlation coefficient of carbon emission concentration, and the comprehensive correlation coefficient of carbon emission range;
[0123] Secondary training module: It is used to input the comprehensive carbon emission correlation coefficient, the comprehensive carbon emission concentration correlation coefficient, and the comprehensive carbon emission range correlation coefficient as additional features into the trained LSTM network model for secondary training to optimize the model parameters until the preset maximum number of iterations is reached or the loss value of the loss function reaches the minimum, and obtain the LSTM network model after secondary training is completed;
[0124] Carbon emission prediction module: It is used to collect real-time environmental data, operation data of transportation vehicles, and operation data of transportation vehicles, input the LSTM network model after secondary training is completed, and predict the carbon emissions, carbon emission concentration, and carbon emission range of transportation vehicles after T time.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting carbon emissions of transportation vehicles operating in high-altitude areas, characterized in that, Including: Step 1: Collect multiple sets of historical data and perform preprocessing; The historical data includes: environmental data at the current moment, operating traffic vehicle data, operating traffic vehicle operation data, carbon emissions after T time, carbon emission concentration after T time, and carbon emission range after T time; Step 2: Use the preprocessed multiple sets of historical data to train the LSTM network model to obtain a trained LSTM network model; Step 3: Analyze the environmental data at the current moment, operating traffic vehicle data, operating traffic vehicle operation data, carbon emissions after T time, carbon emission concentration after T time, and carbon emission range after T time to obtain the comprehensive correlation coefficient of carbon emissions, the comprehensive correlation coefficient of carbon emission concentration, and the comprehensive correlation coefficient of carbon emission range; Step 4: Use the comprehensive correlation coefficient of carbon emissions, the comprehensive correlation coefficient of carbon emission concentration, and the comprehensive correlation coefficient of carbon emission range as additional features and input them into the trained LSTM network model for secondary training to optimize the model parameters until the preset maximum number of iterations is reached or the loss value of the loss function reaches the minimum, obtaining the LSTM network model after secondary training; Step 5: Collect real-time environmental data, operating traffic vehicle data, and operating traffic vehicle operation data, input them into the LSTM network model after secondary training, and predict the carbon emissions, carbon emission concentration, and carbon emission range of the operating traffic vehicle after T time.
2. The carbon emission prediction method for operating transportation vehicles in high-altitude areas according to claim 1, characterized in that, The preprocessing includes data cleaning, feature encoding, data standardization, data normalization, and time alignment.
3. A method for predicting carbon emissions of transportation vehicles operating in high-altitude areas according to claim 1, characterized in that, The specific content of Step 2 includes: Step 21: Construct the LSTM model architecture and initialize it; Step 22: Divide the preprocessed multiple sets of historical data into a training set, a validation set, and a test set according to a preset ratio; Step 23: Input the training set data into the initialized LSTM network model, use the backpropagation algorithm, and update the model parameters by minimizing the loss function to obtain a trained LSTM model.
4. A method for predicting carbon emissions of transportation vehicles operating in high-altitude areas according to claim 3, characterized in that, The loss function is the mean squared error loss function.
5. A method for predicting carbon emissions of transportation vehicles operating in high-altitude areas according to claim 1, characterized in that, The specific content of Step 3 includes: Step 31: Group the preprocessed historical data according to the type of operating traffic vehicle, operating scenario, and altitude; Step 32: Perform correlation analysis on each group of classified data and create interaction features within each group; Step 33: Calculate the correlation coefficients between each group of grouped data and the interaction features within the group and the carbon emissions, carbon emission concentration, and carbon emission range after T time; Step 34: Analyze the data and the interaction features within each group, and combine the correlation coefficients obtained in Step 33 to obtain the comprehensive correlation coefficient of carbon emissions, the comprehensive correlation coefficient of carbon emission concentration, and the comprehensive correlation coefficient of carbon emission range within the group; Step 35: According to the proportion of the number of data samples in each group to the total number of samples, assign corresponding weights to the comprehensive correlation coefficients of each group, and perform weighted summation on the comprehensive correlation coefficients of carbon emissions, the comprehensive correlation coefficients of carbon emission concentrations, and the comprehensive correlation coefficients of carbon emission ranges for all groups respectively, to obtain the total comprehensive correlation coefficients of carbon emissions, the total comprehensive correlation coefficients of carbon emission concentrations, and the total comprehensive correlation coefficients of carbon emission ranges.
6. A method for predicting carbon emissions of transportation vehicles operating in high-altitude areas according to claim 5, characterized in that, Step 32: Conduct a correlation analysis on each group of classified data, and create interaction features within each group, specifically including: Step 321: Calculate the correlation coefficients of the environmental data, the operation data of the transportation vehicles, and the operation data of the transportation vehicles at the current moment within each group of the classified data. Step 322: Preset a correlation coefficient threshold, and create interaction features for the data with correlation coefficients greater than or equal to the correlation coefficient threshold.
7. A method for predicting carbon emissions of transportation vehicles operating in high-altitude areas according to claim 5, characterized in that, The said Step 34 specifically includes: Step 341: Analyze the importance of the data and the interaction features within each group, and obtain the importance scores of each data and interaction feature within the group. Step 342: According to the importance scores, assign weights to the data and the interaction features within each group. Step 343: According to the weights of the data and the interaction features within each group and the correlation coefficients obtained in Step 33, obtain the comprehensive correlation coefficients of carbon emissions, the comprehensive correlation coefficients of carbon emission concentrations, and the comprehensive correlation coefficients of carbon emission ranges within the group.
8. A method for predicting carbon emissions of transportation vehicles operating in high-altitude areas according to claim 7, characterized in that, In Step 341, using the environmental data, the operation data of the transportation vehicles, the operation data of the transportation vehicles, and the interaction features within the group as input features, and the carbon emissions, carbon emission concentrations, and carbon emission ranges after T time as output targets respectively, calculate the average impurity reduction amount of each feature, and determine the importance score of each feature according to the average impurity reduction amount.
Citation Information
Patent Citations
Public institution-oriented LSTM recurrent neural network carbon emission prediction method
CN116882580A
Multi-dimensional carbon emission data acquisition and accounting system based on enterprise data
CN117592666A