A method and system for monitoring carbon emissions in the Yellow River Basin using big data
By constructing a multi-source data acquisition network and a carbon emission accounting model, the timeliness and accuracy of carbon emission monitoring in the Yellow River Basin have been addressed, enabling real-time dynamic monitoring and early warning of carbon emissions, and improving monitoring accuracy and data authenticity.
Patent Information
- Application Number
- CN202510239886.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-03-03
AI Technical Summary
Traditional carbon emission monitoring methods cannot achieve timely, comprehensive and accurate monitoring of the Yellow River Basin, and have problems such as long data acquisition cycles, limited sampling scope and inconsistent data formats.
A multi-source data acquisition network is constructed to integrate data from satellite remote sensing, ground monitoring stations, IoT sensors, and government and enterprise information systems. Through data preprocessing and standardization, multi-dimensional fusion feature vectors are formed. A carbon emission accounting model is constructed using neural networks and support vector machines, and the model parameters are optimized through random forest regression algorithm to achieve real-time dynamic monitoring and early warning.
It enables real-time dynamic monitoring of carbon emissions in the Yellow River Basin, improving monitoring accuracy and data authenticity. It can capture instantaneous changes and automatically trigger early warnings, providing high-quality carbon emission information display and analysis.
Smart Images

Figure CN119850229B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of carbon emission monitoring methods, and particularly relates to a dynamic carbon emission monitoring system for the Yellow River Basin using big data. BACKGROUND
[0002] The Yellow River Basin covers many industries and densely populated areas, has large energy consumption, and has complex and diverse carbon emission sources, including industrial production, energy exploitation, transportation, and residential life in various fields. Traditional carbon emission monitoring methods often rely on periodic field research, sampling detection, and enterprise self-reporting data, and these methods have many limitations: first, the data acquisition period is long, and it is difficult to reflect the dynamic changes of carbon emissions in real time; second, the sampling range is limited, and it is difficult to accurately cover all types of emission sources in the entire basin, which may cause data bias; third, the data formats and standards are not unified among different departments and regions, and the integration is difficult, which hinders the accurate grasp of the overall carbon emission panorama of the basin. SUMMARY
[0003] The technical problem to be solved by the present application is to provide a method and system for timely, comprehensive, and accurate monitoring of carbon emissions in the Yellow River Basin.
[0004] To solve the above technical problems, the technical solution adopted by the present application is: a dynamic carbon emission monitoring method for the Yellow River Basin using big data, comprising the following steps:
[0005] S1, data acquisition: a multi-source data acquisition network is constructed to integrate data related to carbon emissions from satellite remote sensing, ground monitoring stations, Internet of Things sensors, and various government and enterprise information systems in the Yellow River Basin;
[0006] S2, data preprocessing: first, the collected multi-source data is cleaned, then data standardization is performed, and the units and dimensions of data from different sources are unified according to a unified data dictionary; finally, data fusion is performed to form a multi-dimensional fusion feature vector representing the relationship between the multi-source data;
[0007] S3, carbon emission accounting model construction: a carbon emission accounting model is constructed by fusing a neural network and a support vector machine, the multi-dimensional fusion feature vector after processing is used as input, and carbon emission information is used as output to train the carbon emission accounting model, and a random forest regression algorithm is used to dynamically adjust the accounting parameters of the carbon emission accounting model to optimize the model parameters;
[0008] S4, dynamic monitoring and early warning: the constructed carbon emission accounting model is used to monitor the carbon emissions in the Yellow River Basin in real time, and when the carbon emission growth rate of a certain region or industry is too fast or the intensity exceeds the standard, the system automatically triggers an early warning;
[0009] S5, data updating and model optimization: continuously update various data in the data acquisition network, and use new data to optimize the multi-dimensional fusion feature vector and the carbon emission accounting model, and then use the optimized carbon emission accounting model to monitor the carbon emission of the Yellow River Basin.
[0010] Correspondingly, the application also discloses a dynamic carbon emission monitoring system for the Yellow River Basin based on big data, which comprises:
[0011] The data acquisition module is used for constructing a multi-source data acquisition network, and integrating data related to carbon emission from satellite remote sensing, ground monitoring stations, Internet of Things sensors and various government and enterprise information systems in the Yellow River Basin.
[0012] The data preprocessing module is used for firstly cleaning the collected multi-source data, then performing data standardization, unifying the units and dimensions of data from different sources according to a unified data dictionary, and finally performing data fusion to form a multi-dimensional fusion feature vector representing the relationship between the multi-source data.
[0013] The carbon emission accounting model construction module is used for constructing a carbon emission accounting model by using a fusion neural network and a support vector machine, taking the processed multi-dimensional fusion feature vector as input and carbon emission information as output, training the carbon emission accounting model, and using a random forest regression algorithm to dynamically adjust the accounting parameters of the carbon emission accounting model and optimize the model parameters.
[0014] The dynamic monitoring and early warning module is used for using the constructed carbon emission accounting model to perform real-time dynamic monitoring on the carbon emission in the Yellow River Basin, and automatically triggering an early warning when the monitoring shows that the carbon emission growth rate of a certain region or industry is too fast or the intensity exceeds the standard.
[0015] The data updating and model optimization module is used for continuously updating various data in the data acquisition network, and using new data to optimize the multi-dimensional fusion feature vector and the carbon emission accounting model, and then using the optimized carbon emission accounting model to monitor the carbon emission in the Yellow River Basin.
[0016] The above technical scheme has the following beneficial effects: 1) multiple different types of data sources such as satellite remote sensing, ground monitoring stations, Internet of Things sensors, government and enterprise information systems are integrated, which greatly expands the data breadth and depth of carbon emission monitoring, and different data sources have their own suitable fast update frequency, which can track the dynamic changes of carbon emission in the Yellow River Basin, capture instantaneous fluctuations such as sudden equipment failure of a factory leading to sharp increase in energy consumption and carbon emission soaring during urban traffic peak hours, and make the carbon emission monitoring data more real.
[0017] 2) Using geographic information system GIS, data fusion is carried out on the basis of the map of the Yellow River basin, so that the multi-source data can be accurately positioned in space, the carbon emission related information at each geographical position in the basin is integrated, and the evolution trend of carbon emission in the basin space with time is intuitively presented through three-dimensional space-time presentation, so as to provide high-quality input data for subsequent carbon emission accounting model.
[0018] 3) The accounting model is set according to the characteristics of different industries and different fields in the Yellow River basin, which can be fitted to the actual situation of the basin and greatly improve the accuracy of carbon emission accounting. The machine learning algorithm is introduced to mine the implicit influencing factors, dynamically adjust the accounting parameters of the model, and adaptively learn the change of carbon emission law, so as to realize real-time optimization of precision and accurately reflect the actual carbon emission of the Yellow River basin.
[0019] 4) Using the constructed carbon emission accounting model, the overall carbon emission intensity of the basin and the distribution of each region and industry can be calculated at a time, and the visual chart is displayed on the monitoring screen. The high-frequency calculation update and intuitive chart presentation can convert the dynamic change of carbon emission in the Yellow River basin into visual information in real time, and set multi-level carbon emission warning thresholds. Once the regional or industrial carbon emission growth rate is too fast or the intensity exceeds the standard, the system will automatically trigger the warning. BRIEF DESCRIPTION OF DRAWINGS
[0020] The application will be further described in detail below in combination with the drawings and specific embodiments.
[0021] Figure 1 is the main flowchart of the method described in the embodiment of the application;
[0022] Figure 2 is the flowchart of forming a multi-dimensional fusion feature vector in the method described in the embodiment of the application;
[0023] Figure 3 is the flowchart of attribute fusion and feature extraction in the method described in the embodiment of the application;
[0024] Figure 4 is the flowchart of carbon emission accounting model construction in the method described in the embodiment of the application;
[0025] Figure 5 is the flowchart of dynamic monitoring and warning in the method described in the embodiment of the application;
[0026] Figure 6 is the principle block diagram of the system described in the embodiment of the application. DETAILED DESCRIPTION
[0027] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0028] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced according to the claims without some or all of these specific details. In other instances, well-known methods, procedures and components have not been described in detail so as not to unnecessarily obscure aspects of the present application.
[0029] As shown in the Figure 1 embodiments of the present application disclose a kind of Yellow River Basin carbon emission dynamic monitoring method using big data, comprising the following steps:
[0030] S1, data acquisition: build multi-source data acquisition network, integrate the data related to carbon emission from satellite remote sensing, ground monitoring station, Internet of Things sensor and various government, enterprise information system in Yellow River Basin;Satellite remote sensing data is used to obtain the large-area land use change and vegetation cover information of Yellow River Basin which indirectly reflects carbon emission factors;Ground monitoring station is set in the city, industrial park and energy production area of Yellow River Basin, real-time monitoring of atmospheric pollutant concentration and meteorological parameter information;Internet of Things sensor is deployed in carbon emission microcosmic source, including factory workshop, traffic artery and residential area, for the running power of carbon emission microcosmic source energy consumption equipment, vehicle mileage and oil consumption, household electricity and gas data are collected, and real-time transmission to data center;At the same time, the information system of environmental protection, energy and industry and information system of large enterprises in the basin are connected, and the production scale, energy consumption structure and product yield data of enterprises are obtained, and daily synchronization is updated.
[0031] S2, data preprocessing: first, the collected multi-source data is cleaned, then data standardization is performed, and the units and dimensions of data from different sources are unified according to a unified data dictionary; finally, data fusion is performed to form a multi-dimensional fusion feature vector representing the relationship between multi-source data; further, after receiving the collected multi-source data, first, the data is cleaned, and for abnormal values caused by sensor failure and transmission interruption, an anomaly detection algorithm based on time series analysis is used, combined with the mean and standard deviation of adjacent period data to set a threshold to identify and eliminate abnormal data points; for satellite remote sensing collected data, clear image data is restored through image restoration and filtering algorithm; second, data standardization is performed, and the units and dimensions of data from different sources are unified according to a unified data dictionary; finally, data fusion is performed, based on geographic information system, taking the Yellow River Basin map as the base, satellite remote sensing data, ground monitoring data, Internet of Things data and industry data after data cleaning and standardization are matched and fused according to geographic location coordinates to form a multi-dimensional fusion feature vector.
[0032] S3, carbon emission accounting model construction: by fusing neural network and support vector machine, a carbon emission accounting model is constructed, the multi-dimensional fusion feature vector after processing is taken as input, carbon emission information is taken as output, the carbon emission accounting model is trained, and a random forest regression algorithm is used to dynamically adjust the accounting parameters of the carbon emission accounting model to optimize the model parameters;
[0033] S4, dynamic monitoring and early warning: using the constructed carbon emission accounting model, the carbon emission of the Yellow River Basin is monitored in real time, and when the carbon emission growth rate of a certain region or industry is too fast or the intensity exceeds the standard, the system automatically triggers an early warning;
[0034] Further, using the constructed carbon emission accounting model, the carbon emission of the Yellow River Basin is monitored in real time, the collected multi-source data is input into the carbon emission accounting model in real time, the overall carbon emission intensity of the Yellow River Basin and the carbon emission distribution of each region and industry are calculated once every set time, and the results are displayed in the form of a visual chart on the display module; a multi-level carbon emission warning threshold is set, and when the carbon emission growth rate or increment of a certain region or industry exceeds the set threshold, the system automatically triggers an early warning, and the early warning information is pushed to the environmental protection supervision department and the relevant enterprise through short message and / or email, and a pop-up window is reminded on the mobile terminal APP.
[0035] S5, data updating and model optimization: continuously update the data in the data collection network, and use the new data to optimize the multi-dimensional fusion feature vector and the carbon emission accounting model, and then use the optimized carbon emission accounting model to monitor the carbon emission of the Yellow River Basin.
[0036] Further, as Figure 2As shown, the method of forming a multi-dimensional fusion feature vector in S2 includes the following steps:
[0037] S2-1) Geographical coordinate unification and calibration:
[0038] 1) Satellite remote sensing data coordinate processing:
[0039] Firstly, through the remote sensing analysis module in the remote sensing image processing software ArcGIS, coordinate registration operation is performed on the satellite remote sensing image, a plurality of known accurate geographical positions of control points (including landmark mountains, large bridges and other fixed landmarks in the basin, the accurate coordinates of which can be obtained by high-precision geodetic survey) in the Yellow River Basin are selected, and the coordinate system of the remote sensing image is calibrated to the coordinate system consistent with the selected Yellow River Basin map reference by least square method, so that the geographical position of the remote sensing data is accurately corresponded to the actual geographical space of the Yellow River Basin, and the error is controlled within the allowable range (sub-meter or the accuracy standard set according to actual demand);
[0040] 2) Coordinate matching of ground monitoring station data:
[0041] When the ground monitoring station is constructed, the accurate longitude and latitude coordinates of the ground monitoring station are recorded by high-precision GPS positioning equipment, the longitude and latitude coordinates are converted to the coordinate system consistent with the Yellow River Basin, and the corresponding geographical position coordinate information of each type of monitoring data (atmospheric pollutant concentration, meteorological parameters, etc.) of each ground monitoring station is associated in the data acquisition system; Figure 1
[0042] 3) Coordinate association of Internet of Things data:
[0043] The Internet of Things sensors of carbon emission micro sources installed and deployed in factories, workshops, traffic arteries and residential areas obtain their geographical position coordinates by means of field surveying tools or mobile positioning methods, and embed these coordinate information into the data transmission protocol collected by the sensors, so that the running power of each group of collected energy consumption equipment, vehicle mileage and oil consumption, and household electricity and gas data are all attached with geographical coordinates, so that the Internet of Things data has corresponding position in the geographical space of the Yellow River Basin;
[0044] 4) Coordinate positioning of industry data:
[0045] Through the geographical position information of the registered address of the enterprise and the location of the production facility, combined with the geographic coding, the corresponding geographical position coordinates are given to the industry data in the government and enterprise information system, so that the industry data is positioned on the geographical space level of the Yellow River Basin and participates in the subsequent data fusion process;
[0046] S2-2) Data layering and spatialization processing
[0047] 1) Data layering organization
[0048] Using GIS software, different data sources after coordinate unification are managed according to categories; satellite remote sensing image data is taken as a separate data layer, ground monitoring station data is constructed into different data layers according to station categories (including city stations, industrial park stations, energy production area stations, etc.), Internet of Things data is layered according to the areas where sensors are located (different categories such as factories, traffic, and residents), and industry data is layered according to the industries or areas where enterprises belong to;
[0049] 2) Spatialization representation
[0050] Each layer of data is converted by spatialization representation to make the data present in the form of intuitive geographic spatial elements, ground monitoring station data is visualized on the map in the form of point elements, each point represents a monitoring station, and its attribute information includes real-time monitoring of atmospheric pollutant concentration and meteorological parameters; satellite remote sensing data is spread over the corresponding area of the Yellow River Basin map (factory workshop range, traffic section range, residential area range, etc.) in the form of grid data, each grid pixel carries land use change and vegetation cover information; Internet of Things data is represented by area elements or point elements according to its deployment location, and is associated with corresponding energy consumption and operation attributes; industry data is represented by point elements or area elements (for large enterprise parks, etc.) at the location of enterprises, and its attributes include production scale information;
[0051] S2-3) Matching and fusion based on geographic spatial position
[0052] Based on the spatial overlay analysis function of GIS, the data layers of each layer are overlaid according to the geographic position coordinates, first, taking the Yellow River Basin map as the base, the point element data layer of the ground monitoring station is overlaid with the grid data layer of the satellite remote sensing, at each monitoring station position, the real-time monitoring data of the station and the surrounding land use and vegetation cover information reflected by the corresponding remote sensing image are obtained at the same time; then the area or point element layer corresponding to the Internet of Things data is overlaid into the above layers, so that the factories, traffic arteries and residential areas at the micro location integrate energy consumption data with station monitoring data and remote sensing macro data; finally, the industry data layer is overlaid, so that the industry-related data of enterprises is combined with other carbon emission-related data at the specific geographic location, realizing the matching and fusion of multi-source data at the geographic spatial position;
[0053] S2-4) Attribute fusion and feature extraction:
[0054] For each spatial position after superimposition, multi-source data is subjected to attribute fusion operation. By writing a GIS script, relevant attribute data from different data sources but corresponding to the same position is integrated. Energy consumption data collected by Internet of Things sensors in a factory area, surrounding atmospheric pollutant concentration data obtained by ground monitoring stations, and production scale data of the factory in an enterprise information system are subjected to fusion processing according to a method such as weighted average, to form a set of new attribute data comprehensively reflecting the carbon emission related characteristics of the position. Then, from these fused attribute data, representative features are extracted according to key factors affecting carbon emission to form a multi-dimensional fusion feature vector. The vector elements include the fused energy consumption comprehensive index, the pollutant emission comprehensive index and the industrial scale index. The feature vector is used to comprehensively and comprehensively reflect the carbon emission related characteristics of the corresponding geographic position.
[0055] Further, as shown in Figure 3 S2-4) attribute fusion and feature extraction includes the following steps:
[0056] S2-4-1, determining the weight coefficient of weighted average:
[0057] The relative importance of different data sources for reflecting the carbon emission characteristics of the position is analyzed. The weight coefficients of energy consumption data, atmospheric pollutant concentration data obtained by ground monitoring stations and enterprise production scale data are set according to the importance. Referring to the past carbon emission accounting cases of the data source, relevant research data of the same type of data source, and consulting industry experts, the actual proportion of the influence of each data on carbon emission in a similar scenario is obtained. According to these information, the weight coefficients set initially are further calibrated, and the correction range is ±0.1;
[0058] S2-4-2, writing GIS script to realize weighted average fusion:
[0059] The attribute fusion operation is realized by using the script writing language supported by GIS software. According to the specific data table structure, field name and real weight, the relevant attribute data of different data sources at the same position is fused according to the weighted average method. New attribute data related to energy consumption, atmospheric pollutant concentration and production scale is generated on each data source element;
[0060] S2-4-3, multi-dimensional fusion feature vector construction:
[0061] Based on our understanding of carbon emission mechanisms and the characteristics of industries and the environment in the Yellow River Basin, we identified key factors influencing carbon emissions at this location, including energy consumption, air pollutant concentrations, and production scale. We then comprehensively processed energy consumption data into a comprehensive energy consumption index reflecting the overall energy consumption level of the region, including energy consumption per unit area and energy consumption per unit output value. For pollutant emission data, we constructed a weighted comprehensive index based on the environmental impact weights of different pollutants. For production scale data, we extracted the proportion of carbon-intensive product output that reflects the correlation between industrial scale and carbon emissions.
[0062] Within the factory area, IoT sensors collect data on the operating power and energy consumption duration of multiple devices. First, the total energy consumption of the area is calculated by summing the energy consumption of each device. Then, combined with the area's footprint, the energy consumption per unit area is calculated. Simultaneously, the factory's output data is acquired, and the energy consumption per unit output is calculated. These two indicators are then combined with different weights to form a final comprehensive energy consumption index, which is incorporated into the multi-dimensional fusion feature vector. For the concentration data of various air pollutants acquired from ground monitoring stations, the concentrations of each pollutant are weighted and summed according to their impact on the greenhouse effect and air quality, resulting in a comprehensive pollutant emission index that reflects the overall intensity of air pollutant emissions and their environmental impact in the area. This index is added to the multi-dimensional fusion feature vector. Data related to the enterprise's production scale and industrial structure are combined, along with the enterprise's product classification and output share, to obtain an industrial scale index, which is also added to the multi-dimensional fusion feature vector. Finally, the comprehensive energy consumption index, comprehensive pollutant emission index, and industrial scale index are systematically combined to form the multi-dimensional fusion feature vector. ],in This represents a comprehensive energy consumption index. Indicates a comprehensive index of pollutant emissions. Indicators representing industry scale.
[0063] Furthermore, the analysis of the relative importance of different data sources for reflecting the carbon emission characteristics of this location includes the following steps:
[0064] 1) Obtain the corresponding data sequence
[0065] The reference sequence is a sequence of carbon emission characteristic values:
[0066] ;
[0067] in, Indicates the first The carbon emissions value obtained at that time. This represents the number of elements in the carbon emission characteristic value sequence;
[0068] The comparison sequences include an energy consumption data sequence, an atmospheric pollutant concentration data sequence, and an enterprise production scale data sequence.
[0069] The energy consumption data sequence:
[0070] ;
[0071] wherein, represents energy consumption related data corresponding to the kth moment;
[0072] The atmospheric pollutant concentration data sequence:
[0073] ;
[0074] wherein, represents atmospheric pollutant concentration data monitored at the kth moment;
[0075] The enterprise production scale data sequence:
[0076] ;
[0077] wherein, represents a production scale quantitative index corresponding to the kth moment of the enterprise;
[0078] n represents the number of elements in the comparison sequence; 2) Calculate the correlation coefficient
[0079] The reference sequence
[0080] and the sequence after initialization are respectively , ; ;
[0081] Calculate the correlation coefficient of each comparison sequence and the reference sequence at each moment :
[0082]
[0083]
[0084] The correlation coefficient ranges from 0 to 1, and the closer to 1, the closer the correlation, wherein, represents the absolute value of the data difference between the initialized reference sequence and the ith comparison sequence at the kth moment, which is used to represent the difference between the two at that moment; represents the minimum value of the absolute difference between the reference sequence and the comparison sequence at all n time points, i.e., the minimum value of the absolute difference between the reference sequence and the comparison sequence at all n time points , find the time point with the minimum difference between the reference sequence and the comparison sequence at n time points; represents the minimum value of the minimum absolute difference between the reference sequence and the comparison sequence at all i comparison sequences, i.e., the minimum value of the minimum difference between the reference sequence and the comparison sequence at all comparison sequences at each time point, indicating the closest degree between the overall data sequences;
[0085] , represents the maximum difference between the reference sequence and the comparison sequence at each time point, i.e., first find the maximum value of the absolute difference between the reference sequence and the comparison sequence at n time points in each comparison sequence, and then select the maximum value from the maximum values, indicating the maximum difference between the overall data sequences;
[0086] wherein the reference sequence is initialized as:
[0087] ;
[0088] represents the carbon emission value at the kth time point divided by the carbon emission value at the first time point, converting the carbon emission characteristic value sequence into a relative change sequence with the first value as the reference;
[0089] the comparison sequence is initialized as:
[0090] , ;
[0091] represents the energy consumption, atmospheric pollutant concentration or enterprise production scale data at the kth time point divided by the energy consumption, atmospheric pollutant concentration or enterprise production scale data at the first time point, converting the energy consumption characteristic sequence, atmospheric pollutant concentration characteristic sequence and enterprise production scale characteristic sequence into a relative change sequence with the first value as the reference;
[0092] is a discrimination coefficient, with a value range of 0 to 1, used to adjust the size of the correlation coefficient;
[0093] 3) Calculate the grey correlation degree
[0094] After obtaining the correlation coefficients at each time point, the grey correlation degree of each comparison sequence and the reference sequence is calculated by the following formula:
[0095] , ;
[0096] Grey correlation degree for each comparison sequence The correlation coefficient of the reference sequence at n time points is averaged, reflecting the close degree of correlation between the comparison sequence and the reference sequence in the entire data interval, and the grey correlation degree The value range of the grey correlation degree is between 0 and 1, and the larger the value, the stronger the correlation between the corresponding energy consumption, atmospheric pollutant concentration or enterprise production scale data and the carbon emission characteristic value, and the higher the relative importance in reflecting the carbon emission characteristics.
[0097] Further, as shown in Figure 4 , the S3 specifically includes the following steps:
[0098] S3-1, neural network model selection and structure design
[0099] According to the characteristics of the carbon emission data of the Yellow River Basin, a multilayer perceptron MLP neural network is selected as the basic architecture. The MLP has strong non-linear mapping ability and can handle complex input-output relationships, making it suitable for mining the implicit carbon emission rules in the multi-dimensional fusion feature vector. The number of input layer nodes corresponds to the dimension of the multi-dimensional fusion feature vector, so that it can completely receive the processed multi-source data feature information. The number of hidden layer neurons is set in a decreasing manner, with the number of neurons in the first hidden layer being 2 / 3 of the number of input layer nodes, and the subsequent hidden layers decreasing in turn. The hidden layer activation function uses the commonly used ReLU function. The number of output layer nodes is determined according to the dimension of the carbon emission information to be predicted, and the output layer activation function uses a linear activation function.
[0100] S3-2, support vector machine SVM model configuration
[0101] The radial basis kernel function RBF is used as the kernel function of the SVM model to map the original input space to a high-dimensional feature space. The penalty parameter and the kernel width parameter related to the RBF kernel function are optimized by cross-validation method. A part of the labeled training data is divided into multiple subsets, and the subsets are used as the validation set in turn, and the remaining subsets are used as the training set. The values of and are constantly adjusted to find the optimal parameter combination that makes the model have the best prediction performance on the validation set.
[0102] S3-3, model fusion: first, the multi-dimensional fusion feature vector is input into the designed neural network, and after nonlinear transformation of each layer of the neural network, high-level abstract features are extracted, and an intermediate feature representation is output, then the intermediate feature is taken as the input of the support vector machine, and the support vector machine is used for final carbon emission information prediction;
[0103] S3-4, training of carbon emission accounting model
[0104] The collected multi-dimensional fusion feature vector data set with carbon emission annotation information is divided into training set, validation set and test set according to a certain proportion, the training set is used for learning and adjusting model parameters, and the validation set is used for evaluating the generalization ability of the model during training; the mean square error MSE is selected as the loss function; for the neural network part, the Adam optimization algorithm is used for parameter updating training, and for the support vector machine part, the training algorithm of the selected optimization software library is used for parameter solving, and the optimal classification or regression hyperplane is found based on the quadratic programming algorithm;
[0105] According to the set training algorithm and parameters, the training set data is repeatedly input into the fused carbon emission accounting model, the weight and bias parameters of the model are continuously adjusted, the loss function value is gradually reduced, and at the same time, after each round of training is completed, the model is evaluated with the validation set data, the change of the loss value and the mean absolute error MAE on the validation set are recorded, and whether the model appears overfitting is observed;
[0106] S3-5, using random forest regression algorithm to adjust the parameters dynamically:
[0107] Through experiments, the prediction error of the model on the validation set is compared when the number of decision trees is different, and the number of decision trees that makes the performance of the model stable is selected; the maximum depth of the decision tree and the minimum number of samples required for node splitting are configured, and a random forest regression model for parameter adjustment is constructed; each feature in the multi-dimensional fusion feature vector is input into the constructed random forest regression model, the average reduction based on node impurity is used to analyze the relative importance of each feature to the prediction result of carbon emissions, and it is determined which features are key factors affecting carbon emissions and which are relatively secondary factors; at the same time, the random forest regression model mines implicit influencing factors, obtains the influencing mechanism by analyzing the branching conditions of the decision tree and the combination of each feature in different decision trees; new labeled carbon emission data and corresponding multi-dimensional fusion feature vectors are collected regularly and input into the trained carbon emission accounting model to obtain the prediction result, and the deviation between the predicted value and the true value is calculated; according to the key features and influencing factors analyzed by the random forest regression model, the parameters related to these factors in the carbon emission accounting model are determined, and the quantitative relationship between the features and carbon emissions obtained by the random forest regression model is used to dynamically adjust these key parameters to optimize the prediction ability of the model for new data.
[0108] Further, as shown in Figure 5 S4 specifically includes the following steps:
[0109] S4-1, real-time dynamic monitoring and calculation:
[0110] Firstly, through the data interface and network communication protocol, the collected multi-source data (satellite remote sensing, ground monitoring stations, Internet of Things sensors, and government and enterprise information systems, etc.) are transmitted in real time to the computing server or cloud computing platform environment according to the established data format, and the real-time data flowing in are preprocessed during the data transmission process; according to the set time interval (for example, every 1 hour, half an hour, etc., which can be adjusted flexibly according to actual demand and computing resources, monitoring accuracy, etc.), the carbon emission accounting model is started for calculation, and the model takes the processed multi-dimensional fusion feature vector input in real time as input, runs the internal neural network and support vector machine fusion calculation logic, and outputs the corresponding carbon emission information;
[0111] When calculating the overall carbon emission intensity of the Yellow River Basin, the estimated values of carbon emissions of all regions and all industries are aggregated, and combined with the geographical range information of the Yellow River Basin, the carbon emission intensity per unit area is obtained by dividing the total carbon emission by the area; for the calculation of the carbon emission distribution of each region, according to the geographical position coordinate information carried in the data, the carbon emission estimation results are divided into different regions (such as according to the administrative division, different ecological function zones or pre-set monitoring grid regions in the Yellow River Basin, etc.), and the carbon emission values and proportion of each region are calculated; for the carbon emission distribution of each industry, according to the industry classification information of the enterprise, the carbon emission of different industries (such as power, steel, chemical industry in industry, transportation industry, service industry, etc.) is obtained, and the total carbon emission of each industry and the proportion in the overall emission are calculated.
[0112] S4-2, visualization chart display:
[0113] According to the characteristics of the carbon emission related indicators to be displayed, the type of visualization chart is selected, for the overall carbon emission intensity trend, a line chart is used, with time (hour, day, month, etc. time dimension, according to the actual display period) as the horizontal axis, and carbon emission intensity value as the vertical axis; for the carbon emission distribution of each region, a heat map is used for display, taking the Yellow River Basin map as the base, different regions are rendered with different color depths according to their carbon emission values (regions with high carbon emission are rendered with red color system and deeper color indicates more emission, regions with low carbon emission are rendered with green color system and shallower color indicates less emission); pie chart is used to display the proportion of carbon emission of each industry, the whole circle represents the total carbon emission of the Yellow River Basin, each sector corresponds to a different industry, and the area of the sector is proportional to the proportion of carbon emission of the industry;
[0114] In addition to the above main chart types, various chart forms such as column chart (comparing the absolute values of carbon emission of different regions or industries in the same period, etc.), stacked chart (showing the composition of different categories of emission sources in the carbon emission of each region or industry, etc.) can be used to comprehensively display the carbon emission related information of the Yellow River Basin from multiple angles, meeting the needs of different users (environmental protection supervision department personnel, enterprise managers, scientific research personnel, etc.) to view and analyze data.
[0115] Build a visual display module, integrate it into the user interface of the monitoring system, write code to bind the data calculated by the carbon emission accounting model to the corresponding visual chart elements in real time, when using Echarts to draw a line chart to show the overall carbon emission intensity, configure the basic style of the chart through JavaScript code, then use the data interface to get the intensity values output by the carbon emission accounting model in each round of calculation, and dynamically add these data to the data series of the line chart according to the time sequence, to realize real-time update and display of the line chart; for the heat map to show the carbon emission distribution of each region, also through code to associate the carbon emission data of each region output by the model with the corresponding regional elements on the map, adjust the regional color rendering effect according to the value, and reflect the changes of regional carbon emission in real time.
[0116] S4-3, multi-level carbon emission warning threshold setting and warning triggering:
[0117] Set thresholds based on historical data and policy standards: analyze the carbon emission history data of the Yellow River Basin in the past 10 years, obtain the normal range and fluctuation of carbon emission of different regions and industries in the Yellow River Basin under normal production and operation and different seasons, and set the basic warning threshold based on this; at the same time, refer to the relevant policies and regulations of carbon emission issued by the state and local governments and industry standards, combined with the ecological protection target of the Yellow River Basin, comprehensively determine the specific value of the warning threshold of different levels;
[0118] Warning triggering mechanism and information push: compare and judge the carbon emission growth rate and increment of each region and each industry calculated by the carbon emission accounting model with the set multi-level warning threshold in real time, through program logic, once the carbon emission growth rate or increment of a certain region or industry exceeds the corresponding threshold level, trigger the warning mechanism immediately, integrate the short message gateway service, email sending service and mobile APP development interface function into the monitoring system, and realize the warning of different information.
[0119] Generally, it can be divided into three or four levels of warning:
[0120] Level 1 warning (mild warning): triggered when the carbon emission growth rate or increment of a certain region or industry is slightly higher than the normal fluctuation range, but has not yet reached the degree that may have a significant impact on the environment or industrial development. For example, the carbon emission growth rate exceeds 10%-20% (the specific value is adjusted according to the actual situation) compared with the average growth rate in the same period, which can be set as the first level warning threshold range, prompting relevant parties to pay attention to the changes in carbon emission and appropriately investigate the possible reasons.
[0121] Secondary warning (moderate warning): triggered when the growth rate or increment of carbon emissions further exceeds the reasonable range, which may begin to put pressure on regional environmental quality, industrial low-carbon transformation, etc. For example, the growth rate exceeds 20%-50% compared with the historical same period, at which time relevant departments and enterprises need to pay more attention, analyze the reasons and prepare to take some preliminary control measures.
[0122] Tertiary warning (severe warning): triggered when the growth rate or increment of carbon emissions seriously exceeds the normal level, which brings significant risks to the ecological environment and sustainable economic development of the Yellow River Basin. For example, the growth rate exceeds 50% or even higher, at which time strong emission reduction and control measures need to be taken immediately, and environmental protection supervision departments may need to intervene in strict supervision and rectification.
[0123] Correspondingly, as shown in Figure 6 The application also discloses a dynamic monitoring system for carbon emissions in the Yellow River Basin using big data, which comprises:
[0124] The data acquisition module 101 is used to build a multi-source data acquisition network, and integrate data related to carbon emissions from satellite remote sensing, ground monitoring stations, Internet of Things sensors, and various government and enterprise information systems in the Yellow River Basin.
[0125] The data preprocessing module 102 is used to first clean the collected multi-source data, then perform data standardization, unify the units and dimensions of data from different sources according to a unified data dictionary, and finally perform data fusion to form a multi-dimensional fusion feature vector representing the relationship between the multi-source data.
[0126] The carbon emission accounting model construction module 103 is used to construct a carbon emission accounting model by fusing a neural network and a support vector machine, take the processed multi-dimensional fusion feature vector as input and carbon emission information as output, train the carbon emission accounting model, and use a random forest regression algorithm to dynamically adjust the accounting parameters of the carbon emission accounting model and optimize the model parameters.
[0127] The dynamic monitoring and early warning module 104 is used to use the constructed carbon emission accounting model to perform real-time dynamic monitoring of carbon emissions in the Yellow River Basin, and automatically trigger a warning when the growth rate or intensity of carbon emissions in a certain region or industry is too high.
[0128] The data updating and model optimization module 105 is used to continuously update various data in the data acquisition network, optimize the multi-dimensional fusion feature vector and the carbon emission accounting model using the new data, and then use the optimized carbon emission accounting model to monitor the carbon emissions in the Yellow River Basin.
[0129] It should be noted that the implementation method of the corresponding module in the system can refer to the aforementioned dynamic monitoring method, which will not be described here.
[0130] The method and system can timely, comprehensively and accurately monitor the carbon emissions in the Yellow River Basin.
Claims
1. A method for dynamic monitoring of carbon emissions in the Yellow River Basin using big data, characterized by Comprise the following steps: S1, data acquisition: build a multi-source data acquisition network, integrate satellite remote sensing, ground monitoring stations, Internet of Things sensors and various government, enterprise information systems related to carbon emissions data from the Yellow River Basin; S2, data preprocessing: first, clean the collected multi-source data, then standardize the data, unify the units and dimensions of data from different sources according to a unified data dictionary, and finally perform data fusion. Based on geographic information system GIS, the satellite remote sensing data, ground monitoring data, Internet of Things data and industry data after data cleaning and standardization are matched and fused according to geographic location coordinates on the Yellow River Basin map to form a multi-dimensional fusion feature vector; S3, carbon emission accounting model construction: through the fusion of neural network and support vector machine, a carbon emission accounting model is constructed, the processed multi-dimensional fusion feature vector is taken as input, and the carbon emission information is taken as output. The carbon emission accounting model is trained, and the random forest regression algorithm is used to dynamically adjust the accounting parameters of the carbon emission accounting model to optimize the model parameters; S4, dynamic monitoring and early warning: using the constructed carbon emission accounting model, the carbon emissions in the Yellow River Basin are monitored in real time. When the monitoring shows that the carbon emission growth rate of a certain region or industry is too fast or the intensity is out of standard, the system automatically triggers an early warning; S5, data updating and model optimization: continuously update various data in the data acquisition network, and optimize the multi-dimensional fusion feature vector and the carbon emission accounting model using the new data, and then use the optimized carbon emission accounting model to monitor the carbon emissions in the Yellow River Basin; Step S3 specifically comprises the following steps: S3-1, neural network model selection and structure design: Select a multi-layer perceptron MLP neural network as the basic architecture, wherein the number of input layer nodes corresponds to the dimension of the multi-dimensional fusion feature vector; the number of hidden layer neurons is set in a decreasing manner, the number of neurons in the first hidden layer is 2 / 3 of the number of input layer nodes, and the subsequent hidden layers are sequentially decreased, and the hidden layer activation function is selected as ReLU function; the number of output layer nodes is determined according to the dimension of the carbon emission information to be predicted, and the output layer activation function is selected as linear activation function; S3-2, support vector machine SVM model configuration: Using radial basis kernel function RBF as the kernel function of SVM model, the penalty parameter and the kernel width parameter of RBF kernel function are optimized by cross-validation method, a part of labeled training data is divided into multiple subsets, and subsets are used as validation set and the remaining subsets are used as training set in turn, and the values of and are adjusted constantly to find the best combination of parameters that make the model have the best prediction performance on the validation set; S3-3, model fusion: first, input the multi-dimensional fusion feature vector into the designed neural network, perform nonlinear transformation through each layer of the neural network, extract high-level abstract features, and output an intermediate feature representation. Then, the intermediate feature representation is taken as the input of the support vector machine, and the support vector machine performs final carbon emission information prediction. 2.The method of claim 1, wherein, In the S1 data acquisition step: Satellite remote sensing data is used to obtain large-area land use change information in the Yellow River Basin, which indirectly reflects carbon emission factors and vegetation cover information Ground monitoring sites are set up in cities, industrial parks and energy production areas in the Yellow River Basin to monitor atmospheric pollutant concentrations and meteorological parameter information in real time Internet of Things sensors are deployed at carbon emission micro sources, including factory workshops, traffic arteries and residential areas, to collect and transmit to the data center in real time the operating power data of energy consumption equipment , vehicle mileage and fuel consumption data , household electricity and gas consumption data At the same time, it is connected with the information systems of environmental protection, energy and industrial and information departments and large enterprises in the basin to obtain enterprise production scale, energy consumption structure and product output data . 3.The method of claim 1, wherein, The method for forming a multi-dimensional fusion feature vector comprises the following steps: S2-1) geographic coordinate unification and calibration: 1) satellite remote sensing data coordinate processing: First, through the remote sensing analysis module in the remote sensing image processing software ArcGIS, the satellite remote sensing image is subjected to coordinate registration operation, a plurality of known accurate geographic position control points in the Yellow River Basin are selected, the coordinate system of the remote sensing image is calibrated to the coordinate system consistent with the selected Yellow River Basin map datum by the least square method, so that the geographic position of the remote sensing data corresponds to the actual geographic space of the Yellow River Basin; 2) Coordinate matching of ground monitoring station data: When building the layout of ground monitoring stations, record the accurate longitude and latitude coordinates of the stations by GPS positioning equipment, convert the coordinates to the coordinate system consistent with the Yellow River Basin map, and associate the geographic position coordinate information of each type of monitoring data of each ground monitoring station during data collection; 3) Coordinate association of Internet of Things data: The Internet of Things sensors installed and deployed for carbon emission micro sources use field surveying and mapping tools or mobile positioning methods to obtain their geographic position coordinates, and embed these coordinate information into the data transmission protocol collected by the sensors, so that each set of collected energy consumption equipment operation power, vehicle mileage and fuel consumption, and household electricity and gas data are accompanied by geographic coordinates, so that the Internet of Things data has a corresponding position in the geographic space of the Yellow River Basin; 4) Coordinate positioning of industry data: Through the geographic position information of the registered address and production facility location of enterprises, combined with geographic coding, the industry data in the government and enterprise information system is given corresponding geographic position coordinates, so that the industry data is positioned on the geographic space level of the Yellow River Basin; S2-2) Data layering and spatialization processing: 1) Data layering organization Use GIS software to manage different data sources data according to categories after coordinate unification processing; satellite remote sensing image data is taken as a separate data layer, ground monitoring station data is constructed into different data layers according to station categories, Internet of Things data is layered according to the area where the sensors are located, and industry data is layered according to the industry or area where the enterprises belong; 2) Spatialization representation Spatialization representation conversion is performed on each layer of data to present the data in the form of intuitive geographic space elements, ground monitoring station data is visualized and displayed on the map in the form of point elements, each point represents a monitoring station, and its attribute information includes real-time monitoring of atmospheric pollutant concentration and meteorological parameters; satellite remote sensing data is spread over the corresponding area of the Yellow River Basin map in the form of grid data, each grid pixel carries land use change information and vegetation cover information; Internet of Things data is represented in the form of area elements or point elements according to its deployment location, and is associated with the corresponding energy consumption and operation attributes; industry data is represented in the form of point elements or area elements at the location of enterprises, and its attributes include production scale information; S2-3) Matching and fusion based on geographic space position The GIS-based spatial overlay analysis function will overlay the data layers of each layer according to the geographical position coordinates. First, the Yellow River Basin map is taken as the base, and the point feature data layer of the ground monitoring station is overlaid with the satellite remote sensing grid data layer. At each monitoring station position, the real-time monitoring data of the station and the surrounding land use and vegetation cover information reflected by the corresponding remote sensing image are obtained simultaneously. Then, the region or point feature layer corresponding to the Internet of Things data is overlaid into the above layers, so that the factories, traffic arteries and residential areas at the micro location are integrated with the energy consumption data, the station monitoring data and the remote sensing macro data. Finally, the industry data layer is overlaid, so that the industry-related data of the enterprise are combined with other carbon emission-related data at the specific geographical location, realizing the matching and fusion of multi-source data in geographical space position. S2-4) Attribute fusion and feature extraction: For the multi-source data at each spatial position after overlay, attribute fusion operation is performed. By writing GIS scripts, the relevant attribute data from different data sources but corresponding to the same position are integrated. According to the weighted average method, the fusion processing is performed to form a new set of attribute data that comprehensively reflects the carbon emission-related characteristics of the position. Then, from these fused attribute data, representative features are extracted according to the key factors that affect carbon emission, and combined into a multi-dimensional fusion feature vector. The vector elements include the fused energy consumption comprehensive index, the pollution emission comprehensive index and the industry scale index. 4.The method of claim 3, wherein, The step S2-4) attribute fusion and feature extraction includes the following steps: S2-4-1, determining the weight coefficient of weighted average: The relative importance of different data sources for reflecting the carbon emission characteristics of the position is analyzed. According to the importance, the weight coefficients of energy consumption data, atmospheric pollutant concentration data obtained by ground monitoring stations and enterprise production scale data are set; Referring to the past carbon emission accounting cases of this data source, relevant research data of the same type of data source, and consulting industry experts, the actual proportion of the influence of each data on carbon emission in a similar scenario is obtained. According to these information, the weight coefficients set initially are further calibrated, and the correction range is ±0.1; S2-4-2, writing GIS scripts to realize weighted average fusion: Using the script writing language supported by GIS software to realize attribute fusion operation, according to the specific data table structure, field name and real weight, the related attribute data of different data sources at the same position are fused according to the weighted average method. New attribute data of energy consumption, atmospheric pollutant concentration and production scale related to each data source element are generated; S2-4-3, multi-dimensional fusion feature vector construction: Based on the understanding of the carbon emission mechanism and the industry and environmental characteristics of the Yellow River Basin, the key factors affecting the carbon emission of the position are determined, including energy consumption, atmospheric pollutant concentration and production scale. The energy consumption data are integrated into an energy consumption comprehensive index reflecting the overall energy consumption level of the region, including unit area energy consumption and unit output value energy consumption. For pollutant emission data, a weighted comprehensive index based on the environmental impact weight of different pollutants is constructed; for production scale data, the proportion of carbon-intensive product output reflecting the correlation between industrial scale and carbon emissions is extracted; In the factory area, the Internet of Things sensors collect the running power and energy consumption time data of multiple equipment. First, the total energy consumption of the area is calculated by adding up the energy consumption of each device, and then the unit area energy consumption value is calculated combined with the floor area of the area. At the same time, the output value data of the factory is obtained, and the unit output value energy consumption is calculated. These two indicators are combined in a way that different weights are assigned to form the final energy consumption comprehensive index, which is included in the multi-dimensional fusion feature vector. For the various atmospheric pollutant concentration data obtained by the ground monitoring station, the pollutant concentration is weighted and summed according to the influence weight of each pollutant on the greenhouse effect and air quality, to obtain a comprehensive index reflecting the overall intensity of atmospheric pollutant emission and the degree of environmental impact in the area, which is added to the multi-dimensional fusion feature vector. The data related to the production scale of the enterprise and the industrial structure data are combined and fused with the product classification and yield proportion of the enterprise to obtain the industrial scale index, which is added to the multi-dimensional fusion feature vector. The fusion energy consumption comprehensive index, the pollutant emission comprehensive index and the industry scale index are sequentially combined to form a multi-dimensional fusion feature vector ]wherein represents the energy consumption comprehensive index, represents the pollutant emission comprehensive index, represents the industry scale index. 5.The method of claim 4, wherein, The analysis of the relative importance of different data sources for reflecting the carbon emission characteristics of the location includes the following steps: 1) Obtain the corresponding data sequence The reference sequence is the carbon emission characteristic value sequence: ; wherein, represents the carbon emission value obtained at the first time point, is the number of elements in the carbon emission characteristic value sequence. The comparison sequence includes the energy consumption data sequence, atmospheric pollutant concentration data sequence, and enterprise production scale data sequence; Energy consumption data sequence: ; wherein, represents the energy consumption related data corresponding to the time point; Atmospheric pollutant concentration data sequence: ; wherein, represents the atmospheric pollutant concentration data monitored at the first time point; Enterprise production scale data sequence: ; wherein, represents the production scale quantitative index corresponding to the enterprise at the first time. To compare the number of elements in a sequence; 2) Calculate the correlation coefficient Reference sequence and comparison sequence The sequences after initialization are respectively , ; calculating a correlation coefficient of each comparison sequence with the reference sequence at each time instant : correlation coefficient The value range is 0 to 1, the closer to 1, the closer the correlation, wherein, represents the initialized reference sequence at the kth moment and the ith comparison sequence corresponding to the absolute value of the data difference, used to represent the difference between the two at that moment; represents the minimum value of the absolute value of the difference between the reference sequence and the comparison sequence among all n moments, that is, for a particular comparison sequence , find the minimum value of the absolute value of the difference between the reference sequence at the moment of the smallest difference among the n moments; represents the minimum value of the minimum difference absolute value among all i comparison sequences, that is, the minimum value of the minimum difference between all comparison sequences and the reference sequence at each moment, representing the closest degree between the overall data sequences. represents the maximum difference between all comparison sequences and the reference sequence at each time, i.e., first find the maximum value of the absolute difference between each comparison sequence and the reference sequence at n times, and then select the maximum value from these maximum values to represent the maximum difference between the overall data sequences. The reference sequence is initialized: ; represents the carbon emission value at the kth moment divided by the carbon emission value at the first moment, and the carbon emission characteristic value sequence is converted into a relative change sequence with the first value as the benchmark; The comparison sequence is initialized: , ; indicates the energy consumption, atmospheric pollutant concentration or enterprise production scale data at the kth moment divided by the energy consumption, atmospheric pollutant concentration or enterprise production scale data at the first moment, and converts the energy consumption feature sequence, atmospheric pollutant concentration feature sequence and enterprise production scale feature sequence into a relative change sequence with the first value as the benchmark; For the resolution coefficient, the value range is between 0 and 1, which is used to adjust the size of the correlation coefficient; 3) Calculate the gray correlation degree After obtaining the correlation coefficient at each time, the gray correlation degree of each comparison sequence and the reference sequence is calculated by the following formula: , ; Grey correlation degree for each comparison sequence The correlation coefficient of the reference sequence at n time points The correlation coefficient of the reference sequence at n time points The value range of the grey correlation degree is between 0 and 1. The larger the value, the stronger the correlation between the corresponding energy consumption, atmospheric pollutant concentration or enterprise production scale data and the carbon emission characteristic value, and the higher the relative importance in reflecting the carbon emission characteristics. 6.The method of claim 1, wherein, The S3 further includes the following steps: S3-4, training of the carbon emission accounting model: The multi-dimensional fusion feature vector data set collected with carbon emission annotation information is divided into training set, validation set and test set according to a certain proportion. The training set is used for learning and adjusting the model parameters, and the validation set is used to evaluate the generalization ability of the model during training. The mean square error MSE is selected as the loss function. For the neural network part, the Adam optimization algorithm is used for parameter update training. For the support vector machine part, the training algorithm provided by the selected optimization software library is used for parameter solving, and the quadratic programming algorithm is used to find the optimal classification or regression hyperplane. According to the set training algorithm and parameters, the training set data is repeatedly input into the fused carbon emission accounting model, and the weight and bias parameters of the model are continuously adjusted, so that the loss function value gradually decreases. At the same time, after each round of training is completed, the model is evaluated with the validation set data, and the changes of the loss value and the mean absolute error MAE on the validation set are recorded to observe whether the model is overfitting. S3-5, dynamic adjustment of the accounting parameters using the random forest regression algorithm: The model performance is stabilized by comparing the prediction error of the model on the validation set with different numbers of decision trees through experiments; the maximum depth of the decision tree and the minimum number of samples required for node splitting are configured to build a random forest regression model for parameter adjustment; each feature in the multi-dimensional fusion feature vector is input into the built random forest regression model, the average reduction based on node impurity is used to analyze the relative importance of each feature to the carbon emission prediction result, and it is determined which features are the key factors affecting carbon emissions and which are the relatively secondary factors; at the same time, the random forest regression model mines the implicit influencing factors, obtains the influencing mechanism by analyzing the branch conditions of the decision tree and the combination of each feature in different decision trees; new labeled carbon emission data and corresponding multi-dimensional fusion feature vectors are collected regularly and input into the trained carbon emission accounting model to obtain the prediction result, and the deviation between the predicted value and the true value is calculated; according to the key features and influencing factors analyzed by the random forest regression model, the parameters related to these factors in the carbon emission accounting model are determined, and the quantitative relationship between the features and carbon emissions obtained by the random forest regression model is used to adjust these key parameters to optimize the prediction ability of the model for new data.
7. A dynamic monitoring system for carbon emissions in the Yellow River Basin using big data, wherein the system runs the dynamic monitoring method for carbon emissions in the Yellow River Basin according to any one of claims 1-6. The system comprises: a data acquisition module for building a multi-source data acquisition network, integrating carbon emission-related data from satellite remote sensing, ground monitoring stations, Internet of Things sensors, and various government and enterprise information systems in the Yellow River Basin; a data preprocessing module for first cleaning the collected multi-source data, then standardizing the data, unifying the units and dimensions of data from different sources according to a unified data dictionary, and finally fusing the data to form a multi-dimensional fusion feature vector representing the relationship between the multi-source data; a carbon emission accounting model construction module for constructing a carbon emission accounting model using a fusion neural network and a support vector machine, taking the processed multi-dimensional fusion feature vector as input and carbon emission information as output, training the carbon emission accounting model, and using a random forest regression algorithm to dynamically adjust the accounting parameters of the carbon emission accounting model and optimize the model parameters; a dynamic monitoring and early warning module for using the constructed carbon emission accounting model to monitor the carbon emissions in the Yellow River Basin in real time, and automatically triggering an early warning when the carbon emission growth rate of a certain region or industry is too fast or the intensity exceeds the standard; a data updating and model optimization module for continuously updating various data in the data acquisition network, optimizing the multi-dimensional fusion feature vector and the carbon emission accounting model using the new data, and then monitoring the carbon emissions in the Yellow River Basin using the optimized carbon emission accounting model.
Citation Information
Patent Citations
Dual-carbon park carbon emission monitoring display method and system, storage medium and equipment
CN118569878A
Regional carbon emission accounting system
CN119443534A