Carbon emission surplus and deficit prediction method and system based on data mining
Through data mining methods, data collection and cleaning from carbon emission sources are constructed, feature sets are established and prediction models are established, which solves the accuracy and control efficiency of carbon emission prediction, and accurately captures and efficient management of dynamic changes in carbon emissions.
Patent Information
- Application Number
- CN202411598183.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2044-11-11
AI Technical Summary
The existing carbon emissions profit and shortage prediction technology is difficult to accurately capture dynamic changes in real time, resulting in insufficient accuracy and inefficient management and control efficiency.
Through a data mining method, historical data is collected from multiple carbon emission sources, carbon emission data sets are established, data cleaning and feature extraction are carried out, feature sets are constructed, carbon emission prediction models are established, and trigger thresholds are set for management and control.
The accuracy and control efficiency of carbon emissions profit and shortage prediction have been improved, and the accurate capture and effective management of real-time dynamic changes in carbon emissions have been achieved.
Smart Images

Figure CN119539165B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field related to carbon emission monitoring, and specifically to a method and system for predicting carbon emission surplus or deficit based on data mining. Background Art
[0002] In order to respond to global climate change and reduce greenhouse gas emissions, carbon emission prediction and management have become key links. With the accelerated development of industrialization and urbanization, accurate prediction and effective control of carbon emissions are particularly important. However, traditional carbon emission surplus and deficit prediction faces many challenges in practical applications. First, carbon emission prediction often relies on simple linear regression models or time series analysis methods. By fitting historical carbon emission data, it attempts to reveal its trend over time and predict future carbon emissions accordingly. However, changes in carbon emissions are affected by many complex factors, including the efficiency of production equipment, adjustments to energy structure, changes in the external environment, etc. The interaction between these factors often shows highly nonlinear characteristics. Moreover, with the continuous accumulation of carbon emission data, the amount of data has exploded, resulting in limited real-time and accuracy of the prediction results, which cannot meet the real-time prediction and rapid response requirements of carbon emission processing.
[0003] Therefore, at the current stage, the technologies related to carbon emission surplus and deficit prediction have the problem of being unable to accurately and real-time capture the dynamic change characteristics of carbon emissions in the face of complex and changeable carbon emission data, which leads to insufficient accuracy of carbon emission surplus and deficit prediction and low efficiency of carbon emission control. Summary of the Invention
[0004] This application provides a carbon emission surplus and deficit prediction method and system based on data mining, which solves the technical problems in the existing technology that it is difficult to accurately and real-time capture the dynamic change characteristics of carbon emissions in the face of complex and changeable carbon emission data, which in turn leads to insufficient accuracy of carbon emission surplus and deficit prediction and low efficiency of carbon emission control, and achieves the technical effect of improving the accuracy of carbon emission surplus and deficit prediction and control efficiency.
[0005] The present application provides a carbon emission surplus and shortage prediction method based on data mining, which includes: collecting historical carbon emission data from multiple carbon emission sources to establish a carbon emission data set, wherein the carbon emission data set includes processing data of production equipment and energy-using equipment, energy consumption, carbon emission data, and external environment data; after data cleaning the carbon emission data set, feature extraction is performed on the cleaned carbon emission data set based on a fusion feature network to establish a first feature set and a second feature set; training samples are constructed based on the first feature set and the second feature set and surplus and shortage records, and a carbon emission prediction model is established using the training samples; real-time carbon emission feature data is used as input data and input into the carbon emission prediction model to generate surplus and shortage prediction results; a trigger threshold is set, and a trigger analysis is performed based on the surplus and shortage prediction results and the trigger threshold, and carbon emission control is performed based on the trigger analysis results.
[0006] In a possible implementation, the feature extraction of the carbon emission dataset after data cleaning is performed based on the fusion feature network, and the following processing is also performed: calling the equipment feature extraction sub-network within the fusion feature network, performing processing data analysis of the carbon emission dataset after data cleaning, and obtaining equipment operating parameters, including operating time, start and shutdown frequency, load rate, and operating speed; extracting energy consumption characteristics, operating efficiency characteristics, and fault characteristics of the equipment operating parameters to establish a basic feature set; using the basic feature set to fuse and associate production rhythm characteristics, and optimizing redundant features based on the dimensionality reduction of principal component features to establish key influencing factors; and constructing a first feature set based on the key influencing factors.
[0007] In a possible implementation, the basic feature set is used to fuse and associate production rhythm features, and redundant features are optimized based on the principal component feature dimensionality reduction, and the following processing is also performed: variance analysis is performed on the fused and associated features to establish a variance calculation result; feature screening is performed using a preset threshold and variance calculation result to establish a preliminary screening feature set; a predetermined number of principal components are selected from the preliminary screening feature set, and cumulative variance analysis of the current principal component is performed to establish a cumulative variance analysis result; the component contribution in the current principal component is obtained, and judgment is made based on the component contribution and explanation rate threshold; if the cumulative variance analysis result fails to reach the cumulative threshold, a reconstruction principal component instruction is generated, and principal component optimization is performed based on the reconstruction principal component instruction and the judgment result of the component contribution and explanation rate threshold; and principal component feature dimensionality reduction is completed using the principal component optimization result.
[0008] In a possible implementation, the method performs feature extraction on the carbon emission data set after data cleaning based on the fusion feature network, and also performs the following processing: calling the external environment feature extraction subnetwork within the fusion feature network, performing external environment data analysis based on the external environment feature extraction subnetwork, and obtaining meteorological data, geographic data, and policy data; after feature extraction based on the meteorological data, geographic data, and policy data, performing multi-feature cross analysis, and establishing a second feature set based on the cross analysis results.
[0009] In a possible implementation, the method performs feature extraction on the carbon emission dataset after data cleaning based on the fusion feature network, and further performs the following processing: obtaining a real-time data stream, configuring the initial feature weights of the fusion feature network based on the real-time data stream; performing feature extraction on the carbon emission dataset based on the initial feature weights, and establishing a feature gradient of the initial feature weights based on the feature intensity; adaptively updating the initial feature weights using the feature gradient, performing attention optimization of the fusion feature network; and re-extracting attention features based on the attention optimization results.
[0010] In a possible implementation, the real-time carbon emission characteristic data is used as input data and input into the carbon emission prediction model to generate a surplus or shortage prediction result, and the following processing is also performed: an abnormality identification center is established, and the abnormality identification center is configured with an abnormality identification threshold; before the real-time carbon emission characteristic data is input, the real-time carbon emission characteristic data is triggered to determine the abnormality identification threshold; if the trigger determination result is a trigger result, an early warning notification is triggered.
[0011] In a possible implementation, the method constructs training samples based on the first feature set, the second feature set and the surplus and shortage records, establishes a carbon emission prediction model based on the training samples, and further performs the following processing: establishes a dynamic data update mechanism, and the dynamic data update mechanism is used to perform update accumulation judgment on the real-time data stream; if the trust and data volume of the real-time data stream synchronously meet the dynamic data update mechanism, a feedback instruction is generated; based on the feedback instruction, the real-time data stream is controlled to perform the update and supplement of the carbon emission data set, and the carbon emission prediction model is reconstructed based on the update and supplement results.
[0012] The present application also provides a carbon emission surplus and shortage prediction system based on data mining, including: a carbon emission data set establishment module, which is used to collect historical carbon emission data from multiple carbon emission sources and establish a carbon emission data set, wherein the carbon emission data set includes processing data of production equipment and energy-using equipment, energy consumption, carbon emission data, and external environment data; a data set feature extraction module, which is used to perform data cleaning on the carbon emission data set, extract features from the cleaned carbon emission data set based on a fusion feature network, and establish a first feature set and a second feature set; a carbon emission prediction model establishment module, which is used to construct training samples based on the first feature set and the second feature set and surplus and shortage records, and establish a carbon emission prediction model based on the training samples; a surplus and shortage prediction result generation module, which is used to input real-time carbon emission feature data as input data into the carbon emission prediction model to generate surplus and shortage prediction results; a carbon emission control module, which is used to set a trigger threshold, perform trigger analysis based on the surplus and shortage prediction results and the trigger threshold, and perform carbon emission control based on the trigger analysis results.
[0013] The carbon emission surplus and shortage prediction method and system based on data mining proposed in this application is intended to collect historical carbon emission data from multiple carbon emission sources and establish a carbon emission data set; perform feature extraction on the carbon emission data set after data cleaning based on a fusion feature network; construct training samples based on the first feature set and the second feature set and surplus and shortage records, and establish a carbon emission prediction model based on the training samples; input real-time carbon emission feature data into the carbon emission prediction model to generate surplus and shortage prediction results; perform trigger analysis based on the surplus and shortage prediction results and the trigger threshold, and use the trigger analysis results to control carbon emissions. This solves the technical problem in the existing technology that it is difficult to accurately capture the dynamic change characteristics of carbon emissions in real time when faced with complex and changeable carbon emission data, which leads to insufficient accuracy in carbon emission surplus and shortage prediction and low efficiency in carbon emission control, and achieves the technical effect of improving the accuracy of carbon emission surplus and shortage prediction and the control efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments of the present disclosure are briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in precise order. Instead, various steps may be processed in reverse order or simultaneously as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0015] Figure 1 A flow chart of a method for predicting carbon emissions surplus or deficit based on data mining provided in an embodiment of the present application;
[0016] Figure 2 Schematic diagram of the structure of the carbon emission surplus and deficit prediction system based on data mining provided in an embodiment of the present application.
[0017] Explanation of the reference numerals: carbon emission data set establishment module 10 , data set feature extraction module 20 , carbon emission prediction model establishment module 30 , surplus / shortage prediction result generation module 40 , carbon emission control module 50 . DETAILED DESCRIPTION
[0018] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.
[0019] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0020] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict, and the terms “first\second” involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. The terms “including” and “having” and any variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application only.
[0021] The present application embodiment provides a carbon emission surplus and shortage prediction method based on data mining, such as Figure 1 As shown, the method includes:
[0022] Step S100 , collecting historical carbon emission data from multiple carbon emission sources to establish a carbon emission data set, wherein the carbon emission data set includes processing data of production equipment and energy-using equipment, energy consumption, carbon emission data, and external environment data.
[0023] Preferably, historical carbon emission data are obtained from multiple carbon emission sources (equipment or places that may generate carbon emissions). Specifically, carbon emission sources may include industrial facilities (such as factories, power stations, steel mills, chemical plants, etc.), commercial buildings (such as office buildings, large shopping malls), transportation vehicles (such as ships, airplanes, heavy trucks) and energy facilities (such as natural gas stations, heating stations), etc. Historical carbon emission data may come from monitoring instruments, sensors installed on the equipment or reports from environmental monitoring stations, such as real-time monitoring sensors installed on carbon emission equipment or chimneys, which continuously monitor the concentration and flow of emissions, and regular emission reports generated by enterprises and environmental protection agencies, including monthly or annual total carbon emissions. These carbon emission data are used to construct a comprehensive carbon emission data set that can be used for analysis and prediction, including processing data of production equipment and energy-using equipment, energy consumption, carbon emission data, and external environmental data. Among them, the processing data of production equipment refers to the operating parameters of production equipment, such as temperature, pressure, speed, etc., which reflect the equipment The working status under different production conditions may also include the operation of industrial boilers under different loads, the output power of generators, the production rhythm in the process, etc.; the processing data of energy-using equipment refers to the data of equipment that consumes energy during production and operation, including the operating time, load level, adjustment status of heaters, coolers, air-conditioning systems, etc., to help understand the energy consumption of the equipment and its impact on carbon emissions; energy consumption refers to specific energy usage data, such as electricity consumption, gas consumption, fuel consumption, etc., which are used to quantify the input and conversion of energy in the production process; carbon emissions data refers to actual carbon emissions monitoring data, usually derived from sensors and monitoring systems installed on emission sources, reflecting the emissions of greenhouse gases such as CO2 and NOx within a specific time; external environment data refers to external conditions related to carbon emissions, such as temperature, humidity, wind speed, air pressure, etc. For example, changes in temperature may cause an increase or decrease in energy consumption of the air-conditioning system, and wind speed may affect the emission diffusion.
[0024] In step S200 , after cleaning the carbon emission dataset, feature extraction is performed on the cleaned carbon emission dataset based on a fusion feature network to establish a first feature set and a second feature set.
[0025] Preferably, data cleaning refers to cleaning and organizing the erroneous, incomplete or redundant information in the original carbon emission data set before carbon emission data analysis to ensure the accuracy and consistency of the data, including identifying and removing or correcting abnormal data points that are obviously inconsistent with the actual situation, such as abnormal emission values during equipment failure; filling data gaps through interpolation, average substitution and other methods to avoid analysis bias caused by missing data; unifying data formats, such as timestamps, units, etc., to ensure that all data have consistent structure and standards. Then, feature extraction is performed on the carbon emission data set after data cleaning based on the fusion feature network, wherein the fusion feature network is used to extract relevant features from multi-source heterogeneous data, and can integrate data from multiple dimensions and generate high-value features. Specifically, the fusion feature network combines different types of information in the carbon emission data (such as equipment operating parameters, energy consumption, external environmental data, etc.) to create a more comprehensive feature space, and extracts and identifies potential nonlinear and interactive relationships between equipment operation and carbon emissions, thereby establishing a first feature set and a second feature set. The first feature set mainly includes relevant features of equipment operation extracted directly from the carbon emission data set, and the second feature set mainly includes relevant features of the external environment extracted from the carbon emission data set.
[0026] Furthermore, step S200 also includes step S210, calling the equipment feature extraction sub-network within the fusion feature network, performing processing data analysis of the carbon emission data set after data cleaning, and obtaining equipment operating parameters, including operating time, start-up and shutdown frequency, load rate, and operating speed; step S220, extracting energy consumption characteristics, operating efficiency characteristics, and fault characteristics of the equipment operating parameters to establish a basic feature set; step S230, using the basic feature set to fuse and associate production rhythm characteristics, and optimizing redundant features based on the dimensionality reduction of principal component features to establish key influencing factors; step S240, constructing a first feature set based on the key influencing factors.
[0027] Preferably, the device feature extraction subnetwork within the fusion feature network is called. The device feature extraction subnetwork is a machine learning unit used to extract the operating characteristics of the device from the data set, process the carbon emission data set that has undergone data cleaning, parse it into data features related to the device operation, and extract key parameters related to the device operating status, including operating time, start-up and shutdown frequency, load rate and operating speed. Among them, operating time refers to the cumulative working time of the device in a specific period; start-up and shutdown frequency refers to the frequency of starting and stopping the device, which is used to analyze its operating stability and the impact of startup on energy consumption; load rate refers to the ratio of the actual working load of the equipment to its rated load, reflecting the utilization efficiency of the equipment; operating speed refers to the operating speed of the equipment, which affects energy consumption and emissions.
[0028] Preferably, energy consumption characteristics, operating efficiency characteristics, and fault characteristics are extracted from the equipment operating parameters. Specifically, the energy consumption of the equipment under different operating conditions is analyzed, such as energy consumption fluctuations under high load, and the efficiency of the equipment under different load rates and operating speeds is measured, such as its performance in different production rhythms. Through data pattern analysis, possible equipment failures or performance degradation are identified, such as frequent start-up and shutdown, which may indicate potential failures. These extracted features are integrated into a basic feature set; then, the data related to the equipment production rhythm in the basic feature set are fused and correlated to analyze how the equipment operates under different production loads and production cycles, for example, the matching of the equipment operating parameters with production demand in different time periods. Specifically, principal component analysis (PCA) is used to reduce unnecessary redundant features in the feature set, extract key features that best explain data changes, and then identify and determine the key influencing factors of carbon emission prediction through the feature set after dimensionality reduction, characterize the association between equipment operating characteristics and carbon emissions, such as the energy consumption and fault trends of equipment under high load; finally, a first feature set is constructed based on the key influencing factors, which includes equipment operation-related features, such as energy consumption patterns, efficiency performance, and fault indications, to ensure that the prediction model can more accurately and effectively analyze carbon emission trends.
[0029] Furthermore, step S230 also includes step S231, performing variance analysis on the fused and associated features to establish a variance calculation result; step S232, performing feature screening through a preset threshold and variance calculation result to establish a preliminary screening feature set; step S233, selecting a predetermined number of principal components for the preliminary screening feature set, and performing cumulative variance analysis on the current principal component to establish a cumulative variance analysis result; step S234, obtaining the component contribution in the current principal component, and making a judgment based on the component contribution and the explanation rate threshold; step S235, if the cumulative variance analysis result fails to reach the cumulative threshold, generating a reconstructed principal component instruction, and performing principal component optimization based on the reconstructed principal component instruction and the judgment result of the component contribution and the explanation rate threshold; step S236, completing the principal component feature dimensionality reduction based on the principal component optimization result.
[0030] Preferably, variance analysis is performed on the fused and associated features to quantify the contribution of each feature to the overall variation of the data, that is, the variance of each feature is calculated, its variation range is recorded, and which features have a larger variation range in the entire feature set are identified, so as to screen out features with higher recognition and information content, thereby generating variance calculation results, and screening features according to a preset variance threshold. Features with variances lower than the threshold are regarded as redundant features and eliminated, and features with variances greater than the threshold are retained to establish a preliminary screening feature set; then a certain number of principal components are selected from the preliminary screening feature set, and cumulative variance analysis is performed on these principal components to calculate the overall data variation they explain, wherein cumulative variance analysis refers to the accumulation of the explanation rate (component contribution percentage) of multiple principal components, which is used to determine whether the current principal component set meets expectations, and then analyze the component contribution of each principal component. , and compare it with the preset explanation rate threshold to determine whether the explanation ability requirements are met. Component contribution refers to the proportion of each principal component in describing the data variation. The explanation rate threshold refers to the set minimum cumulative explanation rate to ensure that the principal component can fully reflect the variation of the original feature set; if the cumulative variance analysis result does not reach the preset cumulative threshold, an instruction to reconstruct the principal component is generated, that is, a different feature combination is reselected for reorganization to improve the cumulative explanation rate of the principal component set so that it meets the set cumulative threshold. Optimization is achieved by adjusting the structure and component contribution of the principal component set to ensure that the principal component reduces redundancy while maintaining as much original data information as possible. The selected principal component is used for dimensionality reduction, and the high-dimensional feature set is converted into a low-dimensional feature set to retain the most important information, thereby improving the performance and stability of the carbon emission prediction model in the prediction task.
[0031] Furthermore, step S200 also includes step S250, calling the external environment feature extraction subnetwork within the fusion feature network, performing external environment data analysis based on the external environment feature extraction subnetwork, and obtaining meteorological data, geographic data, and strategy data; step S260, after feature extraction based on meteorological data, geographic data, and strategy data, performing multi-feature cross analysis, and establishing a second feature set based on the cross analysis results.
[0032] Preferably, an external environment feature extraction subnetwork within the fusion feature network is used, wherein the external environment feature extraction subnetwork is a machine learning module for extracting features from external environmental data, and works in conjunction with the device feature extraction subnetwork to help obtain information related to carbon emissions from external factors. Specifically, the external environment feature extraction subnetwork parses and extracts features from the external environmental data, obtains factors that have a greater impact on carbon emissions, and obtains meteorological data, geographic data, and policy data. Meteorological data includes temperature, humidity, wind speed, rainfall, air pressure, etc., which affect the operating efficiency and carbon emissions of the equipment. For example, rising temperatures may lead to increased energy consumption of cooling equipment; geographic data includes geographic location, altitude, terrain features, etc. Different geographic locations and terrain conditions may affect the diffusion and detection of carbon emissions. For example, the air flow in valleys is different from that in plains; policy data includes emission-related policy restrictions, emission reduction measures, etc. of environmental protection agencies, enterprises, etc., as well as corresponding time points.
[0033] Preferably, various types of external environmental data are analyzed to extract key features that can affect carbon emissions, such as the impact of high temperature and high humidity environments on equipment operation, the impact of thin air in high altitude areas on emission diffusion, and the adjustment method of equipment use when emission restriction policies take effect. The extracted features of different categories are then cross-analyzed to discover the interactions and influences between different features. Specifically, meteorological, geographical and policy features are combined and analyzed using multivariate regression, decision trees, random forests, etc. to find potential correlations. For example, how temperature changes affect carbon emissions under specific geographical locations and policies is analyzed to identify the associations between multiple features, such as the joint impact of high temperature and policy restriction strategies on carbon emissions. Based on the results of the multi-feature cross-analysis, a second feature set is established, which contains the extracted and analyzed key external environmental features to capture the impact of the external environment on carbon emissions, thereby supporting more accurate and comprehensive carbon emission forecasts and enhancing the model's ability to predict carbon emission trends and external environmental impacts.
[0034] Furthermore, step S200 also includes step S270, obtaining real-time data stream, and configuring the initial feature weights of the fusion feature network based on the real-time data stream; step S280, performing feature extraction on the carbon emission data set based on the initial feature weights, and establishing a feature gradient of the initial feature weights based on the feature intensity; step S290, using the feature gradient to adaptively update the initial feature weights, and performing attention optimization of the fusion feature network; step S2100, re-extracting the attention features based on the attention optimization results.
[0035] Preferably, the latest data is continuously acquired from the device and the external environment as a real-time data stream, and this data is input into the fusion feature network to configure its initial feature weights, that is, according to the current performance and importance of different features in the real-time data, an initial weight is assigned to each feature, and the carbon emission data set is analyzed and feature extracted according to the initial feature weights, that is, the initial configured feature weights are used to determine the extraction priority and strength of each feature. Specifically, the feature gradient of the initial feature weight is established based on the feature strength, which means that the gradient generated by the feature weight in the data extraction process is calculated to reflect the importance of the feature and its adjustment direction in the optimization process, thereby generating a feature gradient, wherein the feature gradient reflects the trend of feature weight change, indicating that each feature is important in extraction. The contribution strength to the overall data in the extraction process; based on the feature gradient, the initial feature weights are dynamically adjusted through the adaptive algorithm, so that the feature network pays more attention to the features that have a greater impact on the prediction results. For example, the gradient descent algorithm is used to adjust the weight of each feature according to the feature gradient, improve the model's response to key features, and enhance the accuracy of feature extraction and prediction; finally, based on the optimized feature weights, new features are extracted from the data set, that is, feature extraction is performed in combination with a deep learning network structure (such as a convolutional neural network or an attention mechanism), and the data is re-analyzed based on the optimization results to ensure that the extracted feature set contains the features that contribute most to the prediction effect, so that the fusion feature network can maintain efficient feature extraction and model optimization when the data is constantly updated.
[0036] Step S300: constructing a training sample based on the first feature set, the second feature set and the surplus / deficient records, and establishing a carbon emission prediction model using the training sample.
[0037] Preferably, training samples are created by combining the first feature set, the second feature set and surplus and shortage records, and each training sample contains input features (such as equipment operating parameters and external environmental data) and target outputs (such as surplus and shortage status or carbon emissions). Specifically, the first feature set refers to data features related to equipment operation extracted from carbon emission sources, describing the impact of the equipment's operating mode under different conditions on carbon emissions, including equipment operating parameters, such as boiler temperature, pressure, power, combustion efficiency, etc., energy consumption indicators, such as fuel consumption, electricity usage, etc., production load information, such as the equipment's operating status, startup and shutdown frequency under different production loads; the second feature set refers to data related to external environmental factors. The data characteristics of the equipment, including ambient temperature, humidity, wind speed and pressure, indirectly affect the operation of the equipment and carbon emissions; surplus and shortage records refer to the surplus or shortage records in the historical carbon emission data, which indicate the difference between the actual carbon emissions and the target emissions during the forecast period; then a prediction model is established based on machine learning or deep learning algorithms (such as regression analysis, random forests, and deep neural networks), and the prediction model is trained using the constructed training samples. Specifically, the training samples are input into the model for training to learn the relationship between the equipment operation characteristics and external environmental characteristics and carbon emissions, and to establish a carbon emission prediction model that can predict future carbon emissions and possible surpluses and shortages based on the real-time input equipment operation and environmental data.
[0038] Furthermore, step S300 also includes step S310, establishing a dynamic data update mechanism, which is used to perform update accumulation judgment on the real-time data stream; step S320, if the trust and data volume of the real-time data stream synchronously meet the dynamic data update mechanism, then a feedback instruction is generated; step S330, based on the feedback instruction, the real-time data stream is controlled to update and supplement the carbon emission data set, and the carbon emission prediction model is reconstructed based on the update and supplement results.
[0039] Preferably, a dynamic data update mechanism is constructed to monitor and determine whether real-time data needs to be updated and accumulated to ensure the timeliness and integrity of the data set. Specifically, real-time data streams are continuously received and analyzed to determine whether the data is new enough or has sufficient trust to update the data set. Based on preset conditions (such as data volume and trust), it is decided whether to add real-time data to the existing data set. Among them, the update accumulation judgment is used to evaluate whether the real-time data stream meets the conditions for updating the data set, such as the quantity and quality (trust) of the data. The data volume refers to whether the accumulation of real-time data meets the minimum standard for updating. The trust is used to evaluate the reliability and accuracy of the real-time data. If the data If the preset trust and quantity standards are met at the same time, it is determined that the update conditions are met, and a feedback instruction is generated to control the real-time data flow to update and supplement the carbon emission data set, that is, to merge the real-time data that meets the conditions into the existing carbon emission data set to maintain the timeliness and relevance of the data, so as to ensure that the carbon emission data set can reflect the latest equipment operation and environmental status; after the data set is updated, the new data is used to retrain or optimize the carbon emission prediction model, such as retraining the model according to the updated data set, calibrating the model parameters with the latest data, and verifying the reconstructed model to ensure its carbon emission prediction ability, and further improve the prediction accuracy and reliability of the carbon emission prediction model.
[0040] Step S400: inputting the real-time carbon emission characteristic data as input data into the carbon emission prediction model to generate a surplus or shortage prediction result.
[0041] Preferably, the carbon emission related data collected in real time is input into the carbon emission prediction model to generate a carbon emission surplus or shortage prediction result. Specifically, the real-time carbon emission characteristic data refers to the latest data collected from the equipment and environment at the current time or in a short period of time, such as the temperature, pressure, power, fuel consumption data collected in real time by the sensors installed on the equipment, as well as external environmental data such as temperature and humidity, as well as operating parameters such as the working status, load conditions, and energy usage of the current equipment; these data are dynamically changing and reflect the current status of the equipment operation and environment; the collected characteristic data set (equipment operating parameters and external environmental data) is input into the carbon emission prediction model as input data, and the carbon emission prediction model analyzes and predicts it, and outputs the carbon emission surplus or shortage prediction result, including surplus, carbon emissions are lower than the target value, indicating that the current situation meets or even exceeds the environmental protection requirements, and shortfall, carbon emissions are higher than the target value, indicating that measures need to be taken to reduce emissions to avoid exceeding the standard; using real-time data to predict carbon emission surplus or shortage can enable early warning control, optimize the production process and maintain efficient production under environmental protection requirements.
[0042] Furthermore, step S400 also includes step S401, establishing an abnormality identification center, and the abnormality identification center is configured with an abnormality identification threshold; step S402, before the real-time carbon emission characteristic data is input, triggering the abnormality identification threshold of the real-time carbon emission characteristic data; step S403, if the triggering judgment result is a triggering result, triggering an early warning notification.
[0043] Preferably, an anomaly identification center is set up in the carbon emission management system to identify data anomalies in advance and issue an early warning when necessary. Specifically, the anomaly identification center is a submodule responsible for monitoring and identifying anomalies in carbon emission data to ensure the reliability of the input carbon emission data. The anomaly identification center is configured with an anomaly identification threshold, wherein the anomaly identification threshold is a preset value or range used to determine whether the real-time input data is abnormal. It is usually set based on the historical carbon emission data. The anomaly identification center continuously monitors the real-time input carbon emission characteristic data and checks the data according to the set anomaly identification threshold, that is, the real-time collected carbon emission characteristic data is compared with the preset anomaly identification threshold. If the data is within the normal range, it is allowed to be input into the prediction model; if the data exceeds the set anomaly identification threshold range, it is marked as abnormal, triggering an early warning notification for processing, thereby further ensuring that the real-time input carbon emission data has been verified and the abnormal data has been cleared before entering the prediction model, thereby achieving more efficient and accurate carbon emission prediction and control.
[0044] Step S500: setting a trigger threshold, performing a trigger analysis based on the surplus / shortage forecast result and the trigger threshold, and performing carbon emission control based on the trigger analysis result.
[0045] Preferably, the trigger threshold is a pre-set carbon emission value used to determine whether carbon emissions are exceeding the standard. The trigger threshold can be a quantitative target or limit value for carbon emissions. The trigger threshold is usually set according to the carbon emission regulations of the environmental protection agency, and the surplus and shortage prediction results generated by the model are compared with the set trigger threshold to determine whether countermeasures need to be taken. Specifically, when the prediction results show that carbon emissions will exceed the set threshold, the trigger analysis mechanism will be activated, marking this situation as a potential risk, and generating a trigger analysis result to determine whether intervention or adjustment is needed. If the prediction result is within the threshold, it is judged to be normal; if the prediction result is close to the threshold, an early warning is issued; if the prediction result exceeds the threshold, an automatic control program is started to control carbon emissions, such as reducing the power of production equipment or reducing operating time to reduce carbon emissions, starting carbon capture equipment or switching to low-carbon emission energy sources, thereby achieving real-time carbon emission management and improving the efficiency of carbon emission control.
[0046] In the above, refer to Figure 1 The carbon emission surplus and shortage prediction method based on data mining according to an embodiment of the present invention is described in detail. Figure 2 A carbon emission surplus and deficit prediction system based on data mining according to an embodiment of the present invention is described.
[0047] The data mining-based carbon emissions surplus and shortage prediction system according to an embodiment of the present invention is used to address the technical problem in the prior art of facing complex and changing carbon emissions data, making it difficult to accurately and real-time capture the dynamic changes in carbon emissions, which in turn leads to insufficient accuracy in carbon emissions surplus and shortage predictions and inefficient carbon emissions management and control. This achieves the technical effect of improving the accuracy of carbon emissions surplus and shortage predictions and the efficiency of carbon emissions management and control. The data mining-based carbon emissions surplus and shortage prediction system includes: a carbon emissions dataset establishment module 10, a dataset feature extraction module 20, a carbon emissions prediction model establishment module 30, a surplus and shortage prediction result generation module 40, and a carbon emissions management and control module 50.
[0048] A carbon emission data set establishment module 10 is used to collect historical carbon emission data from multiple carbon emission sources and establish a carbon emission data set, wherein the carbon emission data set includes processing data of production equipment and energy-using equipment, energy consumption, carbon emission data, and external environment data; a data set feature extraction module 20 is used to perform data cleaning on the carbon emission data set, extract features from the cleaned carbon emission data set based on a fusion feature network, and establish a first feature set and a second feature set; a carbon emission prediction model establishment module 30 is used to construct training samples based on the first feature set and the second feature set and surplus and shortage records, and establish a carbon emission prediction model with the training samples; a surplus and shortage prediction result generation module 40 is used to input real-time carbon emission feature data as input data into the carbon emission prediction model to generate surplus and shortage prediction results; a carbon emission control module 50 is used to set a trigger threshold, perform trigger analysis based on the surplus and shortage prediction results and the trigger threshold, and perform carbon emission control based on the trigger analysis results.
[0049] The specific configuration of the dataset feature extraction module 20 will be described in detail below. The dataset feature extraction module 20 further includes: invoking the device feature extraction subnetwork within the fusion feature network to perform data parsing on the cleaned carbon emission dataset to obtain equipment operating parameters, including operating time, start / stop frequency, load rate, and operating speed; extracting energy consumption characteristics, operating efficiency characteristics, and fault characteristics from the equipment operating parameters to establish a basic feature set; utilizing the basic feature set to fuse and correlate production rhythm features, optimizing redundant features based on principal component feature dimensionality reduction, and establishing key influencing factors; and constructing a first feature set based on the key influencing factors.
[0050] The specific configuration of the dataset feature extraction module 20 will be described in detail below. The dataset feature extraction module 20 further includes: performing variance analysis on the fused and associated features to establish a variance calculation result; performing feature screening using a preset threshold and the variance calculation result to establish a preliminary screening feature set; selecting a predetermined number of principal components from the preliminary screening feature set, and performing cumulative variance analysis on the current principal components to establish a cumulative variance analysis result; obtaining the component contribution in the current principal component and performing a discrimination based on the component contribution and the explanatory rate threshold; if the cumulative variance analysis result fails to reach the cumulative threshold, generating a reconstructed principal component instruction, performing principal component optimization based on the reconstructed principal component instruction and the discrimination result of the component contribution and the explanatory rate threshold; and completing principal component feature dimensionality reduction using the principal component optimization result.
[0051] The specific configuration of dataset feature extraction module 20 will be described in detail below. This module further includes: invoking the external environment feature extraction subnetwork within the fusion feature network; parsing external environment data based on the external environment feature extraction subnetwork to obtain meteorological data, geographic data, and policy data; performing multi-feature cross-analysis after feature extraction based on the meteorological data, geographic data, and policy data; and establishing a second feature set based on the cross-analysis results.
[0052] The specific configuration of the dataset feature extraction module 20 will be described in detail below. The dataset feature extraction module 20 further includes: acquiring a real-time data stream, configuring initial feature weights for the fused feature network based on the real-time data stream; extracting features from the carbon emission dataset based on the initial feature weights, and establishing feature gradients for the initial feature weights based on feature strengths; adaptively updating the initial feature weights using the feature gradients, performing attention optimization on the fused feature network; and re-extracting attention features based on the attention optimization results.
[0053] The specific configuration of the surplus / deficit forecasting result generation module 40 will be described in detail below. The surplus / deficit forecasting result generation module 40 further includes: establishing an anomaly identification center configured with an anomaly identification threshold; performing a trigger determination on the anomaly identification threshold on the real-time carbon emission characteristic data before inputting the real-time carbon emission characteristic data; and triggering an early warning notification if the trigger determination result is a trigger result.
[0054] The specific configuration of the carbon emission prediction model establishment module 30 will be described in detail below. The carbon emission prediction model establishment module 30 further includes: establishing a dynamic data update mechanism for determining updates and accumulation of real-time data streams; generating feedback instructions if the trustworthiness and data volume of the real-time data streams simultaneously meet the dynamic data update mechanism; controlling the real-time data streams to update and supplement the carbon emission dataset based on the feedback instructions; and reconstructing the carbon emission prediction model based on the updated and supplemented results.
[0055] The carbon emission surplus and deficit prediction system based on data mining provided by the embodiment of the present invention can execute the carbon emission surplus and deficit prediction method based on data mining provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0056] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules may be used and run on the user terminal and / or server, and the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention.
[0057] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A carbon emission surplus and deficit prediction method based on data mining, characterized by: The method comprises: Collect historical carbon emission data from multiple carbon emission sources to establish a carbon emission data set, wherein the carbon emission data set includes processing data of production equipment and energy-using equipment, energy consumption, carbon emission data, and external environmental data; After cleaning the carbon emission data set, feature extraction is performed on the cleaned carbon emission data set based on the fusion feature network to establish a first feature set and a second feature set; Building a training sample based on the first feature set, the second feature set and the surplus and shortage records, and establishing a carbon emission prediction model using the training sample; Using real-time carbon emission characteristic data as input data into the carbon emission prediction model to generate surplus or deficit prediction results; Set trigger thresholds, conduct trigger analysis based on surplus / deficit forecast results and trigger thresholds, and use the trigger analysis results to control carbon emissions; The feature extraction of the carbon emission data set after data cleaning based on the fusion feature network also includes: Call the device feature extraction sub-network within the fusion feature network to perform data analysis on the cleaned carbon emission dataset and obtain equipment operating parameters, including operating time, start and stop frequency, load rate, and operating speed; Extract energy consumption characteristics, operating efficiency characteristics, and fault characteristics of equipment operating parameters to establish a basic feature set; Use the basic feature set to fuse and correlate production rhythm features, and optimize redundant features based on principal component feature dimensionality reduction to establish key influencing factors; Construct the first feature set based on key influencing factors; The feature extraction of the carbon emission data set after data cleaning based on the fusion feature network also includes: Call the external environment feature extraction sub-network within the fusion feature network, analyze the external environment data based on the external environment feature extraction sub-network, and obtain meteorological data, geographic data, and strategy data; After feature extraction based on meteorological data, geographic data, and strategic data, multi-feature cross analysis is performed, and the second feature set is established based on the cross analysis results.
2. The carbon emission surplus and deficit prediction method based on data mining according to claim 1, characterized in that: The method of using the basic feature set to fuse and associate production rhythm features and optimizing redundant features based on the principal component feature dimensionality reduction further includes: Perform variance analysis on the fused and associated features to establish the variance calculation results; Perform feature screening using preset thresholds and variance calculation results to establish a preliminary screening feature set; Select a predetermined number of principal components for the initial screening feature set, perform cumulative variance analysis on the current principal components, and establish cumulative variance analysis results; Obtain the component contribution in the current principal component and make judgments based on the component contribution and explanation rate threshold; If the cumulative variance analysis result fails to reach the cumulative threshold, a principal component reconstruction instruction is generated, and principal component optimization is performed based on the principal component reconstruction instruction and the component contribution and explanation rate threshold judgment results; The principal component feature dimensionality reduction is completed using the principal component optimization results.
3. The carbon emission surplus and deficit prediction method based on data mining according to claim 1, characterized in that: The feature extraction of the carbon emission data set after data cleaning based on the fusion feature network also includes: Obtain real-time data streams and configure the initial feature weights of the fusion feature network based on the real-time data streams; Extract features from the carbon emission dataset based on the initial feature weights, and establish a feature gradient of the initial feature weights based on the feature strengths; Use feature gradients to adaptively update initial feature weights and perform attention optimization of the fused feature network; Re-extract attention features based on the attention optimization results.
4. The carbon emission surplus and deficit prediction method based on data mining according to claim 1 is characterized in that: The method further comprises: Establishing an anomaly identification center, wherein the anomaly identification center is configured with an anomaly identification threshold; Before the real-time carbon emission characteristic data is input, the abnormal identification threshold triggering judgment is performed on the real-time carbon emission characteristic data; If the trigger judgment result is a trigger result, an early warning notification is triggered.
5. The carbon emission surplus and deficit prediction method based on data mining according to claim 1, characterized in that: The method further comprises: Establishing a dynamic data update mechanism for performing update accumulation determination of real-time data streams; If the trust level and data volume of the real-time data stream synchronously meet the dynamic data update mechanism, a feedback instruction is generated; Based on feedback instructions, the real-time data flow is controlled to update and supplement the carbon emission data set, and the carbon emission prediction model is reconstructed based on the update and supplement results.
6. The carbon emission surplus and deficit prediction system based on data mining is characterized by: The system is used to implement the carbon emission surplus and deficit prediction method based on data mining according to any one of claims 1 to 5, and the system includes: A carbon emission data set establishment module is used to collect historical carbon emission data from multiple carbon emission sources and establish a carbon emission data set, wherein the carbon emission data set includes processing data of production equipment and energy-using equipment, energy consumption, carbon emission data, and external environmental data; A data set feature extraction module is used to perform data cleaning on the carbon emission data set, extract features from the cleaned carbon emission data set based on a fusion feature network, and establish a first feature set and a second feature set; A carbon emission prediction model establishment module is used to construct a training sample based on the first feature set, the second feature set and the surplus and shortage records, and establish a carbon emission prediction model using the training sample; A surplus / deficit prediction result generation module is configured to input real-time carbon emission characteristic data as input data into the carbon emission prediction model to generate a surplus / deficit prediction result; The carbon emission control module is used to set trigger thresholds, perform trigger analysis based on surplus and shortage forecast results and trigger thresholds, and use the trigger analysis results to perform carbon emission control.
Citation Information
Patent Citations
Carbon quota surplus and deficiency prediction method and device, electronic equipment and storage medium
CN115169678A
Comprehensive energy system multi-element load prediction method and system based on deep learning
CN115936218A