Energy carbon emission monitoring method based on multi-sensor data

By constructing a split evaluation function for the decision tree and combining the correlation and uncertainty measurement of sensor data, the problems of adaptability and noise sensitivity of traditional decision trees in carbon emission monitoring are solved, and the accuracy and stability of monitoring are improved.

CN120764863AActive Publication Date: 2025-10-10BEIJING SHAANXI COAL NEW ENERGY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511278851.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-10-10
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Traditional decision trees cannot adapt to the complex, high-dimensional and interrelated data collected by multiple sensors in carbon emission monitoring. They have low classification accuracy and are sensitive to noise, which affects monitoring accuracy.

Method used

By introducing a variety of information measurement methods and combining the correlation weights and uncertainty measures of each sensor data, a split evaluation function of the decision tree is constructed. Taking into account the correlation weights, uncertainty measures and discrimination, a decision tree model is constructed and trained.

Benefits of technology

The classification accuracy and monitoring accuracy of the decision tree are improved, the impact of noise on the model is reduced, the adaptability and stability of the model are enhanced, and more reliable carbon emission data support is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120764863A_ABST
    Figure CN120764863A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to an energy carbon emission monitoring method based on multi-sensor data, and the method comprises the steps: obtaining the carbon emission concentration at the current moment after energy consumption and the carbon emission related data of each sensor; recording the carbon emission related data of any sensor as target data, recording the current moment and a plurality of previous moments as a backtracking period, and determining the correlation weight of the target data in the backtracking period; determining an uncertainty metric value of the target data in the backtracking period; based on the target data, performing splitting by using a decision tree to obtain a plurality of subsets, and determining the distinction degree between the subsets after splitting based on the target data; constructing a split evaluation function of the decision tree; and constructing and training a decision tree model to realize monitoring of energy carbon emission. According to the method, the split evaluation function is constructed by comprehensively considering the multi-sensor data, so that the method can better adapt to the multi-sensor data, and the monitoring result can more accurately reflect the actual carbon emission condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to an energy carbon emission monitoring method based on multi-sensor data. Background Art

[0002] In the energy sector, carbon emission monitoring technology is evolving with the widespread adoption of multi-sensor technology. By deploying various sensors, the system can collect multi-dimensional carbon emission-related data in real time, including temperature, humidity, energy consumption rate, and gas concentration. Decision tree algorithms, due to their intuitive classification and strong interpretability, have become a common data processing tool for processing and analyzing carbon emission-related data. Carbon emission monitoring requires not only accurate identification of the type and extent of carbon emissions but also the ability to rationally interpret and analyze the monitoring results to provide strong support for energy conservation and emission reduction decisions. Therefore, it is necessary to combine multi-dimensional data for subsequent classification and understanding.

[0003] However, traditional decision tree node splitting is based on a single information metric, such as information gain, information gain ratio, or Gini index. This single metric cannot fully exploit the rich information contained in complex, high-dimensional, and interrelated data collected by multiple sensors. For example, when monitoring carbon emissions from different energy production processes, the correlation between different sensor data changes dynamically with changes in production processes and environmental conditions. Traditional splitting methods are unable to adapt to these changes, resulting in low classification accuracy of decision trees. Furthermore, traditional decision trees do not fully consider the uncertainty and noise of multi-sensor data when splitting nodes. In actual monitoring, sensors may be interfered with by various factors, resulting in errors or outliers in the data. Traditional splitting methods are more sensitive to this noise, which can easily cause the generated decision tree model to learn incorrect patterns, thereby affecting the accuracy of carbon emissions monitoring. Summary of the Invention

[0004] In order to solve the problem that in the process of energy carbon emission monitoring, the node splitting of the traditional decision tree is based on a single information measurement method and cannot adapt to the complex, high-dimensional and interrelated data collected by multiple sensors, resulting in low classification accuracy of the decision tree. At the same time, the traditional decision tree is affected by noise, causing the decision tree model to learn incorrect patterns, thereby affecting the accuracy of carbon emission monitoring. The present invention proposes an energy carbon emission monitoring method based on multi-sensor data, which includes the following steps: The carbon emission concentration at the current moment after energy consumption and the carbon emission-related data of each sensor are obtained; the carbon emission-related data of any sensor is recorded as the target data, and the current moment and several previous moments are recorded as the lookback period. The correlation weight of the target data in the lookback period is determined based on the correlation coefficient between the target data and the carbon emission concentration at each moment in the lookback period, and the time difference between the current moment and each moment in the lookback period; the uncertainty measurement value of the target data in the lookback period is determined based on the information entropy and the coefficient of variation of the target data in the lookback period; based on the target data, a decision tree is used to split to obtain several subsets, and the discrimination between the subsets after the target data is split is determined based on the mean carbon emission concentration corresponding to the moment included in each subset and the mean carbon emission concentration corresponding to the moment included in the target data; a split evaluation function of the decision tree is constructed based on the correlation weight, the uncertainty measurement value, the discrimination, and the preset correction coefficient of each sensor; a decision tree model is constructed based on the split evaluation function and trained, and the newly obtained carbon emission-related data of each sensor is input into the trained decision tree model to realize energy carbon emission monitoring.

[0005] The present invention introduces multiple information measurement methods and combines the correlation weights and uncertainty measurements of each sensor data to more comprehensively reflect the characteristics of the data, thereby improving the classification accuracy of the decision tree and enhancing the accuracy of carbon emission monitoring; it can effectively process high-dimensional, complex and interrelated data collected by multiple sensors, overcome the traditional decision tree's reliance on a single information measurement when splitting nodes, and make the model more adaptable; by introducing correlation weights and uncertainty measurements within the lookback period, it can effectively reduce the impact of noise on the decision tree model, reduce the risk of the model learning incorrect patterns, and thus improve the reliability of the monitoring results; the constructed split evaluation function comprehensively considers the correlation weights, uncertainty measurements and discrimination, making the splitting process of the decision tree more scientific and reasonable, and can more effectively divide the data set and improve the generalization ability of the model; by improving the accuracy of carbon emission monitoring, it can provide more reliable data support for enterprises and governments, help formulate more effective carbon emission reduction policies, and promote the realization of sustainable development goals.

[0006] Furthermore, the carbon emission concentration and carbon emission related data of each sensor are the carbon emission concentration and carbon emission related data of each sensor after data cleaning and normalization processing.

[0007] Furthermore, the normalization process adopts Z-score standardization.

[0008] Furthermore, the carbon emission related data includes energy consumption data, temperature data, humidity data and gas concentration data.

[0009] Further, the correlation coefficient is obtained by: the difference between the target data and the carbon emission concentration in the backtracking period containing each time before and after the Pearson correlation coefficient is recorded as the correlation coefficient of the target data and the carbon emission concentration at each time in the backtracking period.

[0010] Further, the correlation weight satisfies: ; In the formula, is the carbon emission related data of the first sensor in the correlation weight of the backtracking period, is the current time, is the first time in the backtracking period, is the number of times in the backtracking period, is the first sensor, is the correlation coefficient of the carbon emission related data and the carbon emission concentration at the first time in the backtracking period, is a preset time decay coefficient, is a proportional normalization function, is a natural exponential function.

[0011] The calculation of the correlation weight of the present application considers the time factor, and by introducing the time decay coefficient and the time difference, the influence degree of the data at different time points on the current monitoring result can be effectively reflected, so that the data weight of the recent time is higher, and the sensitivity of the model to the latest information is enhanced; by combining the correlation coefficients of different sensors and time decay, the information of multi-source data can be better integrated, the capture ability of the model to the carbon emission characteristics is improved, and the monitoring result is more comprehensive and accurate; by weighting the historical data, the data fluctuation caused by accidental factors can be smoothed, the sensitivity of the model to abnormal values is reduced, and the stability and robustness of the model are enhanced.

[0012] Further, the uncertainty measure value satisfies: ; In the formula, is the uncertainty measure value of the carbon emission related data of the first sensor in the backtracking period, is the information entropy of the carbon emission related data of the first sensor in the backtracking period, is the coefficient of variation of the carbon emission related data of the first sensor in the backtracking period, is a preset weight adjustment parameter, is a proportional normalization function.

[0013] The uncertainty measurement value of the present invention combines information entropy and coefficient of variation, which can comprehensively reflect the performance of sensor carbon emission-related data within the retrospective period, comprehensively consider the distribution characteristics and volatility of the data, and thus more accurately measure the uncertainty of the data; by introducing weight adjustment parameters, the relative influence of information entropy and coefficient of variation can be flexibly adjusted according to actual conditions, so that the uncertainty measurement can better adapt to scenarios with different data characteristics, thereby increasing the scope of application of the method; through the application of the proportional normalization function, the influence of data at different scales on the uncertainty assessment can be eliminated, thereby ensuring the comparability of measurement values ​​between different sensors and improving the reliability of monitoring data; by quantifying uncertainty, potential sources of error can be better identified and located in the monitoring results, which helps to improve the overall accuracy and credibility of the monitoring system and provide a basis for formulating more effective carbon emission reduction measures.

[0014] Furthermore, the discrimination satisfies: Where, Based on the The degree of distinction between the subsets after splitting the carbon emission related data of the sensors, Based on the The number of subsets after splitting the carbon emission related data of each sensor, Based on the The carbon emission related data of each sensor is split into The mean value of carbon emission concentration corresponding to the time contained in the subsets, Based on the The carbon emission related data of each sensor contains the average carbon emission concentration corresponding to the moment, Based on the The carbon emission related data of each sensor is split into The number of moments included in the subset.

[0015] The discriminant method of the present invention quantifies the differences between subsets after the sensor-based carbon emission-related data is split, which can effectively reflect the degree of discreteness of the mean carbon emission concentration of each subset and help identify efficient feature splitting points; by calculating the difference between each subset and the overall mean, it can provide a more intuitive understanding of the model's decision-making process, help identify which data features have a greater influence on carbon emission monitoring, and thus improve the transparency of the model.

[0016] Furthermore, the split evaluation function satisfies: Where, is the split evaluation function of the decision tree, is the number of sensors participating in the split evaluation function calculation, For the The correlation weight of the carbon emission related data of each sensor in the retrospective period, For the The uncertainty measurement value of the carbon emission related data of each sensor in the retrospective period, Based on the The degree of distinction between the subsets after splitting the carbon emission related data of the sensors, It is Preset correction factors for each sensor.

[0017] The splitting evaluation function of the decision tree of the present invention comprehensively considers the correlation weight, uncertainty measurement value, discrimination and preset correction coefficient, and can comprehensively evaluate the effectiveness of feature splitting from multiple dimensions, thereby improving the scientific nature of the splitting strategy; through comprehensive evaluation of the data of each sensor, it can more accurately identify the most valuable features for carbon emission monitoring, thereby effectively splitting in the tree structure, and ultimately improving the classification performance and accuracy of the decision tree.

[0018] Furthermore, the training method adopts Fold cross validation.

[0019] The present invention has the following beneficial effects: (1) By comprehensively considering the correlation, uncertainty and other factors of multi-sensor data to construct a split evaluation function, the potential patterns in the data can be more accurately mined. Compared with the traditional single information measurement method, the present invention can better adapt to the complex, high-dimensional and interrelated data collected by multiple sensors, thereby improving the classification accuracy of the decision tree for carbon emission status and making the monitoring results more accurately reflect the actual carbon emission situation.

[0020] (2) The information entropy and coefficient of variation of the data are taken into account when calculating the uncertainty metric, which effectively handles the uncertainty and noise problems of the data. This makes the improved decision tree more resistant to noise and outliers in the sensor data, reduces the learning of erroneous patterns, improves the stability and reliability of the model, and can output more reliable monitoring results even in the presence of interference.

[0021] (3) When constructing the split evaluation function, the preset correction coefficients of each sensor were introduced, and the business logic of carbon emission monitoring was combined, such as the relationship between different energy types and carbon emissions, the impact of production processes on carbon emissions, etc.; this makes the construction of the decision tree more in line with actual monitoring needs, enhances the practicality of the monitoring results, and can provide more effective support for energy conservation and emission reduction decisions.

[0022] (4) By combining multi-sensor data to build a decision tree model and subsequently training the decision tree model, it is possible to clearly explain how the decision tree model classifies and monitors carbon emission data, making it easier for users to understand and analyze the monitoring results and assist in decision making. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a flowchart of the steps of the energy carbon emission monitoring method based on multi-sensor data in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The following is a clear and complete description of the technical solutions in the embodiments of the present invention. The described embodiments are part of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of the present invention.

[0025] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0026] See also Figure 1 , which shows a flowchart of a method for monitoring energy carbon emissions based on multi-sensor data according to an embodiment of the present invention, the method comprising the following steps: S1: Obtain the carbon emission concentration at the current moment after energy consumption and the carbon emission related data of each sensor.

[0027] It should be noted that the carbon emission concentration is collected by a non-dispersive infrared analyzer, and data from multiple sensors (including temperature sensors, humidity sensors, energy consumption sensors, and gas concentration sensors, etc.) are obtained at the same time. For example, the collection interval is 30 minutes per time.

[0028] Specifically, the carbon emission concentration and the carbon emission related data of each sensor are the carbon emission concentration and the carbon emission related data of each sensor after data cleaning and normalization processing.

[0029] Specifically, the normalization process adopts Z-score standardization.

[0030] Specifically, the carbon emission related data includes energy consumption data, temperature data, humidity data and gas concentration data.

[0031] S2: Record the carbon emission related data of any sensor as target data, record the current moment and several previous moments as the lookback period, and determine the relevance weight of the target data in the lookback period.

[0032] It's important to note that the correlation between a sensor's carbon emission data and carbon emission concentration isn't static but rather changes dynamically over time. In carbon emission monitoring scenarios, factors such as production activities at different stages and changes in environmental conditions can affect the correlation between each sensor's carbon emission data and carbon emission concentration. For example, during peak industrial production periods, the correlation between carbon emission data from energy consumption sensors and carbon emission concentrations may increase. During equipment maintenance, carbon emission data from environmental sensors like temperature and humidity may show a more pronounced correlation with carbon emission concentrations. Furthermore, in actual monitoring, recent production processes and environmental conditions are more closely correlated with current carbon emissions. Longer-term data may no longer be representative due to changes in various factors. Therefore, incorporating a time decay factor can dynamically and accurately capture these time-varying data correlation characteristics.

[0033] The correlation weight of the target data in the retrospective period is determined based on the correlation coefficient between the target data and the carbon emission concentration at each moment in the retrospective period, as well as the time difference between the current moment and each moment in the retrospective period.

[0034] Specifically, the correlation coefficient is obtained as follows: The difference between the Pearson correlation coefficients of the target data and the carbon emission concentration before and after each moment in the retrospective period is recorded as the correlation coefficient between the target data and the carbon emission concentration at each moment in the retrospective period.

[0035] Specifically, the relevance weight satisfies: ; Where, For the The correlation weight of the carbon emission related data of each sensor in the retrospective period, For the current moment, The first A moment, is the number of moments in the lookback period, For the The carbon emission related data of each sensor and the carbon emission concentration in the first The correlation coefficient at each moment, is the preset time attenuation coefficient, is the scale normalization function, is the natural exponential function.

[0036] Implementers can use the total length of the specific backtracking period and the time decay coefficient. For example, the total length of the backtracking period is 7 days and the time decay coefficient is 0.1.

[0037] in, The larger the value, the stronger the correlation between the carbon emission-related data of the sensor and the carbon emission concentration at that moment, and the greater the correlation weight, and vice versa; the closer the data is to the current moment, the larger the corresponding time decay factor The larger the value of , the greater the weight that the correlation coefficient between the recent sensor carbon emission-related data and the carbon emission concentration will be given when calculating the correlation weight, which can better reflect the actual correlation between the sensor data and the carbon emission concentration at the current moment. In the carbon emission monitoring scenario, factors such as production processes and environmental conditions may change rapidly over time. Recent data can better reflect the current actual situation. Highlighting the importance of recent data through the time attenuation factor helps to make the calculated correlation weight more in line with the current actual situation.

[0038] For example: For the convenience of calculation, assume that the backtracking period is set to the latest 4 moments (i.e., data from the past 1.5 hours), the number of sensors is 2 (energy consumption sensor and temperature sensor), and the collected data is shown in Table 1. The correlation coefficient As shown in Table 2 (the Pearson correlation coefficient is a prior art and will not be described in detail here), the calculation is as follows: Table 1

[0039] Table 2

[0040] The weighted sum of carbon emission related data from energy consumption sensors is: : , : , : , : , Then the sum is equal to: ; The weighted sum of the carbon emission-related data of the temperature sensor is: : , : , : , : , Then the sum is equal to: ; Then the relevance weight (normalized by proportion) is: , .

[0041] S3: Determine the uncertainty measure of the target data in the lookback period.

[0042] It's important to note that in actual multi-sensor carbon emissions monitoring, data uncertainty comes from a variety of sources. On the one hand, factors such as sensor accuracy limitations and environmental interference can lead to significant data dispersion. On the other hand, the disordered nature of data distribution can also affect its reliability. Single measures of information entropy or dispersion cannot fully capture these complexities. While information entropy can reflect the degree of data disorder—the chaotic state of the data distribution—it cannot directly reflect the magnitude of data fluctuations. The coefficient of variation focuses on describing the dispersion of data relative to the mean but lacks consideration of the overall shape of the data distribution. Therefore, combining these two measures to calculate uncertainty measures can more comprehensively and accurately quantify data uncertainty.

[0043] According to the information entropy and coefficient of variation of the target data in the retrospective period, the uncertainty measurement value of the target data in the retrospective period is determined.

[0044] Specifically, the uncertainty measure satisfies: ; Where, For the The uncertainty measurement value of the carbon emission related data of each sensor in the retrospective period, For the The information entropy of the carbon emission related data of each sensor in the retrospective period, For the The coefficient of variation of the carbon emission related data of each sensor within the retrospective period, is the preset weight adjustment parameter, is the scale normalization function.

[0045] The calculation of information entropy and coefficient of variation is an existing technology and will not be described in detail here.

[0046] Implementers can set weight adjustment parameters according to specific implementation circumstances, for example, 0.6.

[0047] in, To measure the The disorder degree of sensor data is When it increases, it indicates The distribution of carbon emission related data of each sensor is more dispersed and disordered. In carbon emission monitoring, if the value range of a sensor data is very wide and the probability of occurrence of each value is relatively uniform, then its information entropy will be larger, which indicates that the sensor may have more noise or the regularity of the data is poor, then the uncertainty measurement value of the sensor data will be larger. On the contrary, when When it decreases, it means that the distribution of data is more concentrated and orderly, which means that the data of the sensor is relatively more reliable, and the uncertainty measurement value of the sensor data is smaller; Reflects the The dispersion of the carbon emission related data of each sensor relative to its mean value is Increases, indicating that the dispersion of the carbon emission related data of the sensor becomes larger relative to the mean, which means that the stability of the sensor data is poor, and the uncertainty measurement value of the sensor data is larger. On the contrary, when When it decreases, it means that the discreteness of the carbon emission-related data of the sensor becomes smaller relative to the mean, the data is more stable, and the uncertainty measurement value of the sensor data is smaller.

[0048] Example: Energy consumption sensor data According to the above data sequence, the coefficient of variation can be calculated to be 0.3258 and the information entropy is 1.95, so the sum is equal to: ; Temperature sensor data , according to the above data sequence, we can calculate the coefficient of variation to be 0.0117 and the information entropy to be 1.25, so the sum is equal to: ; Then the uncertainty measure (normalized to scale) is: , .

[0049] S4: Based on the target data, use the decision tree to split and obtain several subsets, and determine the discrimination between the subsets after the target data is split.

[0050] It should be noted that when evaluating the effectiveness of a decision tree split, one of the criteria is the degree of difference between the subsets after the split. A high degree of discrimination between subsets indicates that the sensor data is effective in classifying different carbon emission characteristic data, helping the decision tree to more accurately classify and predict carbon emission status. Traditional decision tree splitting methods pay less attention to this subset-based discrimination consideration, but the present invention can more comprehensively evaluate the effectiveness of the split by quantifying the discrimination. When the subsets formed after a sensor data split have large differences in the mean carbon emission data, it indicates that the sensor data is valuable in distinguishing different carbon emission levels or categories.

[0051] The discrimination between the subsets after splitting based on the target data is determined according to the mean of the carbon emission concentration corresponding to the moment included in each subset and the mean of the carbon emission concentration corresponding to the moment included in the target data.

[0052] Specifically, the discrimination satisfies: ; Where, Based on the The degree of distinction between the subsets after splitting the carbon emission related data of the sensors, Based on the The number of subsets after splitting the carbon emission related data of each sensor, Based on the The carbon emission related data of each sensor is split into The mean value of carbon emission concentration corresponding to the time contained in the subsets, Based on the The carbon emission related data of each sensor contains the average carbon emission concentration corresponding to the moment, Based on the The carbon emission related data of each sensor is split into The number of moments included in the subset.

[0053] in, The greater the difference, the more critical the subset is in distinguishing different carbon emission levels or categories, and the greater the degree of differentiation between the subsets after splitting the carbon emission-related data based on the sensor, and vice versa; The larger the value is, the higher the weight of the subset with more data is in the discrimination calculation, which can more effectively reflect the difference between the subset and the whole, thereby affecting the discrimination of different subsets of sensor data.

[0054] For example: when splitting the data of the energy consumption sensor, the split point is: energy consumption < 120, then subset 1 (< 120): energy consumption value is , corresponding to the carbon concentration of , the number is 2, then the mean ; Subset 2 (≥120): Energy consumption value is , corresponding to the carbon concentration of , the number is 2, then the mean ;but ; When splitting the data of the temperature sensor, the split point is: temperature < 25, then subset 1 (< 25): the temperature value is , corresponding to the carbon concentration of , the number is 2, then the mean ; Subset 2 (≥25): Temperature value is , corresponding to the carbon concentration of , the number is 2, then the mean ;but ; It should be noted that since the above data is simplified data under specific circumstances, the subsets are exactly the same after splitting, so the obtained discrimination is also the same, but there will be differences in actual applications.

[0055] S5: Construct split evaluation function for decision tree.

[0056] It's important to note that traditional decision tree node splitting assessments often rely solely on single metrics, such as information gain, information gain ratio, or the Gini index, making it difficult to fully exploit the complex information relationships within multi-sensor data. Furthermore, in the field of carbon emissions monitoring, more accurate and effective splitting decisions require consideration not only of the data's inherent characteristics but also of business logic. The correlation weight reflects the closeness of the correlation between sensor data and carbon emissions data, the uncertainty measure reflects the data's reliability, the discrimination measures the degree of difference between subsets after splitting the sensor data, and the business logic correction coefficient incorporates actual carbon emissions monitoring experience. Combining these factors to construct a splitting assessment function provides a more comprehensive and accurate basis for node splitting, meeting the complex and ever-changing needs of carbon emissions monitoring.

[0057] A split evaluation function of a decision tree is constructed according to the correlation weight, the uncertainty measure, the discrimination, and a preset correction coefficient of each sensor.

[0058] Specifically, the split evaluation function satisfies: ; Where, is the split evaluation function of the decision tree, is the number of sensors participating in the split evaluation function calculation, For the The correlation weight of the carbon emission related data of each sensor in the retrospective period, For the The uncertainty measurement value of the carbon emission related data of each sensor in the retrospective period, Based on the The degree of distinction between the subsets after splitting the carbon emission related data of the sensors, It is Preset correction factors for each sensor.

[0059] Implementers can set the preset correction coefficient of the sensor according to the specific implementation situation. For example, the energy consumption sensor is 0.8 and the temperature sensor is 0.4. Taking the energy consumption sensor as an example, since energy consumption and carbon emissions are closely related, for large chemical companies, changes in energy consumption in the production process have a rapid and significant impact on carbon emissions. Large-scale energy consumption is often the direct cause of increased carbon emissions. Therefore, the preset correction coefficient of the energy consumption sensor is 0.8.

[0060] in, Reflects the The correlation between the carbon emission data of each sensor and the carbon emission concentration, The larger the value of , the more important the sensor data is to carbon emission monitoring, the greater its proportion in the split evaluation function, and the more significant its impact on the result of the split evaluation function. It reflects the The reliability of carbon emission-related data from each sensor, The smaller, The larger it is, the more reliable the sensor data is and the greater its contribution to the result of the split evaluation function is. Reflects the The degree of difference between the subsets after splitting the carbon emission related data of each sensor, The larger the value is, the better it is at distinguishing different carbon emission situations by splitting the sensor data, and the greater its contribution to the result of the split evaluation function. , indicating that after the sensor data is split, the mean values ​​of the carbon emission data in each subset are the same, and the split evaluation result of this node is also 0; It is based on business logic The importance of each sensor is evaluated. When it is larger, it indicates that the sensor data has a higher importance in the carbon emission monitoring business, which will correspondingly increase its impact on the results of the split evaluation function.

[0061] Example: The evaluation results of the energy consumption sensor are: ; The evaluation results of the temperature sensor are: ; Comparing two evaluation results ,Therefore, the decision tree will select the energy consumption sensor and its corresponding split point (energy consumption < 120) to split the current node of the decision tree, and repeat the above operation for each node.

[0062] S6: Build and train a decision tree model to monitor energy carbon emissions.

[0063] A decision tree model is constructed based on the split evaluation function and trained. The newly acquired carbon emission related data of each sensor is input into the trained decision tree model to realize the monitoring of carbon emissions.

[0064] Starting from the root node, the carbon emission-related data of the sensor after data cleaning and normalization is input into the decision tree. For each node, all possible features (i.e., the carbon emission data of each sensor) and the results of the splitting evaluation function corresponding to the splitting point are calculated. The feature and splitting point with the largest result of the splitting evaluation function are selected for splitting to generate a new child node. The above splitting operation is repeated on the newly generated child node until the node data meets the stopping condition, such as the node data belongs to the same category, the amount of node data is less than a specific threshold (for example, 10), the gain after splitting is less than a certain threshold (for example, 0.1), etc.

[0065] Reuse fold cross validation (illustratively, ) Train the decision tree model and divide the data set into mutually non-overlapping subsets, one of which is used as the test set in turn, and the rest The subsets are used as training sets, the decision tree model is trained and its accuracy on the test set is calculated. By adjusting the decision tree parameters (for example, maximum depth, minimum number of samples, etc.), the decision tree model with the highest accuracy in cross-validation is selected as the final model (that is, the trained decision tree model).

[0066] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for monitoring carbon emissions based on multi-sensor data, characterized in that: include: Obtain the carbon emission concentration at the current moment after energy consumption and the carbon emission related data of each sensor; The carbon emission-related data of any sensor is recorded as the target data, and the current moment and several previous moments are recorded as the lookback period. The correlation weight of the target data in the lookback period is determined based on the correlation coefficient between the target data and the carbon emission concentration at each moment in the lookback period, as well as the time difference between the current moment and each moment in the lookback period. Determine the uncertainty measurement value of the target data in the backtracking period based on the information entropy and coefficient of variation of the target data in the backtracking period; Based on the target data, a decision tree is used to split the data into several subsets. The discrimination between the subsets after the target data is split is determined based on the mean carbon emission concentration corresponding to the time period included in each subset and the mean carbon emission concentration corresponding to the time period included in the target data. Constructing a split evaluation function of a decision tree according to the correlation weight, the uncertainty measure, the discrimination, and a preset correction coefficient of each sensor; A decision tree model is constructed based on the split evaluation function and trained. The newly acquired carbon emission related data of each sensor is input into the trained decision tree model to realize the monitoring of carbon emissions.

2. The method for monitoring energy carbon emissions based on multi-sensor data according to claim 1, characterized in that: The carbon emission concentration and the carbon emission related data of each sensor are the carbon emission concentration and the carbon emission related data of each sensor after data cleaning and normalization processing.

3. The method for monitoring energy carbon emissions based on multi-sensor data according to claim 2, characterized in that: The normalization process adopts Z-score standardization.

4. The method for monitoring energy carbon emissions based on multi-sensor data according to claim 2, characterized in that: The carbon emission related data includes energy consumption data, temperature data, humidity data and gas concentration data.

5. The method for monitoring energy and carbon emissions based on multi-sensor data according to claim 1, characterized in that: The correlation coefficient is obtained as follows: The difference between the Pearson correlation coefficients of the target data and the carbon emission concentration before and after each moment in the retrospective period is recorded as the correlation coefficient between the target data and the carbon emission concentration at each moment in the retrospective period.

6. The method for monitoring energy carbon emissions based on multi-sensor data according to claim 1, characterized in that: The relevance weight satisfies: ; Where, For the The correlation weight of the carbon emission related data of each sensor in the retrospective period, For the current moment, The first A moment, is the number of moments in the lookback period, For the The carbon emission related data of each sensor and the carbon emission concentration in the first The correlation coefficient at each moment, is the preset time attenuation coefficient, is the scale normalization function, is the natural exponential function.

7. The method for monitoring energy carbon emissions based on multi-sensor data according to claim 1, characterized in that: The uncertainty measure satisfies: ; Where, For the The uncertainty measurement value of the carbon emission related data of each sensor in the retrospective period, For the The information entropy of the carbon emission related data of each sensor in the retrospective period, For the The coefficient of variation of the carbon emission related data of each sensor within the retrospective period, is the preset weight adjustment parameter, is the scale normalization function.

8. The method for monitoring energy and carbon emissions based on multi-sensor data according to claim 1, characterized in that: The discrimination satisfies: ; Where, Based on the The degree of distinction between the subsets after splitting the carbon emission related data of the sensors, Based on the The number of subsets after splitting the carbon emission related data of each sensor, Based on the The carbon emission related data of each sensor is split into The mean value of carbon emission concentration corresponding to the time contained in the subsets, Based on the The carbon emission related data of each sensor contains the average carbon emission concentration corresponding to the moment, Based on the The carbon emission related data of each sensor is split into The number of moments included in the subset.

9. The method for monitoring energy and carbon emissions based on multi-sensor data according to claim 1, characterized in that: The split evaluation function satisfies: ; Where, is the split evaluation function of the decision tree, is the number of sensors participating in the split evaluation function calculation, For the The correlation weight of the carbon emission related data of each sensor in the retrospective period, For the The uncertainty measurement value of the carbon emission related data of each sensor in the retrospective period, Based on the The degree of distinction between the subsets after splitting the carbon emission related data of the sensors, It is Preset correction factors for each sensor.

10. The method for monitoring energy and carbon emissions based on multi-sensor data according to claim 1, characterized in that: The training method adopts Fold cross validation.

Citation Information

Patent Citations

  • Supply chain carbon data credible management method based on block chain

    CN120069336A

  • Carbon measurement method for unorganized emissions of greenhouse gases in industrial park

    US20250013809A1