Method for analyzing root causes of carbon emission exceeding based on dynamic bayesian network

By constructing a root cause analysis method for carbon emission exceedances based on dynamic Bayesian networks, combining carbon emission knowledge graphs and expert knowledge, and using normalized flow to accelerate probabilistic reasoning, the method solves the problem of low efficiency in traditional methods and achieves accurate identification and analysis of the causes of carbon emission exceedances across multiple time scales.

CN121684067BActive Publication Date: 2026-04-24YUNNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YUNNAN UNIV
Filing Date
2026-02-10
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively quantify the causal contribution of carbon emission exceedances across multiple time scales, and traditional dynamic Bayesian networks are inefficient to construct, failing to achieve accurate root cause analysis.

Method used

We construct a root cause analysis method for carbon emission exceedance based on dynamic Bayesian networks. We extract influencing factors through carbon emission knowledge graphs, combine carbon emission expert knowledge-constrained DAG structure learning, use maximum likelihood estimation to learn conditional probability parameters, and utilize normalized flow to accelerate probabilistic inference, thereby achieving quantification of causal effects across multiple time scales.

Benefits of technology

It enables accurate identification of the causes of excessive carbon emissions, improves the efficiency of root cause analysis, and provides technical support for enterprises' carbon management and energy structure optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121684067B_ABST
    Figure CN121684067B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of electric digital data processing, in particular to a carbon emission exceeding standard root cause analysis method based on a dynamic Bayesian network. The method comprises the following steps: constructing a carbon emission knowledge graph according to multi-modal carbon emission data, and extracting carbon emission influencing factors to construct a time-series carbon emission data set in the carbon emission knowledge graph; constructing a carbon emission rule graph according to carbon emission expert knowledge, and constructing a carbon emission dynamic Bayesian network based on the carbon emission rule graph, the time-series carbon emission data set and a maximum likelihood estimation method; simulating intervention scenarios of different carbon emission influencing factors based on the carbon emission dynamic Bayesian network, calculating a change amount of carbon emission exceeding standard probabilities before and after intervention, and identifying carbon emission exceeding standard reasons according to the change amount of carbon emission exceeding standard probabilities before and after intervention. The method realizes multi-time scale analysis of carbon emission exceeding standard reasons.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital data processing technology, and in particular to a method for root cause analysis of carbon emission exceedances based on dynamic Bayesian networks. Background Technology

[0002] The core task of root cause analysis of carbon emission exceedances is to identify the dependencies between factors such as enterprise production equipment and production processes from carbon emission monitoring data, production logs and maintenance records from multiple time periods and multiple modes, thereby locating the source of carbon emission exceedances and providing a basis for implementing precise energy-saving and carbon-reduction measures.

[0003] Because the impact of carbon emission factors on emissions exceeding limits has multi-timescale characteristics, root cause analysis needs to consider both intra-time slice and inter-time slice dimensions. Intra-time slice analysis focuses on the direct impact of various factors on carbon emission exceeding limits within the current period (e.g., the current month); while inter-time slice analysis traces back to how equipment anomalies or cumulative effects in historical periods (e.g., previous months) gradually led to the current carbon emission exceeding limits. For example, the current month's exceeding limit may be due to a sudden failure of another piece of equipment, or it may be due to the long-term accumulation of historical equipment anomalies.

[0004] Therefore, the key challenge lies in how to effectively quantify the causal contribution of each influencing factor to carbon emission exceedance at different time scales based on multimodal carbon emission data, and to construct a root cause analysis model that can integrate time-series dependence and cross-time period causal mechanisms.

[0005] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main purpose of this application is to provide a method for analyzing the root causes of carbon emission exceedances based on dynamic Bayesian networks, aiming to solve the technical problem of how to construct a root cause analysis model for carbon emission exceedances that integrates time-series dependence and cross-time period causal mechanisms.

[0007] To achieve the above objectives, this application proposes a root cause analysis method for carbon emission exceedance based on dynamic Bayesian networks. This method includes: constructing a carbon emission knowledge graph based on multimodal carbon emission data; extracting carbon emission influencing factors from the carbon emission knowledge graph to construct a time-series carbon emission dataset; the multimodal carbon emission data including carbon emission data from production systems, equipment operation logs, and maintenance records; constructing a carbon emission rule graph based on carbon emission expert knowledge; and analyzing the carbon emission rule graph, the time-series carbon emission dataset, and maximum likelihood estimation. The estimation method involves constructing a dynamic Bayesian network for carbon emissions. This construction includes learning a search space based on a DAG structure constrained by a carbon emission rule graph, and learning conditional probability table parameters based on maximum likelihood estimation. Based on this dynamic Bayesian network, probabilistic inference is used to simulate intervention scenarios for different carbon emission influencing factors. The change in the probability of exceeding carbon emission standards before and after intervention is calculated, and the causes of exceeding carbon emission standards are identified based on this change. The probabilistic inference simulation of the intervention scenarios includes fitting a probability distribution based on normalized flow.

[0008] One or more technical solutions proposed in this application have at least the following technical effects:

[0009] By constructing a carbon emission knowledge graph based on multimodal carbon emission data, including carbon emission data recorded by monitoring equipment in the enterprise's production system, equipment operation logs, and maintenance records, and extracting carbon emission influencing factors to form a time-series carbon emission dataset, a carbon emission knowledge graph is generated. Subsequently, carbon emission expert knowledge is introduced to constrain the search space for DAG structure learning, thereby efficiently learning the DAG structure. Conditional probability parameters are then learned using the maximum likelihood estimation method, completing the construction of a CedBN. Furthermore, to further improve inference efficiency, normalized flow is used to fit complex probability distributions, accelerating probabilistic inference of the CedBN in intervention scenarios. This quantifies the causal effects of influencing factors within and between different time slices on carbon emission exceedances, achieving accurate root cause identification. By combining carbon emission expert knowledge and multimodal carbon emission data, a dynamic Bayesian network for carbon emissions is constructed, and normalized flow is used to accelerate probabilistic inference, enabling multi-timescale analysis of the causes of carbon emission exceedances. This provides technical support for enterprises' carbon management, energy structure optimization, and green and low-carbon transformation. Attached Figure Description

[0010] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0011] Obviously, those skilled in the art can obtain other figures from these figures without any creative effort.

[0012] Figure 1 This is a flowchart illustrating an embodiment of the carbon emission exceedance root cause analysis method based on dynamic Bayesian networks provided in this application.

[0013] Figure 2 A simplified flowchart is provided for an embodiment of the carbon emission exceedance root cause analysis method based on dynamic Bayesian network in this application;

[0014] Figure 3 This is a schematic diagram of a semi-structured equipment operation log in one embodiment of the carbon emission exceedance root cause analysis method based on dynamic Bayesian network in this application;

[0015] Figure 4 This is a schematic diagram of an unstructured maintenance record in one embodiment of the root cause analysis method for carbon emission exceedance based on dynamic Bayesian networks in this application.

[0016] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0017] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0018] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific embodiments.

[0019] It should be noted that the executing entity in this embodiment can be an electronic device with data processing, network communication and program running functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of realizing the above functions.

[0020] Reference Figure 1 In this embodiment, the root cause analysis method for carbon emission exceedance based on dynamic Bayesian networks includes steps S100 to S300:

[0021] Step S100: Construct a carbon emission knowledge graph based on multimodal carbon emission data, and extract carbon emission influencing factors from the carbon emission knowledge graph to construct a time-series carbon emission dataset; the multimodal carbon emission data includes carbon emission data from the production system, equipment operation logs, and maintenance records.

[0022] Multi-source, heterogeneous carbon emission data exhibits multimodal characteristics. Structured carbon emission data, collected periodically by monitoring equipment, records the carbon emission values ​​for each production stage in tabular form. Semi-structured carbon emission data generated by the production system stores key operating parameters such as equipment operating temperature and real-time power consumption in key-value pair format. Unstructured carbon emission data, written by maintenance personnel, describes equipment malfunctions or maintenance events in text form. To learn the Carbon Emission Dynamic Bayesian Network (CedBN), the multimodal carbon emission data is represented as a unified-modality carbon emission knowledge graph. Carbon emission influencing factors are extracted from this knowledge graph to construct a unified-modality time-series carbon emission dataset.

[0023] Step S200: Construct a carbon emission rule graph based on carbon emission expert knowledge, and construct a dynamic Bayesian network for carbon emissions based on the carbon emission rule graph, the time-series carbon emission dataset, and the maximum likelihood estimation method. The construction of the dynamic Bayesian network for carbon emissions includes learning the search space based on the carbon emission rule graph constrained DAG structure and learning the conditional probability parameters based on the maximum likelihood estimation method.

[0024] The construction of CedBN involves two steps: structure learning and parameter learning. To overcome the low computational efficiency of traditional dynamic Bayesian network construction methods, a CedBN construction method constrained by carbon emission expert knowledge is adopted. First, the carbon emission expert knowledge is converted into a carbon emission rule graph. Then, the structure learning of CedBN is constrained by the matching degree between the candidate directed acyclic graph (DAG) structure and the carbon emission rule graph, achieving efficient and interpretable DAG structure learning. Based on this, the maximum likelihood estimation method is used to learn the conditional probability parameters, quantitatively describing the dependencies between various carbon emission influencing factors, thereby constructing CedBN.

[0025] Step S300: Based on the carbon emission dynamic Bayesian network, probabilistic reasoning is used to simulate intervention scenarios for different carbon emission influencing factors, calculate the change in the probability of carbon emission exceeding the standard before and after the intervention, and identify the cause of carbon emission exceeding the standard based on the change in the probability of carbon emission exceeding the standard before and after the intervention. The probabilistic reasoning simulation of the intervention scenario includes fitting the probability distribution based on normalized flow.

[0026] Root cause analysis based on dynamic Bayesian networks for carbon emissions requires using "do-operators" in causal inference to remove edges from the DAG structure, eliminating the influence of other factors on carbon emission exceedances. This simulates the operation of modifying only the values ​​of specific influencing factors in a virtual intervention scenario, thereby measuring the causal effect of specific influencing factors on carbon emission exceedances. However, in carbon emission scenarios, it is necessary to quantify causal effects at different time scales, leading to low efficiency in root cause analysis due to repeated probabilistic inference. To address this, a method based on Normalizing Flow (NF) to accelerate probabilistic inference is adopted. The structural constraints in CedBN are hard-coded into the Normalizing Flow neural network model, constructing a Normalized Flow for Carbon Emission (SNF-CE) that constrains the dependencies between carbon emission influencing factors. This not only follows the conditional independence constraints defined by CedBN but also supports arbitrary modifications to the structural constraints by "do-operators," efficiently completing root cause analysis of carbon emission exceedances within and between time slices.

[0027] In the technical solution provided in this embodiment, a carbon emission knowledge graph is constructed based on multimodal carbon emission data, such as carbon emission data recorded by monitoring equipment in the enterprise's production system, equipment operation logs, and maintenance records. Carbon emission influencing factors are extracted from this graph to form a time-series carbon emission dataset. Subsequently, carbon emission expert knowledge is introduced to constrain the search space for DAG structure learning, thereby efficiently learning the DAG structure. Conditional probability parameters are then learned using the maximum likelihood estimation method to complete the construction of the CedBN. Furthermore, to further improve inference efficiency, normalized flow is used to fit complex probability distributions, accelerating probabilistic inference of the CedBN in intervention scenarios. This quantifies the causal effects of influencing factors within and between different time slices on carbon emission exceedances, achieving accurate root cause identification. By combining carbon emission expert knowledge and multimodal carbon emission data, a dynamic Bayesian network for carbon emissions is constructed, and normalized flow is used to accelerate probabilistic inference, thereby achieving multi-timescale analysis of the causes of carbon emission exceedances. This provides technical support for enterprises' carbon management, energy structure optimization, and green and low-carbon transformation.

[0028] Optionally, step S100 includes steps S111 to S114: Step S111, extracting a set of triples for structured carbon emission data based on the carbon emission data, wherein the set of triples for structured carbon emission data includes a set of triples for the attribution relationship between production links and equipment, a set of triples for the attribute binding between production links and carbon emission amounts, and a set of triples for the discrete semantic tags of carbon emission amounts; Step S112, extracting a set of triples for semi-structured carbon emission data based on the equipment operation logs; Step S113, extracting a set of triples for unstructured carbon emission data based on the maintenance records; Step S114, merging the set of triples for structured carbon emission data, the set of triples for semi-structured carbon emission data, and the set of triples for unstructured carbon emission data to obtain the carbon emission knowledge graph.

[0029] The construction of the carbon emission knowledge graph aims to extract entities such as transformers, wind turbine systems, and carbon emissions from three types of carbon emission data, construct relationships between entities, and form triples such as (Wind Turbine 1, belongs to, Wind Turbine System). This allows multi-source, heterogeneous carbon emission data to be represented as a unified modality carbon emission knowledge graph.

[0030] Specifically, in the triplet set of structured carbon emission data, the structured carbon emission data mainly comes from carbon emission monitoring equipment deployed in various production stages, recording carbon emission values ​​in a timed manner. Each carbon emission record includes a timestamp and the corresponding carbon emission amount. A time slice The carbon emissions of each production stage are recorded as follows: , will the Carbon emissions records for each time slice are represented as follows: including Carbon emission record .

[0031] The set of triples relating production processes and equipment Each production stage and its associated equipment is represented as an entity in a knowledge graph. Based on the actual equipment connection structure, a "belonging" relationship is established between equipment and production stages, and a (…) is constructed. ,belong, The triplet of ). Among them, Indicates any stage of production. This represents any production equipment. Such triples constitute a set. It is used to characterize the hierarchical structure of a physical system.

[0032] The set of triples that link the production process to carbon emissions For each carbon emission observation This is defined as a dynamic attribute of the device at a specific point in time. To preserve numerical information about the recording time and carbon emissions, carbon emission events are introduced. As an intermediate entity, construct triples ( ,occur, ), ( The value is, )and( Occurred in, ).in, Indicates the first The first time slice The equipment monitored in the production process is number one. Each event that records carbon emissions carries both temporal and numerical information. These triples constitute a set. This is used to describe the carbon emissions of this production process.

[0033] For the discretized semantic tag triplet set of carbon emissions To further enhance semantic interpretability, the original carbon emissions were discretized. Specifically, each production stage was discretized. Historical carbon emission sequence Normalization to Intervals, and based on predefined ten-level intervals. Map it to discrete levels , recorded as This discrete level can be interpreted as a "carbon emission intensity level," thus constructing a semantic tag triple ( Carbon emission intensity level This type of triple constitutes a set. This is used to preserve the relative magnitudes of carbon emission values. Finally, the above sets of triplets are combined to obtain... Structured carbon emission data is represented as a set of triples. .

[0034] In the extraction of triples from semi-structured carbon emission data, semi-structured carbon emission data such as equipment operation logs are typically recorded in key-value pair format, including "equipment operating temperature": The system collects key operating parameters for each device, such as "Real-time Power Consumption": "1kW", and device-related information such as "Device Name": "Boiler" and "Recording Time": "August 1, 2025". Each device is represented as an entity in a carbon emission knowledge graph. Feature vectors representing the device's operating status are extracted from the key operating parameters. K-Means clustering is used to identify different operating states of the devices, constructing triples such as (Fan 1, Owned, Normal Operating Status).

[0035] Will A time slice The operation log of each device is represented as follows: , No. The first time slice One device The runtime log is represented as ,Include records Each record This includes key-value pairs of data, including device name, recording time, and all key operating parameters. The first... individual devices The key operating parameters are represented as follows: .

[0036] First, in order to obtain the operating parameter vector that expresses the operating status of the device, for each device... Each time slice Each record and each key parameter To record Centered on, take the front and back together The records form a sliding window. For time slices missing at the window boundaries, zero-padding is used to obtain a window of length [length missing]. The runtime parameter vector: .in, Indicates the first The device in the The first time slice The first record The values ​​of several key parameters are determined. Then, to extract feature vectors representing the device's operating state, a one-dimensional convolutional neural network is used. For each runtime parameter vector Encoding is performed using two convolutional layers and one global average pooling layer to obtain the embedding vector for each key operating parameter of the device. Further, the first The device in the In the time frame The key operating parameters are embedded sequentially and concatenated, then compressed into a unified state representation using MLP to obtain the operating state of the device in the current time slice. 3D feature vector Next, in order to automatically identify different operating states of the device, the feature vectors of the device's operating states in each time slice are... Perform K-Means clustering. The specific steps include steps A1 to A4:

[0037] Step A1, assuming the device exists Various running states, initialization Cluster centers .in, This indicates the normal operating status of the equipment; the rest... Each cluster center represents an abnormal operating state of the equipment. The feature vector of the equipment operating state for each time slice... Calculate the cosine similarity between it and each cluster center: Step A2, given hyperparameters and Define the feature vector Predicted distribution of different equipment operating states With the model on feature vectors The target distribution of the most reliable clustering assignment KL divergence is used to reflect the differences between distributions, and the clustering objective is optimized by minimizing KL divergence. Step A3, iterate through... Repeat steps A1 and A2 to obtain There are 10 clusters, and the number of clusters in each cluster is calculated. The corresponding average profile coefficient: .in, For feature vectors The average intra-cluster distance of the cluster, For feature vectors The average distance to the nearest other cluster. Step A4, take... ,get Clusters .

[0038] Finally, based on the above equipment operating status Each cluster generates semantic labels for each device state. And construct triples ( ,have, The set of triples representing different operating states of each device is denoted as . This clearly expresses the actual meaning of each cluster. For the normal state cluster... Calculate each parameter mean With variance For each cluster of anomalous states Calculate its mean Deviation from standardization ,like The mark deviates from the direction for ,like The mark deviates from the direction for This generates semantic labels for the cluster of abnormal states: Among them, normal cluster labels .

[0039] In the extraction of triplets from unstructured carbon emission data, unstructured data recorded in text format, such as maintenance records written by maintenance personnel, describes events such as equipment malfunctions and incomplete combustion of raw materials. The data extracted from these maintenance records... The maintenance record for each time slice is represented as follows: , will the Within a time slice A maintenance record is represented as Each record It includes the writing time and text content. This invention defines each maintenance record and event as an entity, constructing a triple such as (text content 1, description, wind turbine failure).

[0040] First, all abnormal events recorded in the maintenance logs are defined as events that may lead to excessive carbon emissions. To extract textual descriptions of potential events from the maintenance logs, each maintenance record... text content This is called maintenance text, and TextRank is used to evaluate the maintenance text. Perform keyword extraction steps B1 to B4:

[0041] Step B1: Use the BERT-wwm-ext Chinese pre-trained model to maintain the text. Perform word segmentation to obtain A set of candidate keywords Step B2, based on the above Given candidate keywords, construct an undirected weighted graph. In this context, nodes represent keywords. ,side Keywords and Adjacent elements in the maintenance text do not exceed Each keyword, edge weight This represents the co-occurrence frequency of the two. Step B3: Based on the above undirected weighted graph, calculate the frequency of each keyword node. Importance rating: .in, Representation and keywords A set of connected keyword nodes, initially Step B4: Repeat step B3 until the scores for each keyword remain unchanged, then take the score before the change. The keywords are used to describe the abnormal event in this maintenance record. .

[0042] Then, since anomalous events often lead to abnormal operating states of equipment, clustering results based on the triplet set of semi-structured carbon emission data are used to construct a method for matching anomalous events. Abnormal operating status of equipment training sample set The first The device in the The cluster index to which a time slice belongs is denoted as The embedded information corresponding to the device's operating status is as follows: .like And maintenance records exist. Then the text will be embedded in the pair As a positive sample; if However, if maintenance text exists, it is treated as a negative sample. Positive and negative sample classification is performed for each device's operating status to obtain the training sample set. .

[0043] Subsequently, each maintenance record was... text content Perform semantic vectorization encoding on the text to obtain feature vectors of the text describing the abnormal events in the maintenance records. The specific steps include steps C1 to C3: Step C1, constructing the input sequence for each maintenance text: Step C2, will The feature vector of the maintained text is obtained by inputting it into the BERT-wwm-ext model: Step C3, take The hidden state of the location serves as a semantic representation of this maintenance record. .

[0044] Next, the representation of the text will be maintained. Mapped to a representation of the device's operating state In a semantic space of the same dimension, the projected text embedding is obtained as follows: Based on the training sample set The InfoNCE contrastive loss function is used to narrow the distance between positive sample pairs and widen the distance between negative sample pairs. .in, The cosine similarity function is used. It is the set of all device state embeddings in the current training batch.

[0045] Furthermore, in order to identify abnormal events from maintenance records that are semantically consistent with abnormal equipment states, a set of triples for maintenance records is constructed. For each maintenance record Calculate its projection embedding And calculate its relationship with the time slice. Similarity of embedded state of all devices Take the maximum similarity. and corresponding device index Based on similarity threshold Regarding label validity, construct a set of triples in two cases. Case 1, if and Clustered into clusters Then construct a triple ( ,describe, Case 2, if but If a data point is not clustered into any cluster, then a triplet is constructed. ,describe, ).

[0046] Finally, the above triplet , and The resulting carbon emission knowledge graph is obtained by merging the data, unifying the multi-source heterogeneous data into a semantically aligned carbon emission knowledge graph.

[0047] Further, step S100 above also includes steps S121 to S124: Step S121, converting the carbon emission records in the triplet set of the structured carbon emission data into corresponding carbon emission intensity levels to obtain first structured carbon emission influencing factor data; Step S122, assigning values ​​to the abnormal state clusters to which the carbon emission data records in the triplet set of the semi-structured carbon emission data belong to to obtain second structured carbon emission influencing factor data; Step S123, converting the maintenance records corresponding to candidate events in the triplet set of the unstructured carbon emission data into values ​​of corresponding carbon emission influencing factors to obtain third structured carbon emission influencing factor data; Step S124, constructing the time-series carbon emission dataset based on the first structured carbon emission influencing factor data, the second structured carbon emission influencing factor data, and the third structured carbon emission influencing factor data.

[0048] For the set of triples of structured carbon emission data in the carbon emission knowledge graph, the triples are ( ,occur, ), ( Occurred in, )and( Carbon emission intensity level This involves considering the carbon emissions of each production stage as a factor influencing carbon emissions. A time slice Carbon emission records for each production stage Each carbon emission record Convert to the corresponding carbon emission intensity level The first structured data on carbon emission influencing factors was obtained. .

[0049] For the set of triplets of semi-structured carbon emission data in the carbon emission knowledge graph, the triplets are ( ,have, ),Will The operating status of each device is used as a factor influencing carbon emissions, and carbon emission data is used as a basis for analysis. Each record in the database is determined according to its corresponding record number. Each abnormal state cluster is assigned a value. This will transform the original semi-structured carbon emission data. Converted into second-structured carbon emission impact factor data ,in This indicates the value of each record after the conversion.

[0050] For the set of triplets of unstructured carbon emission data in the carbon emission knowledge graph, the triplets are ( ,describe, ), extract Candidate events are represented as For each candidate event The corresponding maintenance records The data is converted into values ​​for the corresponding carbon emission influencing factors. If a candidate event occurs at the corresponding time point, the binary carbon emission influencing factor value corresponding to that event is recorded as 1; otherwise, it is recorded as 0, thus obtaining the third-structured carbon emission influencing factor data. .

[0051] The above data on factors affecting carbon emissions , and In China, there are a total of One factor influencing carbon emissions. To support the root cause analysis of carbon emission exceedances, the first factor will be... Whether carbon emissions in a given time slice exceed the standard is represented by a binary random variable. A value of 1 indicates that carbon emissions exceed the limit, and a value of 0 indicates that they do not exceed the limit. This allows us to analyze the data on the factors influencing carbon emissions. , Records and The records were merged to obtain the first Time-series carbon emission data for each time slice ,Include Each record contains: .in, This indicates whether the corresponding event has occurred; a value of 1 indicates that it has occurred, and 0 indicates that it has not occurred. This indicates whether carbon emissions exceed the standard; a value of 1 indicates exceeding the standard, and 0 indicates not exceeding the standard.

[0052] The above The record, consisting of various carbon emission influencing factors and whether carbon emissions exceed the standard, is abbreviated as follows: Thus, the first Carbon emission data for each time slice Represented as This yields the overall time-series carbon emission dataset. .

[0053] In one feasible implementation, to overcome the computational inefficiency caused by the exponential growth of the search space with the number of nodes in traditional dynamic Bayesian network learning, carbon emission expert knowledge is first formalized into first-order logic (FOL) rules, and then a carbon emission rule graph is constructed to measure the matching degree between candidate DAG structures and carbon emission expert knowledge. Subsequently, the discrete DAG structure search problem is transformed into a continuous adjacency matrix optimization problem. By designing a differentiable scoring function that integrates data fit, expert knowledge matching degree, and DAG constraints, carbon emission expert knowledge is embedded into the scoring function of structure learning, and gradient descent is used to achieve efficient and interpretable DAG structure learning. Step S200 may include steps S210 to S240: Step S210, representing the carbon emission expert knowledge as a set of first-order logic rules; Step S220, transforming the set of first-order logic rules into the carbon emission rule graph, wherein the node set of the carbon emission rule graph is composed of predicate instances of first-order logic rules, and the edge set of the carbon emission rule graph is constructed from the implication relations of first-order logic rules; Step S230, performing an alignment operation on the candidate DAG structure and the carbon emission rule graph based on a graph neural network to obtain the overall matching degree function between the candidate DAG structure and the carbon emission rule graph; Step S240, constraining the search space of the DAG structure learning according to the overall matching degree function, and performing DAG structure learning operation to obtain the target DAG structure of the carbon emission dynamic Bayesian network.

[0054] First, the carbon emission expert knowledge is represented as A set of FOL rules Its form is .in, This represents the set of devices involved in all factors affecting carbon emissions. This is a time-slice index. For example, "Abnormal rise in boiler temperature leads to a decrease in combustion efficiency, which in turn causes carbon emissions to exceed standards" can be represented as a rule. and : ;

[0055] Then, in order to uniformly represent the knowledge of all carbon emission experts, the above-mentioned FOL rule set is... Convert to carbon emission rule map Among them, the node set From all those appearing in the FOL rules, such as , , Equal predicate instances constitute; edge set Then, based on the implication relationship of the FOL rules, each rule is constructed. Add a directed edge from the premise predicate node to the conclusion predicate node in the carbon emission rule graph. For example, rule and Edges will be generated respectively.

[0056] ,and

[0057] .

[0058] Then, to measure the fit between the candidate DAG structures and carbon emission expert knowledge, a graph neural network (GNN) was used to analyze the candidate DAG structures. With carbon emission rules map Alignment is performed to quantify the semantic matching degree and obtain candidate DAG structures. With carbon emission rules map The overall matching degree function. The current candidate DAG structure can be represented as... ,in, , Let be the adjacency matrix of the DAG structure, representing the dependencies between various carbon emission influencing factors. Subsequently, in order to constrain the structure learning of CedBN using the aforementioned overall matching degree function, the discrete DAG structure search problem is transformed into a differentiable optimization problem of a continuous adjacency matrix. By using the data fit degree, carbon emission expert knowledge matching degree, and a differentiable scoring function of the DAG constraint, the aforementioned matching degree scoring function is embedded into the scoring function of DAG structure learning, thereby achieving efficient and interpretable CedBN structure learning.

[0059] Finally, the DAG structure within each time slice is... Dependencies across time slices By merging, a complete dynamic Bayesian network structure for carbon emissions is formed. This refers to the target DAG structure of the carbon emission dynamic Bayesian network, which lays the structural foundation for subsequent efficient and accurate multi-timescale root cause analysis.

[0060] Further, step S230 includes steps S231 to S234: Step S231, performing an initialization operation on the semantic embedding vectors of the carbon emission rule graph and the candidate DAG structure to obtain embedding vectors; Step S232, based on the message passing of the graph neural network, performing neighbor information aggregation on the embedding vectors of the carbon emission rule graph and the candidate DAG structure respectively to obtain node embedding vectors; Step S233, traversing the directed edges of the candidate DAG structure and calculating the semantic similarity score; Step S234, calculating the overall matching degree function based on the semantic similarity score.

[0061] Specifically, first, initialize the carbon emission rule map. With candidate DAG structure Each node The semantic embedding vector is For data from device operation logs The nodes of carbon emission influencing factors, and their embedding vectors Initialize as described above For data from maintenance records The nodes of carbon emission influencing factors, and their embedding vectors Initialize as described above .

[0062] Then, in order to map the DAG structure and carbon emission expert knowledge to the same embedding space, separate steps were taken in the candidate DAG structure. And carbon emission rules map Above, neighbor information is aggregated through message passing in a GNN to obtain node embedding vectors. For any graph Nodes in Its neighbor set is defined as the union of the neighbors of the incoming edge and the neighbors of the outgoing edge. Then the node Embedded Calculated as follows: .in, This represents vector concatenation. It is a learnable projection matrix.

[0063] Then, for any regular edge Traverse candidate DAG structures All possible directed edges Calculate the semantic similarity score between the two: .in, Let be the cosine similarity.

[0064] Take all middle The maximum value is obtained from the edge of the rule. Match score: .like (i.e., CedBN has no borders), then .

[0065] Finally, based on the above rule edges Matching score calculation of candidate DAG structure With carbon emission rules map Overall matching degree function: .

[0066] Overall matching degree The larger the value, the more the candidate structure conforms to the knowledge of carbon emission experts.

[0067] Further, step S240 includes steps S241 to S246: Step S241, performing an initialization operation on the DAG structure to be learned to obtain a continuous adjacency matrix; Step S242, measuring the fit of the continuous adjacency matrix to the time-series carbon emission dataset based on the Bayesian information criterion score; Step S243, calculating the DAG constraint loss based on the trace of the continuous adjacency matrix and the number of nodes in the carbon emission rule graph; Step S244, calculating a differentiable loss function based on the Bayesian information criterion score and the DAG constraint loss; Step S245, repeatedly performing the steps of calculating the gradient of the differentiable loss function with respect to the continuous adjacency matrix and updating the continuous adjacency matrix until the change of the differentiable loss function is less than a preset threshold, or the change of the differentiable loss function reaches a preset maximum number of iterations, to obtain a target continuous adjacency matrix; Step S246, performing a threshold discretization operation on the target continuous adjacency matrix to obtain the DAG structure in each time slice.

[0068] Specifically, the DAG structure to be learned is first initialized. A continuous adjacency matrix , among which, element Represents a directed edge The confidence level of existence is then determined. Next, the Bayesian Information Criterion (BIC) is used to measure the current structure. Time-series carbon emission datasets Goodness of fit: .in, The total number of valid samples. Let be the number of non-zero elements (i.e., the number of edges) in the adjacency matrix. To make this expression differentiable, we will... The norm is approximately equal to norm .

[0069] Then, in order to ensure a continuous adjacency matrix The corresponding graph is a directed acyclic graph, considering that the adjacency matrix... traces Equal to the number of nodes in the graph If the constraint is satisfied, calculate the DAG constraint loss: .in, For matrix exponents approximated by truncated Taylor expansion Then, based on the BIC score and DAG-constrained loss mentioned above, the differentiable loss function is calculated. The ability to fit carbon emission data to the joint optimization structure, the consistency of expert knowledge, and the legality of the graph structure are all considered. .in, and Let be the hyperparameter for the penalty intensity. Then, calculate the differentiable loss function. right gradient and update ,in The learning rate is used. This step is repeated until the change in the differentiable loss function is less than a preset threshold or the preset maximum number of iterations is reached, yielding the target continuous adjacency matrix. Then, for the target continuous adjacency matrix... Perform threshold discretization and set a threshold. , will satisfy Set the elements to 1 and the rest to 0 to obtain the DAG structure within each time slice. ,in .

[0070] Furthermore, to construct a complete dynamic Bayesian network, it is necessary to learn dependencies across time slices. This involves dividing the time slices... and The nodes are defined as two sets of independent variables. and And introduce a cross-time adjacency matrix ,express Time variable pair The influence of time-varying variables. This matrix is ​​also learned through steps S241 to S246 above, and its loss function form is the same as... Consistent, but the training data is replaced with joint sample pairs from adjacent time slices. The set of edges across time is obtained. .

[0071] Optionally, step S200 further includes step S250:

[0072] Step S250: Based on the maximum likelihood estimation formula, calculate the maximum likelihood estimate of the conditional probability of the carbon emission influencing factors or the node where carbon emissions exceed the standard within each time slice, and calculate the maximum likelihood estimate of the conditional probability of the nodes involved in the edge across time slices. P。

[0073] The maximum likelihood estimation formula is: ;in, Indicating carbon emission data In the middle, node Values And the combination of values ​​of its parent node is The number of samples, This indicates that the combination of values ​​for the parent node is The total number of samples, the set of parent nodes is .

[0074] In this embodiment, based on the constructed DAG structure The maximum likelihood estimation method is used to learn the conditional probability parameters of each node to quantify the dependence strength among various carbon emission influencing factors. For any node... Let all its possible values ​​be . The set of parent nodes is The conditional probability table (CPT) parameters of this node include any combination of values ​​from its parent node. Under the given conditions, the value of this node is... conditional probability For any carbon emission influencing factor or carbon emission exceeding the limit within each time slice. The maximum likelihood estimate of its conditional probability is: .

[0075] For nodes involved in edges spanning time slices, the conditional probability is also calculated using the maximum likelihood estimation formula, but at this time... This indicates the overall carbon emission data. In the middle, node Values And the combination of values ​​of its parent node is The number of samples.

[0076] Finally, the maximum likelihood estimation formula was used to calculate the values ​​of all CPT parameters in CedBN, completing the CPT parameter learning of CedBN and laying the probabilistic reasoning foundation for subsequent efficient and accurate multi-timescale root cause analysis.

[0077] In one feasible implementation, to achieve efficient probabilistic inference for CedBN and support modifications to structural constraints by "do-operations," NF is used as the core of probabilistic inference. NF is a class of invertible generative models that efficiently maps simple probability distributions to complex data distributions through a series of differentiable and invertible transformations. However, standard NF cannot represent structural constraints, causing the model to be unable to support modifications by "do-operations." To address this, a mask is added to each layer of the NF model to satisfy the structural constraints of CedBN, constructing SNF-CE to ensure that the model maintains high fitting ability while adhering to the conditional independence defined by the CedBN structure. Step S300 may include steps S311 to S313: Step S311, based on a greedy binary matrix factorization algorithm, an inter-layer connection mask is generated according to the adjacency matrix of the carbon emission dynamic Bayesian network; Step S312, based on the inter-layer connection mask, an SNF-CE model based on an affine coupling layer is constructed; Step S313, based on the time-series carbon emission dataset, the SNF-CE model is trained.

[0078] First, connections in the neural network that violate the structural constraints of CedBN are masked, based on the fact that CedBN contains... Adjacency matrix of carbon emission influencing factors and nodes exceeding carbon emission standards A series of inter-layer connection masks are generated using a greedy binary matrix factorization algorithm. Suppose that SNF-CE includes There are 1 hidden layer, each with a width of [missing information]. The mask generation process specifically includes steps D1 to D3:

[0079] Step D1: Initialize the first layer mask Repeat the filling of each row as The non-zero rows, i.e., the rows corresponding to the parent variables that have child nodes, are counted until the row is filled. Okay, we get the first layer mask. .

[0080] Step D2, for the first Layer mask According to the output mask of the previous layer Dependency structure with target Based on the logical relationships, construct the input mask for the current layer. Specifically, for Each line, if it is in The effect may indirectly affect If a variable is prohibited from being depended upon, then the corresponding join is set to 0; all other positions are set to 1, thus obtaining the first... layer mask .

[0081] Step D3 finally yields the mask for each layer. .

[0082] Subsequently, to progressively map the simple Gaussian distribution to the probability distribution in CedBN, an SNF-CE model based on an affine coupling layer is constructed. This is done by topological ordering. The first in the arrangement Variables The transformation function of the affine coupling layer is defined as: .in, For standard Gaussian latent variables, For scaling function, The displacement function, the scaling function, and the displacement function are all masked. The implementation uses multiple masked neural network layers. Based on this, and The output depends only on the variable The set of direct parent nodes in CedBN By stacking A single affine coupling layer constitutes a complete reversible transformation. This refers to the SNF-CE model.

[0083] Finally, in order to learn SNF-CE, using Time-series carbon emission dataset The SNF-CE is trained using a loss function that maximizes the likelihood of the carbon emission data to best fit the carbon emission data. .in, It follows a standard Gaussian distribution.

[0084] After training, SNF-CE can accurately fit the joint probability distribution of CedBN, supporting not only marginal probabilities. Conditional probability It enables rapid computation and provides efficient support for subsequent simulation of intervention scenarios using "do-operations," laying a computational foundation for measuring causal effects.

[0085] In another feasible implementation, step S300 further includes steps S321 to S322: Step S321, based on Monte Carlo sampling, estimate the conditional probability of each intervention to obtain an estimated value of the probability of carbon emission exceeding the standard under each intervention scenario; Step S322, based on the estimated value of the probability of carbon emission exceeding the standard under each intervention scenario, calculate the average causal effect corresponding to each intervention scenario.

[0086] In causal inference, the "do-operation" is used to simulate a virtual experimental scenario where a variable is artificially forced to take a specific value, thereby eliminating interference from other influencing factors and accurately assessing the causal effect of that variable on the outcome. In CedBN, Indicates the variable Delete all edges connecting the parent node to it and force it to have a value of 0. Meanwhile, the remaining edges remain unchanged. By modifying the sampling process of SNF-CE to implement the "do-operation", and based on the "do-operation", the conditional probability of carbon emission exceeding the standard under the intervention scenario is calculated, thereby measuring the causal effect of carbon emission influencing factors on carbon emission exceeding the standard, and providing theoretical and computational support for the root cause analysis of carbon emission exceeding the standard.

[0087] First, in order to assess the factors affecting carbon emissions The conditional probability of exceeding carbon emission limits after intervention, i.e., assessing the intervention. Variables related to whether carbon emissions exceed standards The impact was assessed using Monte Carlo sampling to estimate the probability of carbon emissions exceeding limits under various intervention scenarios. The specific steps include steps E1 to E3: Step E1, starting from the standard Gaussian prior... Independent sampling One latent variable sample Step E2, for each According to topological order Use the formulas sequentially Calculate the value of the next variable in the calculation of factors affecting carbon emissions. When ignoring formulas Scale parameters output from the middle With displacement parameters , directly Set the value as the intervention value Thus, the observed samples after intervention were obtained. Step E3: Statistical analysis of the observed samples. China's carbon emission exceeding standards The frequency of occurrence is used as an estimate of the probability of carbon emissions exceeding the standard under the current intervention scenario: .in, Indicates the first The value of the variable indicating whether carbon emissions exceed the standard in the sample obtained from the second sampling. This is an indicator function that takes the value 1 when the condition is true and 0 otherwise.

[0088] Then, for any carbon emission influencing factor Calculate its impact on carbon emissions exceeding standards The average causal effect (ACE). Let... This indicates abnormal conditions such as equipment malfunction or incomplete combustion. This indicates its normal state, then the carbon emission influencing factor Excessive carbon emissions The ACE is: This value reflects the situation when forced to... The net increase in the probability of carbon emission exceeding the standard when switching from a normal state to an abnormal state. The higher the ACE value, the stronger the causal effect of the factor on the exceeding event; if the ACE is close to 0, it means that although the factor may be related to carbon emission exceeding the standard, it has no significant causal effect. By traversing all carbon emission influencing factors and calculating their ACE, it is possible to quantitatively rank and screen potential root causes, providing a reliable basis for subsequent decision-making.

[0089] Optionally, since traditional root cause analysis methods are usually limited to static causal inference within a single time slice, they are difficult to effectively capture the lagged or cumulative effects of historical conditions on current carbon emission exceedance events. The constructed SNF-CE joint modeling integrates carbon emission influencing factors and carbon emission exceedance variables across all time slices, supporting unified probabilistic inference across time slices, thereby realizing root cause analysis within and between time slices. Specifically, step S300 may also include steps S331 to S334: Step S331, based on the average causal effect, perform root cause analysis within the time slice to identify the immediate root cause leading to the current month's carbon emission exceedance; Step S332, based on the average causal effect, perform root cause analysis between time slices to identify historical root causes with lagged or cumulative effects; Step S333, merge the immediate and historical root causes to obtain a set of carbon emission exceedance root causes; Step S334, based on the average causal effect, sort the set of carbon emission exceedance root causes in descending order and output a structured list of carbon emission exceedance root causes.

[0090] In this embodiment, firstly, root cause analysis is performed within the time slice. This is done for the time slice of interest. Iterate through all carbon emission influencing factors within that time frame. and for each factor Calculate the intervention probability under abnormal and normal conditions respectively. and Therefore, the formula can be used. Calculate its carbon emissions exceeding the standard for the current time slice. Average causal effect Set a threshold Carbon emission influencing factors with ACE values ​​greater than a threshold are represented as a set of root causes within a time slice. This process can accurately identify the immediate root causes of exceeding the carbon emission standard for the current month, such as sudden equipment failure or a sharp drop in combustion efficiency.

[0091] Then, root cause analysis was performed across time slices. Each carbon emission influencing factor DAG structure based on CedBN Backtracking from the root cause node Historical carbon emission influencing factors at the starting point And the historical time slices in this tracing path. All factors affecting carbon emissions Similarly, perform the "do-operation" and calculate its impact on carbon emissions exceeding the limit for the current time slice. The ACE value, i.e. Set a threshold The carbon emission influencing factors with ACE values ​​greater than the threshold are represented as a set of root causes across time slices. This process can effectively identify historical root causes with lag or cumulative effects, such as a decrease in energy efficiency this month due to long-term minor wear and tear on equipment, or continued incomplete combustion caused by quality problems in previous batches of raw materials.

[0092] Finally, by merging the root cause sets within and between time slices, we obtain the root cause set for carbon emission exceedances. The data is then sorted in descending order by ACE value, and a structured list of root causes of carbon emission exceedances is output. Each record contains the following four-tuple information: .in, Indicates the nodes of factors affecting carbon emissions; The time slice to which this carbon emission influencing factor belongs In order to distinguish between immediate root causes and historical root causes; This indicates the causal effect value of the carbon emission influencing factor on the current carbon emission exceedance, in order to quantify the intensity of its impact on the current exceedance event; This represents the DAG structure in CedBN. In this context, from the perspective of the nodes affecting carbon emissions Current results exceeding the standard The tracing path It provides explainable causal attribution. This output can be directly used for precise carbon management, providing enterprises with root cause analysis from phenomenon to essence, from the present to the past. For example, it can initiate emergency response to the immediate root cause of high ACE, and optimize equipment maintenance cycles or supply chain management for the historical root causes of high ACE. This enables efficient, accurate, and explainable root cause analysis of carbon emission exceedance events, providing strong technical support for the green and low-carbon transformation of industrial systems.

[0093] For example, to help understand the implementation process of the carbon emission exceedance root cause analysis method based on dynamic Bayesian networks obtained in this embodiment combined with the above embodiments, please refer to... Figures 2 to 4 Specifically, this includes constructing a time-series carbon emission dataset, learning dynamic Bayesian networks for carbon emissions, and analyzing the root causes of carbon emission exceedances. Structured carbon emission records are shown in Table 1, and semi-structured equipment operation logs in JSON format are shown in... Figure 3 As shown, unstructured maintenance logs are as follows: Figure 4 As shown.

[0094] Table 1: Carbon Emissions Record in 2024

[0095]

[0096] The construction of the time-series carbon emission dataset includes the construction of a carbon emission knowledge graph and the construction of a time-series carbon emission dataset based on the carbon emission knowledge graph. In the construction of the carbon emission knowledge graph, entities and relations are extracted from the three types of data mentioned above to construct a set of triples.

[0097] For triple extraction from structured carbon emission data, each month is treated as a time slice, and a set of triples is extracted from the structured data: a set of triples relating production processes to equipment. Based on the actual connection relationship between equipment and production processes, a ternary set is constructed: (fan system, belonging to the raw material production process), (main blower A, belonging to the fan system), (LED lighting group B, belonging to the lighting equipment), and (main transformer 2, belonging to the transformer). The production process and carbon emission attributes are linked to this ternary set. Taking the raw material production process as an example, we introduce... This refers to the event in which carbon emissions are first recorded by monitoring equipment during the raw material production stage in the first time slice. Construct a ternary set (raw material production stage, occurrence, ... ), ( The value is 30) and ( , occurs in, 1). Discretized semantic tag triples set of carbon emissions The carbon emission data recorded in Table 1 were normalized to Interval, and based on the interval range Map it to Based on data from January 1, 2024 For example, for the former The column is normalized to obtain And map it to the range of intervals. For example, 0.5 is located in The corresponding 6th interval is mapped to 6; 0.22 is located in The corresponding third interval is mapped to 3; 0.07 is located in The first interval corresponds to 1; 0.25 is located in... The third interval corresponds to 3, which is mapped to 3. This constructs a semantic label triple ( Carbon emission intensity level, 6), ( Carbon emission intensity level, 3), ( Carbon emission intensity levels, 1) and ( Carbon emission intensity level, 1). Finally, the above three types of triplet sets will be merged. Structured carbon emission data is represented as a set of triples. .

[0098] Triple extraction from semi-structured carbon emission data. Extracting triples from semi-structured carbon emission data. For example... Figure 3 The operation logs of each device shown are represented as follows: 12 months of operation logs for the three devices. Each month is considered a time slice, resulting in a year-long operational log. Each device has 366 operational log entries. Taking the operational log of a lighting device as an example, the corresponding triples are extracted. First, to obtain the operational parameter vector representing the device's operating status, let's take the first total power operational parameter of the lighting device in the first time slice as an example... Centered on the record, take the previous and next records. The total power parameters in each record constitute an operating parameter vector of length 7: Then, in order to extract feature vectors representing the operating state of the lighting equipment, a one-dimensional convolutional neural network is used. Encode the operating parameter vector to obtain the embedding vector of the total power parameter of the lighting equipment. Further embedding three key operating parameters of the lighting equipment , and The features of the current operating status of the lighting equipment are obtained by sequentially concatenating the data and compressing it using MLP. Next, in order to automatically identify different operating states of the lighting equipment, the feature vectors of the operating states of the lighting equipment in each time slice are analyzed. Perform K-Means clustering. The specific steps are as follows: Step A1, assume that the lighting equipment exists. The system is in various operating states, and four cluster centers are initialized. .in, The first cluster represents the normal operating state of the equipment, and the other three cluster centers represent the abnormal operating state of the equipment, and are randomly initialized. The feature vector for each time slice of the lighting equipment... Calculate the cosine similarity between it and each cluster center: Step A2, given hyperparameters and Define the feature vector Predicted distributions belonging to different operating states With the model on feature vectors The target distribution of the most reliable clustering assignment Minimize KL divergence to optimize clustering objective: Step A3: To automatically determine the optimal number of clusters, the silhouette coefficient is used to determine the number of clusters. Traversal Repeat steps A1 and A2 to obtain There are 10 clusters, and the number of clusters in each cluster is calculated. The corresponding average profile coefficient: Step A4, take ,get Clusters .

[0099] Finally, to clarify the actual meaning of each cluster, semantic labels for the four operating states of the lighting equipment were generated based on the four clusters representing the operating states of the lighting equipment. For normal state clusters Its label Calculate each key operating parameter of the lighting equipment mean With variance For each cluster of anomalous states Semantic labels are generated using offsets. This is based on the first anomalous state cluster. For example, calculate its mean. Deviation from standardization ,because Marker off-direction for Generate semantic labels for this abnormal state cluster: Therefore, based on the semantic tags of the four operating states of the lighting equipment, four triples are constructed (lighting equipment, ownership, ...). (Lighting equipment, owned,) (Lighting equipment, owned,) ) and (lighting equipment, owned, The above operation is performed on each device to obtain triples under different operating states. The set of triples under different operating states of all devices is denoted as . .

[0100] Triple extraction from unstructured carbon emission data. Extracting triples from unstructured carbon emission data. For example... Figure 4 The maintenance log shown is written by maintenance personnel, representing the maintenance records for the 12 months from January to December 2024. Each maintenance record includes a timestamp and a natural language description of the event. Figure 4 Taking the maintenance record from August 2024 as an example, ternaries are extracted from unstructured carbon emission data.

[0101] First, to extract textual descriptions of potential events from maintenance records, TextRank is used for keyword extraction to reveal keywords associated with captured anomalous events. Figure 4 The first maintenance record text content For example:

[0102] Step B1, use the vocabulary of the BERT-wwm-ext Chinese pre-trained model. Perform sub-word segmentation , construct containing A set of candidate keywords Step B2: Construct an undirected weighted graph. Among them, nodes ,side Keywords and Adjacent elements in the maintenance text do not exceed Each keyword, edge weight This represents the co-occurrence frequency of the two. Step B3: Calculate the frequency of each keyword node. The rating, in For example: Step B4: Repeat step B3 until the scores for each keyword remain unchanged, then take the score before the change. The keywords are used as the abnormal events described in this maintenance record. .

[0103] Then, a training sample set for semantic alignment is constructed based on the clustering results of the triple set of semi-structured carbon emission data to maintain records. For example, the cluster index of the main blower A in the 8th month (8th time slice) is denoted as The embedded information corresponding to the device's operating status is as follows: .at this time However, maintenance records exist. Embed the text into the pair As negative samples, they are treated as positive samples.

[0104] Subsequently, the maintenance records were reviewed. text content Perform semantic vectorization encoding of the text. The specific steps are as follows: Step C1, construct the input sequence. Step C2, will The feature vector of the maintained text is obtained by inputting it into the BERT-wwm-ext model. Step C3, take The hidden state of the location serves as a semantic representation of this maintenance record. .

[0105] Next, a learnable projection layer is used to maintain the representation of the text. Mapped to a representation of the device's operating state In a semantic space of the same dimension, the projected text embedding is obtained as follows: Training was performed using the InfoNCE contrastive loss function: .

[0106] Furthermore, for each maintenance record Calculate its projection embedding And calculate its relationship with the time slice. Similarity of embedded state of all devices Take the maximum similarity. and corresponding device index Based on similarity threshold Regarding label validity, two cases are handled, and the created set of triples is denoted as... :

[0107] like and Clustered into clusters Then create a triple ( ,describe, ).

[0108] like but If a data point is not clustered into any cluster, a triple is created. ,describe, ).

[0109] Finally, the above triplet , and The resulting carbon emission knowledge graph is obtained by merging the data, unifying the multi-source heterogeneous data into a semantically aligned carbon emission knowledge graph.

[0110] In the construction of a time-series carbon emission dataset based on a carbon emission knowledge graph, carbon emission influencing factors are extracted from the carbon emission knowledge graph, and a time-series carbon emission dataset is constructed.

[0111] For the triplet in the carbon emissions knowledge graph (raw material production stage, occurrence, ...), ), ( , occurs in, 1) and ( Carbon emission intensity level, 3), taking the raw material production process as a factor influencing carbon emissions. Each record in the 12-month carbon emission data in Table 1 is converted to its corresponding carbon emission intensity level, resulting in the structured carbon emission influencing factor data shown in Table 2. .

[0112] Table 2: Normalized and Discretized Carbon Emission Data

[0113]

[0114] For the triplet in the carbon emission knowledge graph ( ,have, The operating status of the three devices was used as a factor influencing carbon emissions, and the carbon emission data was used as a basis for analysis. Each record in the database is determined according to its corresponding record number. Each abnormal state cluster is assigned a value. .by Figure 3 Taking a record from the lighting equipment operation log as an example, it belongs to the first abnormal state cluster, and its record is converted to a value of 1. After converting all records in the equipment operation log, the carbon emission influencing factor data is obtained. Table 3 provides data on factors influencing carbon emissions. A diagram of the first month (the first time slice).

[0115] Table 3: Carbon Emission Data for January 2024

[0116]

[0117] For the triplet in the carbon emission knowledge graph ( ,describe, ), extract Candidate events are represented as For each candidate event The corresponding maintenance records The data is converted into values ​​corresponding to carbon emission influencing factors. If a candidate event occurs at the corresponding time point, the binary carbon emission influencing factor value corresponding to that event is recorded as 1; otherwise, it is recorded as 0, thus obtaining structured carbon emission influencing factor data. Table 4 shows the... Figure 4 The maintenance records shown have been converted to August carbon emission impact factor data. The illustration.

[0118] Table 4: Carbon Emission Data for August 2024

[0119]

[0120] For the carbon emission influencing factors data shown in Tables 2, 3, and 4 , , Treating the carbon emission influencing factors as discrete random variables, there are a total of Each carbon emission factor is considered. The carbon emission data from the three tables are then horizontally concatenated to obtain a time-series carbon emission dataset. .

[0121] In carbon emission dynamics Bayesian network learning, this includes carbon emission knowledge-guided DAG structure learning and conditional probability parameter learning. The carbon emission knowledge-guided DAG structure learning is based on learning the DAG structure of a CedBN (Cyber-Emission Boundary Network). First, the carbon emission expert knowledge is formalized into first-order logic rules. For example, an expert states: "Wearing of wind turbine bearings will lead to abnormal operation of the wind turbine system" and "Abnormal operation of the wind turbine system will directly lead to an increase in the carbon emissions of the wind turbine system." These two pieces of knowledge are represented as FOL (First-Order Logic) rules: ; .

[0122] Next, the above FOL rules are converted into a carbon emission rule diagram. The graph contains nodes. And add the following directed edges based on the rule implication relationship: ; .

[0123] Furthermore, to measure the degree of fit between candidate DAG structures and expert knowledge, a GNN-based approach is used to analyze the candidate structures. And carbon emission rules map Alignment is performed to quantify the semantic matching degree. Taking the DAG structure learning in January 2024 (the first time slice) as an example, The CCP includes 11 factors affecting carbon emissions. And whether carbon emissions exceed the standard variable In this embodiment, Indicates the carbon emissions of lighting equipment, Indicates the carbon emissions of raw materials, Indicates the carbon emissions of transformers, This indicates the carbon emissions of the wind turbine system (the first four contain 10 possible values). Indicates the operating status of the lighting equipment (including 5 possible values). Indicates the transformer's operating status (including 7 possible values). This indicates the operating status of the wind turbine system (including 6 possible values). Indicates whether the fan bearings are worn. Indicates whether the transformer heat sink is dusty. Indicates whether the driver power supply of the lighting equipment is aging. Indicates whether the gas pressure sensor is faulty. This indicates whether carbon emissions exceed the limit (the last five variables can take values ​​of 0 and 1). Then, the candidate structure... With carbon emission rules map The steps for calculating semantic matching degree include: First, initializing the carbon emission rule graph. With candidate structure Each node The semantic embedding vector is For the three carbon emission influencing factor nodes representing the equipment's operating status, their embedding vectors are pre-trained... For nodes representing carbon emission impact factors of maintenance events, their embedding vectors are also pre-trained vectors located in the same semantic space as the equipment operating status. The embedding vectors of the remaining nodes are randomly initialized. Next, in the candidate structure... And carbon emission rules map Above, neighbor information is aggregated through message passing in a GNN, enabling node embedding to simultaneously encode its own semantics and its role in the graph structure. For any graph... Nodes in Its neighbor set is defined as the union of the neighbors of the incoming edge and the neighbors of the outgoing edge. Then the node Embedded Update Next, for any regular edge... Traverse all possible directed edges in the DAG structure. Calculate their semantic similarity score And then take all. middle The maximum value is used as the edge of the rule. Match score Finally, the candidate DAG structures of CedBN are calculated. With carbon emission rules map Overall matching degree .

[0124] In order to use the above The scoring function constrains the structure learning of CedBN, relaxing the discrete DAG search problem into a problem involving continuous adjacency matrices. The optimization problem involves the following steps: First, initialize the adjacency matrix. The value is a random small value. Next, the goodness of fit between the current DAG structure and carbon emission data is calculated. Assuming the current adjacency matrix has 50 edges and the effective sample size is 31 days, the calculation yields... Next, according to the candidate structure With carbon emission rules map The steps for calculating semantic matching degree yield expert knowledge matching degree. This is done for each candidate structure. and rule diagram Perform embedding. Calculate the regular edges. exist Corresponding element The semantic similarity score is 0.92; the regular edge Corresponding element The score is 0.88. Calculate the overall match score. Then, calculate the DAG constraint terms. This indicates that the current structure is close to acyclic. Subsequently, the penalty strength hyperparameter is set... , Calculate the differentiable loss function: Finally, let the learning rate... Gradient descent method The process involves iterative updates. After 200 iterations, the loss function remains essentially unchanged, resulting in the final continuous adjacency matrix. Set a threshold. For continuous adjacency matrices Threshold discretization is performed to obtain an adjacency matrix with values ​​of 0 or 1. That is, the DAG structure within the first time slice. .

[0125] Using the same method, the adjacency matrix across time slices is learned by utilizing joint sample pairs from adjacent time slices (such as data pairs from month 1 and month 2). The set of edges across time is obtained. For example, the model might learn that "the aging of the lighting equipment driver power supply in the first month" will affect "the operating status of the lighting equipment in the second month".

[0126] Finally, the DAG structure within each time slice is... Dependencies across time slices The combined structures form a complete CedBN structure, laying a reliable structural foundation for subsequent parameter learning and root cause analysis.

[0127] In conditional probability parameter learning, the CPT parameters of CedBN are learned. Maximum likelihood estimation is used to learn the CPT parameters of each node in the DAG structure of CedBN. (Taking the example of whether carbon emissions exceed the limit within the first time slice), assuming its parent node is... Then its CPT parameters include , , and .by For example, its maximum likelihood estimate is: .in, Indicating carbon emission data In the middle, node and The number of samples; Indicates the parent node The total number of samples. For edges spanning time slices, the nodes involved, for example, child nodes... The parent node is Then its conditional probability ,at this time This indicates the overall carbon emission data. China satisfies and The number of samples, Indicating carbon emission data China satisfies The number of samples is determined. Finally, the CPT parameters of all nodes are calculated to complete the construction of CedBN.

[0128] The root cause analysis of carbon emission exceedances includes the construction of the SNF-CE, causal effect measurement based on the SNF-CE, and multi-timescale root cause analysis and output. In the construction of the SNF-CE, an SNF-CE for accelerating probabilistic inference is built and trained. There are 12 time slices from January to December 2024, each containing 11 carbon emission influencing factors and 1 variable indicating whether carbon emissions exceed the limit, totaling... There are 144 discrete random variables. The goal of SNF-CE is to jointly model the complete joint probability distribution of these 144 variables. .

[0129] To hard-encode the full-time unfolded graph structure constraints of CedBN into the neural network, inter-layer connection masks are generated based on the adjacency relationships of CedBN. The full graph structure of CedBN is a static DAG containing 144 nodes, and its overall adjacency matrix is... .based on Generate mask sequences for a 5-layer neural network of SNF-CE The specific steps for mask construction include: performing a global topological sort on 144 variables. Ensure that any variable All ancestors (including those in historical time slices) are placed before it; for the ordered... Variables Define its position in the whole graph The set of direct parent nodes in the data is To ensure the mask conforms to the structural constraints of CedBN, subsequent masks are constructed layer by layer. For the first layer mask... Its input dimension is 144 (the number of original variables), and its output dimension is also 144. For the corresponding variables in the output dimension... The Okay, just... Set the column position corresponding to the variable in the middle to 1, and the rest to 0. Ensure that each output neuron of the first layer MLP only receives input from its valid parent node. For subsequent layers... Layer mask ( The input is the 144-dimensional activation vector from the previous layer. To maintain structural consistency, the... Layer mask Adopted and The same sparse pattern, i.e. This design guarantees the composite mapping function of the entire network. The Jacobian matrix strictly maintains the same Consistent sparse lower triangular structure (according to) (After sorting).

[0130] Construct an SNF-CE model based on an affine coupling layer. (The model is then processed according to topological order.) The first in the arrangement Variables The transformation function of this affine coupling layer is defined as follows: Among them, the scaling function and displacement function All are masked by the above. A masked 5-layer MLP implementation. This ensures that the generation of each variable depends only on its set of valid parent nodes in the full-time CedBN, thus strictly adhering to the conditional independence assumption of dynamic Bayesian networks. By stacking five such affine coupling layers, a complete invertible transformation is constructed. Using time-series carbon emission datasets The SNF-CE model is trained. Each training sample is a complete 12-month sequence, i.e., a 144-dimensional vector. The training objective is to maximize the log-likelihood of the data, and the loss function is calculated as follows: .

[0131] In this embodiment, the Adam optimizer is used, and the learning rate is set to... The batch size is 64, and the training run is 300 epochs. After training, the SNF-CE model can accurately fit the full-time joint probability distribution of CedBN. It not only supports the rapid calculation of conditional / marginal probabilities within or across any time slice, but also provides efficient computational support for simulating interventions on variables in any historical or current time slice using "do-operations". This lays a solid computational foundation for measuring causal effects across multiple time scales.

[0132] In the causal effect measurement based on SNF-CE, the causal effect of the carbon emission exceedance event that occurred in April 2024 (i.e., the 4th time slice) is measured using the constructed CedBN and trained SNF-CE. The carbon emission influencing factor is analyzed using "whether the wind turbine bearings are worn". For example, SNF-CE can be used to efficiently calculate the impact of carbon emission exceedance variables. The average causal effect. Specifically, to simulate a scenario where the "wind turbine bearing wear" event is forcibly intervened in, a "do-operation" is performed on the SNF-CE model. Specifically, during the forward sampling process of SNF-CE, when the variable representing the "wind turbine bearing wear" event is generated... When (i.e., whether "wind turbine bearing wear" occurs in time slice 4), perform the following modifications: ignore the scale parameters and displacement parameters calculated from its parent node; directly... The value is forcibly set as the intervention value. To calculate ACE, it needs to be set as an abnormal wear condition. and normal condition without wear To assess the conditional probability of carbon emissions exceeding the limits under the two interventions mentioned above, Monte Carlo sampling was used for estimation. The specific steps are as follows: Step E1, from the standard Gaussian prior distribution... Independent sampling One latent variable sample Step E2, for each latent variable sample Using the modified sampling process (i.e., forced sampling) or Generate the corresponding complete observation samples Step E3: Count the number of carbon emission exceedance events in all samples. The frequency of occurrence, and the probability of intervention. Unbiased estimation: Step E4: Repeat steps E1 through E3, but set the intervention value to... ,get .

[0133] Calculate the causal effect of "whether the wind turbine bearings are worn" on excessive carbon emissions: The results indicate that if the "wind turbine bearing wear" event occurs, the probability of exceeding carbon emission limits increases by a net 55%, demonstrating a strong causal effect of this carbon emission influencing factor. Similarly, all 11 carbon emission influencing factors within the fourth time slice were examined, and their respective ACE values ​​were calculated. For example, the carbon emissions from "lighting equipment" were calculated... The ACE of "Transformer Operating Status" is 0.30. The ACE of the sample was 0.05. In this way, the causal effects of all potential root causes were quantified, providing a precise numerical basis for subsequent root cause ranking and screening.

[0134] In the multi-timescale root cause analysis and output, taking the carbon emission exceedance event that occurred in April 2024 (i.e., the 4th time slice) as an example, based on the constructed CedBN and SNF-CE models, root cause analysis was performed within and between time slices, and a structured root cause list was output. Specifically, the root cause analysis within the time slice was conducted, targeting all carbon emission influencing factors within the 4th time slice. The "do-operation" intervention was performed sequentially, and the SNF-CE model was used to efficiently calculate the carbon emission exceedance variables after the intervention. The conditional probabilities were calculated. The ACE values ​​for 11 carbon emission influencing factors were obtained. Thresholds were set. Factors with ACE values ​​greater than a certain threshold were selected as candidate root causes within the time slice. (Assuming the wind turbine bearings are worn...) The ACE value is Carbon emissions from lighting equipment The ACE value is Then the set of root causes within the time slice And, perform root cause analysis across time slices. For each carbon emission influencing factor, further trace its potential upstream influencing factors in historical time slices to identify historical root causes with lag or cumulative effects. For example, regarding whether the wind turbine bearings are worn in time slice 4... Assuming it has no parent node. Regarding the carbon emissions of lighting equipment in the fourth time slice. Assuming that in the DAG structure of CedBN, its source path is Therefore, its historical parent node is whether the lighting equipment driver power supply in the third time slice is aging. Historical variables need to be calculated. Regarding the current exceeding standards The complete path of the cross-time-slice causal effect is as follows: For evaluation right The causal effect was calculated, and its ACE value was 0.252. A causal effect threshold was set. The root cause set between time slices is obtained. The root causes within and between time slices are merged to obtain the complete root cause set. Sort by ACE value in descending order to generate a structured list of root causes of carbon emission exceedances, as shown in Table 5.

[0135] Table 5: List of root causes of carbon emissions exceeding standards in the fourth time segment

[0136]

[0137] The list clearly reveals that the primary cause of the carbon emission exceedances in April 2024 was the "wind turbine bearing wear" event that occurred during the current period; followed by "excessive carbon emissions from lighting equipment" during the current period; in addition, the historical maintenance issue of "aging power supplies for lighting equipment" that occurred in the previous period (March 2024) also played a significant role in contributing to the current exceedances through the continued impact of equipment condition. This provides enterprises with a comprehensive decision-making basis, from "immediate action" to "historical tracing."

[0138] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the root cause analysis method for carbon emission exceedance based on dynamic Bayesian networks. Any simple modifications based on this technical concept are within the scope of protection of this application. The above descriptions are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

[0139] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for analyzing the root causes of carbon emission exceedances based on dynamic Bayesian networks, characterized in that, The root cause analysis method for carbon emission exceedance based on dynamic Bayesian networks includes: A carbon emission knowledge graph is constructed based on multimodal carbon emission data, and carbon emission influencing factors are extracted from the carbon emission knowledge graph to construct a time-series carbon emission dataset. The multimodal carbon emission data includes carbon emission data from the production system, equipment operation logs, and maintenance records. The steps of constructing a carbon emission rule graph based on carbon emission expert knowledge, and then constructing a dynamic Bayesian network for carbon emissions based on the carbon emission rule graph, the time-series carbon emission dataset, and the maximum likelihood estimation method include: The carbon emission expert knowledge is represented as a set of first-order logic rules; The set of first-order logic rules is transformed into the carbon emission rule graph, the set of nodes of the carbon emission rule graph is composed of predicate instances of first-order logic rules, and the set of edges of the carbon emission rule graph is constructed from the implication relations of first-order logic rules. Based on a graph neural network, an alignment operation is performed on the candidate DAG structure and the carbon emission rule graph to obtain the overall matching degree function between the candidate DAG structure and the carbon emission rule graph; The search space for DAG structure learning is constrained by the overall matching degree function, and DAG structure learning operation is performed to obtain the target DAG structure of the carbon emission dynamic Bayesian network. The construction of the carbon emission dynamic Bayesian network includes learning the search space based on the carbon emission rule graph constrained DAG structure and learning the conditional probability table parameters based on the maximum likelihood estimation method. Based on the aforementioned dynamic Bayesian network of carbon emissions, the SNF-CE model is used to accelerate probabilistic inference. Probabilistic inference simulates intervention scenarios for different carbon emission influencing factors, calculates the change in the probability of carbon emission exceeding standards before and after intervention, and identifies the causes of carbon emission exceeding standards based on the change in the probability of carbon emission exceeding standards before and after intervention. The steps include: Based on the greedy binary matrix factorization algorithm, an inter-layer connection mask is generated according to the adjacency matrix of the carbon emission dynamic Bayesian network. Based on the interlayer connection mask, construct an SNF-CE model based on an affine coupling layer; The SNF-CE model is trained based on the time-series carbon emission dataset. The probabilistic reasoning simulation of the intervention scenario includes fitting the probability distribution based on normalized flow.

2. The method for analyzing the root causes of carbon emission exceedances based on dynamic Bayesian networks as described in claim 1, characterized in that, The steps of constructing a carbon emission knowledge graph based on multimodal carbon emission data, and extracting carbon emission influencing factors from the carbon emission knowledge graph to construct a time-series carbon emission dataset include: Based on the carbon emission data, a set of triples for structured carbon emission data is extracted. The set of triples for structured carbon emission data includes a set of triples for the attribution relationship between production links and equipment, a set of triples for the attribute binding between production links and carbon emission amounts, and a set of triples for the discrete semantic tags of carbon emission amounts. Based on the equipment operation log, extract the set of triplets of semi-structured carbon emission data; Based on the maintenance records, extract the set of triplets of unstructured carbon emission data; The carbon emission knowledge graph is obtained by merging the sets of triples from the structured carbon emission data, the sets of triples from the semi-structured carbon emission data, and the sets of triples from the unstructured carbon emission data.

3. The method for analyzing the root causes of carbon emission exceedances based on dynamic Bayesian networks as described in claim 2, characterized in that, The step of constructing a carbon emission knowledge graph based on multimodal carbon emission data, and extracting carbon emission influencing factors from the carbon emission knowledge graph to construct a time-series carbon emission dataset, further includes: The carbon emission records in the triplet set of the structured carbon emission data are converted into corresponding carbon emission intensity levels to obtain the first structured carbon emission influencing factor data. The abnormal state cluster to which the carbon emission data record in the triplet set of the semi-structured carbon emission data belongs is assigned a value to obtain the second structured carbon emission influencing factor data. The maintenance records corresponding to candidate events in the triplet set of the unstructured carbon emission data are converted into the values ​​of the corresponding carbon emission influencing factors to obtain the third structured carbon emission influencing factor data. The time-series carbon emission dataset is constructed based on the first structured carbon emission influencing factor data, the second structured carbon emission influencing factor data, and the third structured carbon emission influencing factor data.

4. The method for analyzing the root causes of carbon emission exceedances based on dynamic Bayesian networks as described in claim 1, characterized in that, The step of performing an alignment operation between the candidate DAG structure and the carbon emission rule graph based on a graph neural network to obtain the overall matching degree function between the candidate DAG structure and the carbon emission rule graph includes: An initialization operation is performed on the semantic embedding vectors of the carbon emission rule graph and the candidate DAG structure to obtain the embedding vectors; Based on the message passing of the graph neural network, neighbor information is aggregated for the embedding vectors of the carbon emission rule graph and the candidate DAG structure to obtain the node embedding vectors. Traverse the directed edges of the candidate DAG structure and calculate the semantic similarity score; The overall matching degree function is calculated based on the semantic similarity score.

5. The method for analyzing the root causes of carbon emission exceedances based on dynamic Bayesian networks as described in claim 4, characterized in that, The steps of constraining the search space of the DAG structure learning according to the overall matching degree function and performing DAG structure learning operations to obtain the target DAG structure of the carbon emission dynamic Bayesian network include: The DAG structure to be learned is initialized to obtain a continuous adjacency matrix; The goodness of fit of the continuous adjacency matrix to the time-series carbon emission dataset is measured based on the Bayesian information criterion scoring. Calculate the DAG constraint loss based on the trace of the continuous adjacency matrix and the number of nodes in the carbon emission rule graph; Based on the Bayesian information criterion score and the DAG constraint loss, calculate the differentiable loss function; Repeat the steps of calculating the gradient of the differentiable loss function with respect to the continuous adjacency matrix and updating the continuous adjacency matrix until the change of the differentiable loss function is less than a preset threshold, or the change of the differentiable loss function reaches a preset maximum number of iterations, to obtain the target continuous adjacency matrix. A threshold discretization operation is performed on the target continuous adjacency matrix to obtain the DAG structure in each time slice.

6. The method for analyzing the root causes of carbon emission exceedances based on dynamic Bayesian networks as described in claim 1, characterized in that, The step of constructing a carbon emission rule graph based on carbon emission expert knowledge, and then constructing a dynamic Bayesian network for carbon emissions based on the carbon emission rule graph, the time-series carbon emission dataset, and the maximum likelihood estimation method, further includes: Based on the maximum likelihood estimation formula, the maximum likelihood estimate of the conditional probability of carbon emission influencing factors or nodes where carbon emissions exceed the standard is calculated in each time slice, and the maximum likelihood estimate of the conditional probability P of the nodes involved in the edge across time slices is calculated. The maximum likelihood estimation formula is: ; in, Indicating carbon emission data In the middle, node Values And the combination of values ​​of its parent node is The number of samples, This indicates that the combination of values ​​for the parent node is The total number of samples, the set of parent nodes is , Representing the A time slice, and ,and Represents the total number of time slices. ; Representing the One carbon emission influencing factor, and , Represents the total number of factors influencing carbon emissions. .

7. The method for analyzing the root causes of carbon emission exceedances based on dynamic Bayesian networks as described in claim 1, characterized in that, The step of probabilistically simulating intervention scenarios for various carbon emission influencing factors based on the carbon emission dynamic Bayesian network, calculating the change in the probability of carbon emission exceeding the standard before and after the intervention, and identifying the causes of carbon emission exceeding the standard based on the change in the probability of carbon emission exceeding the standard before and after the intervention, further includes: Based on Monte Carlo sampling, the conditional probability of each intervention is estimated to obtain the estimated probability of carbon emission exceeding the standard under each intervention scenario; Based on the estimated probability of carbon emissions exceeding the standard under each intervention scenario, the average causal effect corresponding to each intervention scenario is calculated.

8. The method for analyzing the root causes of carbon emission exceedances based on dynamic Bayesian networks as described in claim 7, characterized in that, The step of using the dynamic Bayesian network for carbon emissions to probabilistically infer and simulate intervention scenarios for different carbon emission influencing factors, calculating the change in the probability of exceeding carbon emission standards before and after intervention, and identifying the causes of exceeding carbon emission standards based on the change in the probability of exceeding carbon emission standards before and after intervention, further includes: Based on the average causal effect, root cause analysis is performed within the time slice to identify the immediate root causes that lead to the carbon emissions exceeding the standard in the current month. Based on the average causal effect, root cause analysis is performed between time slices to identify historical root causes with lag effects or cumulative effects. By combining the immediate and historical root causes, a set of root causes for carbon emission exceedances is obtained; Based on the average causal effect, the set of root causes of carbon emission exceeding the standard is sorted in descending order, and a structured list of root causes of carbon emission exceeding the standard is output.

Citation Information

Patent Citations

  • Causal discovery method combining knowledge graph and automatic variational coding

    CN112949860A

  • Decision model construction method based on big data environment

    CN120258453A