A method for generating an emission source inventory based on upr and process reasoning
By using UPR and process inference methods, emission sources are automatically extracted and supplemented from heterogeneous data, which solves the problem of underestimation of carbon footprint due to manual modeling. This improves the completeness of emission source identification and the accuracy of calculation, and supports carbon footprint accounting and green supply chain management.
Patent Information
- Application Number
- CN202611097904.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-08-25
AI Technical Summary
In existing technologies, the construction of product carbon footprint emission source inventories relies on manual extraction, which leads to high data heterogeneity and easy omission of hidden emission sources, resulting in lower carbon footprint accounting results.
By using a method based on UPR and process reasoning, candidate emission sources are automatically extracted from heterogeneous data, implicit emission sources are supplemented based on process knowledge, a standardized emission source list is generated, and data self-consistency is optimized by using mass conservation constraints and least squares adjustment.
It significantly improves the completeness of emission source identification and the accuracy of carbon footprint accounting, reduces the underreporting rate of hidden emission sources, and provides a high-quality data foundation to support carbon footprint accounting and green supply chain management.
Smart Images

Figure CN122633734A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of carbon emission monitoring technology, and in particular to a method for generating emission source inventories based on UPR and process inference. Background Technology
[0002] With the advancement of global carbon neutrality goals, product carbon footprint accounting has become a crucial aspect of international trade and green supply chain management. Enterprises need to quickly establish carbon emission source inventories compliant with standards such as ISO 14067 for various products in order to calculate carbon footprints and identify emission reduction hotspots.
[0003] In existing technologies, the construction of product carbon footprint emission source inventories mainly relies on two methods: one is the manual modeling method, in which life cycle assessment (LCA) experts read the product's bill of materials, process specifications, energy consumption ledgers, and transportation data, and manually select background processes and enter emission source information item by item based on the process database, generating an inventory for carbon footprint calculation; the other is the rule template-based method, in which fixed emission source templates are pre-set for specific product types (such as plastic parts and metal parts), and activity data is filled from product information through keyword matching.
[0004] However, the original product data is highly heterogeneous. Product bills of materials, process flow diagrams, energy consumption ledgers, etc., often exist in different data formats, and text, tables, and images are mixed together. Manually extracting emission source information is time-consuming and laborious, and it is easy to miss key data. Although existing databases contain a large amount of industrial process input and output data, there is a semantic gap between this data and the prospective modeling requirements of specific products, making it difficult to automatically map them to corresponding items in the product emission source list. Original product data usually only records the main materials and energy consumption, and many auxiliary emission sources that are inevitably present in the process are not included, resulting in the omission of implicit emission sources and a systematic underestimation of carbon footprint accounting results. Summary of the Invention
[0005] This application provides an emission source inventory generation method based on UPR and process reasoning, which solves the problems in the prior art where product carbon footprint modeling relies on manual extraction and omission of hidden emission sources, resulting in low carbon footprint accounting results. It achieves the technical effect of automatically extracting candidate emission sources from heterogeneous data and supplementing hidden emission sources based on process knowledge, significantly improving the completeness of emission source identification and the accuracy of carbon footprint accounting.
[0006] This application provides a method for generating emission source inventories based on UPR and process inference, including: S1: Obtain the original modeling data of the target product and convert it into an intermediate text representation to obtain candidate emission sources; determine the boundary type of the target product according to the life cycle boundary rules; based on the candidate emission sources and boundary types, retrieve process data from the unit process database as a standard exchange flow object; S2: Based on the target product, candidate emission sources, and boundary type, perform process reasoning and complete the implicit emission sources according to the pre-set completion rules; S3: Generate a standardized emission source list for the target product based on candidate emission sources, implicit emission sources, and standard exchange flow objects; S31: Pre-build a product case library, extract the meta-features of the target product and construct a target feature map, calculate the graph edit distance between the target feature map and each case feature map, and generate a similar case candidate set; find the common nodes between the target feature map and each similar case in the similar case candidate set; for non-common nodes, trigger difference rule reasoning to determine the difference emission sources; merge the difference emission sources into the implicit emission sources, combine the candidate emission sources and the standard exchange flow object, and output a comprehensive emission source list for the target product; The original modeling data includes product BOM, process flow diagram, equipment parameter table, energy consumption ledger, transportation documents and product specifications; the meta-features refer to the basic process-related units extracted from the original product modeling data that describe the inherent attributes of the product itself and cannot be further subdivided.
[0007] Furthermore, the method also includes: S4: Selecting mass flow nodes and energy flow nodes in the emission source inventory according to the production process flow, establishing conservation edges, recording the conversion coefficient for each edge and assigning a feasible interval to each node; constructing conservation equations and objective functions based on conservation edges and feasible intervals, solving for the minimum adjustment amount as the adjusted target value; integrating the adjusted target value with the comprehensive emission source inventory fields to form an accurate emission source inventory and calculate the carbon footprint.
[0008] Furthermore, the completion rules are set as follows: determine the product type and identify the corresponding general processing operation flow, identify trigger keywords, and complete the hidden emission sources; pre-set trigger keywords and corresponding hidden emission sources according to different product types.
[0009] Furthermore, constructing the target feature map includes: extracting all identifiable meta-features from the intermediate text representation, treating the meta-features as nodes, and the process relationships between features as edges, to generate a feature topology map of the target product as the target feature map.
[0010] Furthermore, a product case library is pre-built. Each case includes a list of product meta-features and a topological relationship diagram. Each meta-feature independently represents a specific attribute of the product in the corresponding dimension. Multiple meta-features are interconnected through process relationships to jointly constitute a complete description of the product's process knowledge. Meta-features include material type, geometry, connection method, surface treatment, and tolerance level.
[0011] Furthermore, the target feature map is compared with each similar case in the candidate set of similar cases to find the corresponding common subgraph. The feature nodes in all common subgraphs are taken as common nodes and directly inherited using the implicit emission source completion rule. For non-common nodes, the difference rule reasoning is as follows: based on the type of difference features of non-common nodes, identify and determine the extra nodes and missing nodes; search the general rule base to determine the difference emission source. Excess nodes are feature nodes that exist in the target feature map but not in the similar case map; missing nodes are feature nodes that exist in the similar case map but not in the target feature map.
[0012] Furthermore, for multiple nodes, based on the type and connection relationship of the multiple nodes, the implicit emission source completion rules for the corresponding type of node are retrieved from the general rule base, and new emission source entries are automatically generated; for missing nodes, the corresponding rules in similar cases are ignored and no completion is performed; the general rule base is a database containing all meta-features of each product, constructed based on product data literature and industry standards.
[0013] Furthermore, the mass flow node refers to a substance with mass, whose activity data unit can be converted to the SI kilogram; the energy flow node is used to describe energy input and output, and its activity data unit can be converted to the SI joule.
[0014] Furthermore, the boundary type of the target product is automatically determined according to the life cycle boundary rules. If the target product is a final consumer product, a complete boundary is adopted, and the emission sources of the product use stage and the disposal stage need to be supplemented. If the target product is an intermediate industrial product, a production boundary is adopted, and the use stage and the disposal stage are not supplemented.
[0015] One or more technical solutions provided in this application have at least the following technical effects or advantages: By unifying text processing and entity extraction, the efficiency of processing heterogeneous data is significantly improved; through process reasoning and completion rules, the completeness of identifying hidden emission sources is greatly enhanced. By pre-constructing a process knowledge base containing multiple product types, emission sources that are missed in the original data but are inherently present in the process, such as injection molding cooling water, mold release agents, cutting fluids, and wastewater treatment, are automatically completed based on product type and trigger keywords. This significantly reduces the underreporting rate of hidden emission sources and improves the accuracy of carbon footprint accounting.
[0016] This system enhances reasoning capabilities for complex products and hybrid material scenarios through meta-feature graph matching and case transfer reasoning. Breaking away from traditional rule limitations based on fixed product types, it decomposes products into indivisible meta-features such as material, shape, connection method, surface treatment, and tolerance level, and constructs a feature topology graph. It retrieves the most similar cases by computationally calculating graph editing distance, inherits rules from common subgraphs, and handles multiple feature nodes through difference rule reasoning. It can handle hybrid material products such as plastic-coated metal and electronic components, filling gaps in the existing rule base and further improving the generalization ability of emission source identification.
[0017] The physical consistency of activity data is ensured through mass conservation constraint propagation and weighted least squares adjustment. Conservation equations are established by identifying mass flow and energy flow nodes in the inventory, and uncertainty intervals and weights are assigned based on data quality labels. Least squares solutions are used to find the adjusted solution closest to the original observations while satisfying all conservation equations, prioritizing high-precision data and adjusting low-precision estimates. This addresses the imbalance between material input and output, ensuring the reliability and accuracy of carbon footprint calculations. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of a method for generating an emission source inventory based on UPR and process reasoning in an embodiment of the present invention. Detailed Implementation
[0019] To facilitate understanding of the present invention, a more complete description of this application will be given below with reference to the accompanying drawings, which illustrate preferred embodiments of the invention. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to enable a more thorough and complete understanding of the disclosure of the present invention.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to limit the invention; the term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0021] This invention addresses the technical challenges in product carbon footprint modeling, including the difficulty in handling heterogeneous raw data, the easy omission of hidden emission sources, physical inconsistencies in activity data, and inconsistent inventory structures. It proposes an automatic emission source inventory generation method based on a unit process database and process reasoning. First, heterogeneous data from multiple sources, such as product bills of materials, process flow diagrams, and energy consumption ledgers, are uniformly converted into text representations. Natural language processing (NLP) techniques are used to extract structured information of candidate emission sources. By constructing a meta-feature topology graph, the graph edit distance is calculated to retrieve the most similar historical cases. Emission source completion rules are inherited from common subgraphs, and difference rule reasoning is triggered for nodes with multiple features, thereby completing hidden emission sources that are inherently present in the process but not documented. This method is particularly suitable for mixed materials and complex process scenarios. Based on this, mass flow and energy flow nodes are identified, and a mass conservation constraint equation is established. Uncertain intervals and weights are assigned according to data quality labels, and weighted least squares or interval propagation algorithms are used to solve for the optimal adjustment value that satisfies the conservation equation. This ensures consistency between material input and output and prioritizes the protection of high-precision raw data. The final output is a standardized emission source inventory that includes lifecycle stages, activity data, technical specifications, transportation information, allocation tags, data quality labels, and complete knowledge logic, which can be directly connected to a carbon footprint factor library to calculate carbon emissions. This invention automates the entire emission source inventory generation process, significantly improving the completeness of hidden emission source identification, the physical consistency of activity data, and the audit traceability of results, providing a high-quality data foundation for product carbon footprint accounting, supply chain carbon management, and green product design.
[0022] Example 1: As Figure 1 As shown, an emission source inventory generation method based on UPR and process inference is described, the method comprising: S1: Obtain the original modeling data of the target product and convert it into an intermediate text representation to obtain candidate emission sources; determine the boundary type of the target product according to the life cycle boundary rules; and retrieve process data from the unit process database as a standard exchange flow object based on the candidate emission sources and boundary types.
[0023] S2: Based on the target product, candidate emission sources, and boundary type, perform process reasoning and, according to pre-set completion rules, complete the implicit emission sources that were missed by the candidate emission sources extracted from the intermediate text representation.
[0024] S3: Generate a standardized emission source list for the target product based on candidate emission sources, implicit emission sources, and standard exchange flow objects.
[0025] The original modeling data includes the product BOM (Bill of Materials), process flow diagram, equipment parameter table, energy consumption ledger, transportation documents, and product specifications. At least three types of original modeling data for the target product are obtained, including the product BOM, process flow diagram, equipment parameter table, energy consumption ledger, transportation documents, and product specifications. These original modeling data are then converted into a unified text representation. All converted text content is merged into an intermediate text representation in the order it is read, preserving paragraph boundaries and table row and column relationships. A BOM (Bill of Materials) is a computer-readable product structure data file that describes all raw materials, components, their quantities, and hierarchical relationships required for product production. A UPR (Unit Process Record) is a standardized data unit in the life cycle assessment database that records the input and output information of a single industrial process.
[0026] Based on natural language processing technology, entity and relation extraction is performed on the intermediate text representation to obtain structured data of candidate emission sources. Specific extracted fields include: target product name, life cycle stage, emission category, emission source name, standardized label, activity data value, activity data unit, technical specifications, process step description, whether transportation is involved, transportation mode, transportation distance, fuel type, carbon emission result value, whether it is allocated data, data quality label, and knowledge logic fields. Only explicitly stated content in the text is extracted. For data missing in the text but necessary for modeling, it is supplemented through subsequent process reasoning steps. Specifically, when the same text line or table row describes both material consumption and corresponding transportation information, the line is split into two independent candidate entities: one is a material entity, recording the material name, consumption amount, and unit; the other is a transportation service entity, recording the transportation mode, transportation distance, and vehicle / ship type. The splitting is based on a pre-defined keyword list, including transportation, transportation distance, vehicle, ship, train, and logistics.
[0027] The boundary type of the target product is automatically determined based on the life cycle boundary rules. If the target product is a final consumer product (such as home appliances, clothing, or automobiles), a complete boundary is used, and the emission sources of the product's use and disposal stages need to be completed. If the target product is an intermediate industrial product (such as steel, chemical raw materials, or electronic components), a production boundary is used, and the use and disposal stages are not completed.
[0028] For the extracted transportation information, further supplement the implicit emissions in the transportation link: when the transportation description involves multimodal transport, it is automatically broken down into multiple transportation segments, and each transportation segment records the transportation mode, distance and default fuel type.
[0029] Based on candidate emission sources and boundary types, process data is retrieved from at least one unit process database. The unit process database is a specialized database that stores a large amount of industrial process input and output data. Different products correspond to different databases, and a general database format can be used for storage. Carbon emission factors and technical specifications are provided for each emission source per unit activity, so that activity data (such as 300kWh) can be converted into carbon emissions.
[0030] Each record in the parsing unit process data is used to extract the process name, geographical region, functional unit, input stream list, output stream list, environmental emission stream list, annotation information, classification information, and data time information.
[0031] All unit process data from all sources are uniformly converted into standard exchange stream objects. The parsing module reads the raw data from each database, and after field mapping, unit unification, and semantic normalization, standard exchange stream objects are generated. Their function is to unify and abstract heterogeneous data from different databases into an intermediate format. Each standard exchange stream object contains at least the following fields: input / output direction, emission category, emission source name, activity data reference value, unit, lifecycle stage, comments, geographic region, and data quality indicators.
[0032] Perform a cleaning operation on all generated standard exchange stream objects. After cleaning, sort all exchange streams according to their life cycle stage and emission category.
[0033] The process reasoning involves using the obtained target product name, product type, material type, and process keywords identified from the intermediate text representation. It then calls upon the process knowledge base to perform reasoning, and according to pre-set completion rules, completes the implicit emission sources omitted from the candidate emission sources extracted from the intermediate text representation. These implicit emission sources are those not explicitly stated in the intermediate text representation but inherently present in the process (e.g., cooling water power consumption).
[0034] The completion rules are set as follows: determine the product type and identify the corresponding general processing operation flow, identify trigger keywords, and complete the hidden emission sources. Trigger keywords and corresponding hidden emission sources are pre-set according to different product types. For example: if the product type is plastic or rubber products, and the trigger keyword is that the material contains "PP / PE / ABS / rubber" or the process contains "injection molding / vulcanization", then the following hidden emission sources are completed: power consumption of injection molding or vulcanization process, power consumption of mold cooling water circulation pump, power consumption of cooling tower fan, mold release agent consumption, and power consumption of waste crusher; among which, cooling water power consumption is estimated at 10% of the rated power of the injection molding machine. If the product type is metal processing parts, and the trigger keyword is that the process contains "turning / milling / drilling / grinding / stamping", then the following hidden emission sources are completed: power consumption of cutting or stamping equipment, cutting fluid consumption, and waste metal shavings generation and treatment; among which, cutting fluid consumption is estimated at two liters of cutting fluid per ton of processed parts. If the product type is electronic assembly, please complete the following hidden emission sources: power consumption of surface mount equipment, power consumption of reflow soldering, solder paste consumption, and cleaning agent consumption; among which, reflow soldering power consumption is estimated at 0.05 kWh per square centimeter of circuit board. If the product type is textile printing and dyeing products, please complete the following hidden emission sources: steam consumption of pretreatment equipment, steam consumption of dyeing machine, steam consumption of setting machine, consumption of dyes and auxiliaries, and power consumption of wastewater treatment.
[0035] Based on the extracted candidate and latent emission sources, and combined with standard exchange flow objects from the unit process database, a standardized emission source list for the target product is generated. For each emission source, a matching and filling operation is performed: First, semantic matching is performed between the emission source names and the emission source names in the standard exchange flow object. If a match is successful, information such as technical specifications, geographical region, and data quality indicators are inherited from the standard exchange flow object and populated into the corresponding fields of the emission source list. Second, based on the process step and life cycle stage of the emission source, typical activity data reference values are extracted from the annotation information of the standard exchange flow object to verify whether the extracted activity data is within a reasonable range.
[0036] Finally, for emission sources that cannot be matched with standard exchange flow objects (such as brand new processes not included in the database), they are marked as UPRs to be supplemented, and the product and process context in which the emission source appears is recorded, and the database is updated.
[0037] A standardized emission source inventory should include at least the following fields: life cycle stage, emission source category, emission source name, activity data value, activity data unit, technical specifications, process step description, transport distance and mode of transport, allocation tag, data quality label, and knowledge logic field. The knowledge logic field records the origin and reasoning basis for each emission source, and includes the following possible values: direct document extraction, process rule reasoning, UPR database matching, and default value estimation.
[0038] The generated standardized emission source inventory is output in a structured data format. Using this standardized emission source inventory as input, the carbon footprint factor library calculation service is invoked. Based on fields such as emission source name, technical specifications, life cycle stage, geographical region, and transportation mode, corresponding carbon emission factors are matched. Activity data is multiplied by the factors to obtain the carbon emissions of each emission source. These are then aggregated by life cycle stage, and the final carbon footprint result for the target product is output.
[0039] By unifying text processing and entity extraction, the efficiency of processing heterogeneous data is significantly improved; through process reasoning and completion rules, the completeness of identifying hidden emission sources is greatly enhanced. By pre-constructing a process knowledge base containing multiple product types, emission sources that are missed in the original data but are inherently present in the process, such as injection molding cooling water, mold release agents, cutting fluids, and wastewater treatment, are automatically completed based on product type and trigger keywords. This significantly reduces the underreporting rate of hidden emission sources and improves the accuracy of carbon footprint accounting.
[0040] Example 2: The above examples, through unified text processing, entity extraction, process completion, and UPR matching, achieve automated generation of a standardized emission source inventory from heterogeneous product data, improving modeling efficiency and inventory completeness. However, the process reasoning in Example 1 relies on hard-coded rules based on product type. For hybrid material products, special processes, or new materials, hidden emission sources may still be omitted or incorrectly included. This example further improves upon the above examples.
[0041] The method further includes: S31: Pre-constructing a product case library, extracting meta-features of the target product and constructing a target feature map, calculating the graph edit distance between the target feature map and each case feature map, generating a similar case candidate set, specifically selecting the K cases with the smallest distance as the similar case candidate set; finding common nodes between the target feature map and each similar case in the similar case candidate set; for non-common nodes, triggering difference rule reasoning to determine the difference emission sources; merging the difference emission sources into the implicit emission sources, and combining the candidate emission sources and the standard exchange flow object to output a comprehensive emission source list for the target product. The standardized emission source list output in Example 1 is based on the results of document extraction and hard-coded process rule completion. Based on the above, Example 2 refines and completes the implicit emission sources through meta-feature map matching and case transfer reasoning, outputting a more complete comprehensive emission source list.
[0042] In some embodiments, a product case library is pre-built. Each case includes a list of meta-features and a topological relationship diagram of the product. The meta-features refer to the basic, indivisible, process-related units extracted from the original product modeling data to describe the inherent attributes of the product itself. Each meta-feature independently represents a specific attribute of the product in a corresponding dimension. Multiple meta-features are interconnected through process relationships to jointly constitute a complete description of the product's process knowledge. Graph edit distance is defined as a metric for feature graph similarity. Meta-features specifically include: material type (the basic material used in the product or its components, such as PP, ABS, 45# steel, etc.), geometry (the product's shape or structural features, such as plate, rod, shell, etc.), connection method (the way the product is joined internally or between components, such as snap-fit connection, bonding, welding, etc.), surface treatment (the type of post-processing technology on the product surface, such as electroplating, spraying, anodizing, etc.), and tolerance level (the product's processing accuracy requirements, such as IT6, IT8, free tolerance).
[0043] Constructing a target feature map includes: extracting all identifiable meta-features from the intermediate text representation, treating the meta-features as nodes, and the spatial or technological relationships between features as edges, to generate a feature topology map of the target product as the target feature map.
[0044] Calculate the graph edit distance (GED) between the target feature map and the feature map of each case in the case library (i.e., the minimum operation cost of transforming one graph into another by adding or deleting nodes and edges). Select the K cases with the smallest distance (K values [3,5]) as the candidate set of similar cases. As K increases, the similarity between the Kth nearest case and the target feature decreases sharply. Statistical analysis of a typical product feature library shows that the GED of the 3rd nearest case is usually 1.5-2 times that of the 1st nearest case, the GED of the 5th nearest case is 2-3 times that of the 1st nearest case, while for the 6th nearest case and beyond, the GED quickly exceeds the irrelevant threshold, and introducing them will introduce noise from irrelevant rules. Therefore, K=5 is the balance point between effectiveness and efficiency.
[0045] Specifically, Graph Edit Distance (GED) measures the degree of difference between two feature graphs. It is the sum of the minimum operation costs required to completely transform a source graph into a target graph through editing operations such as inserting, deleting, and replacing nodes, inserting, deleting, and replacing edges. In this invention, the smaller the GED between the feature graph of the target product and the feature graphs of cases in the case library, the more similar the feature structures of the two products are, and the more transferable the rules for process inference are.
[0046] Let the source graph be The target image is Where V represents the set of nodes and E represents the set of edges. From arrive The graph edit distance is defined as: ; in, This is the i-th edit operation; k is the length of the edit operation sequence; This is the cost of performing the operation. All possible sequences of edit operations form a set, and the sequence with the minimum total cost is taken as the graph edit distance. In actual computation, edit operations are limited to the following six categories, each assigned a fixed cost (the cost can be adjusted according to semantic importance): Deleting node v: Cost Inserting node v: Cost Replace node v with v': Cost Semantic similarity based on node labels: 0 for completely identical labels, 0.2 for semantically similar labels (e.g., "injection molding" and "injection molding"), and 1 for semantically completely different labels (e.g., "metal" and "ceramic"); Deleting edge (u,v): cost Inserting edge (u,v): Cost Replace edge (u,v) with (u',v'): cost 0.1. The above values can be obtained through industry expert calibration or historical data statistics. Those skilled in the art can adjust them according to the actual scenario, but this will not affect the comparison and sorting of graph editing distances.
[0047] In some embodiments, for simplicity, the replacement cost of all nodes is based on the semantic similarity of the node labels; if the labels are the same (e.g., “PP material” and “PP material”), the cost is 0; if the labels are different but semantically similar (e.g., “injection molding” and “injection molding”), the cost is 0.2; if the labels are completely different (e.g., “metal” and “ceramic”), the cost is 1. The cost of deleting or inserting a node is fixed at 1 (if the node should exist but is missing) or weighted by node type (e.g., the cost of a core material node is higher than that of an auxiliary feature node). The cost of edge operations is generally much smaller than the cost of node operations; for example, the cost of deleting or inserting an edge is 0.1. The graph edit distance is essentially a minimum-cost edit path length, with units consistent with the cost unit.
[0048] Set a threshold based on empirical statistics. To determine whether two feature maps are significantly different and therefore considered unrelated, the threshold is calculated as follows: ; in, Let |V1| represent the number of nodes in the node sets of graphs G1 and G2, respectively. For example, if G1 has 10 nodes, then |V1| = 10. This represents the number of edges contained in the edge set of graphs G1 and G2; and These are the upper bound coefficients for the costs of nodes and edges, respectively. In typical process feature graphs, the number of nodes (meta-features) is usually between 5 and 20, and the number of edges does not exceed twice the number of nodes; [Setting...] (Maximum cost of node replacement) (Maximum cost of edge operations), then when both graphs have 10 nodes and 15 edges, the maximum possible GED is approximately In practical applications, when the normalized GED (GED divided by the maximum possible cost) is greater than 0.6, it is considered irrelevant, i.e.: ;in, The maximum edit distance is assumed to be in the worst-case scenario, set as the total cost of the edit operations required to delete and then insert all nodes. For a specific implementation, a fixed absolute threshold is preset, for example... When GED > 5.0, it is considered unrelated. This value is determined through experimental statistics. Under typical product characteristic dimensions, the GED of similar products is generally between 0 and 3, while the GED of different product categories exceeds 6.
[0049] Find common nodes between the target feature map and each similar case in the candidate set of similar cases. Compare the target feature map with each similar case in the candidate set to find the corresponding common subgraph. Take the feature nodes in all common subgraphs as common nodes and directly inherit and use the implicit emission source completion rule (i.e., the implicit emission source supplemented by the completion rule in the above embodiment). For non-common nodes, i.e. feature nodes other than the common nodes, identify and determine the extra nodes and missing nodes, triggering difference rule reasoning: extra nodes are feature nodes that exist in the target feature map but not in the similar case map (e.g., metal-plastic interface, surface spraying requirements, presence of adhesive interface); missing nodes are feature nodes that exist in the similar case map but not in the target feature map (e.g., the similar case has a "heat treatment process", but the target product does not).
[0050] For non-public nodes, the difference rule reasoning is as follows: based on the type of difference characteristics of non-public nodes, identify and determine additional and missing nodes; search the general rule base to determine the differential emission sources. Specifically, for additional nodes, based on the type and connection relationship of the additional nodes, retrieve the implicit emission source completion rules for the corresponding type of node from the general rule base, and automatically generate new emission source entries; for missing nodes, ignore the corresponding rules in similar cases and do not complete them. The general rule base is a database containing all meta-features of each product, constructed based on product data literature and industry standards. It can also be summarized from historical cases or manually entered by experts. Building a database is a conventional technique, and this application does not impose specific restrictions here.
[0051] Specifically, extract node type labels, such as interface type as adhesive, surface treatment as electroplating, and shape feature as deep hole. Search for rule entries matching this type label in the general rule base. The matching method can be exact matching, label-level matching, or semantic matching, without specific limitations. If a general rule is found, output the predefined list of implicit emission sources and the default estimation formula in the general rule. If no rule is found, mark it as an unknown feature, record the unknown feature and its context, and do not complete the relevant content in the list. For each missing feature node, directly ignore the rules associated with it in similar cases without any completion.
[0052] The original implicit emission sources and the differential emission sources derived from the differential rule reasoning are merged, and after deduplication, a comprehensive emission source list for the target product is output. Each implicit emission source is accompanied by an initial activity data estimate and the estimation method, which are written into the knowledge logic field. The reasoning results are stored in the case library.
[0053] Identify all emission source nodes in the inventory. An emission source node refers to a single record in the comprehensive emission source inventory, i.e., a specific emission source, including: input material nodes (such as polypropylene granules), output product nodes (such as finished injection molded parts), waste nodes (such as waste materials and scrap metal), and energy consumption nodes (such as electricity consumption for injection molding and cooling water). Each node includes a name, activity data, unit, life cycle stage, etc.
[0054] The method further includes: S4: Selecting mass flow nodes and energy flow nodes in the emission source inventory according to the production process flow, establishing conservation edges, recording the conversion coefficient for each edge, and assigning a feasible interval to each node; based on the conservation edges and feasible intervals, constructing conservation equations and objective functions, and solving to minimize the adjustment amount as the adjusted target value; integrating the adjusted target value with the fields of the comprehensive emission source inventory to form a precise emission source inventory and calculate the carbon footprint. Through mass conservation constraint adjustment, the activity data in the comprehensive emission source inventory is corrected to self-consistent values, ultimately forming a precise emission source inventory. The standardized emission source inventory, the comprehensive emission source inventory, and the precise emission source inventory are in a progressive relationship.
[0055] The mass flow nodes refer to substances with mass, whose activity data units can be converted to kilograms or tons in the International System of Units (SI). In the lifecycle phase, mass flow nodes belong to the material input, product output, or waste output stages of raw material acquisition, manufacturing, and waste disposal; for example, polypropylene granules, steel, cutting fluid, packaging cartons, scrap metal, wastewater, and volatile organic compounds (in mass form). The energy flow nodes describe energy input and output, whose activity data units can be converted to joules or kilowatt-hours in the SI. In the lifecycle phase, they belong to the energy consumption input of the manufacturing stage; for example, electricity consumption for injection molding, electricity consumption for cooling water circulation pumps, natural gas combustion, and steam consumption. Specifically, they are identified and distinguished based on the emission source names in the inventory: emission sources containing "electricity consumption," "natural gas," or "steam" belong to energy flows; those containing "granules," "sheets," or "waste" belong to mass flows.
[0056] Conservation edges are established by adding directed edges between mass flow nodes and energy flow nodes according to the process flow sequence and energy conversion relationships. Mass conservation edge: Input material mass = Product output mass + Waste mass + Loss (loss can be considered an implicit node). Energy conservation edge: Input electrical energy / fuel calorific value = Product internal energy increment + Heat dissipation + Energy carried away by exhaust gas. Isolated nodes where a conservation relationship cannot be directly established (such as transportation emissions) do not participate in constraint propagation and retain their original values. Each edge records a conversion coefficient and assigns a feasible interval to each node. The conversion coefficient is set according to product industry standard rules, such as 1kg of polypropylene granules being converted into 0.95kg of product and 0.05kg of waste through injection molding.
[0057] If the node data is directly extracted from a document and the data quality label is high, then the interval width is set to 0, i.e., a fixed value. If it comes from process inference estimation but the confidence level is above 80%, the data quality label is medium, and the interval width is [missing value]. The remaining data quality is labeled as low, with a range width of [missing information]. A range is set based on the estimation method in the knowledge logic. The specific range is set according to the industry standards and data quality labels corresponding to the product. This application does not impose specific restrictions here. For example, if the power consumption of cooling water is estimated at 10% of the power consumption of injection molding, the range is set to [8.5%, 11.5%]. If there is no clear basis, the default range is [original value × 0.7, original value × 1.3]. If it comes from the default value of UPR, the variance information provided in UPR data is used to set the range. If there is a unit inconsistency (such as the input node unit is ton, and the product node unit is kilogram), it is automatically converted to the same unit and unified to kilogram; the energy data unit is unified to joule.
[0058] In some embodiments, based on the established conservation boundary, each production process corresponds to two conservation equations (mass and energy). All mass flow nodes and energy flow nodes participating in the conservation are uniformly numbered as i = 1, 2, ..., N, where N is the total number of nodes. The original activity data of each node i is denoted as... The adjusted target value is denoted as (Unknown quantity).
[0059] For the conservation of mass, each equation takes the form: ; in, This is the set of indices for the input material nodes in the k-th process; Let be the set of indices for the output product nodes and output waste nodes in the k-th process; It is the target quality value of the input material at the i-th node; The j-th node outputs the target quality values for both products and waste. The loss amount (in mass) of the k-th process is treated as an independent variable node, with its original value... Initially set to 0, and a feasible interval is assigned, which is the total input mass [0, 5%]. When calculating the loss variable, if the calculation result is negative, it is automatically zeroed and the adjustment information is recorded.
[0060] For the conservation of energy, each energy equation takes the form: ; in, Let be the energy value (joules) of the e-th energy input node. This is the total number of energy input nodes; The increase in the product's internal energy is estimated based on the product's specific heat capacity and mass change, and is considered a known constant. Heat loss is considered as a variable. The heat loss at the f-th energy output node, i.e. the energy carried away by the waste gas and wastewater, is used as a variable. This refers to the total number of energy output nodes. When heat loss and heat dissipation cannot be directly obtained, the closest product and hidden emission sources are retrieved based on the product name. Historical values are obtained through statistical analysis of historical similar data records; this will not be elaborated upon further here. For simplicity, energy conservation can be analogized to mass conservation.
[0061] Solve all the conservation equations simultaneously and write them in matrix form. , where N is a column vector consisting of the adjusted values of all nodes (including loss nodes), A is a coefficient matrix, each equation corresponds to one row, the input node coefficient is +1, the output node coefficient is -1, and the loss node coefficient is -1.
[0062] In the law of conservation of mass, input materials entering the process should be marked with a positive sign (increasing system mass); output products and waste leaving the process should be marked with a negative sign (decreasing system mass); losses are also considered as mass lost from the system, and therefore are also marked with a negative sign. The equation indicates that the mass entering equals the mass leaving plus losses, satisfying the law of conservation of mass. The coefficients +1 and -1 only indicate the direction (increase or decrease), not the magnitude.
[0063] Define a weighted least squares objective function to minimize the weighted sum of squares of the adjustment amounts: ; in, It is the weighted sum of squares of the adjustments at all emission source nodes, and is a scalar value; This is the raw activity data for each node i. It is the adjusted target value, and the specific value is calculated by solving the constraint system composed of conservation equations; The weight of the i-th node is inversely proportional to the data quality label. When the data quality label is high, the node's value is considered unadjustable, equivalent to... In actual calculations, such nodes are directly fixed as... Removed from the variable space.
[0064] The weighted sum of squares is the square of the adjustment amount (adjusted value minus the original value) for each emission source node, multiplied by the node's weight (higher data quality means higher weight). By minimizing the objective function, a set of adjusted values with the smallest overall deviation can be found, while satisfying all conservation equations. This set of values best respects the original high-precision data (due to its large weight and small adjustment range), while placing the main adjustment pressure on the low-precision estimated data. The result is physically consistent activity data with minimal overall deviation, improving the accuracy of carbon emission identification and carbon footprint analysis.
[0065] Assuming the interval corresponds to a confidence interval of over 95%, when the data quality label is medium, the weights are... ,in, When the data quality label is low, the weight... ,in, .
[0066] For loss variables Its initial value is set to 0, and the standard deviation is set to 5% of the total input mass of this process, as the starting parameter (based on specific industry experience values; in conventional manufacturing industries (such as injection molding, machining, casting, and assembly), process losses (including flash, overflow, volatiles, spillage, scrap, oxidation loss, etc.) typically account for 1%-10% of the total input material mass). In the mass conservation equation, the loss variable acts as an independent slack variable to absorb mass imbalances caused by incomplete data or estimation errors (e.g., material losses such as flash, volatiles, and spillage that are not separately recorded). By assigning a reasonable initial value and standard deviation to the loss, the mass of the loss is automatically estimated during the solution process, thus ensuring the mass conservation equation holds true. The calculated loss value can also serve as a reference for activity data of hidden emission sources (such as volatile organic compounds), improving the completeness of carbon emission identification.
[0067] In some embodiments, when the number of equations is small and interval information is of greater interest (e.g., the user wants to know the range of all possible values rather than a single-point solution, thus deriving an emission range), an interval propagation algorithm can be used as an alternative, performing an iterative process, specifically: initializing each variable. The feasible interval for each conservation equation, for example Calculate the new interval for the third variable based on the known intervals of the two variables: ; and with Find the intersection of the original intervals. If the intersection is empty, mark the equation as conflicted. Repeat the traversal of all equations until all variable intervals no longer shrink or the maximum number of iterations (e.g., 100) is reached. Output the final feasible interval for each variable. If a single value is needed, the midpoint of the interval can be used as the adjustment value. Interval propagation can explicitly express the range of uncertainty. This embodiment uses the least squares method by default, and switches to interval propagation when there are high conflicts or user requirements.
[0068] After obtaining the adjusted target value, the adjustment results are inspected for quality, and all adjustment information is written into the knowledge logic field of the emission source inventory to ensure that the entire adjustment process is traceable and auditable.
[0069] For each emission source node i involved in the adjustment, calculate the relative adjustment magnitude: ; in, It is a relative adjustment range; It is the initial value; It is the adjusted target value; in particular, when This formula is not applicable when the loss node is initially 0; in such cases, the absolute adjustment amount should be used directly. As a basis for judgment.
[0070] Set the following warning thresholds: If or If 10% of the total mass is input (for loss nodes), a significant warning will be triggered, indicating that the reliability of the original data for that node is extremely low, or that there are excessively strong constraints in the conservation equations, requiring the user to manually review the data.
[0071] like If it is, it will be marked as a moderate adjustment, automatically recorded without warning, and used for auditing reference only.
[0072] like This is considered within the normal adjustment range.
[0073] If the intersection of feasible intervals for any equation is found to be empty (i.e., all constraints cannot be satisfied simultaneously), an irreconcilable conflict is identified. In this case, no numerical adjustments are performed; the original values of all nodes remain unchanged. A data consistency conflict flag is added to the metadata section of the manifest, and all conflicting equations are listed, including the node names involved in the equation, their original values, and data quality labels. For each conflicting equation, possible causes of the conflict are automatically suggested (e.g., in the mass conservation equation, if the total mass of the input material is much less than the sum of the mass of the product and waste, the possible cause is that the waste node is omitted or the unit of activity data in the document is incorrect).
[0074] The value after conservation adjustment ( The data is integrated with other fields of the original list to form a complete, self-consistent, and computable accurate emission source list, and the carbon footprint is calculated based on the emission source list.
[0075] After outputting the final list in a structured data format, users can directly use it as input for the product carbon footprint calculation service. The calculation service matches the corresponding carbon emission factors in the factor library based on fields such as emission source name, technical specifications, life cycle stage, geographical region, and transportation mode in the list. It then multiplies the activity data (adjusted values) of each emission source by the factor to obtain the carbon emissions of that emission source. Finally, the data is summarized by life cycle stage to output the total carbon footprint of the target product.
[0076] In this embodiment, the reasoning capability for complex products and hybrid material scenarios is enhanced through meta-feature graph matching and case transfer reasoning. Breaking through the limitations of traditional rules based on fixed product types, the product is decomposed into indivisible meta-features such as material, shape, connection method, surface treatment, and tolerance level, and a feature topology graph is constructed. The most similar cases are retrieved by computing graph editing distance, rules of common subgraphs are inherited, and multiple feature nodes are processed through difference rule reasoning. This approach can handle hybrid material products such as plastic-coated metal and electronic components, filling gaps in the existing rule base and further improving the generalization capability of emission source identification.
[0077] The physical consistency of activity data is ensured through mass conservation constraint propagation and weighted least squares adjustment. Conservation equations are established by identifying mass flow and energy flow nodes in the inventory, and uncertainty intervals and weights are assigned based on data quality labels. Least squares solutions are used to find the adjusted solution closest to the original observations while satisfying all conservation equations, prioritizing high-precision data and adjusting low-precision estimates. This addresses the imbalance between material input and output, ensuring the reliability and accuracy of carbon footprint calculations.
[0078] To make the technical solution of the present invention clearer, the above specific embodiments may involve several specific numerical values, thresholds, parameters, formulas, and examples. Those skilled in the art should understand that these specific contents are merely exemplary descriptions given for ease of implementation and are not essential technical features necessary for realizing the present invention, nor do they constitute a limitation on the scope of protection of the present invention. The mathematical formulas, parameter values, model structures, calculation rules, constraints, and data acquisition methods described in the specific embodiments section of this application are preferred examples for realizing the technical concept of the present invention, not the only implementation methods. Based on the core logic and inventive concept disclosed in this application, those skilled in the art can make reasonable equivalent substitutions, parameter adjustments, or logical optimizations to the parameters, formulas, and algorithm models according to actual conditions, industry standards, or conventional experimental methods, without producing substantial differences. Any changes that do not depart from the inventive concept of this application fall within the scope of this application. All mathematical expressions in this application should ensure that their physical meaning is clear when applied; parameters such as weights and thresholds can be calibrated using statistical or machine learning methods known in the art.
[0079] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for generating an emission source inventory based on UPR and process inference, characterized in that, include: S1: Obtain the original modeling data of the target product and convert it into an intermediate text representation to obtain candidate emission sources; determine the boundary type of the target product based on the life cycle boundary rules; Based on candidate emission sources and boundary types, process data in the unit process database is retrieved as a standard exchange flow object; S2: Based on the target product, candidate emission sources, and boundary type, perform process reasoning and complete the implicit emission sources according to the pre-set completion rules; S3: Generate a standardized emission source list for the target product based on candidate emission sources, implicit emission sources, and standard exchange flow objects; S31: Pre-build a product case library, extract the meta-features of the target product and construct the target feature map, calculate the graph edit distance between the target feature map and each case feature map, and generate a similar case candidate set; Find common nodes between the target feature map and each similar case in the candidate set of similar cases; for non-common nodes, trigger difference rule reasoning to determine the difference emission sources; merge the difference emission sources into the implicit emission sources, combine the candidate emission sources and the standard exchange flow object, and output a comprehensive emission source list for the target product; The original modeling data includes product BOM, process flow diagram, equipment parameter table, energy consumption ledger, transportation documents and product specifications; the meta-features refer to the basic process-related units extracted from the original product modeling data that describe the inherent attributes of the product itself and cannot be further subdivided.
2. The emission source inventory generation method based on UPR and process inference according to claim 1, characterized in that, The method further includes: S4: Selecting mass flow nodes and energy flow nodes in the emission source inventory according to the production process flow, establishing conservation edges, recording the conversion coefficient for each edge and assigning a feasible interval to each node; constructing conservation equations and objective functions based on conservation edges and feasible intervals, solving for the minimum adjustment amount as the adjusted target value; integrating the adjusted target value with the comprehensive emission source inventory fields to form an accurate emission source inventory and calculate the carbon footprint.
3. The emission source inventory generation method based on UPR and process reasoning according to claim 1, characterized in that, The completion rules are set as follows: determine the product type and identify the corresponding general processing operation flow, identify trigger keywords, and complete the hidden emission sources; pre-set trigger keywords and corresponding hidden emission sources according to different product types.
4. The emission source inventory generation method based on UPR and process inference according to claim 1, characterized in that, The standardized emission source inventory includes at least: life cycle stage, emission source category, emission source name, activity data value, activity data unit, technical specifications, process step description, transportation distance and mode of transportation, allocation tag, data quality tag, and knowledge logic field; wherein, the knowledge logic field is used to record the source and reasoning basis of each emission source, including direct document extraction, process rule reasoning, UPR database matching, and default value estimation.
5. The emission source inventory generation method based on UPR and process inference according to claim 1, characterized in that, Constructing the target feature map includes: extracting all identifiable meta-features from the intermediate text representation, treating the meta-features as nodes, and the technological relationships between features as edges, to generate a feature topology map of the target product as the target feature map.
6. The emission source inventory generation method based on UPR and process inference according to claim 1, characterized in that, A product case library is pre-built. Each case includes a list of product meta-features and a topological relationship diagram. Each meta-feature independently represents a specific attribute of the product in the corresponding dimension. Multiple meta-features are interconnected through process relationships to jointly constitute a complete description of product process knowledge. Meta-features include material type, geometry, connection method, surface treatment, and tolerance level.
7. The emission source inventory generation method based on UPR and process inference according to claim 1, characterized in that, The target feature map is compared with each similar case in the candidate set of similar cases to find the corresponding common subgraph. The feature nodes in all common subgraphs are taken as common nodes, and the implicit emission source completion rules are directly inherited and used. For non-common nodes, the difference rule reasoning is as follows: based on the type of difference features of non-common nodes, the extra nodes and missing nodes are identified and determined; the general rule base is searched to determine the difference emission source. Excess nodes are feature nodes that exist in the target feature map but not in the similar case map; missing nodes are feature nodes that exist in the similar case map but not in the target feature map.
8. The emission source inventory generation method based on UPR and process inference according to claim 7, characterized in that, For multiple nodes, based on the type and connection relationship of the multiple nodes, the implicit emission source completion rules for the corresponding type of node are retrieved from the general rule base, and new emission source entries are automatically generated; for missing nodes, the corresponding rules in similar cases are ignored and no completion is performed; the general rule base is a database containing all meta-features of each product, built based on product data literature and industry standards.
9. The emission source inventory generation method based on UPR and process inference according to claim 2, characterized in that, The mass flow node refers to a substance with mass, and its activity data unit can be converted to the SI kilogram; the energy flow node is used to describe energy input and output, and its activity data unit can be converted to the SI joule.
10. The emission source inventory generation method based on UPR and process inference according to claim 1, characterized in that, The boundary type of the target product is automatically determined based on the life cycle boundary rules. If the target product is a final consumer product, a complete boundary is used, and the emission sources of the product use stage and the disposal stage need to be supplemented. If the target product is an intermediate industrial product, a production boundary is used, and the use stage and the disposal stage are not supplemented.