Carbon footprint acquisition method and system based on knowledge graph

By using a knowledge graph-based approach, we have achieved automatic standardization and probability distribution estimation of multi-source data, which solves the problems of data uncertainty and difficulty in reflecting causal relationships in carbon footprint calculation, and improves the accuracy and adaptability of carbon footprint calculation.

CN122114389BActive Publication Date: 2026-07-31JIAXING UNIV +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIAXING UNIV
Filing Date
2026-04-24
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies are unable to fully reflect the diversity, uncertainty, and complex causal relationships between parameters of multi-source heterogeneous data in carbon footprint calculation, resulting in limited accuracy of carbon footprint estimation and difficulty in achieving dynamic adjustment and scientific reasoning.

Method used

A knowledge graph-based approach is adopted to generate standard word roots by semantic recognition and normalization of multi-source parameter groups. Data weighting is performed using the analytic hierarchy process and statistical methods to construct a probability distribution model, and carbon footprint probability distribution is estimated by combining causal relationships.

Benefits of technology

It improves the efficiency and accuracy of data fusion, quantifies differences in data quality, enhances the robustness and adaptability of carbon footprint calculation, better reflects the uncertainty of real data, and improves the interpretability and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114389B_ABST
    Figure CN122114389B_ABST
Patent Text Reader

Abstract

This invention provides a knowledge graph-based carbon footprint acquisition method and system, belonging to the field of carbon footprint acquisition technology. By constructing a semantic normalization mechanism, this invention achieves standardized processing of multi-source parameter groups, improving the efficiency and accuracy of data fusion. Combining the analytic hierarchy process (AHP) with multiple evaluation dimensions, it analyzes the reliability of parameters, quantifies and integrates data quality differences, ensuring high credibility and stability of the fused data. Statistical methods are used for repeated sampling to achieve non-parametric estimation of data uncertainty, and error control between the fused data and the sampling mean enables dynamic optimization of the probability distribution model, enhancing the robustness and scientific rigor of carbon footprint calculation. The combination of knowledge graph and probability distribution enables joint probabilistic reasoning of carbon footprint, allowing it to better reflect the uncertainty of real data, thereby improving the overall interpretability and adaptability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of carbon footprint acquisition technology, specifically to a carbon footprint acquisition method and system based on knowledge graphs. Background Technology

[0002] With increasing global environmental awareness, the accurate calculation and management of carbon footprints has become a key focus for enterprises. Carbon footprint data typically originates from multiple heterogeneous systems and platforms, including Enterprise Resource Planning (ERP) systems, IoT devices, third-party reports, and publicly available official data. These multi-source data sources exhibit significant differences in format, naming, collection methods, and data quality, posing a substantial challenge to accurate carbon footprint measurement. Furthermore, existing technologies largely rely on deterministic models and manual rules for data processing and calculation, which struggle to fully reflect the diversity, uncertainty, and complex causal relationships between parameters. This limits the accuracy of carbon footprint estimation and hinders dynamic adjustments and scientific reasoning.

[0003] In the prior art, CN120429838A discloses a method and apparatus for intelligent fusion of multi-source heterogeneous data for carbon footprint calculation. This method includes: acquiring multi-source heterogeneous data of a product, including basic product information and carbon footprint data throughout the product's lifecycle; parsing the data types of the multi-source heterogeneous data and selecting an algorithm matching the data types to convert the multi-source heterogeneous data into structured data, obtaining a structured normalized dataset; performing semantic extraction on the structured normalized dataset based on a loaded carbon footprint domain ontology, obtaining a semantically normalized triplet dataset; performing quality assessment on the semantically normalized triplet dataset and constructing an initial carbon footprint dataset including quality information based on the quality assessment results; and performing conflict detection and conflict ablation on the initial carbon footprint dataset to obtain the final carbon footprint dataset of the product. Although this method achieves the fusion of multi-source heterogeneous data by constructing a triplet dataset, it focuses on semantic extraction of the data ontology, and the generated dataset is still a fixed value, failing to fully consider data uncertainty. This results in insufficient simulation and reasoning ability for real data, and overall weak accuracy and adaptability.

[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for obtaining carbon footprints based on knowledge graphs, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: The carbon footprint acquisition method based on knowledge graphs includes the following steps: S1: Collect multi-source parameter groups related to the carbon footprint of the product to be analyzed, perform semantic recognition on the parameter names of each sub-parameter in the multi-source parameter group, generate several sets of standard word roots based on the industry standards of the product to be analyzed, and classify each sub-parameter after semantic recognition into the standard word root with the highest semantic similarity. S2: Based on the multi-source parameter group, the original data of each sub-parameter is obtained. The original data of each sub-parameter under the same standard root word are preprocessed to generate standard data. The source credibility, historical stability and collection method of the original data are used as evaluation dimensions. The analytic hierarchy process is used to generate the credibility weight of different sub-parameters. The credibility weight is used to weight the standard data and generate fused data. S3: Randomly select several sets of standard word roots to construct a validation set. For each set of standard word roots in the validation set, use statistical methods to repeatedly sample all standard data under the standard word root to generate the probability distribution of the theoretical data of the standard word root. Calculate the corresponding sampling mean based on the distribution to estimate the probability distribution of the standard word root. S4: Using the fused data as a reference benchmark, calculate the relative error and mean error between the sampling mean and the fused data. Based on the mean error, determine whether the probability distribution setting corresponding to the theoretical data of the standard word roots is reasonable. Then, based on the judgment result, estimate the probability distribution of all standard word roots. S5: Construct a knowledge graph about the carbon footprint of the product to be analyzed by using each standard word root as a node, and input the probability distribution and distribution parameters of the theoretical data of each standard word root into the knowledge graph to expand the node attributes, so as to realize the probability distribution estimation of the carbon footprint of the product to be analyzed.

[0007] Preferably, the specific process of step S1 is as follows: S101: Collect multi-source parameter groups, use a pre-trained natural language model to perform semantic recognition and normalization on the parameter names of each sub-parameter, and map the normalized parameter names into high-dimensional semantic vectors. S102: Based on carbon footprint industry standards, several sets of standard word roots are generated, and a pre-trained natural language model is used to map the standard word roots into high-dimensional semantic vectors. S103: For each sub-parameter, calculate the semantic similarity between the parameter name and all standard word roots based on its parameter name and the high-dimensional semantic vector of the standard word root, and classify the sub-parameter into the standard word root with the highest semantic similarity.

[0008] Preferably, the logic for the data preprocessing is as follows: For each standard word root, the original data of each sub-parameter is converted to the same dimension in turn, and the original data of each sub-parameter is aligned based on the time axis. For each sub-parameter under the standard word root, the mean and standard deviation of its original data in the time axis direction are calculated sequentially, and statistical methods are used to identify and remove outliers in the mean. For each sub-parameter under the standard word root, the K-nearest neighbor method is used to impute missing values, and the mean value after imputation is used as the standard data for that sub-parameter.

[0009] Preferably, the credibility weight includes a first weight and a second weight, and the generation logic of the fused data is as follows: For each standard word root, the data source and collection method of each sub-parameter are obtained in sequence. The collection method includes automatic and manual methods, and is labeled as a binary label, with a value of 0 indicating manual collection and a value of 1 indicating automatic collection. For each sub-parameter under the standard root, scores are given to different data sources to quantify source credibility, the normalized standard deviation of its standard data in the time axis direction is calculated to quantify historical stability, and the collection method is quantified based on the binary label of the collection method. For each sub-parameter under the standard word root, the three evaluation dimensions of source credibility, historical stability and collection method are compared in pairs based on the analytic hierarchy process to construct a judgment matrix. Based on the judgment matrix, feature weights are calculated to obtain the first weight of different evaluation dimensions. Then, the three evaluation dimensions are weighted based on the first weight to obtain the second weight corresponding to the sub-parameter. For the standard word root, the second weight of each sub-parameter is normalized, and then the standard data of each sub-parameter is weighted and summed using the normalized second weight. The weighted sum is used as the fusion data of the standard word root.

[0010] Preferably, step S3 includes: S301: Randomly select several groups from all standard word roots to construct a validation set. For each standard word root in the validation set, use the Bootstrap resampling method to sample the standard data of each sub-parameter with replacement. S302: For each sample with replacement, the number of samples is equal to the number of sub-parameters under the standard root word, and the frequency of the sampled data is used to weight and sum them to generate weighted sampled data; S303: Repeatedly perform sampling with replacement to generate several sets of weighted sampling data. Use a preset distribution model to fit the several sets of weighted sampling data to obtain their probability distribution, so as to realize the probability distribution estimation of the theoretical data of the standard word root. And calibrate the mean of all weighted sampling data as the sampling mean.

[0011] Preferably, the logic for determining whether the probability distribution setting corresponding to the standard word root is reasonable is as follows: Along the time axis, the relative error between the downsampled mean and the fused data at each time point is calculated, the corresponding mean error is calculated, and the mean error is compared with a preset error threshold. If the mean error does not exceed the error threshold, the probability distribution of the standard word root is considered to be reasonable. If the mean error exceeds the error threshold, the probability distribution of the standard word root is considered to be unreasonable.

[0012] Preferably, the logic for using the judgment result as a reference to estimate the probability distribution of all standard word roots is as follows: For each probability distribution in the validation set with an unreasonable standard word root, change the fitting model used when fitting the distribution and re-estimate the probability distribution until the mean error does not exceed the error threshold. The fitting models used for all standard word roots in the statistical validation set are sorted in descending order of frequency of use. When estimating the probability distribution of all standard word roots, the distributions are fitted sequentially in descending order of the sorting list to ensure that the mean error does not exceed the error threshold.

[0013] Preferably, the distribution model includes normal distribution, kernel density estimation, Poisson distribution, and empirical distribution.

[0014] Preferably, step S5 includes: S501: Using a multi-source parameter set as input to the knowledge graph, each standard word root is used as a node of the knowledge graph, and the probability distribution and distribution parameters of the standard word root are used as extended attributes of the node. S502: Based on the calculation process of the carbon footprint of the product to be analyzed, define the causal relationship between nodes and form the edges of the knowledge graph; S503: Introduces a reasoning algorithm based on a probabilistic graphical model, combines the causal relationships between nodes, generates the probability distribution of theoretical data on the carbon footprint of the product to be analyzed, and uses it as the output of the knowledge graph.

[0015] A knowledge graph-based carbon footprint acquisition system, wherein the carbon footprint acquisition system is used to execute the above-described carbon footprint acquisition method, specifically including: The data acquisition module is used to collect multi-source parameter groups related to the carbon footprint of the product to be analyzed, perform semantic recognition and normalization on the parameter names of each sub-parameter in the multi-source parameter group, generate several sets of standard word roots based on industry standards, and classify each sub-parameter into the standard word root with the highest semantic similarity. The data processing module is used to preprocess the raw data of each sub-parameter under the same standard root word to generate standard data. The module uses source credibility, historical stability and collection method as evaluation dimensions, and generates credibility weights for different sub-parameters based on the analytic hierarchy process. The credibility weights are then used to weight the standard data and generate fused data. The distribution fitting module is used to randomly sample and construct a validation set from all standard word roots. For each group of standard word roots in the validation set, the probability distribution is estimated sequentially. Statistical methods are used to repeatedly sample the standard data to generate the probability distribution, and the corresponding sampling mean is obtained simultaneously. The data judgment module is used to calculate the relative error between the sample mean and the fused data using the fused data as a reference benchmark, and to judge whether the probability distribution setting corresponding to the set of standard word roots is reasonable based on the calculation result. Then, the judgment result is used as a reference to estimate the probability distribution of all standard word roots. The graph construction module is used to construct a knowledge graph about carbon footprint by using each standard word root as a node, and inputs the probability distribution and distribution parameters of the theoretical data of each standard word root into the knowledge graph to expand the node attributes, so as to realize the probability distribution estimation of the carbon footprint of the product to be analyzed.

[0016] Compared with the prior art, the beneficial effects of the present invention are: This invention achieves automatic standardization of multi-source parameter sets by constructing a semantic normalization mechanism based on a natural language model, greatly improving the efficiency and accuracy of data fusion. It combines the analytic hierarchy process (AHP) to analyze parameter reliability from multiple evaluation dimensions, effectively quantifying and integrating data quality differences, ensuring high credibility and stability of the fused data. Statistical methods are used for repeated sampling to achieve nonparametric estimation of data uncertainty, and error control between the fused data and the sampling mean enables dynamic optimization of the probability distribution model, enhancing the robustness and scientific rigor of carbon footprint calculation. Based on the combination of knowledge graphs and probability distributions, it deeply characterizes the causal relationships and statistical dependencies between parameters, realizing joint probabilistic inference of carbon footprint, enabling it to better reflect the uncertainty of real data, thereby improving the overall interpretability and adaptability of the system. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the overall method flow of the present invention; Figure 2 This is a flowchart illustrating step S1 in this invention; Figure 3 This is a flowchart illustrating step S3 in this invention; Figure 4This is a flowchart illustrating step S5 in this invention; Figure 5 This is a schematic diagram of the module structure of the carbon footprint acquisition system in this invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0020] Example: Please see Figures 1-4 The present invention provides a technical solution: The carbon footprint acquisition method based on knowledge graphs includes the following steps: S1: Collect multi-source parameter sets related to the carbon footprint of the product to be analyzed, perform semantic recognition and normalization on the parameter names of each sub-parameter in the multi-source parameter set, generate several sets of standard word roots based on industry standards, and classify each sub-parameter into the standard word root with the highest semantic similarity.

[0021] The specific process of step S1 is as follows: S101: Collect multi-source parameter sets, use a pre-trained natural language model to perform semantic recognition and normalization on the parameter names of each sub-parameter, and map the normalized parameter names into high-dimensional semantic vectors.

[0022] The natural language models used here include Word2Vec, BERT, and SentenceTransformer. The steps for semantic recognition and normalization of the parameter names of each sub-parameter include format correction of the parameter names, removal of irrelevant characters, and conversion of synonyms, abbreviations, and spelling variations, such as converting "CO2 emissions" to "CO2 emissions". As for the dimensional range of the high-dimensional semantic vector, for the Word2Vec model, it is usually set between 100 and 300 dimensions; for the BERT model, it is usually set between 768 and 1024 dimensions; and for the SentenceTransformer model, it is usually set between 384 and 768 dimensions.

[0023] S102: Based on the industry standard of carbon footprint, several sets of standard word roots are generated, and a pre-trained natural language model is used to map the standard word roots into high-dimensional semantic vectors.

[0024] The standard roots generated here are for the purpose of unifying the parameter names of sub-parameters from different sources. The standard roots can be set as: "electricity consumption", "natural gas consumption", "carbon dioxide emissions", etc.

[0025] S103: For each sub-parameter, calculate the semantic similarity between the parameter name and all standard word roots based on its parameter name and the high-dimensional semantic vector of the standard word root, and classify the sub-parameter into the standard word root with the highest semantic similarity.

[0026] The semantic similarity here can be calculated using cosine similarity, and the specific formula is as follows: In the formula , They represent the first The parameter name of the sub-parameter and the first sub-parameter A standard word root, , These are indices for parameter names and standard word roots, respectively. Indicates the first The parameter name of the sub-parameter and the first sub-parameter Cosine similarity between standard word roots , They represent the first The parameter name of the sub-parameter and the first sub-parameter A high-dimensional semantic vector of a standard word root, , These represent the L2 norms of the two vectors, respectively. In the final output, each standard word root will have a mapping table reflecting the mapping relationship between the standard word root and the parameter names of its sub-parameters. This can be represented in set form. For example, the mapping table for the standard word root "electricity consumption" can be represented in set form as: {"electricity consumption", "electricity consumption", "electricity usage", ...}. If the parameter name of a sub-parameter has the same cosine similarity as multiple standard word roots, it will be randomly assigned to one of the standard word roots.

[0027] Understandably, due to the different sources and acquisition methods of data in multi-source parameter groups, parameter names vary and contain synonyms, ambiguities, abbreviations, and spelling differences. Therefore, even for the same type of product to be analyzed, there will be much data with similar semantics. For example, data sources may include enterprise ERP systems, IoT devices, publicly available official data, and third-party reports. Even when referring to the parameter "electricity consumption," different sources may use different parameter names (e.g., energy consumption, power consumption), and the data format and units may also differ (e.g., kW / h, MW / h). Therefore, in this step, by standardizing the parameter names of each sub-parameter in the multi-source parameter group, we not only avoid data conflicts and statistical biases caused by parameter heterogeneity, facilitating subsequent data fusion and cross-system integration, but also eliminate the hassle of manual review and identification, greatly improving overall work efficiency.

[0028] S2: The raw data of each sub-parameter under the same standard root word are preprocessed to generate standard data. Using source credibility, historical stability, and collection method as evaluation dimensions, the credibility weights of different sub-parameters are generated based on the analytic hierarchy process (AHP). These credibility weights are then used to weight the standard data and generate fused data. The raw data for each sub-parameter is the unprocessed data from the directly collected multi-source parameter groups.

[0029] The logic for data preprocessing is as follows: For each standard word root, the original data of each sub-parameter is converted to the same dimension in turn, and the original data of each sub-parameter is aligned based on the time axis. Taking the standard term "electricity consumption" as an example, the original data in the sub-parameters are converted into parameters in the same dimension, kW / h. Accordingly, those originally in the dimension of MW / h need to be multiplied by 1000. Data alignment is performed on the time axis to ensure consistency between different original data.

[0030] For each sub-parameter under the standard root, the mean and standard deviation of its original data along the time axis are calculated sequentially, and statistical methods are used to identify and remove outliers in the mean.

[0031] The statistical methods used here can be "3" The "outlier rule" means that if the original data exceeds the range of three standard deviations above or below the mean, it is considered an outlier, thus avoiding its impact on the overall data.

[0032] For each sub-parameter under this standard terminology, the K-nearest neighbor method is used to impute missing values, and the mean of the imputed values ​​is used as the standard data for that sub-parameter. Taking daily electricity consumption over 30 days as an example, after the above data processing, the resulting standard data can be represented in set form, i.e.: This indicates the first word of the standard root word. Standard data for each sub-parameter, That is, the set elements within this standard data. This represents the index of a set element, equivalent to the index of a time node on a timeline. This represents the total number of elements in the set, equivalent to the total number of time points. Here, the set elements represent the electricity consumption from day 1 to day 30, and the total number of elements in the set is 30.

[0033] The logic for generating fused data is as follows: For each standard word root, the data source and collection method of each sub-parameter are obtained in sequence. The collection method includes automatic and manual collection, and is labeled as a binary label, with a value of 0 indicating manual collection and a value of 1 indicating automatic collection.

[0034] As mentioned earlier, the data sources for sub-parameters are diverse, including enterprise ERP systems, IoT devices, official public data, and third-party reports. The degree of automation in the collection method affects the accuracy of the data. Automated collection is better than manual entry. Therefore, binary labels are used for differentiation, and the scores corresponding to automated collection are higher, which means that the corresponding sub-parameters will also have higher weights in subsequent calculations.

[0035] For each sub-parameter under the standard root, scores are assigned to different data sources to quantify source credibility, the normalized standard deviation of the standard data in the time axis direction is calculated to quantify historical stability, and the collection method is quantified based on the binary label of the collection method.

[0036] It's understandable that the standard root and its corresponding sub-parameters essentially belong to the same type of data; however, due to different data sources, there may be data discrepancies, thus requiring integration. Furthermore, since different data sources have varying degrees of authority—for example, official public data or third-party reports are generally more credible than self-reported data from companies—when scoring and quantifying different data sources, expert experience can be used to assess the authority of the source and assign corresponding scores. For instance, a company's ERP system might score 0.25, IoT devices 0.5, third-party reports 0.75, and official public data 1. The more authoritative the data source, the higher the score and the higher its credibility, with the overall score controlled between 0 and 1. The normalized standard deviation along the time axis can be obtained through minimax normalization, calculated as follows: In the formula Indicates the first The normalized standard deviation of each sub-parameter along the time axis (i.e., the value of historical stability). Indicates the first The standard deviation of each sub-parameter along the time axis , This represents the maximum and minimum standard deviations of different sub-parameters under this standard terminology. The smaller the standard deviation, the closer the normalized standard deviation is to 1, indicating that the data is more stable.

[0037] For each sub-parameter under the standard word root, the three evaluation dimensions of source credibility, historical stability and collection method are compared in pairs based on the analytic hierarchy process (AHP) to construct a judgment matrix. Based on the judgment matrix, feature weights are calculated to obtain the first weight of different evaluation dimensions. Then, the three evaluation dimensions are weighted based on the first weight to obtain the second weight corresponding to the sub-parameter.

[0038] The values ​​of the three evaluation dimensions are labeled as follows: ~ The judgment matrix constructed using the analytic hierarchy process (AHP) Represented as: In the formula This indicates the relative importance of the first evaluation dimension compared to the second evaluation dimension. , Similarly, the importance value can be determined using the Saaty scaling system. By solving the judgment matrix using the eigenvalue method, the first weights of the three evaluation dimensions can be obtained. Using these first weights to weight the three evaluation dimensions yields the second weights corresponding to the sub-parameters. The formula for calculating the second weights is: In the formula Indicates the first The second weight corresponding to each sub-parameter ~ These represent the first weight of the three evaluation dimensions.

[0039] For the standard word root, the second weight of each sub-parameter is normalized, and then the standard data of each sub-parameter is weighted and summed using the normalized second weight. The weighted sum is used as the fusion data of the standard word root.

[0040] The normalized second weight is the ratio of the second weight of a certain sub-parameter to the sum of the second weights of all sub-parameters under that standard word root. The fused data of the standard word root can be represented as: In the formula This represents the fusion data of the standard word root. This indicates the number of sub-parameters under this standard root word. This indicates its second weight. Subscripts are added to the fused data of standard word roots for indexing. , That is to represent the first The fusion data of standard word roots.

[0041] In this step, the data quality of the sub-parameters is comprehensively evaluated from three dimensions: source credibility, historical stability, and collection method. This encompasses data authority, data stability, and the degree of automation in collection. At the same time, AHP is used to compare the three evaluation dimensions in pairs to scientifically obtain the weights of each indicator, avoiding subjective and arbitrary assignment and ensuring the mathematical rationality of the weight calculation. This ensures that the final fused data can effectively reflect the optimal estimate of the true parameters (i.e., theoretical data) of the standard word root under the current data and knowledge conditions, thus providing a reasonable central position for subsequent probability distribution estimation.

[0042] S3: Randomly select several sets of standard word roots to construct a validation set. For each set of standard word roots in the validation set, use statistical methods to repeatedly sample all standard data under that standard word root to generate the probability distribution of the theoretical data of the standard word root. Calculate the corresponding sampling mean based on this distribution to estimate the probability distribution of the standard word root.

[0043] Step S3 includes: S301: Randomly select several groups from all standard word roots to construct a validation set (e.g., select 30%). For each standard word root in the validation set, use the Bootstrap resampling method to sample the standard data of its sub-parameters with replacement.

[0044] S302: For each sample with replacement, the number of samples is equal to the number of sub-parameters under the standard root word, and the frequency of the sampled data is used to weight and sum them to generate weighted sampled data; S303: Repeatedly perform sampling with replacement to generate several sets of weighted sampling data. Use a preset distribution model to fit the several sets of weighted sampling data to obtain their probability distribution, so as to realize the probability distribution estimation of the theoretical data of the standard word root. And calibrate the mean of all weighted sampling data as the sampling mean.

[0045] Understandably, for a given standard word root, its data is calculated from its corresponding sub-parameters. Essentially, it's a calculated, unknown "theoretical data," and the fused data calculated above reflects the most probable value of this theoretical data. Even for the same parameter (i.e., the standard word root), each sub-parameter can differ due to various factors such as collection errors, measurement biases, and environmental influences. Setting this parameter to a fixed value directly would fail to reflect these random variations in actual data collection. Estimating the probability distribution of the standard word root is essentially estimating the probability distribution of this unknown theoretical data, thereby reflecting the true fluctuation range and probabilistic characteristics of the parameter and enhancing the reliability of the data results.

[0046] For a given standard word root, it is based on a dataset. The form can be expressed as: When sampling with replacement, first define the number of sampling repetitions (e.g., 1000 times), and collect samples with replacement each time. Next, generate a new dataset. That is, sampling is performed sequentially with replacement each time. We take 1000 samples and repeat the process 1000 times. The sampled dataset can be represented as: subscript in formula An index representing the number of repetitions. Indicates the index of the sampled data. They represent the first The first time sampling with replacement Each sample data point and the number of times it appears. The weighted sample data of this dataset. The calculation formula is: Repeating sampling with replacement 1000 times yields 1000 sets of weighted sampling data, i.e.: The probability distribution can be obtained by fitting these weighted sample data.

[0047] Understandably, the Bootstrap resampling method obtains data by sampling with replacement from standard data. Therefore, the sampled data is not independently newly collected data, but rather a simulated sample constructed based on the standard data to reflect sampling errors and uncertainties. It is a randomly generated "virtual sample" used to estimate the distribution of statistics. When calculating the sample mean, unlike the previous calculation of the "mean over a period of time" along the time axis, the mean calculated here is essentially the "mean of different weighted sampled data at the same moment," which is equivalent to the central position of the probability distribution formed by several sets of weighted sampled data.

[0048] In this step, by using Bootstrap resampling to generate multiple virtual samples and simulating sampling errors, the statistical characteristics and variation range of the standard root word data can be more comprehensively characterized. This allows for subsequent carbon footprint analysis using knowledge graphs, enabling not only mean estimation but also quantification of errors and confidence intervals. The final output parameter, "carbon footprint," is transformed into a probability distribution, fully reflecting uncertainties related to data sources, collection methods, and historical stability. This not only enhances the scientific rigor and interpretability of the analysis results but also provides quantitative evidence for the dynamic changes and risk analysis of carbon footprints. This allows decision-makers to gain a more comprehensive understanding of the actual fluctuation range of carbon emissions, formulate more reasonable management and control strategies, and better adapt to various complex scenarios.

[0049] S4: Using the fused data as a reference, calculate the relative error between the sample mean and the fused data. Based on the calculation results, determine whether the probability distribution setting corresponding to the standard word roots is reasonable. Then, use the judgment results as a reference to estimate the probability distribution of all standard word roots.

[0050] The calculation formulas for the fused data and the weighted sampled data after resampling reveal that their underlying calculation logic is the same. The difference lies in the central tendency estimate: the fused data is a fixed value, representing the central tendency estimate with the highest confidence level among the multi-source parameter values, considering factors such as data credibility, source weights, and quality differences. The weighted sampled data, on the other hand, consists of multiple sets used to form a statistical distribution, reflecting the volatility and uncertainty of the actual data. In short, the fused data should theoretically be the central point of the statistical distribution. Therefore, it can be used as a reference for probability distribution estimation. If the fused data is the same as or close to the sampled mean, the probability distribution estimate of the standard word root can be considered reasonable.

[0051] The logic for determining whether the probability distribution setting corresponding to the standard word root is reasonable is as follows: Along the time axis, the relative error between the downsampled mean and the fused data at each time point is calculated, the corresponding mean error is calculated, and the mean error is compared with a preset error threshold. The formula for calculating the mean error is as follows: In the formula Indicates the first The mean error of each standard word root, Represents the sample mean. , These represent the sample mean and the set elements within the fused data, respectively. The smaller the error mean, the more similar the trends of the sample mean and the fused data are over time, and the closer they are to each other.

[0052] If the mean error does not exceed the error threshold, the probability distribution of the standard word root is considered to be reasonable. If the mean error exceeds the error threshold, the probability distribution of the standard word root is considered to be unreasonable.

[0053] The logic for using the judgment result as a reference to estimate the probability distribution of all standard word roots is as follows: For each probability distribution in the validation set with an unreasonable standard word root, change the fitting model used when fitting the distribution and re-estimate the probability distribution until the mean error does not exceed the error threshold. The fitting models used for all standard word roots in the statistical validation set are sorted in descending order of frequency of use. When estimating the probability distribution of all standard word roots, the distributions are fitted sequentially in descending order of the sorting list to ensure that the mean error does not exceed the error threshold.

[0054] Distribution models include normal distribution, kernel density estimation, Poisson distribution, and empirical distribution.

[0055] In this step, by comparing the error between the fused data and the sampling mean, problems in the probability distribution fitting can be identified in a timely manner, ensuring the scientific nature of the model output. When the error exceeds the standard, the distribution model can be refitted, promoting iterative updates of the model and improving the robustness and adaptability of the overall solution.

[0056] S5: Construct a knowledge graph about carbon footprint by treating each standard term as a node. Input the probability distribution and distribution parameters of the theoretical data for each standard term into the knowledge graph to expand node attributes, thereby achieving probability distribution estimation of the carbon footprint of the product being analyzed. Different distribution models include different distribution parameters. Taking the normal distribution as an example, its corresponding distribution parameters include mean, standard deviation, etc. The specific types of distribution parameters are determined by the distribution model used when estimating the probability distribution of the standard term.

[0057] Step S5 includes: S501: Using a multi-source parameter set as input to the knowledge graph, each standard word root is used as a node of the knowledge graph, and the probability distribution and distribution parameters of the standard word root are used as extended attributes of the node. S502: Based on the calculation process of the carbon footprint of the product to be analyzed, define the causal relationship between nodes and form the edges of the knowledge graph; S503: Introducing inference algorithms based on probabilistic graphical models (such as Bayesian networks and Markov random fields), and combining the causal relationships between nodes, a probability distribution of the carbon footprint of the product to be analyzed is generated and used as the output of the knowledge graph. Since using knowledge graphs to generate product carbon footprints is a mature existing technology, such as using tools like CiteSpace for visualization, the specific calculation principles will not be elaborated here.

[0058] In existing technologies, carbon footprint calculations are typically deterministic models, which struggle to capture data uncertainty and parameter dependencies, resulting in limited accuracy. This step, by combining probability distribution attributes and probabilistic graphical model reasoning with causal edges to form a joint probability distribution, scientifically characterizes uncertainty and dependencies, thereby improving computational accuracy and robustness.

[0059] Please see Figure 5 The present invention also provides a carbon footprint acquisition system based on knowledge graphs, for executing the carbon footprint acquisition method described above, specifically including: a data acquisition module, a data processing module, a distribution fitting module, a data judgment module, and a graph construction module.

[0060] The data acquisition module is used to collect multi-source parameter sets related to the carbon footprint of the product to be analyzed, perform semantic recognition and normalization on the parameter names of each sub-parameter in the multi-source parameter set, generate several sets of standard word roots based on industry standards, and classify each sub-parameter into the standard word root with the highest semantic similarity. The data processing module is used to preprocess the raw data of each sub-parameter under the same standard root word to generate standard data. It uses source credibility, historical stability and collection method as evaluation dimensions, and generates credibility weights for different sub-parameters based on the analytic hierarchy process. The credibility weights are then used to weight the standard data and generate fused data. The distribution fitting module is used to randomly sample and construct a validation set from all standard word roots. For each group of standard word roots in the validation set, the probability distribution is estimated in turn. Statistical methods are used to repeatedly sample the standard data to generate the probability distribution, and the corresponding sample mean is obtained simultaneously. The data judgment module is used to calculate the relative error between the sample mean and the fused data using the fused data as a reference benchmark. Based on the calculation results, it judges whether the probability distribution setting corresponding to the standard word roots is reasonable, and then uses the judgment results as a reference to estimate the probability distribution of all standard word roots. The graph construction module is used to construct a knowledge graph about carbon footprint by using each standard word root as a node. The probability distribution and distribution parameters of the theoretical data of each standard word root are input into the knowledge graph to expand the node attributes, so as to realize the probability distribution estimation of the carbon footprint of the product to be analyzed.

[0061] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0062] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0063] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0064] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A carbon footprint acquisition method based on knowledge graphs, characterized in that, The specific steps include: S1: Collect a multi-source parameter set related to the carbon footprint of the product to be analyzed within a certain time interval, perform semantic recognition on the parameter names of each sub-parameter in the multi-source parameter set, generate several sets of standard word roots based on the industry standards of the product to be analyzed, and classify each sub-parameter after semantic recognition into the standard word root with the highest semantic similarity. S2: The original data of each sub-parameter under the same standard root word are preprocessed to generate standard data. The source credibility, historical stability and collection method of the original data are used as evaluation dimensions. The analytic hierarchy process is used to generate credibility weights for different sub-parameters. The credibility weights are used to weight the standard data and generate fused data. S3: Randomly select several standard word roots to construct a validation set. For each standard word root in the validation set, use statistical methods to repeatedly sample all standard data under that standard word root to generate the probability distribution of the theoretical data of the standard word root, and calculate the corresponding sampling mean based on the probability distribution. Step S3 includes: S301: Randomly select several standard word roots from all standard word roots to construct a validation set. For each standard word root in the validation set, use the Bootstrap resampling method to sample the standard data of each sub-parameter of the standard word root with replacement. S302: For each sample with replacement, the number of samples is equal to the number of sub-parameters under the standard root word, and the sampled data is weighted and summed using the frequency of occurrence of the sampled data to generate weighted sampled data; S303: Repeatedly perform sampling with replacement to generate several sets of weighted sampling data. Use a preset distribution model to fit all sets of weighted sampling data to obtain the probability distribution of the weighted sampling data, so as to estimate the probability distribution of the theoretical data of the standard word root. The mean of all weighted sampling data is calibrated as the sampling mean. S4: For each standard word root, using the fused data of each standard word root as a reference benchmark, calculate the relative error and mean error between the sampling mean and the fused data. Based on the mean error, determine whether the probability distribution setting corresponding to the theoretical data of the standard word root is reasonable. Then, based on the judgment result, estimate the probability distribution of all standard word roots. S5: Construct a knowledge graph about the carbon footprint of the product to be analyzed by using each standard word root as a node, and input the probability distribution and distribution parameters of the theoretical data of each standard word root into the knowledge graph to expand the node attributes, so as to realize the probability distribution estimation of the carbon footprint of the product to be analyzed.

2. The carbon footprint acquisition method based on knowledge graphs according to claim 1, characterized in that: The specific process of step S1 is as follows: S101: Collect multi-source parameter groups, use a pre-trained natural language model to perform semantic recognition and normalization on the parameter names of each sub-parameter in the multi-source parameter groups, and map the normalized parameter names into high-dimensional semantic vectors. S102: Based on carbon footprint industry standards, several sets of standard word roots are generated, and a pre-trained natural language model is used to map the standard word roots into high-dimensional semantic vectors. S103: For each sub-parameter, calculate the semantic similarity between the parameter name and each standard word root based on the parameter name and the high-dimensional semantic vector of the standard word root, and classify the sub-parameter into the standard word root with the highest semantic similarity.

3. The carbon footprint acquisition method based on knowledge graphs according to claim 1, characterized in that: The logic for the data preprocessing is as follows: For each standard word root, the original data of each sub-parameter is converted to the same dimension in turn, and the original data of each sub-parameter is aligned based on the time axis. For each sub-parameter under the standard word root, the mean and standard deviation of the original data of each sub-parameter in the time axis direction are calculated in turn, and statistical methods are used to identify and remove outliers in the original data. For each sub-parameter under the standard word root, the K-nearest neighbor method is used to fill in missing values, and the results are used as the standard data for that sub-parameter after the filling is completed.

4. The carbon footprint acquisition method based on knowledge graphs according to claim 3, characterized in that: The credibility weight includes a first weight and a second weight, and the logic for generating the fused data is as follows: For each standard word root, the data source and collection method of each sub-parameter are obtained in sequence. The collection method includes automatic and manual collection. The collection method is labeled as a binary label, with a value of 0 indicating manual collection and a value of 1 indicating automatic collection. For each sub-parameter under the standard root, scores are assigned to different data sources to quantify source credibility, the normalized standard deviation of the standard data for each sub-parameter in the time axis direction is calculated to quantify historical stability, and the collection method is quantified based on the binary label of the collection method. For each sub-parameter under the standard word root, the three evaluation dimensions of source credibility, historical stability and collection method are compared in pairs based on the analytic hierarchy process to construct a judgment matrix. Based on the judgment matrix, feature weights are calculated to obtain the first weight of different evaluation dimensions. Then, the three evaluation dimensions are weighted based on the first weight to obtain the second weight corresponding to the sub-parameter. For each standard word root, the second weight of each sub-parameter belonging to that standard word root is normalized, and then the standard data of each sub-parameter is weighted and summed using the normalized second weight. The weighted sum is used as the fusion data of that standard word root.

5. The carbon footprint acquisition method based on knowledge graphs according to claim 4, characterized in that: The logic for determining whether the probability distribution setting corresponding to the standard word root is reasonable is as follows: Along the time axis, the relative error between the downsampled mean and the fused data at each time point is calculated, the corresponding mean error is calculated, and the mean error is compared with a preset error threshold. If the mean error does not exceed the error threshold, the probability distribution of the standard word root is considered to be reasonable. If the mean error exceeds the error threshold, the probability distribution of the standard word root is considered to be unreasonable.

6. The carbon footprint acquisition method based on knowledge graphs according to claim 5, characterized in that: The logic for using the judgment result as a reference to estimate the probability distribution of all standard word roots is as follows: For each probability distribution in the validation set with an unreasonable standard word root, change the fitting model used when fitting the standard word root distribution and re-estimate the probability distribution until the mean error does not exceed the error threshold. The fitting models used for all standard word roots in the statistical validation set are sorted in descending order of frequency of use. When estimating the probability distribution of all standard word roots, the distribution is fitted sequentially in descending order of the sorting table to ensure that the mean error corresponding to the standard word root does not exceed the error threshold.

7. The carbon footprint acquisition method based on knowledge graphs according to claim 6, characterized in that: The distribution models include normal distribution, kernel density estimation, Poisson distribution, and empirical distribution.

8. The carbon footprint acquisition method based on knowledge graphs according to claim 6, characterized in that: Step S5 includes: S501: Using a multi-source parameter set as input to the knowledge graph, each standard word root is used as a node of the knowledge graph, and the probability distribution and distribution parameters of the standard word root are used as extended attributes of the node. S502: Based on the calculation process of the carbon footprint of the product to be analyzed, define the causal relationship between nodes and form the edges of the knowledge graph; S503: Introduces a reasoning algorithm based on a probabilistic graphical model, combining the causal relationships between nodes to generate a probability distribution of the theoretical data of the carbon footprint of the product to be analyzed and output it as a knowledge graph.

9. A carbon footprint acquisition system based on knowledge graphs, characterized in that: The carbon footprint acquisition system is used to execute the carbon footprint acquisition method as described in any one of claims 1-8, specifically including: The data acquisition module is used to collect multi-source parameter groups related to the carbon footprint of the product to be analyzed, perform semantic recognition and normalization on the parameter names of each sub-parameter in the multi-source parameter group, generate several sets of standard word roots based on industry standards, and classify each sub-parameter into the standard word root with the highest semantic similarity. The data processing module is used to preprocess the raw data of each sub-parameter under the same standard root word to generate standard data. The module uses source credibility, historical stability and collection method as evaluation dimensions, and generates credibility weights for different sub-parameters based on the analytic hierarchy process. The credibility weights are then used to weight the standard data and generate fused data. The distribution fitting module is used to randomly sample and construct a validation set from all standard word roots. The probability distribution of each set of standard word roots in the validation set is estimated sequentially. The standard data of each set of standard word roots is repeatedly sampled using statistical methods to generate a probability distribution, and the sampling mean corresponding to each set of standard word roots is obtained simultaneously. The data judgment module is used to calculate the relative error between the sample mean and the fused data using the fused data as a reference benchmark, and to judge whether the probability distribution setting corresponding to the set of standard word roots is reasonable based on the calculation result. Then, the judgment result is used as a reference to estimate the probability distribution of all standard word roots. The graph construction module is used to construct a knowledge graph about carbon footprint by using each standard word root as a node, and inputs the probability distribution and distribution parameters of the theoretical data of each standard word root into the knowledge graph to expand the node attributes, so as to realize the probability distribution estimation of the carbon footprint of the product to be analyzed.