A method for quality assessment and correction of industrial material product life cycle carbon emission data
Patent Information
- Application Number
- CN202511310944.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2045-09-15
AI Technical Summary
[0013] The beneficial effects of this invention are: it can simply and efficiently locate and correct abnormal data in the carbon emission database, achieve data cleaning, maintain the integrity of the carbon emission database, and ensure the accuracy of carbon emission calculation results.
Smart Images

Figure CN121387867B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a technology for processing carbon emission data throughout the life cycle of industrial materials products, and more particularly to a method for quality assessment and correction of carbon emission data throughout the life cycle of industrial materials products. Background Technology
[0002] Industrial manufacturing is one of the major sources of carbon emissions today, and global regulation of carbon emissions from industrial products is receiving increasing attention. Lifecycle carbon emission calculations for industrial products are of great significance to the economic benefits and social responsibility of various industries. However, industrial product manufacturing processes are complex, defining carbon emission boundaries for products and involving product processes, materials, and equipment. Therefore, a product lifecycle carbon emission database is needed as the foundation for carbon emission calculations. The construction of this database requires the integration of a large amount of data, which may include a certain amount of outlier data, affecting the accuracy of carbon emission calculations. Therefore, a data correction and quality assessment method for industrial product lifecycle data is needed. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method for quality assessment and correction of carbon emission data throughout the life cycle of industrial materials products.
[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for quality assessment and errata correction of carbon emission data throughout the life cycle of industrial materials products, the method comprising the following steps: Step 1: Product lifecycle data collection and feature building Collect and organize industrial material product lifecycle data, and construct carbon emission characteristics of industrial material product lifecycle to describe key data at each stage of the lifecycle, providing a foundation for subsequent carbon emission calculations; Step 2: Data Quality Inspection and Assessment Data statistics and feature learning methods are used to detect the completeness, consistency, timeliness, uniqueness, and reliability of product lifecycle data, and to locate abnormal data. Step 3: Data Cleaning and Correction By utilizing known data characteristics and identified anomalies, a composite method consisting of data cleaning techniques such as data modeling is employed to perform real-time data cleaning and correction.
[0005] Furthermore, in step 2, the integrity check is to check whether the data at each lifecycle stage is complete to ensure that there are no missing values; the consistency check is to check whether the data from different sources are consistent to avoid data conflicts; the timeliness check is to check the timestamp and update time of the data to ensure that the data is not outdated; the uniqueness check is to check for duplicate records; and the reliability check is to verify the source and collection method of the data to ensure that the data source is reliable.
[0006] Furthermore, in step 2, a neural network model is trained on the data, and the data quality is evaluated using a score. Further, a combination of Isolation Forest and DBSCAN is used to perform data cleaning and correction.
[0007] Furthermore, the steps for data cleaning and correction using a combination of isolated forest and DBSCAN are as follows: 1) Use isolated forests for preliminary anomaly detection Train an isolated forest model and calculate an anomaly score for each data point; mark points with anomaly scores higher than a threshold as anomalies; 2) Density analysis using DBSCAN DBSCAN was used to perform density analysis on the dataset to obtain the category label for each data point, including core points, boundary points, and noise points. Points marked as noise points were considered as anomalous data. 3) Combining the results of Isolation Forest and DBSCAN If a data point is marked as an anomaly in both the Isolation Forest and DBSCAN, the probability of that point being an anomaly is high, and the point is deleted directly. For data points marked as anomalies in the Isolation Forest but as boundary points in DBSCAN, interpolation or regression methods are used for correction. Data points considered normal by both the Isolation Forest and DBSCAN are left unprocessed. 4) Final revision Remove outlier data points; correct boundary points using interpolation or regression; use KNN, regression, or interpolation methods to correct the remaining outlier data. 5) Re-verify Reassess outliers: Improve data quality by re-examining the corrected data using methods such as Isolation Forest, DBSCAN, or others.
[0008] Furthermore, the specific implementation method of the isolated forest is as follows: Isolation forests construct random trees and use path length to evaluate the anomaly of each data point. The anomaly scoring formula is as follows: Formula 1, in, It is the anomaly score of data point x, ranging from [0,1]. The closer it is to 1, the higher the probability of an anomaly. It is the average path length of data point x across all trees; : represents the path length of data point x in a separate isolated tree. For flexible model adjustments; It is the normalization coefficient, defined as: Formula 2, Where H(n): the harmonic number of the nth term is approximately... ,for Euler's constant is approximately 0.577. The specific implementation method of DBSCAN is as follows: Potential outliers identified by the isolated forest are further analyzed using DBSCAN and classified into core points, boundary points, and noise points. For each point in the dataset, a neighborhood N(p) is defined for each data point p, and its neighborhood is defined as follows: Formula 3, This is the distance between points p and q, which can be expressed using Euclidean distance or other suitable distance metrics. The following formula applies to the classification of data points: Formula 4 The neighborhood N(p) of a point is defined as satisfying The set of points; Define a point p as the center and a radius of . The area; The judgment criteria are: Formula 5 MinPts is the core point density standard, and its value is selected as the data dimension + 1; Core point determination: If Then p is the core point; Boundary point determination: If If a core point exists in the neighborhood, then it is a boundary point; Noise point determination: If p is neither a core point nor a boundary point, then it is a noise point; Then, neighborhood mean interpolation is used to correct the data: Formula Six Where: N(x): the neighborhood of point x, that is, the set of points that are within a certain distance from x; is the value of the data point in the neighborhood; |N(x)| is the number of neighborhood points; for The interpolation correction results.
[0009] Furthermore, step 1 of the product lifecycle includes a raw material acquisition stage, in which basic information data of the raw materials are collected and features are constructed.
[0010] Furthermore, step 1 of the product lifecycle includes a manufacturing stage, in which key data from the manufacturing process are collected and features are constructed.
[0011] Furthermore, step 1 of the product lifecycle includes a transportation process stage, during which product transportation data is collected and features are constructed.
[0012] Furthermore, step 1 of the product lifecycle includes a recycling process phase, during which data is collected during the recycling process.
[0013] The beneficial effects of this invention are: it can simply and efficiently locate and correct abnormal data in the carbon emission database, achieve data cleaning, maintain the integrity of the carbon emission database, and ensure the accuracy of carbon emission calculation results. Attached Figure Description
[0014] Figure 1 This is a flowchart illustrating the execution process of the present invention; Figure 2 This is a flowchart of the isolated forest method of the present invention; Figure 3 This is a flowchart of the DBSCAN method of the present invention; Figure 4 This is the result of running the isolated forest method in the implementation example; Figure 5 This is the result of DBSCAN processing abnormal data in the example. Figure 6 This is the result of the example using linear interpolation to fill in missing data. Detailed Implementation
[0015] The present invention will be further described in detail below with reference to embodiments, such as... Figure 1 As shown, a method for quality assessment and errata correction of carbon emission data throughout the life cycle of industrial materials products is provided, the method comprising the following steps: Step 1: Product lifecycle data collection and feature building Collect and organize industrial material product lifecycle data, and construct carbon emission characteristics of industrial material product lifecycle to describe key data at each stage of the lifecycle, providing a foundation for subsequent carbon emission calculations; Step 2: Data Quality Inspection and Assessment Data statistics and feature learning methods are used to detect the completeness, consistency, timeliness, uniqueness, and reliability of product lifecycle data, and to locate abnormal data. Step 3: Data Cleaning and Correction By utilizing known data characteristics and identified anomalies, a composite method consisting of data cleaning techniques such as data modeling is employed to perform real-time data cleaning and correction.
[0016] In step 2, the integrity check is to check whether the data at each lifecycle stage is complete and ensure that there are no missing values; the consistency check is to check whether the data from different sources are consistent and avoid data conflicts; the timeliness check is to check the timestamp and update time of the data to ensure that the data is not outdated; the uniqueness check is to check for duplicate records; and the reliability check is to verify the source and collection method of the data to ensure that the data source is reliable.
[0017] The five standardized features described above constitute a multi-dimensional input vector, serving as the input to the neural network evaluation model. The model learns the weights of each feature in the overall quality score through training and outputs the overall data quality score. During training, mean squared error (MSE) is used as the loss function. The MSE is calculated by comparing the ideal scores of each indicator with the actual scores of the data after neural network training, and then used to train the neural network. This approach not only achieves a close integration between the five data evaluation indicators and the model structure but also ensures that the model has good adaptability and interpretability to different types of data features.
[0018] Taking steel products as an example, the material data for each stage is first obtained as shown in Table 1 below: .
[0019] The collected data was normalized, and the database was jointly detected using the Isolation Forest and DBSCAN methods. Missing, anomalous, and duplicate data are shown below. For the Isolation Forest method, refer to Formulas 1 and 2. Due to the randomness inherent in the construction process of the Isolation Forest, the results of each run will vary, with some anomalous data not being detected and others being misjudged as anomalous. However, it is guaranteed that the majority of rows in the results will contain anomalous data. The results of the Isolation Forest method are shown below. Figure 4 .
[0020] DBSCAN method: The results are as follows Figure 5 The DBSCAN algorithm defines a radius region. Using the core point density standard MinPts, normal and abnormal data can be distinguished, as shown in formulas three through five. In this application, With a value set to 300 and MinPts set to 5, rows containing anomalous data are marked as -1, which is reflected in the Cluster column of the results. Twenty rows containing anomalous data have been marked with red boxes.
[0021] The entire steel product dataset contained a total of 980 missing data points, 490 outliers, and 1470 duplicate data points. Missing data accounted for approximately 5.5%, and outliers accounted for approximately 2.7%. Subsequently, data cleaning methods were used to correct the errors, such as interpolation imputation. See Formula Six for the results. Figure 6 The areas marked in red were originally empty and were filled using interpolation. By employing a hybrid approach combining Isolation Forest and DBSCAN, 450 outlier data points were successfully detected and corrected, achieving a detection rate of 91.8%. Furthermore, the proposed data cleaning method successfully removed 450 outlier data points, achieving a removal rate of 100%.
[0022] The above content is only used to illustrate the technical solution of the present invention. Simple modifications or equivalent substitutions made by those skilled in the art to the technical solution of the present invention do not depart from the essence and scope of the technical solution of the present invention.
Claims
1. A method for quality assessment and correction of industrial material product lifecycle carbon emission data, characterized in that: The method includes the following steps: Step 1: Product lifecycle data collection and feature building Collect and organize industrial material product lifecycle data, and construct carbon emission characteristics of industrial material product lifecycle to describe key data at each stage of the lifecycle, providing a foundation for subsequent carbon emission calculations; Step 2: Data Quality Inspection and Assessment Data statistics and feature learning methods are used to detect the completeness, consistency, timeliness, uniqueness, and reliability of product lifecycle data, and to locate abnormal data. Step 3: Data Cleaning and Correction Based on data characteristics, abnormal data is located, and a composite method consisting of data modeling and data cleaning techniques is used to clean and correct the data in real time. The method uses a combination of isolated forest and DBSCAN to achieve data cleaning and correction. The specific implementation method of the isolated forest is as follows: Isolation forests construct random trees and use path length to evaluate the anomaly of each data point. The anomaly scoring formula is as follows: Formula 1, in, It is the anomaly score of data point x, ranging from [0,1]. The closer it is to 1, the higher the probability of an anomaly. It is the average path length of data point x across all trees; : This represents the path length of data point x in a separate isolated tree, used for flexible model adjustments; It is the normalization coefficient, defined as: Formula 2, Where H(n): the harmonic number of the nth term is approximately... , is Euler's constant, which is equal to 0.577; The specific implementation method of DBSCAN is as follows: Potential outliers identified by the isolated forest are further analyzed using DBSCAN and classified into core points, boundary points, and noise points. For each point in the dataset, a neighborhood N(p) of data point p is defined, and its neighborhood is defined as follows: Formula 3, is the distance between points p and q. The following formula applies to the classification of data points: Formula 4 The neighborhood N(p) of a point is defined as satisfying The set of points; m represents the number of feature dimensions of the data points; Define a point p as the center and a radius of . area; The judgment criteria are: Formula 5 MinPts is the core point density standard, and its value is selected as the data dimension + 1; Core Point Judgment: If Then p is the core point; Boundary point determination: If If a core point exists in the neighborhood, then it is a boundary point; Noise point determination: If p is neither a core point nor a boundary point, then it is a noise point; Then, neighborhood mean interpolation is used to correct the data: Formula Six Where: N(x): the neighborhood of point x, that is, the set of points that are within a certain distance from x; is the value of the data point in the neighborhood; |N(x)| is the number of neighborhood points; for The interpolation correction results.
2. The method for quality assessment and errata correction of carbon emission data throughout the life cycle of industrial materials products according to claim 1, characterized in that: In step 2, Integrity checks verify the completeness of data at each lifecycle stage, ensuring that no values are missing. Consistency checks verify whether data from different sources are consistent, thus avoiding data conflicts. Timeliness testing involves checking the timestamps and update times of the data to ensure that the data is not outdated. Uniqueness detection checks for duplicate records; Reliability testing involves verifying the source and collection method of the data to ensure its reliability.
3. The method for quality assessment and errata correction of carbon emission data throughout the life cycle of industrial materials products according to claim 1, characterized in that: In step 2, a neural network model is trained on the data, and the data quality is evaluated by a score.
4. The method for quality assessment and errata correction of carbon emission data throughout the life cycle of industrial materials products according to claim 1, characterized in that: The steps for data cleaning and correction using a combination of isolated forest and DBSCAN are as follows: 1) Use isolated forests for preliminary anomaly detection Train an isolated forest model and calculate an anomaly score for each data point; mark points with anomaly scores higher than a threshold as anomalies; 2) Density analysis using DBSCAN DBSCAN was used to perform density analysis on the dataset to obtain the category label for each data point, including core points, boundary points, and noise points. Points marked as noise points were considered as anomalous data. 3) Combining the results of Isolation Forest and DBSCAN If a data point is marked as an anomaly in both the Isolation Forest and DBSCAN, the probability of that point being an anomaly is high, and the point is deleted directly. For data points marked as anomalies in the Isolation Forest but as boundary points in DBSCAN, interpolation or regression methods are used for correction. Data points considered normal by both the Isolation Forest and DBSCAN are left unprocessed. 4) Final revision Remove outlier data points; correct boundary points using interpolation or regression; use KNN, regression, or interpolation to correct remaining outlier data. 5) Re-verify Reassess outliers: Improve data quality by re-examining the corrected data using methods such as Isolation Forest and DBSCAN.
5. The method for quality assessment and errata correction of carbon emission data throughout the life cycle of industrial materials products according to claim 1, characterized in that: Step 1 of the product lifecycle includes the raw material acquisition stage, in which basic information data of raw materials are collected and features are constructed.
6. The method for quality assessment and errata correction of carbon emission data throughout the life cycle of industrial materials products according to claim 1, characterized in that: Step 1 of the product lifecycle includes the manufacturing stage, in which key data from the manufacturing process are collected and features are constructed.
7. The method for quality assessment and errata correction of carbon emission data throughout the life cycle of industrial materials products according to claim 1, characterized in that: Step 1 of the product lifecycle includes the transportation process stage, during which product transportation data is collected and features are constructed.
8. The method for quality assessment and errata correction of carbon emission data throughout the life cycle of industrial materials products according to claim 1, characterized in that: Step 1 of the product lifecycle includes a recycling process phase, during which data is collected during the recycling process.
Citation Information
Patent Citations
Big data anomaly detection method and device, equipment, storage medium and product
CN118643444A
Risk address identification method and apparatus, electronic device, and storage medium
WO2025179836A1