Road engineering carbon emission analysis method based on multi-source heterogeneous data fusion
By using a multi-source integrated sensor system and multi-source heterogeneous data processing technology, the problems of full life-cycle coverage and data fusion in carbon emission analysis of road engineering have been solved, enabling accurate identification and accounting of carbon emission factors and adapting to intelligent carbon emission accounting under complex working conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUNAN COMM RES INST CO LTD
- Filing Date
- 2025-11-06
- Publication Date
- 2026-04-10
AI Technical Summary
Existing carbon emission analysis methods for road engineering lack full life-cycle coverage and the ability to integrate multi-source heterogeneous data, making it difficult to fully explore the potential correlations between data. Furthermore, the methods for identifying and correcting abnormal data of carbon emission factors are limited, data quality is difficult to guarantee, and the dynamic optimization capability of the accounting model is limited.
A multi-source integrated sensor system is used to collect and preprocess multi-source heterogeneous carbon emission influencing factors throughout the entire life cycle of road construction projects. Through anomaly identification and correction of multi-source heterogeneous data, multi-source heterogeneous standard data of carbon emission factors are generated, and cross-modal global connection fusion is performed to establish global fusion feature data of carbon emission factors. Finally, an optimized carbon emission accounting relationship model is generated.
It achieves comprehensive collection and systematic analysis of carbon emission-related influencing factors, improves the fusion efficiency and accuracy of multi-source data, can adapt to intelligent carbon emission accounting under complex operating conditions, and generates highly consistent and reliable carbon emission accounting results.
Smart Images

Figure CN121073006B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of carbon emission accounting, and particularly relates to a road engineering carbon emission analysis method based on multi-source heterogeneous data fusion. BACKGROUND
[0002] With the acceleration of urbanization, the number and scale of road construction projects are increasing, and engineering construction activities not only promote social and economic development, but also bring about significant carbon emissions. Road construction involves construction equipment energy consumption, building material production and transportation, construction site environment and other aspects, each of which directly or indirectly affects the total carbon emissions. However, the existing road engineering carbon emission analysis method lacks comprehensive collection and analysis of the whole life cycle of road engineering, and the fusion ability of multi-source heterogeneous data is insufficient, which is difficult to fully tap the potential correlation between data; the means of abnormal data identification and correction of carbon emission factors is single, the data quality is difficult to guarantee, and the dynamic optimization ability of the accounting model is limited, which is difficult to meet the intelligent carbon emission accounting demand under complex working conditions. SUMMARY
[0003] Therefore, the present application provides a road engineering carbon emission analysis method based on multi-source heterogeneous data fusion to solve at least one of the above technical problems.
[0004] To achieve the above purpose, a road engineering carbon emission analysis method based on multi-source heterogeneous data fusion comprises the following steps:
[0005] Step S1: using a multi-source integrated sensor system to collect and preprocess multi-source heterogeneous carbon emission influencing factors of road construction projects throughout the whole life cycle, to generate multi-source heterogeneous carbon emission factor data;
[0006] Step S2: performing multi-source heterogeneous data anomaly identification and correction processing on the multi-source heterogeneous carbon emission factor data, to generate multi-source heterogeneous standard carbon emission factor data;
[0007] Step S3: performing global connection fusion processing on the multi-source heterogeneous standard carbon emission factor data, to generate global fusion feature data of carbon emission factors;
[0008] Step S4: establishing an optimized relationship model for actual working condition carbon emission accounting based on the global fusion feature data of carbon emission factors, to generate an optimized carbon emission accounting relationship model; and performing carbon emission intelligent accounting operation on road construction projects based on the optimized carbon emission accounting relationship model.
[0009] Further, the multi-source integrated sensor system comprises a construction equipment energy consumption sensor module, a building material monitoring sensor module and a construction site monitoring sensor module, and step S1 comprises the following steps:
[0010] Step S11: Collecting construction equipment energy consumption data of the road construction project by using the construction equipment energy consumption sensor module to obtain construction equipment energy consumption data;
[0011] Step S12: Collecting carbon emission data of the road construction material production and transportation stage of the road construction project by using the building material monitoring sensor module to obtain material production and transportation carbon emission data;
[0012] Step S13: Monitoring and processing the construction site environment data of the road construction project by using the construction site monitoring sensor module to obtain construction environment data;
[0013] Step S14: Processing the construction equipment energy consumption data, material production and transportation carbon emission data, and construction environment data for multi-source heterogeneous integration of carbon emission factors in the whole life cycle to obtain preliminary carbon emission factor multi-source heterogeneous data;
[0014] Step S15: Preprocessing the preliminary carbon emission factor multi-source heterogeneous data to generate carbon emission factor multi-source heterogeneous data.
[0015] Further, step S2 includes the following steps:
[0016] Step S21: Analyzing the heterogeneous type characteristics according to the carbon emission factor multi-source heterogeneous data to generate carbon emission factor heterogeneous type characteristic data;
[0017] Step S22: Establishing a data anomaly identification relationship model of carbon emission factor heterogeneous types based on the carbon emission factor heterogeneous type characteristic data to obtain a carbon emission factor heterogeneous type anomaly identification model;
[0018] Step S23: Transferring the carbon emission factor multi-source heterogeneous data to the carbon emission factor heterogeneous type anomaly identification model to perform anomaly data identification processing of each heterogeneous type, and outputting multi-source heterogeneous first anomaly data and / or multi-source heterogeneous preliminary valid data;
[0019] Step S24: Performing fuzzy quantization feature conversion processing on the multi-source heterogeneous preliminary valid data to obtain multi-source heterogeneous valid fuzzy quantization feature data;
[0020] Step S25: Transferring the multi-source heterogeneous valid fuzzy quantization feature data to the carbon emission factor heterogeneous type anomaly identification model to perform anomaly fuzzy feature similarity analysis processing of each heterogeneous type valid data, and outputting multi-source heterogeneous second anomaly data and / or outputting carbon emission factor multi-source heterogeneous verification data;
[0021] Step S26: Based on the multi-source heterogeneous first abnormal data and the multi-source heterogeneous second abnormal data, each heterogeneous type of abnormal data correction processing is performed to obtain multi-source heterogeneous abnormal correction data, and the multi-source heterogeneous abnormal correction data is fed back to the carbon emission factor multi-source heterogeneous verification data for abnormal data repair feedback adjustment processing of the carbon emission factor, to generate carbon emission factor multi-source heterogeneous standard data.
[0022] Further, step S22 includes the following steps:
[0023] According to the historical data of the carbon emission factor multi-source heterogeneous data, heterogeneous type of data abnormal prior data analysis is performed to generate carbon emission factor heterogeneous type abnormal prior data;
[0024] According to the carbon emission factor heterogeneous type characteristic data, the abnormal feature index analysis of the carbon emission factor heterogeneous type is performed to generate carbon emission factor heterogeneous type abnormal feature index data, and a preliminary carbon emission factor heterogeneous type abnormal recognition model is established based on the carbon emission factor heterogeneous type abnormal feature index data;
[0025] According to the carbon emission factor heterogeneous type abnormal prior data, the preliminary carbon emission factor heterogeneous type abnormal recognition model is trained and the parameter weight adjustment processing of each heterogeneous type abnormal verification model is performed to generate the carbon emission factor heterogeneous type abnormal recognition model.
[0026] Further, step S23 includes: when the carbon emission factor heterogeneous type abnormal recognition model identifies the matching abnormal data in the carbon emission factor multi-source heterogeneous data, the multi-source heterogeneous first abnormal data is output, and / or when the carbon emission factor heterogeneous type abnormal recognition model does not identify the matching abnormal data in the carbon emission factor multi-source heterogeneous data, the multi-source heterogeneous preliminary effective data is output.
[0027] Further, step S24 includes: when the carbon emission factor heterogeneous type abnormal recognition model identifies that the abnormal fuzzy feature similarity score of the multi-source heterogeneous effective fuzzy quantization feature data is greater than the preset abnormal similarity score threshold, the multi-source heterogeneous second abnormal data is output, and / or when the carbon emission factor heterogeneous type abnormal recognition model identifies that the abnormal fuzzy feature similarity score of the multi-source heterogeneous effective fuzzy quantization feature data is not greater than the preset abnormal similarity score threshold, the carbon emission factor multi-source heterogeneous verification data is output.
[0028] Further, step S3 includes the following steps:
[0029] Step S31: The carbon emission factor modal division processing is performed on the carbon emission factor multi-source heterogeneous standard data to generate carbon emission factor modal data, wherein the carbon emission factor modal data includes carbon emission factor time series data, carbon emission factor static structured data, and carbon emission factor spatial type data;
[0030] Step S32: Carbon emission time sequence dynamic feature extraction is performed on the carbon emission factor time sequence type data to generate carbon emission time sequence dynamic feature data;
[0031] Step S33: Static structured interaction feature analysis is performed on the carbon emission factor static structured data to generate carbon emission factor static interaction feature data;
[0032] Step S34: Carbon emission spatial distribution feature analysis is performed on the carbon emission factor spatial type data to generate carbon emission factor spatial distribution feature data;
[0033] Step S35: Based on the Bayesian network algorithm, the carbon emission time sequence dynamic feature data, the carbon emission factor static interaction feature data, and the carbon emission factor spatial distribution feature data are subjected to cross-modal correlation processing of the carbon emission factor to generate carbon emission factor correlation modal data;
[0034] Step S36: Based on the neural network algorithm, the carbon emission factor correlation modal data is subjected to nonlinear feature analysis of the correlation modal to generate carbon emission factor correlation modal nonlinear feature data, and global connection processing of the correlation modal nonlinear feature is performed according to the carbon emission factor correlation modal nonlinear feature data to generate carbon emission factor global fusion feature data.
[0035] Further, step S35 includes the following steps:
[0036] Carbon emission cross-modal causal relationship analysis is performed on the carbon emission time sequence dynamic feature data, the carbon emission factor static interaction feature data, and the carbon emission factor spatial distribution feature data by the Bayesian network algorithm to generate carbon emission cross-modal causal relationship data;
[0037] Carbon emission modal mutual influence correlation feature analysis is performed according to the carbon emission cross-modal causal relationship data to generate carbon emission modal mutual influence correlation feature data, and a carbon emission modal mutual influence attention weight parameter is designed according to the carbon emission modal mutual influence correlation feature data;
[0038] Based on the carbon emission modal mutual influence attention weight parameter, attention weighted cross-modal correlation processing is performed on the carbon emission time sequence dynamic feature data, the carbon emission factor static interaction feature data, and the carbon emission factor spatial distribution feature data to generate carbon emission factor correlation modal data.
[0039] Further, step S4 includes the following steps:
[0040] Step S41: The carbon emission factor global fusion feature data is subjected to link division processing of the whole life cycle to obtain link division carbon emission factor global fusion feature data;
[0041] Step S42: Perform multi-factor coupled carbon emission conversion feature analysis according to the link division carbon emission factor global fusion feature data, and generate link division coupled carbon emission conversion feature data;
[0042] Step S43: Establish the carbon emission factor fusion feature and carbon emission accounting mapping relationship of each link according to the link division coupled carbon emission conversion feature data, and generate a carbon emission accounting relationship model;
[0043] Step S44: Perform carbon emission accounting actual working condition dynamic feature optimization processing on the carbon emission conversion relationship model, and generate an optimized carbon emission accounting relationship model;
[0044] Step S45: Perform carbon emission intelligent accounting operation on the road construction project based on the optimized carbon emission accounting relationship model.
[0045] Further, step S44 includes the following steps:
[0046] Perform uncertainty feature analysis of each link actual working condition according to the link division carbon emission factor global fusion feature data, and generate link division actual working condition uncertainty feature data;
[0047] Based on the random forest algorithm, establish a tree model of each link actual working condition, and map the link division actual working condition uncertainty feature data to the tree model of each link actual working condition to perform uncertainty feature learning processing of each link actual working condition, and generate a link division carbon emission actual working condition uncertainty relationship model;
[0048] Perform carbon emission accounting actual working condition dynamic feature optimization processing on the carbon emission conversion relationship model through the link division carbon emission actual working condition uncertainty relationship model, and generate an optimized carbon emission accounting relationship model.
[0049] The application has the advantages that the application can comprehensively and systematically collect carbon emission related influencing factors by introducing a multi-source integrated sensor system in the whole life cycle of road construction engineering. Not only the construction equipment energy consumption data is considered, but also the carbon emission data of building material production and transportation process and the construction site environment data are introduced, so that the multi-dimension and multi-link coverage of data sources is ensured. Through multi-source heterogeneous integration processing of the construction equipment energy consumption data, material production and transportation carbon emission data and construction environment data, the problems of different data sources, such as different formats, different sampling frequencies and large dimension differences, are effectively solved, and the structured unified expression of heterogeneous data is realized. The integration mechanism significantly improves the fusion efficiency of multi-source data. The preprocessing step is introduced after data collection, which can filter noise, complete missing values and standardize the preliminary multi-source heterogeneous data, effectively improve the integrity and consistency of the data, and reduce the interference of abnormal data and redundant data on the analysis results from the source. Based on the multi-source heterogeneous data, the heterogeneous type characteristic analysis is carried out, so as to provide structured feature support for subsequent anomaly detection, and reveal the essential difference of different types of data. By introducing the abnormal recognition relationship model based on heterogeneous characteristic data and training and weight adjustment through historical prior data, the data abnormal recognition model for different heterogeneous types is generated, which not only has the ability to identify normal abnormal data, but also can adapt to the complex actual situation in road engineering scene, ensuring the accuracy and adaptability of the abnormal recognition result. Through fuzzy quantification feature conversion and fuzzy similarity analysis, the fuzzy analysis method is introduced to carry out secondary verification on the preliminary effective data, and identify the potential abnormality hidden in the boundary state, which is especially suitable for the data scene with continuous deviation in complex environment. Through the two-stage abnormality recognition (preliminary identification + fuzzy similarity analysis), the integrity of abnormal data identification is maximized. Through the feedback adjustment mechanism of abnormal correction data and verification data, dynamic repair and iterative optimization are realized. This feedback correction method can continuously improve the self-adaptability of the model while correcting the data, ensuring that the finally generated multi-source heterogeneous standard data has high consistency and high reliability. In the carbon emission analysis process, the modal difference of multi-source heterogeneous data is significant, for example, the construction equipment energy consumption data is time series data, the material transportation data is mostly static structured data, and the construction environment monitoring data has obvious spatial distribution characteristics. The modal division and feature extraction of the standardized carbon emission factor data effectively realize the fine modeling of time series features, static features and spatial features, and improve the comprehensiveness and pertinence of feature expression. The Bayesian network algorithm is introduced to analyze the causal relationship of cross-modal data, which can not only identify the direct or indirect relationship between different modalities, but also establish the causal chain of modal interaction. Through the attention weighting mechanism, different weights are given to different modal features, which can automatically highlight the modal features that have a greater impact on carbon emission results, and enhance the expression efficiency and robustness of the fused data.The non-linear feature analysis on the cross-modal causal relationship in combination with the neural network algorithm can effectively capture the complex coupling relationship between the carbon emission factors. The global fusion feature data generated by the global connection mechanism can integrate multi-dimensional information to the greatest extent, ensuring the high correlation and high explanatory power of the input data of the subsequent carbon emission accounting model. Not only does it solve the problem of efficient fusion of multi-source heterogeneous data in the prior art, but also significantly improves the depth and breadth of data mining, laying a solid foundation for establishing a precise carbon emission accounting model. By refining and dividing the whole life cycle into multiple key links such as construction equipment use, material transportation and on-site environment, the carbon emission accounting is accurately decomposed into multiple links, thereby realizing the quantitative modeling of each link and avoiding the errors caused by the averaging process in the traditional overall analysis method. Based on the global fusion feature data divided by the links, the carbon emission conversion feature analysis of multi-factor coupling can capture the interaction effects between the links. For example, the non-linear interaction between construction equipment energy consumption and construction environmental conditions is reflected through multi-factor coupling modeling, making the accounting results more consistent with the actual working conditions. The introduction of the random forest algorithm for modeling uncertain factors can improve the adaptability to complex working conditions through the ensemble learning mechanism of the tree model. Especially in the scene of complex construction environment and equipment state of road construction engineering, the uncertain characteristics of each link are effectively identified and learned, thereby optimizing the dynamic response capability of the carbon emission accounting relationship model. The finally generated optimized carbon emission accounting relationship model not only has high precision and robustness, but also can be dynamically adjusted according to the actual working conditions to support the intelligent carbon emission accounting operation of road engineering.
[0050] Therefore, the road engineering carbon emission analysis method based on multi-source heterogeneous data fusion can cover the whole life cycle of road engineering and has the ability to deeply process the fusion of multi-source heterogeneous data to ensure that the potential association between data is fully mined. Through the establishment of the carbon emission factor heterogeneous type anomaly recognition model and the multi-level data anomaly recognition, the abnormal data of the carbon emission factor is accurately identified, and through the analysis of the coupling characteristics of the carbon emission factor and the actual operation condition, the specific carbon emission that conforms to the actual operation condition is calculated, meeting the intelligent carbon emission accounting demand under complex working conditions. BRIEF DESCRIPTION OF DRAWINGS
[0051] Figure 1 The figure is a step flowchart of the road engineering carbon emission analysis method based on multi-source heterogeneous data fusion of the present application;
[0052] Figure 2 The figure is a detailed implementation step flowchart of step S3 in the present application; Figure 1
[0053] The objectives, functional features and advantages of the present application will be further illustrated in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0054] The technical method of the present application will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without any creative work fall within the scope of protection of the present application.
[0055] In addition, the accompanying drawings are only schematic illustrations of the present application, and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated description thereof will be omitted. Some block diagrams shown in the drawings are functional entities, which do not necessarily have to correspond to physically or logically independent entities. The functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0056] To achieve the above-mentioned object, please refer to Figures 1 to 2 The present application provides a road engineering carbon emission analysis method based on multi-source heterogeneous data fusion. In the embodiments of the present application, please refer to Figure 1 Fig. 1 is a step flow diagram of a road engineering carbon emission analysis method based on multi-source heterogeneous data fusion according to the present application. The road engineering carbon emission analysis method based on multi-source heterogeneous data fusion includes the following steps:
[0057] Step S1: Collecting and data preprocessing of multi-source heterogeneous carbon emission influencing factors of the whole life cycle of road construction engineering by using a multi-source integrated sensor system, to generate multi-source heterogeneous carbon emission factor data.
[0058] In the embodiment of the present application, the sensor modules are arranged according to functions at the construction site: the construction equipment energy consumption sensor module includes a fuel flow meter (pulse output, sampling frequency 10 Hz), an electric energy meter (pulse / pulse width output, sampling frequency 1 Hz), an engine speed and working hour meter (CAN bus interface, sampling frequency 1 Hz), and a hydraulic system pressure sensor (sampling frequency 5 Hz); the building material monitoring sensor module includes a load cell / load sensor (weighing accuracy ±0.5%), a truck-mounted GPS locator (frequency 1 Hz), and a passive RFID tag for batch identification; the construction site monitoring sensor module includes a temperature and humidity sensor (minute-level sampling), a wind speed and direction meter (second-level sampling), a particulate matter sensor (PM2.5, PM10, minute-level sampling), and a noise level meter. Before being put into use, each sensor device is calibrated according to a unified calibration procedure: the fuel flow meter is verified by a standard flow source, the electric energy meter is verified by a reference electric meter, and the weighing sensor is verified by a standard weight. The construction equipment energy consumption sensor module, the building material monitoring sensor module, and the construction site monitoring sensor module are integrated into a multi-source integrated sensor system, and each carbon emission factor data collected through the multi-source integrated sensor system is integrated. Time synchronization is implemented by receiving a reference time through satellite positioning (GNSS) and combining a network time protocol (NTP), and all sampling records are provided with accurate time stamps, sensor identifiers, and geographic coordinates. After original sampling, a preprocessing chain is executed: a fourth-order low-pass Butterworth filter is first applied to continuous signals to remove high-frequency oscillation noise, and a sliding median filter is applied to isolated spikes to eliminate pulse anomalies; cubic spline interpolation is used to complete short-time missing data (less than 5 minutes); long-time missing data (more than 5 minutes) are estimated using a model-driven interpolation method based on nearest neighbor time series similarity and the uncertainty is recorded; discrete events (such as material loading batches) are batched and standard record items are generated using event time as an anchor point. Unit standardization is completed in the preprocessing stage: after the fuel volume is converted into mass, it is multiplied by the greenhouse gas equivalent coefficient corresponding to the fuel category, the electric energy is calculated by multiplying the daily time amplitude by the regional power grid time period emission factor, and the material is calculated by multiplying the mass by the material unit equivalent coefficient. The preprocessing output is a structured record set, and the record fields include: time stamp, sensor ID, measurement type, measurement value, unit of measurement, geographic coordinates, operation status identifier, quality flag bit, and initial uncertainty estimation, which provides a unified and traceable data basis for subsequent anomaly identification and cross-modal fusion.
[0059] Step S2: multi-source heterogeneous data anomaly identification and correction processing is performed on the carbon emission factor multi-source heterogeneous data to generate carbon emission factor multi-source heterogeneous standard data;
[0060] In the embodiment of the present application, after obtaining the multi-source heterogeneous data of carbon emission factors, first, the heterogeneous type characteristic library is established according to the data type: the statistical characteristics (mean, variance, skewness, kurtosis, interquartile range), time domain characteristics (maximum slope, start-stop count, duty cycle ratio), frequency domain characteristics (main frequency and bandwidth of power spectral density), information theory characteristics (sample entropy, self-entropy) and autocorrelation coefficient vector of each time series signal in the sliding window (window length 30 minutes, step 5 minutes) are calculated, and a high-dimensional feature description vector of each signal is formed, and the same type of sensor grouping and type labeling are performed according to the vector. Based on the above characteristic data, a hierarchical anomaly recognition relationship model is constructed: a detector based on isolation forest combined with a statistical threshold rule based on a sliding window is constructed for single variable continuous signals, and the threshold rule is determined by the quantile of the historical operation distribution and saved as a type priori; a long short-term memory autoencoder (LSTM autoencoder) is constructed for multivariate time series overall correlation anomaly, and the correlation anomaly is recognized through reconstruction error; the indicators meeting the physical constraints are incorporated into the rule-based logical verification module (for example, fuel consumption cannot be negative, instantaneous current cannot exceed the rated value of the device), and all detectors are cross-validated and grid searched to determine the hyperparameters and record the model performance indicators on the historical labeled anomaly sample set. In the running, the preprocessed output is input into the anomaly recognition model according to the sliding window, and the model outputs the first anomaly data or the preliminary valid data for each time window. The preliminary valid data is further subjected to fuzzy quantization feature conversion: for continuous variables, a fuzzy membership function (triangular membership or trapezoidal membership) is defined, the numerical value is mapped to a membership degree vector, and the gravity method is used to numerize it into a fuzzy quantization feature vector. The fuzzy quantization feature vector is sent to the same anomaly recognition relationship model to perform fuzzy feature similarity analysis, cosine similarity and fuzzy similarity indicators are adopted, and compared with the similarity threshold value established in advance, if the similarity exceeds the threshold value, it is labeled as the second type of anomaly and output; if it does not exceed the threshold value, it is output as multi-source heterogeneous verification data. According to the anomaly type, a hierarchical correction strategy is adopted for the labeled first anomaly and second anomaly: Kalman smoothing correction is applied to short-time noise and isolated spikes, drift compensation based on a state space model is used to estimate the baseline and correct for sensor drift, Gaussian process regression is used to estimate the missing or truncated data and output the prediction uncertainty at the same time, and the items involving multi-source information contradiction are repaired for consistency by cross-source regression (using related sensor features to construct a regressor), and the error measurement before and after correction is recorded, all correction results form multi-source heterogeneous anomaly correction data and are written back as multi-source heterogeneous verification data, which is used to trigger model feedback training. Periodically (for example, once a week), the anomaly recognition model is retrained and weight adjusted based on the corrected data, and the training sample weight and reconstruction error threshold are adjusted according to the precision and recall rate of cross-validation after retraining, so as to form a closed loop of anomaly recognition and correction process and output the verified consistent high carbon emission factor multi-source heterogeneous standard data.
[0061] Step S3: Carbon emission factor multi-source heterogeneous standard data is subjected to global connection fusion processing of carbon emission factor cross-modal, to generate carbon emission factor global fusion feature data;
[0062] In the embodiment of the present application, after obtaining the carbon emission factor multi-source heterogeneous standard data, first, the data is divided according to the data modal: the time series data includes high-frequency time series of construction equipment energy consumption, equipment operation state and sensor continuous monitoring sequence; the static structured data includes material batch attribute, material unit equivalent coefficient, equipment rated parameter and contract data; the spatial data includes GPS track point, material unloading position coordinate and construction site boundary vector. Multi-scale dynamic feature extraction is performed on the time series data: short-time window (5 minutes) calculates instantaneous power, start-stop times and slope, long-time window (24 hours) calculates daily cumulative amount, load curve features and extracts scale coefficient of transient energy consumption pulse through continuous wavelet transform; the static structured data is constructed into an interaction feature matrix, such as the product of material quality and arrival time, the ratio of unit carbon dioxide equivalent of single material to use time, and a fixed-length binary vector encoding is applied to the category variable for subsequent network input; the spatial data is subjected to road section mapping and spatial statistics: the track points are aligned to the road section by using the map matching algorithm, the transportation distance and stay time are summarized by road section, and the spatial distribution grid based on 200-meter grid is generated by Kriging interpolation to express the local emission density. Cross-modal causal relationship analysis is performed based on Bayesian network: the modal features are nodes, the structure learning is performed according to the scoring criterion and conditional independence test method to determine the directed edge, and the conditional probability table is obtained by maximum likelihood estimation, the generated cross-modal causal relationship is used to estimate the influence strength between modes and design the modal interaction attention weight parameter. On this basis, a multi-branch neural network is constructed for nonlinear feature analysis: the time branch is composed of a one-dimensional convolution layer followed by a long short-term memory layer to extract time series deep dynamic features; the static branch is composed of a plurality of fully connected layers to express structured interaction information; the spatial branch adopts a graph convolution network to capture the topological dependence between road sections. The outputs of each branch are weighted and aggregated by the cross-modal attention module, the attention weight is initialized based on the causal influence strength derived from the Bayesian network and updated by the gradient method during the training process. The training objective function adopts the sum of mean square error and L2 regularization term, and the adaptive first-order momentum optimizer is used for parameter update, the early stopping mechanism is implemented during the training process to prevent overfitting, and after the network training is completed, a global fusion feature vector of uniform dimension is output for each time window.
[0063] Step S4: An optimized relationship model for actual working condition carbon emission accounting is established based on the carbon emission factor global fusion feature data, to generate an optimized carbon emission accounting relationship model; and the carbon emission intelligent accounting operation is performed on the road construction project based on the optimized carbon emission accounting relationship model.
[0064] In the embodiments of the present application, after obtaining the global fusion feature of carbon emission factors, the life cycle stages are divided according to the life cycle stages: the material production stage (from the material production time to the loading time), the transportation stage (from the loading time to the unloading time, the road section meeting the average speed greater than 5 km / h and the single trip duration not less than 10 minutes is the transportation section), the construction stage (from the installation to the completion of the construction acceptance, the equipment running hours are greater than zero and the location is in the construction area), the maintenance stage and the operation stage are defined according to the acceptance record time and are respectively summarized. The multi-factor coupled carbon emission conversion feature analysis is performed for each stage: the generalized additive model with interaction term and the polynomial regression are used to factorize the input features and calculate the interaction effect coefficient, the structural equation model is constructed for the nonlinear interaction to quantify the transmission path between the latent variables and the observed variables, and thus the conversion feature parameter matrix of the stage level is obtained. The mapping relationship model is established based on the coupling features of the stage division: the integrated mapper based on the random forest regression and the gradient boosting regression tree is constructed for each stage, the global fusion feature is used as the input and the output is the estimated value of the stage carbon emission, the historical calibrated measured emission data are used for supervised learning in the training process and the generalization error is evaluated by cross-validation; after the model training is completed, the feature importance ranking and the local interpretability index (such as SHAP value) are derived to support the interpretability of the accounting results. In order to realize the dynamic optimization of the accounting under the actual working condition, the uncertainty learning is introduced: the Bootstrap sample of the random forest is used to construct the prediction quantile and output the confidence interval through the quantile regression forest, the threshold is applied to the prediction uncertainty, and when the uncertainty interval exceeds the preset tolerance, the model adaptive adjustment module is triggered; the adaptive adjustment implements the incremental retraining of the mapper through the sliding time window (example: the last 7 days of data) and updates the mapping parameters combined with the latest calibration data to correct the deviation. The final optimized carbon emission accounting relationship model is used to perform the intelligent accounting operation of the road construction project: the carbon dioxide equivalent estimation value of each source is accumulated in the pre-defined time window and the emission curve and the summary report are generated in the time and segment, and the confidence interval and the key driving factors are output, the accounting process records the complete audit log for backtracking. When the confidence interval or the deviation of the real-time accounting result exceeds the set threshold, the abnormal backtracking process is activated to re-verify the related input features and the stage mapping, and the necessary mapper retraining is triggered, so as to realize the closed-loop guarantee of the accounting accuracy and reliability.
[0065] Further, wherein the multi-source integrated sensor system includes a construction equipment energy consumption sensor module, a building material monitoring sensor module, and a construction site monitoring sensor module, step S1 includes the following steps:
[0066] Step S11: collecting construction equipment energy consumption data of the construction period of the road construction project by using the construction equipment energy consumption sensor module to obtain construction equipment energy consumption data;
[0067] In the embodiment of the present application, the construction equipment energy consumption collection is arranged according to the type of the equipment. The mobile vehicle and the internal combustion equipment such as the excavator are equipped with a fuel flow meter (pulse output, sampling frequency 10 Hz), an engine speed sensor (CAN bus interface, sampling frequency 1 Hz), a working time meter (cumulative timer, resolution 1 s) and an oil pressure sensor (sampling frequency 5 Hz); the electric equipment is equipped with an electric energy meter (pulse output or pulse width output, sampling frequency 1 Hz) and voltage and current measuring points. All sensors are subjected to zero point and range verification according to the factory calibration procedure, the fuel flow meter is verified using a standard flow source, and the electric energy meter is calibrated by comparing with a reference meter. The sensors are fixed on the non-vibration sensitive parts of the equipment using anti-vibration brackets, the cables are shielded and sealed at the wiring points through waterproof connectors, the sampling clock is synchronized with the satellite positioning time (GNSS) and the network time protocol (NTP), and the sampling record contains the time stamp, sensor ID, measurement item, unit, equipment working condition label and initial uncertainty estimate. The diagnostic measures include real-time self-checking of fuel flow sudden change, continuous zero value and overrange, abnormal event triggering boundary marker and recording of surrounding environment and equipment working condition for subsequent verification. The energy consumption instantaneous value is accumulated as the energy consumption total according to the sampling frequency, and the volume to mass conversion is performed according to the fuel density and calorific value for subsequent emission conversion.
[0068] Step S12: collecting carbon emission data of the road construction material production and transportation stage of the road construction project by using the building material monitoring sensor module to obtain material production and transportation carbon emission data;
[0069] In the embodiment of the present application, the carbon emission monitoring of the material production and transportation stage adopts a combination of multi-point metering and identification tracking. The production end sets quality measuring devices at the batching and discharging links, including a weighbridge (weighing accuracy ±0.5%), a continuous belt scale (sampling period 1 s) and fuel flow meters and combustion furnace outlet flue gas Concentration detectors are used (periodically calibrated for zero point and range using standard gas). Each batch of materials is affixed with an RFID tag as a batch identifier, and the loading time is cross-confirmed by the timestamp recorded by the weighbridge and the RFID reader. During transportation, vehicle-mounted positioning devices (GNSS, frequency Hz) and vehicle-mounted fuel metering devices (instantaneous fuel flow meters) are installed on the vehicles, recording the start and end times of the journey, mileage, average speed, and loading status. Emissions during the material production stage are calculated by multiplying the material mass by the material unit emission factor, and the greenhouse gas emissions generated by fuel combustion during the process are recorded simultaneously. Transportation emissions are calculated by multiplying the fuel consumption of the transport vehicle by the fuel unit emission factor, and corrected for by the loading rate. Material batches, weighing records, vehicle mileage, and fuel consumption are systematically correlated through timestamps and batch identifiers, forming a traceable production-transportation chain record and preliminary emission estimates.
[0070] Step S13: Use the construction site monitoring sensor module to monitor and process the on-site environmental data of the road construction project during the construction period to obtain construction environmental data;
[0071] In this embodiment of the invention, environmental monitoring at the construction site involves deploying several meteorological and particulate matter monitoring points within a spatial grid within the construction area to reflect the heterogeneity of the site environment. The monitoring point configuration includes temperature and relative humidity sensors (minute-level sampling), wind speed and direction meters (second-level sampling), particulate matter sensors (PM2.5, PM10, minute-level sampling), volatile organic compound detectors (VOCs, hourly sampling), and on-site noise meters. Meteorological monitoring points are installed at a height of 2 meters above ground level and are ensured to be unobstructed. Particulate matter monitoring points are located near the edge of the work area and are equipped with rain covers and a regular filter replacement program to maintain stable measurements. Environmental data is used to calculate evaporation loss correction coefficients, dust resuspension factors, and combustion efficiency correction terms. Specifically, this is achieved by calculating the moving averages of temperature, humidity, and wind speed within a time window and obtaining the evaporation rate coefficient using an empirical formula. Particulate matter observations are correlated with the intensity of construction activities (number of loading operations, vehicle mileage on the road) to estimate dust contribution. Quality flags at the environmental monitoring points reflect sensor malfunction, contamination, or maintenance status to ensure traceability and auditability for subsequent use.
[0072] Step S14: Perform multi-source heterogeneous integration processing of carbon emission factors throughout the entire life cycle of construction equipment energy consumption data, material production and transportation carbon emission data, and construction environment data to obtain preliminary multi-source heterogeneous carbon emission factor data.
[0073] In the embodiment of the present application, multi-source heterogeneous integrated processing is implemented event-level fusion with time and space as anchor points. Time synchronization takes GNSS timestamp as reference, and adopts minimum time difference matching algorithm to align high-frequency time sequence measurement points with low-frequency device records; spatial alignment associates vehicle trajectory points to road segment vectors through map matching algorithm and realizes partition mapping according to construction area boundary. Batch cascade rule is defined as a three-element key: RFID batch ID, weighbridge time interval and vehicle driving trajectory interval, and records meeting the consistency of the three-element key are merged into a single production and transportation event. Kalman filter is applied to redundant measurement items to fuse instantaneous measurement values, and inverse variance weighting method is used to calculate fusion estimation and fusion uncertainty. Quantities with physical consistency constraints between observations are solved by constraint optimization method (for example, material mass conservation constraint is used to correct the deviation between belt scale and weighbridge readings). Low-frequency sources are aligned with high-frequency sources by using the nearest steady-state value expansion method, and missing segments are estimated by using a time sequence regression model based on similar working conditions and recording uncertainty indicators. The fusion output is the preliminary carbon emission factor record at the event level or time window level, and the fields include event identification, time interval, geographic range, contribution source type, original measurement set, preliminary estimated emission and fusion uncertainty estimation, and complete traceability link.
[0074] Step S15: data preprocessing is performed on the preliminary carbon emission factor multi-source heterogeneous data to generate carbon emission factor multi-source heterogeneous data.
[0075] In the embodiment of the present application, strict preprocessing is performed on the preliminary fusion data to generate standardized carbon emission factor records. Noise processing adopts a multi-level filtering strategy: fourth-order low-pass Butterworth filter is applied to high-frequency oscillation, sliding median filter is applied to isolated pulses, and continuous wavelet transform is used for denoising and reconstruction of stationary components. The deviation value determination is based on the three standard deviation criterion discrimination logic, short-time missing (not more than 5 minutes) is completed by linear interpolation, and long-term missing is estimated by Gaussian process regression and the estimation confidence is output. Unit unification and conversion strictly adopt physical quantity conversion formula: fuel quality is equal to fuel volume multiplied by fuel density; electric energy is obtained by multiplying the power consumption (kilowatt-hour) by the grid unit emission factor; material production stage is obtained by multiplying the material quality by the material unit generation emission factor; environmental gas is obtained by monitoring the gas quality related to carbon element. The final standard record fields include: time window, event identification, spatial boundary, emission source category, data quality label and complete traceability information, and all records are saved in audit log format for subsequent verification and model calibration.
[0076] Further, step S2 includes the following steps:
[0077] Step S21: heterogeneous type characteristic analysis is performed according to the carbon emission factor multi-source heterogeneous data to generate carbon emission factor heterogeneous type characteristic data;
[0078] In the embodiment of the present application, when analyzing the heterogeneous type characteristics according to the multi-source heterogeneous data of carbon emission factors, first, the data is classified by modal into time series type, static structured type and spatial type. The time series type calculates the time domain statistics (mean, variance, skewness, kurtosis, interquartile range), instantaneous slope and start-stop ratio, peak value ratio and periodicity characteristics of each signal in the sliding window (window length 30 minutes, step 5 minutes); the power spectrum density estimation is performed on the frequency domain characteristics to extract the main frequency and bandwidth; the sample entropy and approximate entropy are calculated for complex dynamic signals to depict irregularity. The static structured type performs statistical description on category distribution, frequency, mode and missing pattern, and calculates batch arrival interval distribution, batch quality coefficient of variation and batch internal and external difference indicators. The spatial type maps the trajectory points to the road segment through map matching, calculates the road segment level transportation intensity, residence time distribution and local density, and estimates the spatial autocorrelation scale and Moran's I statistics using the semi-variation function. The above features are standardized to form a high-dimensional feature vector cluster, the homogeneity within the cluster is evaluated using intra-class variance, and the distinguishability between clusters is measured using Mahalanobis distance; the unsupervised division is performed based on the density clustering algorithm (DBSCAN) and Gaussian mixture model to identify abnormal sensor groups or abnormal behavior patterns, and finally the heterogeneous type characteristic data classified by type is output, which includes the feature list, statistical distribution parameter, clustering label and type internal baseline model parameter of each type, for subsequent use in training the abnormal identification model.
[0079] Step S22: establishing a data anomaly identification relationship model of carbon emission factor heterogeneous types based on the carbon emission factor heterogeneous type characteristic data, to obtain a carbon emission factor heterogeneous type anomaly identification model;
[0080] In the embodiment of the present application, when the data anomaly recognition relationship model is established based on the heterogeneous type characteristic data of carbon emission factors, a hierarchical detection framework is constructed and a special detector is configured for each heterogeneous type. For single variable numerical signals, the historical quantile rule (lower limit is 0.5 percentile, upper limit is 99.5 percentile) and isolated forest are used for parallel judgment; in the training stage, the isolated forest uses the historical normal section and the labeled abnormal section to calculate the abnormal score, and the threshold is determined by cross-validation to maximize the F1 value on the validation set. For multi-variable time series correlation anomalies, a long short-term memory autoencoder (LSTM Autoencoder) is trained, and the reconstruction error is used as an anomaly indicator. The training set uses multi-condition historical samples and uses the early stopping method to prevent overfitting. For category sequence type, a hidden Markov model is established to judge sequence anomalies by transition probability deviation. Spatial anomalies use spatial scanning statistics and local Moran's I deviation test. The outputs of each detector are fused into the final discrimination score by a weighted voter, and the weight is determined by the detection contribution of the historical labeled sample through parameter estimation based on expectation maximization and recorded as model metadata. The model training process records the performance indicators such as ROC curve, AUC, precision and recall, and saves the model version to realize the auditable recognition relationship model.
[0081] Step S23: transmit the carbon emission factor multi-source heterogeneous data to the carbon emission factor heterogeneous type anomaly recognition model to perform anomaly data recognition processing of each heterogeneous type, and output multi-source heterogeneous first abnormal data and / or multi-source heterogeneous preliminary valid data;
[0082] In the embodiment of the present application, when the carbon emission factor multi-source heterogeneous data is transmitted to the anomaly recognition model for recognition processing, the detection process is performed piece by piece according to events or time windows: first, for each time series signal, the quantile rule and isolated forest are applied within a 30-minute sliding window to determine whether the abnormal score exceeds the set threshold (for example, isolated forest score ≥ 0.7 and value exceeds 99.5 percentile), and if so, it is labeled as multi-source heterogeneous first abnormal data, and the abnormal type (spike, long zero value, drift), timestamp, sensor identifier and recent environmental context are recorded. If single variable anomaly determination is not triggered, the window and adjacent window are sent together to the LSTM autoencoder for multi-variable reconstruction error detection; if the reconstruction error is greater than the set threshold during training, the first abnormality is output. The bidirectional verification of cross-source data is performed by physical consistency rules, for example, if the fuel consumption to engine work time ratio deviates from the historical working condition by ±20%, it is counted as an abnormality; if all detectors are not labeled as abnormal and pass the physical consistency verification, multi-source heterogeneous preliminary valid data is output, and the confidence score and a number of key feature values used for judgment are recorded. All judgment results form an abnormal log in chronological order, which is used as input for subsequent fuzzy conversion and correction strategies, and the threshold and feature basis for triggering judgment are labeled in the model metadata for auditing.
[0083] Step S24: fuzzy quantification feature conversion processing is performed on the multi-source heterogeneous preliminary effective data to obtain multi-source heterogeneous effective fuzzy quantification feature data;
[0084] In the embodiment of the application, when the multi-source heterogeneous preliminary effective data is subjected to fuzzy quantification feature conversion, language membership sets are defined for continuous variables, and the language items are definitely defined as "low", "normal" and "high". The membership functions are triangular or trapezoidal functions and are constructed with historical quantiles as base points: the lower limit Q0 is the 0.1 percentile of the historical data, the first limit point Q1 is the 25 percentile, the median Q2 is the 50 percentile, the third limit point Q3 is the 75 percentile, and the upper limit Q4 is the 99.9 percentile. The "low" membership function is calculated in a linearly increasing or decreasing manner in [Q0, Q1, Q2], the "normal" membership function is dominant in [Q1, Q2, Q3], and the "high" membership function is defined in [Q2, Q3, Q4]. The membership degree is calculated by using a linear interpolation formula, and the output is in the interval [0, 1]. The frequency distribution of the category features is mapped to a membership degree vector, and the batch identification is obtained by calculating the similarity with the historical high-quality batch template. The membership degrees of each variable are concatenated to form a multi-dimensional fuzzy quantification feature vector, and principal component analysis is performed on the vector as necessary to compress the dimension and retain more than 90% of the variance. In order to facilitate subsequent numerical similarity calculation, the barycentric defuzzification method is used to calculate the centroid of some fuzzy sets that need to be numerically valued to obtain a representative numerical value, while retaining the membership degree vector as a fuzzy description.
[0085] Step S25: the multi-source heterogeneous effective fuzzy quantification feature data is transmitted to the carbon emission factor heterogeneous type anomaly recognition model to perform anomaly fuzzy feature similarity analysis processing on each heterogeneous type effective data, and multi-source heterogeneous second anomaly data is output and / or carbon emission factor multi-source heterogeneous verification data is output;
[0086] In the embodiment of the present application, when the multi-source heterogeneous effective fuzzy quantization feature data is transmitted to the anomaly recognition model to perform fuzzy feature similarity analysis, a number of prototype clusters are constructed based on the historical verified fuzzy feature library, and the prototype clusters are obtained by fuzzy C-means clustering. Each cluster center is saved in the form of membership vector. A number of similarity indexes are calculated for the fuzzy vector to be tested: fuzzy cosine similarity, fuzzy Jaccard index, and normalized value of fuzzy Euclidean distance. The similarity threshold is obtained by searching the maximum F1 on the validation set, and the example threshold is set to be less than 0.65 or greater than 0.35. If the similarity is less than 0.65 or the normalized distance is greater than 0.35, it is considered as multi-source heterogeneous second abnormal data. On the contrary, if the similarity is not less than the threshold and the confidence of the historical prototype distribution is not less than 0.7, the multi-source heterogeneous verification data of carbon emission factors is output. In order to enhance the time consistency, the moving average of the similarity of the last 5 time windows is calculated, and the weighted combination of the moving average and the current similarity is used as the final discrimination basis. The similarity discrimination result records the similarity value, the matched cluster identifier and the confidence, and the verification data is accompanied by distribution explanatory information to support the subsequent correction logic.
[0087] Step S26: Perform anomaly data correction processing of each heterogeneous type based on the multi-source heterogeneous first abnormal data and the multi-source heterogeneous second abnormal data, obtain multi-source heterogeneous anomaly correction data, and feed back the multi-source heterogeneous anomaly correction data to the carbon emission factor multi-source heterogeneous verification data for anomaly data repair feedback adjustment processing of the carbon emission factor, and generate carbon emission factor multi-source heterogeneous standard data.
[0088] In the embodiment of the present application, when the first and second abnormal data are corrected based on multi-source heterogeneous data, a hierarchical correction strategy is adopted according to the type of the abnormal data, and multi-source heterogeneous abnormal correction data is output. For transient spikes, a robust smoothing based on median filtering is adopted, a median with a window length of 5 minutes is used to replace the original point, and the residual before and after replacement is recorded. For sensor drift, baseline estimation is realized by comparing the reference sensor, the drift trend is fitted by piecewise linear regression, and the trend is deducted to restore the baseline value. For long-time missing, Gaussian process regression is adopted, the radial basis white noise kernel is selected as the kernel function, the last 7 days of smooth section data are used for training, and the estimated value and prediction variance are output. For cross-source contradictory items, the inverse variance weighted least squares method is used to solve the consistency estimation, and the weight is determined by the reciprocal of the uncertainty of each observation item. Each correction operation simultaneously performs uncertainty propagation, calculates the confidence interval of the corrected carbon emission element by Monte Carlo sampling (for example: 1000 times sampling), and records the uncertainty decomposition. The correction result forms a multi-source heterogeneous abnormal correction data set, and then the correction result is fed back to the multi-source heterogeneous verification data to trigger model adaptive training: based on the corrected labeled samples, the threshold and weight of the abnormal identification model are periodically re-estimated according to the sliding window, and the model performance change is recorded after each retraining. Finally, the multi-source heterogeneous standard data of carbon emission factors are output, each standard record contains original observation, corrected value, correction method code, uncertainty before and after correction, quality mark and complete traceability link, to meet the input requirements of subsequent accounting models.
[0089] Further, step S22 includes the following steps:
[0090] According to the historical data of the multi-source heterogeneous data of the carbon emission factors, the abnormal prior data analysis of the heterogeneous type data is performed, and the heterogeneous type abnormal prior data of the carbon emission factors is generated;
[0091] In the embodiment of the present application, historical observation records of no less than twelve months are used as the prior construction basis. The historical records are first segmented according to sensor categories and working conditions, and the segmentation length is in units of days and at least 360 daily samples are kept to cover seasonal changes. Quality screening is first performed for each segment; the quality screening rules are: missing rate higher than 5% is marked as low quality; continuous zero value duration exceeding 30 minutes is marked as an abnormal segment; measurement distribution skewness is tested by skewness and kurtosis and recorded. Time series stability is evaluated by enhanced Dickey-Fuller test, if non-stationary, it is processed by difference or seasonal decomposition, seasonal decomposition uses local weighted regression (STL) to extract trend and seasonal components, and residual error is used as the modeling object of this stage. The residual error distribution is fitted to a parameterized distribution, and the normal distribution and the lognormal distribution are compared according to experience to minimize the Akaike information criterion (AIC) to determine the distribution type and save the mean and variance as the parameter prior. For the frequency of sudden abnormalities, drift rate and missing distribution, Bernoulli-Beta process and Gamma distribution are used to model the event rate and variance, and the hyperparameters are estimated using empirical Bayes method, and hierarchical Bayesian structure is used at the group level to reflect the population differences of similar sensors and output the prior parameter set of the population level and the individual level. Cross-sensor collaboration characteristics use shrinkage covariance estimation (Ledoit-Wolf) to get a robust covariance matrix and use it as a multivariate prior covariance for regularization of subsequent anomaly detectors. Simulated observations are generated by sampling from the prior distribution and compared with the key statistics quantile of the actual observation, if the difference is significant, the hyperparameters are iteratively adjusted until the fitting degree meets the set test criteria (example: Kolmogorov-Smirnov test p value greater than 0.05). All prior parameters and test results are recorded as "carbon emission factor isomer type anomaly prior data", and archived in the form of time stamp and version number for traceability.
[0092] According to the carbon emission factor isomer type characteristic data, the abnormal feature index data of the carbon emission factor isomer type is generated, and a preliminary carbon emission factor isomer type anomaly recognition model is established based on the carbon emission factor isomer type abnormal feature index data;
[0093] In the embodiment of the present application, the carbon emission factor isomer type characteristic data is taken as input, and feature engineering is first performed to define an abnormality discrimination index set. The time domain features include mean value, standard deviation, maximum and minimum value, maximum slope, start-stop ratio and peak value ratio in a sliding window (window length 30 minutes, step 5 minutes); the frequency domain features are obtained by Welch spectrum estimation, and the window segment length is 256 points and the main frequency energy and bandwidth ratio are extracted; the time-frequency features are extracted by continuous wavelet transform (mother wavelet Morlet) on scales 1 to 128 to extract the transient pulse energy distribution; the complexity features include sample entropy parameters m=2, r=0.2×standard deviation, m represents the embedding dimension for constructing a vector, and r represents the tolerance threshold for matching similar vectors; the multivariate features include Mahalanobis distance, covariance principal component projection and energy conservation residual (for example, the deviation amount of the ratio of fuel consumption to engine working hours from the historical baseline). The class and batch information is constructed to form a statistical frequency and a transition matrix, and the Kullback-Leibler divergence of the transition matrix is used for sequence anomaly measurement. Based on the above features, a preliminary recognition model set is established: single variable uses dynamic z-score combined with historical quantile threshold for rapid screening; density and isolation detection uses isolation forest, and the parameter setting is that the number of base learners is 200, the upper limit of subsample is 1024, and the initial pollution rate is estimated from the history (example range 0.1% to 1%); associated anomaly detection uses LSTM self-encoding structure, and the sequence length is set to 120 steps, the number of encoder layers is 3, the number of hidden units is compressed from 128 to 32 layer by layer, the training target is to minimize the reconstruction error, and L2 regularization coefficient 1e-4 and dropout 0.2 are used to control overfitting. The feature importance is sorted by random forest feature importance score and mutual information evaluation, and the recursive feature elimination method is used to retain the feature subset with the highest explanation degree, and the obtained feature index and its threshold, detector structure and initial hyperparameter jointly constitute the preliminary carbon emission factor isomer type abnormality recognition model, and the performance indicators such as ROC curve, AUC, precision, recall and median detection delay are recorded under the time series segmentation cross-validation framework for verification.
[0094] According to the carbon emission factor isomer type abnormality prior data, the preliminary carbon emission factor isomer type abnormality recognition model is subjected to isomer type abnormality checking model training and parameter weight adjustment processing to generate a carbon emission factor isomer type abnormality recognition model.
[0095] In the embodiment of the present application, the heterogeneous type of carbon emission factor anomaly prior data is used as the regularization and sample weight source to perform supervised and unsupervised parallel training on the preliminary carbon emission factor heterogeneous type anomaly identification model. The historical labeled abnormal samples and normal samples are used for supervised training, the class imbalance is balanced by the bootstrap resampling method, and the abnormal samples are given a higher weight to reduce the risk of false negatives, and the weight ratio is initially set to normal:abnormal=1:5, and adjusted according to the cost function in the verification process. The unsupervised model parameter optimization adopts Bayesian regularization based on prior distribution: the model weight is set to a prior Gaussian distribution with a mean of zero and a variance determined by the prior noise variance, and the corresponding objective function is the negative log posterior. When optimizing, both the data fitting term and the prior penalty term are considered. The time series blocking cross-validation is used for hyperparameter adjustment, which is divided into five forward rolling folds. The hyperparameter grid includes the isolated forest pollution rate range [0.001, 0.01, 0.05], the LSTM learning rate range [1e-4, 1e-3, 1e-2], the encoding layer width range [64, 128, 256], and the regularization coefficient range [1e-6, 1e-3]. The Gaussian process-based Bayesian optimization is used for efficient search of continuous hyperparameters, and the F1 score and false alarm rate are used for comprehensive evaluation. The threshold is set to keep the false alarm rate below 2 times per day and the false negative rate below 5%. The early stopping strategy is embedded in the training process, and the tolerance is to stop and restore the best weight if there is no improvement for 10 consecutive validation periods. The final model requires PR-AUC to exceed 0.75 on an independent posterior test set and meet the minimum recall rate constraint according to the sensor category. The model configuration file, prior parameters during training, performance index table, and version identifier are output at the same time when the model is released, and the feedback mechanism for online calibration is recorded: when the distribution of continuous observations and the prior Kullback-Leibler divergence exceeds the threshold value 0.2 or the number of newly added manually labeled anomalies exceeds 200, the model retraining process is triggered to maintain the consistency and robustness of the identification ability.
[0096] Further, step S23 comprises: when the carbon emission factor heterogeneous type anomaly identification model identifies the matching abnormal data in the carbon emission factor multi-source heterogeneous data, outputting the multi-source heterogeneous first abnormal data, and / or when the carbon emission factor heterogeneous type anomaly identification model does not identify the matching abnormal data in the carbon emission factor multi-source heterogeneous data, outputting the multi-source heterogeneous preliminary effective data.
[0097] In the embodiment of the present application, in the online discrimination stage of the carbon emission factor heterogeneous type anomaly recognition model, a discrimination input vector is constructed according to a 30-minute sliding window and a step of 5 minutes. The input vector includes normalized time domain statistics (mean, variance, skewness, kurtosis), instantaneous slope, start-stop ratio, sample entropy, power spectrum main frequency and bandwidth, device working condition label and spatial neighborhood mean. For each continuous signal, single variable threshold detection and isolated forest-based anomaly score determination are simultaneously performed: the upper and lower thresholds of the single variable threshold are the 99.5th and 0th.5th percentiles of the history, respectively, and any observation exceeding the interval is recorded as a threshold anomaly; the isolated forest anomaly score is not less than 0.70, which is recorded as an isolated anomaly. For multivariate correlation anomalies, long short-term memory autoencoder (LSTM autoencoder) reconstruction error detection is performed, and the reconstruction error exceeding the 95th percentile threshold during training is recorded as a reconstruction anomaly. For category sequence type features, hidden Markov model transition probability deviation detection is performed, and the transition probability deviation degree exceeding 0.30 is recorded as a sequence anomaly. The physical consistency rules are verified in parallel with several working condition constraints: the fuel consumption cannot be negative, the instantaneous fuel consumption and the device rated consumption ratio exceeding 1.5 times are recorded as physical anomalies, and the single material mass conservation deviation exceeding ±2% is recorded as inconsistent mass. The determination logic adopts hierarchical consistency rules: when any two types of detectors trigger an anomaly in the same time window or the same detector repeatedly triggers in three consecutive sliding windows, the multi-source heterogeneous first anomaly data is output; when all detectors do not trigger and all physical consistency checks pass and the comprehensive confidence is not less than 0.85, the multi-source heterogeneous preliminary valid data is output. The comprehensive confidence is calculated according to the formula: Confidence = 1-(w_T Score_T + w_ IF Score _IF+w_AE× Score _AE), wherein Score _T is the threshold deviation degree normalized value, Score_IF is the isolated forest anomaly score, Score_AE is the autoencoder reconstruction error normalized value; the weights are w_T=0.4, w_IF=0.3, w_AE=0.3. The output items include: record type identification (first anomaly or preliminary valid), anomaly category code, trigger detector list, trigger timestamp, sensor identification, adjacent window context vector, physical consistency residual and confidence value. For the data labeled as first anomaly, the correction priority code is automatically attached according to the anomaly type and queued for subsequent repair process; for the data output as preliminary valid, a verification token is generated and passed to the next fuzzy quantization process.
[0098] Furthermore, step S24 includes: when the carbon emission factor heterogeneity anomaly identification model identifies an anomaly fuzzy feature similarity score of multi-source heterogeneous effective fuzzy quantization feature data that is greater than a preset anomaly similarity score threshold, outputting multi-source heterogeneous second anomaly data; and / or, when the carbon emission factor heterogeneity anomaly identification model identifies an anomaly fuzzy feature similarity score of multi-source heterogeneous effective fuzzy quantization feature data that is not greater than a preset anomaly similarity score threshold, outputting carbon emission factor multi-source heterogeneity verification data.
[0099] In this embodiment of the invention, after fuzzy quantization feature transformation of the preliminarily valid multi-source heterogeneous data, three linguistic terms "low," "normal," and "high" are defined for each continuous feature based on historical distribution. The membership function uses a trigonometric or trapezoidal function, with membership parameters constructed using historical quantiles Q0=0.1%, Q1=25%, Q2=50%, Q3=75%, and Q4=99.9%. Continuous values are calculated using linear interpolation to generate membership vectors; categorical features are used to generate membership vectors using frequency mapping, referencing historical distributions of high-quality batches. Based on a historically verified fuzzy vector sample library, fuzzy C-means clustering is applied to obtain prototype clusters. The number of clusters is set to 8, the fuzziness rate m=2, the maximum iteration is 300, and the convergence threshold is 1e-5. Cluster centers are saved as prototype libraries in membership vector format. The current fuzzy vector is compared with both the abnormal and normal prototype clusters using a weighted synthesis of fuzzy cosine similarity and normalized fuzzy Euclidean distance. The similarity score S is calculated using the formula S=0.7. S_cos+0.3 (1-D_norm) is calculated, where S_cos For fuzzy cosine similarity, D_norm is the distance index after normalizing the fuzzy Euclidean distance to its historical maximum value. If the maximum similarity between the current vector and any anomalous prototype cluster is not less than 0.60, then multi-source heterogeneous second anomalous data is output, and the trigger prototype ID, similarity value, membership dimension with the highest contribution, and time window index are written into the output record. If the anomalous prototype similarity is less than 0.60 and the maximum similarity with the normal prototype cluster is not less than 0.75, then carbon emission factor multi-source heterogeneous verification data is output, along with the matching normal prototype ID and matching confidence. For vectors at the fuzzy boundary (abnormal similarity < 0.60 and normal similarity < 0.75), time consistency expansion is performed: the similarity moving average of the most recent 5 sliding windows is calculated, and the similarity threshold is re-evaluated using the moving average; if the normal similarity of the moving average is not less than 0.70, then carbon emission factor multi-source heterogeneous verification data is output; otherwise, second anomalous data is output and marked as a fuzzy boundary anomalous candidate.
[0100] Furthermore, as an embodiment of the present invention, reference is made to... Figure 2 As shown, Figure 1The detailed step flowchart of step S3 is shown in the figure. In this embodiment, step S3 includes the following steps:
[0101] Step S31: Perform modal division processing on the carbon emission factor multi-source heterogeneous standard data to generate carbon emission factor modal data, wherein the carbon emission factor modal data includes carbon emission factor time series data, carbon emission factor static structured data, and carbon emission factor spatial data.
[0102] In the embodiment of the present application, when the carbon emission factor multi-source heterogeneous standard data is divided into modes, first, automatic annotation is performed according to observation attributes and metadata rules to clearly define the time granularity, spatial elements, and structured fields of each record. The criteria are defined as follows: if the record contains a continuous timestamp field and the average sampling frequency is higher than once per hour, it is classified as time series data; if the record is mainly based on batch attributes, material attributes, and specification parameters and is in discrete event units, it is classified as static structured data; if the record contains latitude and longitude or local projection coordinates and needs to be mapped, it is classified as spatial data. All records increase metadata fields during modal annotation, including sampling frequency, coordinate reference system (WGS84 or local UTM projection), measurement unit, measurement point accuracy, and measurement uncertainty estimation. Perform uniform time zone correction on the timestamp and align it according to the UTC reference. For time series data records with different sampling frequencies, apply the sampling rate label for subsequent resampling strategy application. Spatial data is simultaneously projected and transformed to obtain metric distance measurement, and event window boundaries are set for discrete events. The modal division result is output as an index table and metadata list of the three types of data sets.
[0103] Step S32: Extract carbon emission time series dynamic characteristics from the carbon emission factor time series data to generate carbon emission time series dynamic characteristic data.
[0104] In this embodiment of the invention, when extracting time-series dynamic features from carbon emission factor time-series data, a multi-scale sliding window strategy is adopted. The window is divided into a short-term window (5 minutes in length, 1 minute in step), a medium-term window (60 minutes in length, 15 minutes in step), and a long-term window (24 hours in length, 1 hour in step). Within each window, statistical features are calculated, including mean, variance, skewness, kurtosis, interquartile range, maximum value, minimum value, and median. Time-domain dynamic features are calculated, including instantaneous slope, start-stop frequency, duty cycle, and peak frequency. Frequency-domain features are calculated using Welch spectrum estimation to extract the dominant frequency and bandwidth ratio, and continuous wavelet transform (Morlet wavelet) is used to extract scaling coefficients for short-term transients to characterize pulse energy distribution. Seasonal-trend decomposition (STL) is performed on non-stationary trends to separate the trend from the residuals, and sample entropy, approximate entropy, and several lag values of the autocorrelation function are calculated for the residuals. To reduce feature dimensionality, principal component analysis is used to retain 90% of the variance. The output feature vector for each time window includes multi-scale statistics, time-frequency energy spectrum coefficients, complexity measures, and trend residual energy. These features are used to describe the dynamic characteristics of equipment energy consumption, load fluctuations, and transient events, and serve as input for subsequent correlation analysis.
[0105] Step S33: Perform static structured interaction feature analysis on the static structured data of carbon emission factors to generate static interaction feature data of carbon emission factors;
[0106] In this embodiment of the invention, when performing static interaction feature analysis on static structured data of carbon emission factors, categorical variables are first target-coded using the historical average emission amount at the batch level as the coded value, and the coding uncertainty is recorded; numerical variables are normalized and interaction terms with categorical variables are calculated, such as the product of material quality and supplier reputation score, and material units. The ratio of equivalent to arrival time difference and the difference between the rated power of the equipment and the actual load ratio were used. To identify higher-order interaction effects, a second-order polynomial interaction matrix was constructed, and the effect strength of each interaction term was evaluated. The effect strength was measured by partial regression coefficients and confidence intervals. Statistical significance was assessed using hypothesis testing, and interaction terms with p-values less than 0.05 were retained. For high-cardinality categories, binning and hierarchical aggregation strategies were used to control dimensionality expansion, and the information contribution of categories to the target variable was evaluated using mutual information and chi-square statistics. The output is a static interaction feature matrix, including core interaction terms, standardized regression coefficients, confidence indices, and variable importance ranking, providing structured factors and prior knowledge constraints for cross-modal causal analysis.
[0107] Step S34: Perform spatial distribution characteristic analysis on the spatial data of carbon emission factors to generate spatial distribution characteristic data of carbon emission factors;
[0108] In this embodiment of the invention, when performing spatial distribution characteristic analysis on spatial data of carbon emission factors, a map matching algorithm is first used to align transport trajectory points to road segment vectors and generate travel statistics (mileage, dwell time, average speed, and loading rate) for each road segment. At the geographic grid level, transport intensity and emission estimates are summarized using 200m × 200m grid units, and the emission density per unit area of each unit is calculated. Ordinary Kriging interpolation is used to spatially interpolate sparse observation points to obtain a continuous emission surface, and a semi-variogram is calculated to estimate the spatial autocorrelation scale. Moran's I index is used to test global spatial autocorrelation, and the Getis-Ord Gibbs index is applied. High-value clusters (hotspots) are statistically identified, with a significance level of 0.01. Spatial gradients are calculated at the road segment level to measure the derivative of emission intensity with location, and road segment slope, traffic density, and loading frequency are used as explanatory variables for spatial regression. A spatial error model is used to correct estimation biases caused by spatial autocorrelation. Spatial analysis outputs include emission density at both the road segment and grid levels, hotspot indicators, spatial autocorrelation parameters, and local trend vectors, providing geographical constraints and spatial influencing factors for cross-modal correlation.
[0109] Step S35: Based on the Bayesian network algorithm, perform cross-modal correlation processing on carbon emission time series dynamic feature data, carbon emission factor static interaction feature data and carbon emission factor spatial distribution feature data to generate carbon emission factor correlation modal data.
[0110] In this embodiment of the invention, when performing cross-modal association processing on three types of modal features based on the Bayesian network algorithm, a hybrid structure learning process is adopted to balance constraint information and data-driven evidence. First, the temporal dynamic features, static interaction features, and spatial distribution features are discretized or kept continuous according to variable type, and a candidate node set is constructed. Constraint-driven conditional independence tests are performed (Fisher Z test for continuous variables, G test for discrete variables) to eliminate unsupported edges. Then, a score-driven greedy search algorithm is applied to evaluate the network structure using the Bayesian Information Criterion (BIC) and perform local optimization within the candidate space. To resist the randomness of single-learning, multiple learning iterations guided by the bootstrap method are run, and the frequency of edge occurrence is counted. Only edges with a frequency of not less than 0.7 are retained as stable causal relationships. Parameter estimation uses maximum likelihood estimation, and a conditional Gaussian distribution is applied to fit continuous parent nodes. A conditional probability table is established for discrete nodes. The output results include the directed acyclic graph structure, the conditional probability or regression coefficient of each directed edge, and the edge weight value. The influence intensity matrix between modes is calculated from the edge weights, and the attention parameters of the modal interactions are generated accordingly to guide the initial attention configuration and cross-modal weighting strategy of the subsequent neural network.
[0111] Step S36: Based on the neural network algorithm, perform nonlinear feature analysis on the correlation mode data of carbon emission factors to generate nonlinear feature data of correlation mode of carbon emission factors, and perform global connection processing on the nonlinear features of correlation mode of carbon emission factors to generate global fusion feature data of carbon emission factors.
[0112] In this embodiment of the invention, when performing nonlinear feature analysis on carbon emission factor-related modal data, the data is first jointly modeled with carbon emission cross-modal causal relationship data, and the nonlinear correlation patterns are extracted using the multi-layer mapping capability of the neural network structure. Specifically, the input layer is set as a high-dimensional vector composed of a weighted combination of temporal dynamic features, static interaction features, and spatial distribution features, and a causal relationship matrix is appended before the input layer as a constraint. By embedding causal regularization terms in the network structure, the model ensures that the causal dependency logic between modalities is preserved during feature extraction. The hidden layer is configured with multiple nonlinear activation units, each employing a different nonlinear function to capture the nonlinear variation patterns of complex features, such as non-stationary patterns in temporal fluctuations, local clustering effects in spatial distribution, and nonlinear interactions between static features. The network parameters are iteratively updated using a backpropagation training mechanism, and the convergence direction of the parameters is adjusted using a causal constraint loss function, thereby generating nonlinear feature data of carbon emission factor-related modalities that reflects deep nonlinear relationships between modalities. After extracting the nonlinear feature data of the related modalities, it is further processed through global connections to achieve the fusion of multimodal features at the global level. In the global connectivity process, the nonlinear feature vectors of different modalities are first standardized to ensure consistent dimensions. Then, a global connectivity graph is constructed based on causal weight parameters. This graph uses each modal feature as a node and causal weights as edge weights, and performs cross-modal aggregation of feature vectors from different modalities through graph convolution operations. This process combines local features with the global structure, enabling different modal features to not only be superimposed at the numerical level but also to form global connections at the topological level. Finally, a unified-dimensional feature representation is output through a fully connected layer, yielding globally fused feature data of carbon emission factors. This data contains nonlinear coupling information of temporal, static, and spatial modalities, providing higher accuracy and robustness input support for subsequent optimization of the carbon emission accounting relationship model.
[0113] In another embodiment of the present application, a multi-layer neural network structure is constructed to meet the data processing requirements of the nonlinear characteristics of the correlation mode of carbon emission factors. First, in the input layer, the network receives the pre-processed and normalized multi-source feature vector. The dimension of the feature vector is determined according to the number of carbon emission factors and their modal characteristic expansion, such as construction equipment energy consumption factor, building material carbon emission factor, construction site environmental factor, etc. After feature fusion, a 128-dimensional input vector is formed to ensure complete coverage of multi-modal features. In the design of the hidden layer, a multi-layer fully connected structure combined with nonlinear mapping is adopted. The first hidden layer is set to 256 neurons to improve the ability to capture complex feature relationships by increasing the dimension, and the ReLU activation function is used to avoid the problem of gradient disappearance. The second hidden layer is set to 128 neurons to continue to maintain a high feature abstraction capability, and Batch Normalization processing is introduced in this layer to improve the training stability and alleviate the risk of overfitting. The third hidden layer is set to 64 neurons, which uses the tanh activation function to enhance the expression ability of the nonlinear correlation between carbon emission factors, especially suitable for capturing small fluctuations and asymmetric characteristics. In the design of the deep structure of the network, a double-branch parallel subnetwork is also introduced to process the differences between different modal features. One sub-branch focuses on time series features (such as dynamic fluctuations of energy consumption and emissions) and uses a double-layer LSTM structure with 64 hidden units in each layer to capture long-term dependencies. The other sub-branch is for static structured features (such as material carbon emission coefficients and equipment model parameters) and uses a two-layer fully connected structure with 128 and 64 neurons respectively, and uses the ReLU activation function for mapping. Finally, the output results of the two sub-branches are fused through the Concatenate Layer to form a comprehensive high-dimensional feature representation. In the design of the output layer, the fused high-dimensional features are mapped to a single regression node to output the interaction feature information of the correlation mode of the correlation mode data of the carbon emission factors, and the linear activation function is used to ensure the continuity of the numerical value. In addition, to prevent overfitting, a Dropout layer (dropout rate set to 0.3) is added before the output layer to randomly mask some neuron connections and enhance the model generalization ability. This neural network structure takes into account the feature abstraction capability and modal difference processing capability, and through the combination of multi-layer fully connected network and double-branch parallel design, it realizes efficient modeling of the nonlinear characteristic data of the correlation mode of carbon emission factors, making it easier for the interaction features between carbon emission factors in road construction projects to reflect the actual carbon emission accounting status.
[0114] Further, step S35 includes the following steps:
[0115] The carbon emission cross-modal causal relationship data is generated by performing carbon emission cross-modal causal relationship analysis on the carbon emission time series dynamic characteristic data, the carbon emission factor static interaction characteristic data and the carbon emission factor spatial distribution characteristic data through a Bayesian network algorithm;
[0116] The carbon emission modal mutual influence correlation characteristic data is generated by performing carbon emission modal mutual influence correlation characteristic analysis according to the carbon emission cross-modal causal relationship data, and the carbon emission modal mutual influence attention weight parameter is designed according to the carbon emission modal mutual influence correlation characteristic data;
[0117] The carbon emission factor correlation modal data is generated by performing attention weighted cross-modal correlation processing on the carbon emission time series dynamic characteristic data, the carbon emission factor static interaction characteristic data and the carbon emission factor spatial distribution characteristic data based on the carbon emission modal mutual influence attention weight parameter.
[0118] In the embodiment of the present application, when the carbon emission time series dynamic characteristic data, carbon emission factor static interaction characteristic data and carbon emission factor spatial distribution characteristic data are analyzed for cross-modal causal relationship, the three types of modal characteristics are variable standardized and distribution corrected to ensure that data from different sources are comparable in the same probability framework. Based on the Bayesian network structure learning algorithm, a candidate node set is set, wherein the time series dynamic characteristic data node represents the energy consumption rate, fuel consumption rate, instantaneous power fluctuation and the like in the equipment operation process; the static interaction characteristic data node represents the interaction term of material carbon emission factor and supply batch parameter, the coupling term of construction process and equipment parameter; the spatial distribution characteristic data node represents the emission density of transportation route, regional emission hotspot index and spatial autocorrelation statistic. Through conditional independence test (combined with Fisher Z test and chi-square test), the insignificant dependent relationship is gradually removed, and the optimal directed acyclic graph is constructed based on Bayesian information criterion (BIC). The directed edges in the graph structure represent the potential causal relationship, for example, the positive causal effect of the fuel consumption rate of the construction equipment on the emission density of the transportation route. Further, the maximum likelihood estimation is used to fit the conditional probability distribution of the edges, the continuous variable uses the conditional Gaussian model, and the discrete variable uses the conditional probability table, to generate the carbon emission cross-modal causal relationship data, and to clearly define the causal transmission path and influence strength between different modalities. After obtaining the cross-modal causal relationship data, the modal mutual influence relationship is quantitatively analyzed. According to the edge weight and conditional probability output by the Bayesian network, the mutual influence strength value between different modalities is calculated, which is mapped to the [0, 1] interval by normalization processing, and is used to measure the degree of action of different modalities in the overall causal structure. The modal mutual influence matrix is constructed, which records the influence direction and influence strength between any two types of modalities, for example, the causal weight value of the time series dynamic characteristic on the static interaction characteristic is 0.65, and the causal weight value of the spatial distribution characteristic on the time series dynamic characteristic is 0.42. Through row normalization processing of the matrix, the relative influence contribution value between modalities is obtained, and the weighted importance of the modal to the global carbon emission change is calculated combined with the modal characteristic dimension scale. Based on these results, the modal mutual influence attention weight parameter is designed, the attention weight of the high causal strength modality is set in the interval of 0.6-0.8, and the attention weight of the low causal strength modality is set in the interval of 0.2-0.4, thereby establishing a strict weight allocation mechanism. The carbon emission modal mutual influence correlation characteristic data and the corresponding attention weight parameter are generated, ensuring that the contribution of different modalities in the subsequent fusion is consistent with the real causal relationship. In the attention weighted cross-modal association processing, the three types of modal characteristic vectors are aligned according to the time stamp and spatial label, ensuring that the input features have comparability in the time dimension and spatial dimension.The modal interaction attention weight parameter is used to weight each modal feature, for example, in the fusion of time sequence dynamic features and spatial distribution features, if the causal weight shows that the spatial feature has a greater impact on the time sequence feature, then the spatial feature is given a higher weight value in the calculation of the weighted sum. Through this weighting method, not only the core feature information of each type of modal can be retained, but also the real causal contribution between different modes can be highlighted. In the specific calculation process, first, single-modal weighted aggregation is performed to obtain the weighted feature vectors of time sequence, static and space; then cross-modal weighted combination is performed to obtain global associated modal data. The data contains not only the dynamic change law of the time dimension, but also the constraint condition of the static parameter, and reflects the regional difference of the spatial distribution. The finally generated carbon emission factor associated modal data is a set of high-dimensional vectors, which internally contains the weighted coupling results of each modal feature, and can be used as the input of subsequent neural network nonlinear feature analysis to realize the multi-dimensional fusion expression of road engineering carbon emission.
[0119] Further, step S4 comprises the following steps:
[0120] Step S41: performing link division processing on the carbon emission factor global fusion feature data to obtain link division carbon emission factor global fusion feature data;
[0121] In the embodiment of the application, when performing link division processing on the carbon emission factor global fusion feature data for the whole life cycle, first, the link categories and their time and space boundaries are clearly defined, and the link categories are strictly divided into material production stage, transportation stage, construction stage, maintenance stage and acceptance stage. The material production stage is bounded by the material factory delivery time to the loading time, and the judgment basis is the RFID batch timestamp and the factory scale delivery record; the transportation stage is bounded by the loading time to the unloading time, and the judgment basis is the vehicle trajectory start and end point and the loading state label, and the transportation segment needs to meet the travel duration greater than 10 minutes and the average speed greater than 5 kilometers / hour to be included; the construction stage is bounded by the first entry of the equipment into the construction area to the engineering acceptance time, and the judgment basis is the cumulative working hours of the equipment positioning in the construction area and the contract acceptance document timestamp; the maintenance stage is bounded by the equipment maintenance record time window, and the acceptance stage is bounded by the project completion acceptance time. The link division slices the global fusion features in time sequence, and each link slice output includes link identification, time interval, coverage space boundary, global fusion feature matrix in the time period (summarized by minutes or hours), link-level statistics (cumulative energy consumption, material entry quantity, transportation mileage, average load) and uncertainty estimation item. To ensure mapping accuracy, the link division processing also implements consistency verification rules: batch-level material mass conservation verification, transportation mileage and trajectory mileage comparison, and physical consistency verification of equipment working hours and energy consumption ratio.
[0122] Step S42: Perform multi-factor coupled carbon emission conversion feature analysis of link division according to the link division carbon emission factor global fusion feature data, and generate link division coupled carbon emission conversion feature data;
[0123] In the embodiment of the application, when performing multi-factor coupled carbon emission conversion feature analysis of link division according to the link division carbon emission factor global fusion feature data, a hybrid physical-statistical modeling process is adopted. For each link, an initial conversion equation is first constructed based on the conservation of physical quantities and energy balance relationship, for example, the linear and quadratic relationship between fuel consumption and transportation mileage, loading rate, and vehicle rated oil consumption rate in the transportation stage; a second-order polynomial coupling term is established between equipment energy consumption and operation power, start-stop frequency, and working load rate in the construction stage. Subsequently, data-driven parameter fitting is performed on these initial equations: a generalized additive model is used to smooth fit the nonlinear terms, and a polynomial regression with interaction terms is used to evaluate the high-order coupling effect, the regression process introduces L1 and L2 regularization to suppress overfitting and evaluate the generalization error through time series blocking cross-validation. To depict the propagation path of potential variables and observed variables, a structural equation model is constructed to quantify the indirect influence of hidden factors (such as construction intensity or material humidity) on observed emissions, and the maximum likelihood method is used for parameter estimation and the self-help method is used for swing test of parameter confidence interval. Sensitivity analysis is performed on each link, and local elasticity coefficient and global Sobol index are used to measure the contribution rate of each input variable to link emission, and the result generates a link-level coupled feature matrix, the matrix elements are standardized conversion coefficients, interaction term strength and uncertainty decomposition, which are used for subsequent mapping model and physical interpretation.
[0124] Step S43: Establish the carbon emission factor fusion feature and carbon emission accounting mapping relationship of each link according to the link division coupled carbon emission conversion feature data, and generate a carbon emission accounting relationship model;
[0125] In the embodiment of the present application, when establishing the mapping relationship between the carbon emission factor fusion features of each link and carbon emission accounting based on the coupling carbon emission conversion feature data of the links, an integrated mapper architecture is adopted. The mapper of each link is composed of a random forest regressor and a gradient boosting regression tree in parallel. The random forest is set to have 500 base learners, a maximum tree depth of 20, and a minimum number of leaf nodes of 10 to ensure robust fitting of non-linear and outliers. The gradient boosting regression tree is set to have 1000 iterations, a learning rate of 0.01, and an early stopping method on the validation set to control the number of iterations. The mapper takes the link input feature vector (including global fusion features, expanded items of the coupling feature matrix, and uncertainty indicators) as the independent variable, and outputs the link-level carbon dioxide equivalent estimate and prediction quantiles (0.05, 0.5, 0.95). Historical calibrated emission samples are used for training, and blocked cross-validation is used for evaluation. The loss function is weighted mean squared error, and the weights are set according to the reciprocal of the observation uncertainty. After training, the mapper is subjected to local interpretive analysis. The SHAP value is used to calculate the local contribution of each input feature to the output and generate a feature importance ranking report. To support auditability, each mapper also outputs residual distribution, segmented error statistics, and model stability indicators. The mapper parameters and the prior information version on the training day are archived for comparison when the accounting results need to be traced.
[0126] Step S44: Perform dynamic feature optimization processing on the carbon emission conversion relationship model for actual working condition of carbon emission accounting, and generate an optimized carbon emission accounting relationship model;
[0127] In the embodiment of the present application, when the carbon emission conversion relationship model is actually processed and dynamically optimized, the uncertainty learning and online calibration mechanism is introduced. Firstly, the quantile output by the mapper is subjected to uncertainty consistency test, the prediction interval is constructed by using quantile regression forest, and the confidence level is calibrated by using the prediction interval coverage rate; when the Kullback-Leibler divergence of the distribution of historical prediction residual and the distribution of recent residual exceeds the threshold value 0.2 or the average relative deviation of continuous 24 hours exceeds the set tolerance (for example, 10%), the dynamic adjustment mechanism is triggered. The dynamic adjustment mechanism includes two paths: path one is incremental retraining, the mapper is updated by limited iterations using high-quality calibration samples of the last 7 days, and in the update, the old samples are applied with decay weight to retain long-term information and highlight recent working conditions; path two is model weight adaptive redistribution, the parallel models are weighted and reconstructed according to the recent prediction performance, and the weight of the excellent model is redistributed according to the inverse mean square error. In order to ensure stability, the model rollback point and early stop strategy are implemented in the update process, and the consistency test is completed in the independent verification window before and after the update, and the performance indicators are recorded. All dynamic optimization operations record the version number, trigger factor and update parameters, and a new optimized carbon emission accounting relationship model is generated after the update for real-time accounting, and the confidence interval of the model output is continuously monitored to evaluate the optimization effect.
[0128] Step S45: performing carbon emission intelligent accounting work on the road construction project based on the optimized carbon emission accounting relationship model.
[0129] In the embodiment of the present application, when the carbon emission intelligent accounting work is performed on the road construction project based on the optimized carbon emission accounting relationship model, a hierarchical accounting process and traceable result release are implemented. The accounting is performed according to the predetermined time granularity (for example, by hour, by day, by construction phase), first, the link mapper is called to output the point estimate value and confidence interval of each link time slice, then the engineering level is accumulated and summarized, and the slicing display according to source, link and spatial grid is realized. The accounting result includes fields: time interval, link identifier, source type, carbon dioxide equivalent estimate, uncertainty decomposition and traceability link. In addition, the accounting process outputs trend analysis indicators (compared with the same period, compared with the same period, key driving factor change rate) and abnormal alarm records for management decision, the alarm conditions include sudden increase outside the prediction confidence interval, abnormal rise of single driving factor influence rate and spatial hot spot mutation, and the alarm and backtracking conclusion form a closed loop management, ensuring the accuracy, explainability and traceability of intelligent accounting.
[0130] Further, step S44 includes the following steps:
[0131] According to the link division carbon emission factor global fusion feature data, the uncertainty feature analysis of each link actual working condition is performed, and the link division actual working condition uncertainty feature data is generated;
[0132] The actual working condition tree model of each link is established based on a random forest algorithm, and the actual working condition uncertainty characteristic data of the link division is mapped to the actual working condition tree model of each link for uncertainty characteristic learning and processing of the actual working condition of each link, to generate an actual working condition uncertainty relationship model of the link division carbon emission.
[0133] The actual working condition dynamic characteristic optimization processing of the carbon emission accounting conversion relationship model is performed through the actual working condition uncertainty relationship model of the link division carbon emission, to generate an optimized carbon emission accounting relationship model.
[0134] In the embodiment of the present application, when analyzing the uncertain characteristics of the actual working conditions of each link according to the global fusion feature data of the carbon emission factor, a working condition fluctuation detection mechanism is first established. For the material production link, the standard deviation and coefficient of variation of the production batch energy consumption record are used to calculate the working condition stability, and the CUSUM test is introduced to determine the mutation points in the energy consumption curve, and the mutation point interval is taken as the high-risk interval of uncertainty. For the transportation link, the speed standard deviation, instantaneous acceleration root mean square and idling time proportion of the GPS trajectory speed sequence are calculated, which are combined into a transportation working condition fluctuation index. The larger the index, the higher the uncertainty of the transportation link. For the construction link, based on the time series of engine speed, hydraulic pressure and load rate of construction equipment, sample entropy and multi-scale permutation entropy indicators are extracted. The above indicators constitute the actual working condition uncertainty feature vector of each link, and are simultaneously attached with multi-level uncertainty labels. Each link has one-level labels such as "fluctuation type", "stable type" and "mutation type", and two-level labels such as road fluctuation, road stability, or equipment standby running stability, equipment running fluctuation, etc. Finally, the actual working condition uncertainty feature data of the link division is formed, providing an input feature space for subsequent learning. When establishing the tree model of the actual working condition of each link based on the random forest algorithm, 500 decision trees are used to form the forest, the maximum depth of each tree is limited to 15 layers, the minimum leaf node sample number is set to 5, and the Gini index is used as the splitting standard. The input features are the actual working condition uncertainty feature data of the link division, including fluctuation index, energy consumption standard deviation, speed variation coefficient, sample entropy value, permutation entropy value and other multi-dimensional features, and the output is the link-level uncertainty category and its quantitative score. The self-sampling method is introduced in the training process to ensure sample diversity, and the model generalization performance is evaluated by out-of-bag error (OOB Error). The internal feature importance ranking of the model is calculated by the average impurity reduction, for example, the construction equipment speed entropy indicator ranks first in the construction link uncertainty classification, and the transportation speed fluctuation coefficient has a higher weight in the transportation link. The actual working condition uncertainty relationship model of the link division carbon emission generated by the learning process can quantitatively describe the influence degree of different working condition fluctuation modes on the accuracy of carbon emission accounting. When using the actual working condition uncertainty relationship model of the link division carbon emission to dynamically optimize the conversion relationship model of carbon emission, the output of the uncertainty relationship model is first introduced as a dynamic adjustment factor into the conversion relationship model. Specifically: when the uncertainty score of a certain link is higher than 0.7, the system automatically reduces the weight of the conversion coefficient of the link, and relaxes the upper and lower limits of the confidence interval in carbon emission accounting to reflect the higher uncertainty. For example, in the transportation link, when the speed fluctuation is significantly increased, the conversion relationship weight between transportation fuel consumption and mileage is reduced by 10%, and at the same time, the uncertainty adjustment term is increased. If the uncertainty score is less than 0.3, the original accounting relationship is maintained and the narrow confidence interval is maintained.The sliding time window mechanism is introduced in the dynamic optimization process, and the uncertainty trend in the last 72 hours is taken as the reference to smooth the adjustment of the parameters, avoiding the over-sensitivity of the model caused by single-point abnormality. The final output optimization carbon emission accounting relationship model contains dynamic weight correction term, link-level uncertainty confidence factor and updated mapping function. The optimization model realizes the robust correction of carbon emission accounting results under actual working conditions, ensuring that the accounting can truly reflect the emission level under fluctuating working conditions, and has clear uncertainty boundary and interpretability.
[0135] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, the scope of the application being defined by the appended claims and not by the above description, therefore all variations falling within the meaning and scope of the equivalent requirements of the application file are intended to be included in the present application.
[0136] The above description is merely one specific implementation of the application, which enables those skilled in the art to understand or implement the application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the application. Therefore, the application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A road engineering carbon emission analysis method based on multi-source heterogeneous data fusion, characterized in that, The method comprises the following steps: Step S1: collecting and preprocessing multi-source heterogeneous carbon emission influencing factors of road construction projects in the whole life cycle by using a multi-source integrated sensor system to generate multi-source heterogeneous carbon emission factor data; Step S2: performing multi-source heterogeneous data anomaly identification and correction processing on the multi-source heterogeneous carbon emission factor data to generate multi-source heterogeneous standard carbon emission factor data; Step S3: performing global connection fusion processing on the multi-source heterogeneous standard carbon emission factor data to generate global fusion feature data of carbon emission factors; Step S3 comprises the following steps: Step S31: performing modal division processing on the multi-source heterogeneous standard carbon emission factor data to generate carbon emission factor modal data, wherein the carbon emission factor modal data comprises carbon emission factor time series data, carbon emission factor static structured data, and carbon emission factor spatial data; Step S32: extracting carbon emission time series dynamic characteristics from the carbon emission factor time series data to generate carbon emission time series dynamic characteristic data; Step S33: performing static structured interaction feature analysis on the carbon emission factor static structured data to generate carbon emission factor static interaction feature data; Step S34: performing carbon emission spatial distribution feature analysis on the carbon emission factor spatial data to generate carbon emission factor spatial distribution feature data; Step S35: performing cross-modal correlation processing on the carbon emission factor based on the carbon emission time series dynamic characteristic data, the carbon emission factor static interaction feature data, and the carbon emission factor spatial distribution feature data based on a Bayesian network algorithm to generate carbon emission factor correlation modal data; Step S36: performing correlation modal nonlinear feature analysis on the carbon emission factor correlation modal data based on a neural network algorithm to generate carbon emission factor correlation modal nonlinear feature data, and performing global connection processing on the correlation modal nonlinear features based on the carbon emission factor correlation modal nonlinear feature data to generate global fusion feature data of carbon emission factors; Step S4: establishing an optimized relationship model for actual working condition carbon emission accounting based on the global fusion feature data of carbon emission factors to generate an optimized carbon emission accounting relationship model; and performing carbon emission intelligent accounting for road construction projects based on the optimized carbon emission accounting relationship model; The establishment of the optimized relationship model for actual working condition carbon emission accounting based on the global fusion feature data of carbon emission factors comprises: Step S41: performing link division processing on the global fusion feature data of carbon emission factors to obtain link division carbon emission factor global fusion feature data; Step S42: performing multi-factor coupled carbon emission conversion feature analysis based on the link division carbon emission factor global fusion feature data to generate link division coupled carbon emission conversion feature data; Step S43: establishing a carbon emission factor fusion feature and carbon emission accounting mapping relationship for each link based on the link division coupled carbon emission conversion feature data to generate a carbon emission accounting relationship model; Step S44: performing actual working condition dynamic feature optimization processing on the carbon emission conversion relationship model for carbon emission accounting to generate an optimized carbon emission accounting relationship model.
2. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 1, characterized in that, The multi-source integrated sensor system includes a construction equipment energy consumption sensor module, a building material monitoring sensor module, and a construction site monitoring sensor module, and step S1 includes the following steps: Step S11: collecting construction equipment energy consumption data of the road construction project during the construction period using the construction equipment energy consumption sensor module to obtain construction equipment energy consumption data; Step S12: collecting carbon emission data of the road construction project during the road building material production and transportation stage using the building material monitoring sensor module to obtain material production and transportation carbon emission data; Step S13: monitoring and processing site environment data of the road construction project during the construction period using the construction site monitoring sensor module to obtain construction environment data; Step S14: performing multi-source heterogeneous integration processing of carbon emission factors for the entire life cycle of the construction equipment energy consumption data, the material production and transportation carbon emission data, and the construction environment data to obtain preliminary carbon emission factor multi-source heterogeneous data; Step S15: performing data preprocessing on the preliminary carbon emission factor multi-source heterogeneous data to generate carbon emission factor multi-source heterogeneous data.
3. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 1, characterized in that, Step S2 includes the following steps: Step S21: analyzing the heterogeneous type characteristics based on the carbon emission factor multi-source heterogeneous data to generate carbon emission factor heterogeneous type characteristic data; Step S22: establishing a data anomaly identification relationship model for the carbon emission factor heterogeneous type based on the carbon emission factor heterogeneous type characteristic data to obtain a carbon emission factor heterogeneous type anomaly identification model; Step S23: transmitting the carbon emission factor multi-source heterogeneous data to the carbon emission factor heterogeneous type anomaly identification model to perform anomaly data identification processing for each heterogeneous type, outputting multi-source heterogeneous first abnormal data and / or multi-source heterogeneous preliminary valid data; Step S24: performing fuzzy quantization feature conversion processing on the multi-source heterogeneous preliminary valid data to obtain multi-source heterogeneous valid fuzzy quantization feature data; Step S25: transmitting the multi-source heterogeneous valid fuzzy quantization feature data to the carbon emission factor heterogeneous type anomaly identification model to perform anomaly fuzzy feature similarity analysis processing for each heterogeneous type valid data, outputting multi-source heterogeneous second abnormal data and / or outputting carbon emission factor multi-source heterogeneous verification data; Step S26: performing anomaly data correction processing for each heterogeneous type based on the multi-source heterogeneous first abnormal data and the multi-source heterogeneous second abnormal data to obtain multi-source heterogeneous anomaly correction data, and feeding back the multi-source heterogeneous anomaly correction data to the carbon emission factor multi-source heterogeneous verification data for anomaly data repair feedback adjustment processing of the carbon emission factor to generate carbon emission factor multi-source heterogeneous standard data.
4. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 3, characterized in that, Step S22 includes the following steps: Performing data anomaly priori data analysis of the heterogeneous type based on the historical data of the carbon emission factor multi-source heterogeneous data to generate carbon emission factor heterogeneous type anomaly priori data; Performing anomaly feature index analysis of the carbon emission factor heterogeneous type based on the carbon emission factor heterogeneous type characteristic data to generate carbon emission factor heterogeneous type anomaly feature index data, and establishing a preliminary carbon emission factor heterogeneous type anomaly identification model based on the carbon emission factor heterogeneous type anomaly feature index data; According to the carbon emission factor isomerism type anomaly priori data, the preliminary carbon emission factor isomerism type anomaly recognition model is subjected to isomerism type anomaly checking model training and parameter weight adjustment processing, and a carbon emission factor isomerism type anomaly recognition model is generated.
5. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 3, characterized in that, Step S23 includes: when the carbon emission factor isomerism type anomaly recognition model identifies the matching abnormal data in the carbon emission factor multi-source isomerism data, outputting the multi-source isomerism first abnormal data, and / or when the carbon emission factor isomerism type anomaly recognition model does not identify the matching abnormal data in the carbon emission factor multi-source isomerism data, outputting the multi-source isomerism preliminary effective data.
6. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 3, characterized in that, Step S24 includes: when the carbon emission factor isomerism type anomaly recognition model identifies that the abnormal fuzzy feature similarity score of the multi-source isomerism effective fuzzy quantization feature data is greater than the preset abnormal similarity score threshold, outputting the multi-source isomerism second abnormal data, and / or when the carbon emission factor isomerism type anomaly recognition model identifies that the abnormal fuzzy feature similarity score of the multi-source isomerism effective fuzzy quantization feature data is not greater than the preset abnormal similarity score threshold, outputting the carbon emission factor multi-source isomerism checking data.
7. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 1, characterized in that, Step S35 includes the following steps: Through the Bayesian network algorithm, carbon emission cross-modal causal relationship data is generated by analyzing the carbon emission time series dynamic feature data, the carbon emission factor static interaction feature data and the carbon emission factor spatial distribution feature data. According to the carbon emission cross-modal causal relationship data, carbon emission modal mutual influence correlation feature data is generated by analyzing the carbon emission modal mutual influence correlation feature data, and the carbon emission modal mutual influence attention weight parameter is designed according to the carbon emission modal mutual influence correlation feature data. Based on the carbon emission modal mutual influence attention weight parameter, the carbon emission time series dynamic feature data, the carbon emission factor static interaction feature data and the carbon emission factor spatial distribution feature data are subjected to attention weighted cross-modal correlation processing to generate carbon emission factor correlation modal data.
8. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 1, characterized in that, Step S44 includes the following steps: According to the link division carbon emission factor global fusion feature data, link division actual working condition uncertain feature data is generated by analyzing the uncertain features of each link actual working condition. Based on the random forest algorithm, each link actual working condition tree model is established, and the link division actual working condition uncertain feature data is mapped to each link actual working condition tree model for uncertain feature learning processing of each link actual working condition, and a link division carbon emission actual working condition uncertain relationship model is generated. Through the link division carbon emission actual working condition uncertain relationship model, the carbon emission conversion relationship model is subjected to carbon emission accounting actual working condition dynamic feature optimization processing to generate an optimized carbon emission accounting relationship model.
Citation Information
Patent Citations
Highway construction project carbon emission accounting evaluation method and system
CN119624168A
Method and system for dynamically monitoring carbon emission of Yellow River basin by using big data
CN119850229A