Road engineering carbon emission analysis method based on multi-source heterogeneous data fusion

By using a multi-source integrated sensor system and multi-source heterogeneous data processing technology, the problems of full life-cycle coverage and data fusion in carbon emission analysis of road engineering have been solved, achieving efficient and accurate carbon emission accounting and adapting to intelligent analysis under complex working conditions.

CN121073006AActive Publication Date: 2025-12-05HUNAN COMM RES INST CO LTD

Patent Information

Application Number
CN202511612905.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2025-12-05
Estimated Expiration
2045-11-06

AI Technical Summary

Technical Problem

Existing carbon emission analysis methods for road engineering lack full life-cycle coverage and the ability to integrate multi-source heterogeneous data, making it difficult to fully explore the potential correlations between data. The methods for identifying and correcting abnormal data of carbon emission factors are limited, data quality is difficult to guarantee, and the dynamic optimization capability of the accounting model is limited, making it difficult to meet the needs of intelligent carbon emission accounting under complex working conditions.

Method used

A multi-source integrated sensor system is used to collect and preprocess multi-source heterogeneous carbon emission influencing factors throughout the entire life cycle of road construction projects. Through anomaly identification and correction of multi-source heterogeneous data, multi-source heterogeneous standard data of carbon emission factors are generated, and cross-modal global connection fusion processing is performed to establish an optimized carbon emission accounting relationship model and realize intelligent carbon emission accounting.

Benefits of technology

It achieves comprehensive collection and systematic analysis of carbon emission-related influencing factors, improves the efficiency and accuracy of multi-source data fusion, identifies and corrects abnormal data, generates highly consistent and reliable carbon emission accounting models, and can adapt to the intelligent accounting needs under complex working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121073006A_ABST
    Figure CN121073006A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of carbon emission accounting, in particular to a road engineering carbon emission analysis method based on multi-source heterogeneous data fusion. The method comprises the following steps: collecting carbon emission factor multi-source heterogeneous data; performing multi-source heterogeneous data anomaly identification and correction processing on the carbon emission factor multi-source heterogeneous data to generate carbon emission factor multi-source heterogeneous standard data; performing carbon emission factor cross-modal global connection fusion processing on the carbon emission factor multi-source heterogeneous standard data to generate carbon emission factor global fusion feature data; establishing an optimization relation model of actual working condition carbon emission accounting based on the carbon emission factor global fusion feature data, and generating an optimization carbon emission accounting relation model; and performing carbon emission intelligent accounting operation on the road construction project based on the optimized carbon emission accounting relation model. According to the method, the accurate accounting of the carbon emission of the road engineering is realized by carrying out fusion analysis on the multi-source heterogeneous data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of carbon emission accounting, and in particular to a road engineering carbon emission analysis method based on multi-source heterogeneous data fusion. BACKGROUND

[0002] With the acceleration of urbanization, the number and scale of road construction projects are increasing, and engineering construction activities not only promote social and economic development, but also bring about significant carbon emission problems. Road construction involves construction equipment energy consumption, building material production and transportation, construction site environment and other aspects, each of which directly or indirectly affects the total carbon emission. However, the existing road engineering carbon emission analysis method lacks comprehensive collection and analysis of the whole life cycle of road engineering, and the fusion ability of multi-source heterogeneous data is insufficient, making it difficult to fully explore the potential correlation between data; the means of abnormal data identification and correction of carbon emission factors is single, the data quality is difficult to guarantee, and the dynamic optimization ability of the accounting model is limited, which is difficult to meet the intelligent carbon emission accounting demand under complex working conditions. SUMMARY

[0003] Therefore, the present application provides a road engineering carbon emission analysis method based on multi-source heterogeneous data fusion to solve at least one of the above technical problems.

[0004] To achieve the above purpose, a road engineering carbon emission analysis method based on multi-source heterogeneous data fusion comprises the following steps: Step S1: using a multi-source integrated sensor system to collect and preprocess multi-source heterogeneous carbon emission influencing factors of road construction projects throughout their life cycle, to generate multi-source heterogeneous carbon emission factor data; Step S2: performing multi-source heterogeneous data anomaly identification and correction processing on the multi-source heterogeneous carbon emission factor data to generate multi-source heterogeneous standard carbon emission factor data; Step S3: performing global connection fusion processing on the multi-source heterogeneous standard carbon emission factor data to generate global fusion feature data of carbon emission factors; Step S4: establishing an optimized relationship model for actual working condition carbon emission accounting based on the global fusion feature data of carbon emission factors to generate an optimized carbon emission accounting relationship model; and performing carbon emission intelligent accounting operations on road construction projects based on the optimized carbon emission accounting relationship model.

[0005] Further, the multi-source integrated sensor system comprises a construction equipment energy consumption sensor module, a building material monitoring sensor module and a construction site monitoring sensor module, and step S1 comprises the following steps: Step S11: Collecting construction equipment energy consumption data of the road construction project during the construction period using the construction equipment energy consumption sensor module to obtain construction equipment energy consumption data; Step S12: Collecting carbon emission data of the road construction project during the road construction material production and transportation stage using the construction material monitoring sensor module to obtain material production and transportation carbon emission data; Step S13: Monitoring and processing site environment data of the road construction project during the construction period using the construction site monitoring sensor module to obtain construction environment data; Step S14: Processing carbon emission factors of the whole life cycle multi-source heterogeneous data of construction equipment energy consumption data, material production and transportation carbon emission data, and construction environment data to obtain preliminary carbon emission factor multi-source heterogeneous data; Step S15: Preprocessing the preliminary carbon emission factor multi-source heterogeneous data to generate carbon emission factor multi-source heterogeneous data.

[0006] Further, step S2 includes the following steps: Step S21: Analyzing the heterogeneous type characteristics according to the carbon emission factor multi-source heterogeneous data to generate carbon emission factor heterogeneous type characteristic data; Step S22: Establishing a carbon emission factor heterogeneous type data anomaly recognition relationship model based on the carbon emission factor heterogeneous type characteristic data to obtain a carbon emission factor heterogeneous type anomaly recognition model; Step S23: Transferring the carbon emission factor multi-source heterogeneous data to the carbon emission factor heterogeneous type anomaly recognition model to perform anomaly data recognition processing of each heterogeneous type, and outputting multi-source heterogeneous first anomaly data and / or multi-source heterogeneous preliminary valid data; Step S24: Performing fuzzy quantization feature conversion processing on the multi-source heterogeneous preliminary valid data to obtain multi-source heterogeneous valid fuzzy quantization feature data; Step S25: Transferring the multi-source heterogeneous valid fuzzy quantization feature data to the carbon emission factor heterogeneous type anomaly recognition model to perform anomaly fuzzy feature similarity analysis processing of each heterogeneous type valid data, and outputting multi-source heterogeneous second anomaly data and / or outputting carbon emission factor multi-source heterogeneous verification data; Step S26: Performing anomaly data correction processing of each heterogeneous type based on the multi-source heterogeneous first anomaly data and the multi-source heterogeneous second anomaly data to obtain multi-source heterogeneous anomaly correction data, and feeding back the multi-source heterogeneous anomaly correction data to the carbon emission factor multi-source heterogeneous verification data for carbon emission factor anomaly data repair feedback adjustment processing to generate carbon emission factor multi-source heterogeneous standard data.

[0007] Further, step S22 includes the following steps: According to the historical data of the multi-source heterogeneous data of the carbon emission factor, the heterogeneous type of data abnormal prior data analysis is performed, and heterogeneous type of carbon emission factor abnormal prior data is generated; According to the carbon emission factor heterogeneous type characteristic data, the abnormal characteristic index analysis of the carbon emission factor heterogeneous type is performed, and the carbon emission factor heterogeneous type abnormal characteristic index data is generated. Based on the carbon emission factor heterogeneous type abnormal characteristic index data, a preliminary carbon emission factor heterogeneous type abnormal recognition model is established. According to the carbon emission factor heterogeneous type abnormal prior data, the model training and parameter weight adjustment processing of the preliminary carbon emission factor heterogeneous type abnormal recognition model are performed for each heterogeneous type of abnormal verification, and the carbon emission factor heterogeneous type abnormal recognition model is generated.

[0008] Further, step S23 includes: when the carbon emission factor heterogeneous type abnormal recognition model identifies the matching abnormal data in the multi-source heterogeneous data of the carbon emission factor, outputting the multi-source heterogeneous first abnormal data, and / or when the carbon emission factor heterogeneous type abnormal recognition model does not identify the matching abnormal data in the multi-source heterogeneous data of the carbon emission factor, outputting the multi-source heterogeneous preliminary effective data.

[0009] Further, step S24 includes: when the carbon emission factor heterogeneous type abnormal recognition model identifies that the abnormal fuzzy feature similarity score of the multi-source heterogeneous effective fuzzy quantization feature data is greater than the preset abnormal similarity score threshold, outputting the multi-source heterogeneous second abnormal data, and / or when the carbon emission factor heterogeneous type abnormal recognition model identifies that the abnormal fuzzy feature similarity score of the multi-source heterogeneous effective fuzzy quantization feature data is not greater than the preset abnormal similarity score threshold, outputting the carbon emission factor multi-source heterogeneous verification data.

[0010] Further, step S3 includes the following steps: Step S31: The carbon emission factor modal division processing is performed on the carbon emission factor multi-source heterogeneous standard data, and the carbon emission factor modal data is generated, wherein the carbon emission factor modal data includes carbon emission factor time series data, carbon emission factor static structured data, and carbon emission factor spatial data. Step S32: The carbon emission factor time series data is subjected to carbon emission time series dynamic feature extraction, and carbon emission time series dynamic feature data is generated. Step S33: The carbon emission factor static structured data is subjected to static structured interactive feature analysis, and carbon emission factor static interactive feature data is generated. Step S34: The carbon emission factor spatial data is subjected to carbon emission spatial distribution feature analysis, and carbon emission factor spatial distribution feature data is generated. Step S35: Based on the Bayesian network algorithm, the carbon emission time series dynamic characteristic data, the carbon emission factor static interaction characteristic data and the carbon emission factor spatial distribution characteristic data are processed for cross-modal correlation of carbon emission factors, and carbon emission factor correlation modal data is generated; Step S36: Based on the neural network algorithm, the carbon emission factor correlation modal data is analyzed for nonlinear characteristics of the correlation modal, and carbon emission factor correlation modal nonlinear characteristic data is generated. According to the carbon emission factor correlation modal nonlinear characteristic data, global connection processing of the nonlinear characteristics of the correlation modal is performed, and carbon emission factor global fusion characteristic data is generated.

[0011] Further, step S35 includes the following steps: The carbon emission time series dynamic characteristic data, the carbon emission factor static interaction characteristic data and the carbon emission factor spatial distribution characteristic data are analyzed for carbon emission cross-modal causal relationship by the Bayesian network algorithm, and carbon emission cross-modal causal relationship data is generated; According to the carbon emission cross-modal causal relationship data, carbon emission modal mutual influence correlation characteristic analysis is performed, and carbon emission modal mutual influence correlation characteristic data is generated. According to the carbon emission modal mutual influence correlation characteristic data, carbon emission modal mutual influence attention weight parameter is designed; Based on the carbon emission modal mutual influence attention weight parameter, the carbon emission time series dynamic characteristic data, the carbon emission factor static interaction characteristic data and the carbon emission factor spatial distribution characteristic data are processed for attention weighted cross-modal correlation, and carbon emission factor correlation modal data is generated.

[0012] Further, step S4 includes the following steps: Step S41: The carbon emission factor global fusion characteristic data is processed for life cycle division, to obtain life cycle division carbon emission factor global fusion characteristic data; Step S42: According to the life cycle division carbon emission factor global fusion characteristic data, life cycle division multi-factor coupled carbon emission conversion characteristic analysis is performed, and life cycle division coupled carbon emission conversion characteristic data is generated; Step S43: According to the life cycle division coupled carbon emission conversion characteristic data, carbon emission factor fusion characteristic and carbon emission accounting mapping relationship of each link is established, and a carbon emission accounting relationship model is generated; Step S44: The carbon emission conversion relationship model is processed for carbon emission accounting actual working condition dynamic characteristic optimization, and an optimized carbon emission accounting relationship model is generated; Step S45: Based on the optimized carbon emission accounting relationship model, carbon emission intelligent accounting operation is performed for road construction engineering.

[0013] Further, step S44 includes the following steps: According to the link division carbon emission factor global fusion characteristic data, the uncertainty characteristics of each link actual working condition are analyzed, and link division actual working condition uncertainty characteristic data are generated; Based on the random forest algorithm, a tree model of each link actual working condition is established, and the link division actual working condition uncertainty characteristic data are mapped to the tree model of each link actual working condition for uncertainty characteristic learning and processing of each link actual working condition, thereby generating a link division carbon emission actual working condition uncertainty relationship model; Through the link division carbon emission actual working condition uncertainty relationship model, the actual working condition dynamic characteristic optimization processing of the carbon emission conversion relationship model for carbon emission accounting is performed, and an optimized carbon emission accounting relationship model is generated.

[0014] The application has the advantages that the application can comprehensively and systematically collect carbon emission related influencing factors by introducing a multi-source integrated sensor system in the whole life cycle of road construction engineering. Not only the construction equipment energy consumption data is considered, but also the carbon emission data of building material production and transportation process and the construction site environment data are introduced, so that the multi-dimension and multi-link coverage of data sources is ensured. Through multi-source heterogeneous integration processing of the construction equipment energy consumption data, material production and transportation carbon emission data and construction environment data, the problems of different data sources, such as different formats, different sampling frequencies and large dimension differences, are effectively solved, and the structured unified expression of heterogeneous data is realized. The integration mechanism significantly improves the fusion efficiency of multi-source data. The preprocessing step is introduced after data collection, which can filter noise, complete missing values and standardize the preliminary multi-source heterogeneous data, effectively improve the integrity and consistency of the data, and reduce the interference of abnormal data and redundant data on the analysis results from the source. Based on the multi-source heterogeneous data, the heterogeneous type characteristic analysis is carried out, so as to provide structured feature support for subsequent anomaly detection, and reveal the essential difference of different types of data. By introducing the abnormal recognition relationship model based on heterogeneous characteristic data and training and weight adjustment through historical prior data, the data abnormal recognition model for different heterogeneous types is generated, which not only has the ability to identify normal abnormal data, but also can adapt to the complex actual situation in road engineering scene, ensuring the accuracy and adaptability of the abnormal recognition result. Through fuzzy quantification feature conversion and fuzzy similarity analysis, the fuzzy analysis method is introduced to carry out secondary verification on the preliminary effective data, and identify the potential abnormality hidden in the boundary state, which is especially suitable for the data scene with continuous deviation in complex environment. Through the two-stage abnormality recognition (preliminary identification + fuzzy similarity analysis), the integrity of abnormal data identification is maximized. Through the feedback adjustment mechanism of abnormal correction data and verification data, dynamic repair and iterative optimization are realized. This feedback correction method can continuously improve the self-adaptability of the model while correcting the data, ensuring that the finally generated multi-source heterogeneous standard data has high consistency and high reliability. In the carbon emission analysis process, the modal difference of multi-source heterogeneous data is significant, for example, the construction equipment energy consumption data is time series data, the material transportation data is mostly static structured data, and the construction environment monitoring data has obvious spatial distribution characteristics. The modal division and feature extraction of the standardized carbon emission factor data effectively realize the fine modeling of time series features, static features and spatial features, and improve the comprehensiveness and pertinence of feature expression. The Bayesian network algorithm is introduced to analyze the causal relationship of cross-modal data, which can not only identify the direct or indirect relationship between different modalities, but also establish the causal chain of modal interaction. Through the attention weighting mechanism, different weights are given to different modal features, which can automatically highlight the modal features that have a greater impact on carbon emission results, and enhance the expression efficiency and robustness of the fused data.The non-linear feature analysis of the cross-modal causal relationship combined with the neural network algorithm can effectively capture the complex coupling relationship between the carbon emission factors. The global connection mechanism is used to realize the fusion of cross-modal non-linear features, and the generated global fusion feature data can integrate multi-dimensional information to the greatest extent, ensuring the high correlation and high explanatory power of the input data of the subsequent carbon emission accounting model. Not only does it solve the problem of efficient fusion of multi-source heterogeneous data in the prior art, but also significantly improves the depth and breadth of data mining, laying a solid foundation for establishing a precise carbon emission accounting model. By refining and dividing the whole life cycle into multiple key links such as construction equipment use, material transportation and on-site environment, the carbon emission accounting is accurately decomposed into multiple key links, thereby realizing the quantitative modeling of each link and avoiding the errors caused by the averaging process in the traditional overall analysis method. Based on the global fusion feature data divided by the link, the carbon emission conversion feature analysis of multi-factor coupling can capture the interaction between each link. For example, the non-linear interaction between construction equipment energy consumption and construction environmental conditions is reflected through multi-factor coupling modeling, making the accounting results more consistent with the actual working conditions. The random forest algorithm is introduced to model the uncertain factors, which can improve the adaptability to complex working conditions through the ensemble learning mechanism of the tree model. Especially in the scene of road construction engineering with complex and variable construction environment and equipment state, the uncertain characteristics of each link are effectively identified and learned, thereby optimizing the dynamic response capability of the carbon emission accounting relationship model. The finally generated optimized carbon emission accounting relationship model not only has high precision and robustness, but also can be dynamically adjusted according to the actual working conditions to support the intelligent carbon emission accounting operation of road engineering.

[0015] Therefore, the road engineering carbon emission analysis method based on multi-source heterogeneous data fusion can cover the whole life cycle of road engineering and has the ability to deeply process the fusion of multi-source heterogeneous data to ensure that the potential association between data is fully mined. Through the establishment of the carbon emission factor heterogeneous type anomaly recognition model and the multi-level data anomaly recognition, the abnormal data of the carbon emission factor is accurately identified, and through the analysis of the coupling characteristics of the carbon emission factor and the actual operation condition, the specific carbon emission that conforms to the actual operation condition is calculated, meeting the intelligent carbon emission accounting demand under complex working conditions. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 The figure is a step flowchart of the road engineering carbon emission analysis method based on multi-source heterogeneous data fusion of the present application; Figure 2 The figure is a step flowchart of the road engineering carbon emission analysis method based on multi-source heterogeneous data fusion of the present application; Figure 1 The figure is a detailed implementation step flowchart of step S3 in the present application; The implementation, functional characteristics and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0017] The technical method of the present application will be described clearly and completely in combination with the drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0018] In addition, the drawings are only schematic illustrations of the present application, and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated description thereof will be omitted. Some block diagrams shown in the drawings are functional entities, which do not necessarily have to correspond to physically or logically independent entities. The functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.

[0019] To achieve the above-mentioned object, please refer to Figures 1 to 2 The present application provides a road engineering carbon emission analysis method based on multi-source heterogeneous data fusion. In the embodiments of the present application, please refer to Figure 1 The present application provides a road engineering carbon emission analysis method based on multi-source heterogeneous data fusion. In the embodiments of the present application, please refer to Step S1: Collecting and data preprocessing of multi-source heterogeneous carbon emission influencing factors of road construction engineering in the whole life cycle by using a multi-source integrated sensor system, and generating multi-source heterogeneous carbon emission factor data; In the embodiment of the present application, the sensor modules are arranged according to functions at the construction site: the construction equipment energy consumption sensor module includes a fuel flow meter (pulse output, sampling frequency 10 Hz), an electric energy meter (pulse / pulse width output, sampling frequency 1 Hz), an engine speed and work hour meter (CAN bus interface, sampling frequency 1 Hz), and a hydraulic system pressure sensor (sampling frequency 5 Hz); the building material monitoring sensor module includes a load cell / load sensor (weighing accuracy ±0.5%), a truck-mounted GPS locator (frequency 1 Hz), and a passive RFID tag for batch identification; the construction site monitoring sensor module includes a temperature and humidity sensor (minute-level sampling), a wind speed and direction meter (second-level sampling), a particulate matter sensor (PM2.5, PM10, minute-level sampling), and a noise level meter. Before being put into use, each sensor device is calibrated according to a unified calibration procedure: the fuel flow meter is verified by a standard flow source, the electric energy meter is verified by a reference electric meter, and the weighing sensor is verified by a standard weight. The construction equipment energy consumption sensor module, the building material monitoring sensor module, and the construction site monitoring sensor module are integrated into a multi-source integrated sensor system, and each carbon emission factor data collected through the multi-source integrated sensor system is integrated. Time synchronization is implemented by receiving a reference time through satellite positioning (GNSS) and combining a network time protocol (NTP), and all sampling records are provided with accurate time stamps, sensor identifiers, and geographic coordinates. After original sampling, a preprocessing chain is executed: a fourth-order low-pass Butterworth filter is first applied to continuous signals to remove high-frequency oscillation noise, and a sliding median filter is applied to isolated spikes to eliminate pulse anomalies; cubic spline interpolation is used to complete short-time missing data (less than 5 minutes); long-time missing data (more than 5 minutes) are estimated using a model-driven interpolation method based on nearest neighbor time series similarity and recording the uncertainty; discrete events (such as material loading batches) are batched and standard record items are generated using event time as an anchor point. Unit standardization is completed in the preprocessing stage: after the fuel volume is converted into mass, it is multiplied by the greenhouse gas equivalent coefficient corresponding to the fuel category, the electric energy is calculated by multiplying the daily time amplitude by the regional power grid time period emission factor, and the material is calculated by multiplying the mass by the material unit equivalent coefficient. The preprocessing output is a structured record set, and the record fields include: time stamp, sensor ID, measurement type, measurement value, unit of measurement, geographic coordinates, operation status identifier, quality flag bit, and initial uncertainty estimation, which provides a unified and traceable data basis for subsequent anomaly identification and cross-modal fusion.

[0020] Step S2: multi-source heterogeneous data anomaly identification and correction processing is performed on the carbon emission factor multi-source heterogeneous data to generate carbon emission factor multi-source heterogeneous standard data; In the embodiment of the present application, after obtaining the multi-source heterogeneous data of carbon emission factors, first, the heterogeneous type characteristic library is established according to the data type: the statistical characteristics (mean, variance, skewness, kurtosis, interquartile range), time domain characteristics (maximum slope, start-stop count, duty cycle ratio), frequency domain characteristics (main frequency and bandwidth of power spectral density), information theory characteristics (sample entropy, self-entropy) and autocorrelation coefficient vector of each time series signal are calculated in the sliding window (window length 30 minutes, step 5 minutes), forming a high-dimensional feature description vector of each signal, and the same type of sensor grouping and type labeling are carried out accordingly. Based on the above characteristic data, a hierarchical anomaly recognition relationship model is constructed: for single variable continuous signals, a detector based on isolation forest combined with statistical threshold rules based on sliding window is constructed, the threshold rule is determined by the quantile of the historical running distribution and saved as a type priori; for the overall correlation anomaly of multi-variable time series, a long short-term memory autoencoder (LSTM autoencoder) is constructed, and the correlation anomaly is recognized through the reconstruction error; for the indicators meeting the physical constraints, a rule-based logical verification module is added (for example, fuel consumption cannot be negative, instantaneous current cannot exceed the rated value of the device), and all detectors are cross-validated and grid searched to determine the hyperparameters and record the model performance indicators on the historical labeled anomaly sample set. In the running, the preprocessed output is input into the anomaly recognition model according to the sliding window, and the model outputs the first abnormal data or the preliminary valid data for each time window. The preliminary valid data is further subjected to fuzzy quantization feature conversion: for continuous variables, define fuzzy membership functions (triangular membership or trapezoidal membership), map the numerical values to membership degree vectors and use the barycentric method to numerize them into fuzzy quantization feature vectors. The fuzzy quantization feature vector is sent to the same anomaly recognition relationship model to perform fuzzy feature similarity analysis, cosine similarity and fuzzy similarity indicators are adopted, and compared with the similarity threshold value established in advance, if the similarity exceeds the threshold value, it is labeled as the second type of anomaly and output; if it does not exceed the threshold value, it is output as multi-source heterogeneous verification data. For the first anomaly and the second anomaly labeled, a hierarchical correction strategy is adopted according to the anomaly type: Kalman smoothing correction is applied to short-time noise and isolated spikes, drift compensation based on state space model is used to estimate the baseline and correct for sensor drift, Gaussian process regression is used to estimate the missing or truncated data and output the prediction uncertainty at the same time, and the items involving multi-source information contradiction are repaired for consistency by cross-source regression (using related sensor features to construct a regressor), and the error measurement before and after correction is recorded. All the correction results form the multi-source heterogeneous anomaly correction data and are written back as multi-source heterogeneous verification data to trigger model feedback training. Periodically (for example, once a week), retrain and weight adjustment are performed on the anomaly recognition model based on the corrected data, and the training sample weight and reconstruction error threshold are adjusted according to the precision and recall rate of cross-validation after retraining, so as to form a closed loop of anomaly recognition and correction process and output the consistent high carbon emission factor multi-source heterogeneous standard data verified.

[0021] Step S3: carbon emission factor multi-source heterogeneous standard data is processed by carbon emission factor cross-modal global connection fusion to generate carbon emission factor global fusion feature data; In the embodiment of the present application, after obtaining the carbon emission factor multi-source heterogeneous standard data, first, the data is divided according to the data modal: the time series data includes high-frequency time series of construction equipment energy consumption, equipment operation state and sensor continuous monitoring sequence; the static structured data includes material batch attribute, material unit equivalent coefficient, equipment rated parameter and contract data; the spatial data includes GPS track point, material unloading position coordinate and construction site boundary vector. Multi-scale dynamic feature extraction is performed on the time series data: short-time window (5 minutes) calculates instantaneous power, start-stop times and slope, long-time window (24 hours) calculates daily cumulative amount, load curve features and extracts scale coefficient of transient energy consumption pulse through continuous wavelet transform; the static structured data is constructed into an interaction feature matrix, such as the product of material quality and arrival time, the ratio of unit carbon dioxide equivalent of single material to use time, and a fixed-length binary vector encoding is applied to the category variable for subsequent network input; the spatial data is mapped and spatially counted: the track points are aligned to the road section by using the map matching algorithm, the transportation distance and stay time are aggregated by road section, and the spatial distribution grid based on 200-meter grid is generated by Kriging interpolation to express the local emission density. Cross-modal causal relationship analysis is performed based on Bayesian network: the modal features are nodes, the structure learning is performed according to the scoring criterion and conditional independence test method to determine the directed edge, and the conditional probability table is obtained by maximum likelihood estimation, the generated cross-modal causal relationship is used to estimate the influence strength between modes and design the modal interaction attention weight parameter. On this basis, a multi-branch neural network is constructed for nonlinear feature analysis: the time branch is composed of a one-dimensional convolution layer followed by a long short-term memory layer to extract time series deep dynamic features; the static branch is composed of several fully connected layers to express structured interaction information; the spatial branch adopts a graph convolution network to capture the topological dependence between road sections. The outputs of each branch are weighted and aggregated by the cross-modal attention module, the attention weight is initialized based on the causal influence strength derived from the Bayesian network and updated by the gradient method during the training process. The training objective function uses the sum of mean square error and L2 regularization term and updates the parameters with an adaptive first-order momentum optimizer, the early stopping mechanism is implemented during the training process to prevent overfitting, and after the network training is completed, a global fusion feature vector of uniform dimension is output for each time window.

[0022] Step S4: based on the carbon emission factor global fusion feature data, an optimized relationship model for actual working condition carbon emission accounting is established to generate an optimized carbon emission accounting relationship model; based on the optimized carbon emission accounting relationship model, the carbon emission intelligent accounting operation is performed on the road construction project.

[0023] In the embodiment of the present application, after obtaining the global fusion characteristics of carbon emission factors, the life cycle stages are divided: the stage category is clearly defined as the material production stage (from the material production time to the loading time), the transportation stage (from the loading time to the unloading time, the road section meeting the conditions of average speed greater than 5 km / h and single trip duration not less than 10 minutes is the transportation section), the construction stage (from the installation to the completion of the construction acceptance period, the equipment with running hours greater than zero and located in the construction area is included), and the maintenance stage and the operation stage are defined and summarized respectively according to the acceptance record time. Perform multi-factor coupled carbon emission conversion characteristic analysis on each stage: use generalized additive model with interaction term and multinomial regression to factorize the input characteristics and calculate the interaction effect coefficient, construct a structural equation model for nonlinear interaction to quantify the transmission path between latent variables and observed variables, and obtain the conversion characteristic parameter matrix of the stage level. Establish a mapping relationship model based on the coupling characteristics of the stage division: for each stage, construct an integrated mapper based on random forest regression and gradient boosting regression tree, the mapper takes the global fusion characteristics as the input and output as the stage carbon emission estimation value, and uses historical calibrated measured emission data for supervised learning during the training process and uses cross-validation to evaluate the generalization error; After model training, the feature importance ranking and local interpretability indicators (such as SHAP value) are exported to support the interpretability of the accounting results. In order to realize the dynamic optimization of the accounting under actual working conditions, uncertainty learning is introduced: the Bootstrap sample of the random forest is used to construct the prediction quantile and output the confidence interval through the quantile regression forest, and the threshold is applied to the prediction uncertainty, when the uncertainty interval exceeds the preset tolerance, the model adaptive adjustment module is triggered; adaptive adjustment implements incremental retraining of the mapper through a sliding time window (example: the last 7 days of data) and updates the mapping parameters combined with the latest calibration data to correct the bias. The final optimized carbon emission accounting relationship model is used to perform intelligent accounting operations for road construction projects: in the pre-defined time window, the carbon dioxide equivalent estimation values of each source are accumulated in units of stages to generate time- and segment-specific emission curves and summary reports, while the confidence interval and key driving factors are output, and the accounting process records complete audit logs for backtracking. When the confidence interval or bias of the real-time accounting result exceeds the set threshold, the abnormal backtracking process is activated to re-verify the related input characteristics and stage mapping, and trigger the necessary mapper retraining, thereby realizing the closed-loop guarantee of accounting accuracy and reliability.

[0024] Further, wherein the multi-source integrated sensor system includes a construction equipment energy consumption sensor module, a building material monitoring sensor module, and a construction site monitoring sensor module, step S1 includes the following steps: Step S11: using the construction equipment energy consumption sensor module to collect construction equipment energy consumption data of the road construction project during the construction period to obtain construction equipment energy consumption data; In the embodiment of the present application, the construction equipment energy consumption collection is arranged according to the type of the equipment. The mobile vehicles and internal combustion equipment such as excavators are equipped with fuel flow meters (pulse output, sampling frequency 10 Hz), engine speed sensors (CAN bus interface, sampling frequency 1 Hz), working time meters (cumulative timer, resolution 1 s) and oil pressure sensors (sampling frequency 5 Hz); the electric equipment is equipped with electric energy meters (pulse output or pulse width output, sampling frequency 1 Hz) and voltage and current measuring points. All sensors are subjected to zero point and range verification according to the factory calibration procedure, the fuel flow meter is verified using a standard flow source, and the electric energy meter is calibrated by comparing with a reference electric meter. The sensors are fixed on the non-vibration sensitive parts of the equipment using anti-vibration brackets, the cables are shielded and sealed at the wiring points through waterproof connectors, the sampling clock is synchronized with the satellite positioning time (GNSS) and network time protocol (NTP), and the sampling record includes time stamp, sensor ID, measurement item, unit, equipment working condition label and initial uncertainty estimate. The diagnostic measures include real-time self-checking of fuel flow sudden change, continuous zero value and overrange, and abnormal event triggering boundary marker and recording of surrounding environment and equipment working condition for subsequent verification. The energy consumption instantaneous value is accumulated as the total energy consumption according to the sampling frequency, and the volume to mass conversion is performed according to the fuel density and heat value for subsequent emission conversion.

[0025] Step S12: collecting carbon emission data of the road construction material production and transportation stage of the road construction project by using the building material monitoring sensor module to obtain material production and transportation carbon emission data; In the embodiment of the present application, the carbon emission monitoring of the material production and transportation stage adopts a combination of multi-point metering and identification tracking. The production end sets quality measurement devices at the batching and discharging links, including a weighbridge (weighing accuracy ±0.5%), a continuous belt scale (sampling period 1 s), a fuel flow meter of the combustion system and a flue gas concentration detector (zero point and range verification is performed periodically using standard gas) at the outlet of the combustion furnace. Each batch of materials is attached with an RFID label as batch identification, and the time stamp recorded by the weighbridge at the loading time is cross-confirmed by the RFID reader and writer; the transportation link installs a vehicle positioning device (GNSS, frequency Hz) and a vehicle fuel metering device (instantaneous fuel flow meter) on the vehicle, and records the start and end time, mileage, average speed and loading state of the journey. The material production stage emission is calculated by multiplying the material quality by the material unit emission factor, and the greenhouse gas metering value generated by the process fuel combustion is recorded at the same time; the transportation emission is calculated by multiplying the fuel consumption of the transportation vehicle by the fuel unit emission factor and is modified in combination with the loading rate. The batch of materials, weighing record, vehicle mileage and fuel consumption are regularly associated through time stamp and batch identification, forming traceable production-transportation link record and preliminary emission estimation items.

[0026] Step S13: The construction site monitoring sensor module is used for monitoring and processing the field environment data of the construction period of the road construction project to obtain construction environment data. In the embodiment of the present application, the construction site environment monitoring arranges a plurality of meteorological and particulate matter measuring points in the construction area according to a spatial grid to reflect the heterogeneity of the site environment. The measuring point configuration includes temperature and relative humidity sensors (minute-level sampling), wind speed and direction measuring instruments (second-level sampling), particulate matter sensors (PM2.5, PM10, minute-level sampling), volatile organic compound detectors (VOCs, hour-level sampling), and on-site noise meters. The meteorological measuring point is installed at a height of 2 meters above the ground and is ensured to be unobstructed, and the particulate matter measuring point is close to the edge of the work area and is provided with a rain cover and a regular filter core replacement program to maintain stable metering. The environmental data is used to calculate the evaporation loss correction coefficient, the dust resuspension factor and the combustion efficiency correction term, which is specifically realized by calculating the moving average of temperature, humidity and wind speed in a time window and obtaining the evaporation rate coefficient according to the empirical formula, and the particulate matter observation value is associated with the construction activity intensity (loading frequency, road vehicle driving mileage) to estimate the dust contribution. The quality flag of the environmental measuring point reflects the sensor failure, pollution or maintenance state to ensure the traceability and auditability in subsequent use.

[0027] Step S14: The construction equipment energy consumption data, material production and transportation carbon emission data, and construction environment data are subjected to multi-source heterogeneous integration processing of carbon emission factors in the whole life cycle to obtain preliminary multi-source heterogeneous data of carbon emission factors; In the embodiment of the present application, the multi-source heterogeneous integration processing realizes event-level fusion with time and space as anchor points. Time synchronization takes GNSS timestamp as reference, and minimum time difference matching algorithm is adopted to align high-frequency time series measuring points with low-frequency device records; spatial alignment associates vehicle trajectory points to road segment vectors and realizes partition mapping according to construction area boundaries through map matching algorithm. Batch cascade rule is defined as a ternary key: RFID batch ID, weighbridge time interval and vehicle driving trajectory interval, and records meeting the consistency of the ternary key are merged as a single production and transportation event. Kalman filter is applied to fuse instantaneous measurement values for redundant measurement items, and inverse variance weighting method is used to calculate fusion estimation and fusion uncertainty, and quantities with physical consistency constraints between observations are solved by constraint optimization method (for example, material mass conservation constraint is used to correct the deviation between belt scale and weighbridge readings). Low-frequency sources are expanded to high-frequency alignment according to the nearest steady-state value, and missing segments are estimated by a time series regression model based on similar working conditions and an uncertainty index is recorded. The fusion output is preliminary carbon emission factor records at event level or time window level, and the fields include event identification, time interval, geographic range, contribution source type, original measurement set, preliminary estimated emission amount and fusion uncertainty estimation, and complete traceability link.

[0028] Step S15: data preprocessing is performed on the preliminary carbon emission factor multi-source heterogeneous data to generate carbon emission factor multi-source heterogeneous data.

[0029] In the embodiment of the present application, strict preprocessing is performed on the preliminary fusion data to generate standardized carbon emission factor records. The noise processing adopts a multi-stage filtering strategy: a fourth-order low-pass Butterworth filter is applied to high-frequency oscillation, a sliding median filter is applied to isolated pulses, and continuous wavelet transform is used to denoise and reconstruct stationary components for non-stationary noise. The bias value determination is based on a three-standard-deviation criterion discrimination logic, short-time missing (not more than 5 minutes) is completed by linear interpolation, and long-term missing is estimated by Gaussian process regression and the estimation confidence is output. The unit unification and conversion strictly adopts the physical quantity conversion formula: the fuel quality is equal to the fuel volume multiplied by the fuel density; the electric energy is obtained by multiplying the electricity consumption (kilowatt-hour) by the unit emission factor of the power grid; the material production stage is obtained by multiplying the material quality by the material unit generation emission factor; the environmental gas is obtained by monitoring the gas quality related to carbon elements. The final standard record field includes: time window, event identification, spatial boundary, emission source category, data quality mark and complete traceability information, and all records are saved in the format of audit log for subsequent verification and model calibration.

[0030] Further, step S2 includes the following steps: Step S21: heterogeneous type characteristic analysis is performed on the carbon emission factor multi-source heterogeneous data to generate carbon emission factor heterogeneous type characteristic data; In the embodiment of the present application, when analyzing the heterogeneous type characteristics according to the multi-source heterogeneous data of carbon emission factors, first, the data is classified by modal into time series type, static structured type and spatial type. The time series type calculates the time domain statistics (mean, variance, skewness, kurtosis, interquartile range), instantaneous slope and start-stop ratio, peak value ratio and periodicity characteristics of each signal in the sliding window (window length 30 minutes, step 5 minutes); the power spectral density is estimated and the main frequency and bandwidth are extracted for the frequency domain characteristics; the sample entropy and approximate entropy are calculated for complex dynamic signals to depict irregularity. The static structured type statistically describes the category distribution, frequency, mode and missing pattern, and calculates the batch arrival interval distribution, batch quality coefficient of variation and batch internal and external difference indicators. The spatial type maps the trajectory points to the road segment through map matching, calculates the road segment level transportation intensity, residence time distribution and local density, and estimates the spatial autocorrelation scale and Moran's I statistics using the semi-variant function. The above-mentioned features are standardized to form a high-dimensional feature vector cluster, the intraclass variance is used to evaluate the homogeneity within the cluster, and the Mahalanobis distance is used to measure the distinguishability between clusters; the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) and Gaussian Mixture Model are used to perform unsupervised division to identify abnormal sensor groups or abnormal behavior patterns, and finally the heterogeneous type characteristic data classified by type is output, which includes the feature list, statistical distribution parameters, clustering label and type internal baseline model parameters of each type, so as to be used for subsequent abnormal identification model training.

[0031] Step S22: establishing a data anomaly identification relationship model of carbon emission factor heterogeneous types based on the carbon emission factor heterogeneous type characteristic data, to obtain a carbon emission factor heterogeneous type anomaly identification model; In the embodiment of the present application, when establishing a data anomaly identification relationship model based on the carbon emission factor heterogeneous type characteristic data, a hierarchical detection framework is constructed and a special detector is configured for each heterogeneous type. For single variable numerical signals, the historical quantile rule (lower limit is 0.5 percentile, upper limit is 99.5 percentile) and the isolation forest are used for parallel discrimination; the isolation forest uses the historical normal segment and the labeled abnormal segment to calculate the abnormal score in the training stage, and the threshold is determined through cross-validation to maximize the F1 value on the validation set. For multi-variable time series correlation anomaly, a long short-term memory autoencoder (LSTM Autoencoder) is trained, and the reconstruction error is used as an abnormal index. The training set uses multi-condition historical samples and uses the early stopping method to prevent overfitting. For category sequence type, a hidden Markov model is established to discriminate sequence anomaly by using transition probability deviation. Spatial anomaly uses spatial scanning statistics and local Moran's I deviation test. The outputs of each detector are fused into a final discrimination score by a weighted voter, and the weights are obtained by parameter estimation based on expectation maximization according to the detection contribution of the historical labeled samples and are recorded as model metadata. The model training process records performance indicators such as ROC curve, AUC, precision and recall, and saves the model version to realize an auditable identification relationship model.

[0032] Step S23: transmitting the carbon emission factor multi-source heterogeneous data to the carbon emission factor heterogeneous type anomaly identification model to perform anomaly data identification processing of each heterogeneous type, and outputting multi-source heterogeneous first abnormal data and / or multi-source heterogeneous preliminary valid data; In the embodiment of the present application, when the carbon emission factor multi-source heterogeneous data is transmitted to the anomaly identification model for identification processing, the detection process is performed piece by piece according to events or time windows: firstly, the quantile rule and the isolation forest discrimination are applied to each time sequence signal in a 30-minute sliding window, if the abnormal score exceeds the set threshold (for example, the isolation forest score is greater than or equal to 0.7 and the value exceeds the 99.5 percentile), it is marked as multi-source heterogeneous first abnormal data, and the abnormal type (spike, long-time zero value, drift), timestamp, sensor identifier and the latest environmental context are recorded. If the single-variable abnormality determination is not triggered, the window and the adjacent window are sent to the LSTM autoencoder for multivariate reconstruction error detection; if the reconstruction error is greater than the set threshold during the training period, the first abnormality is output. The cross-source data is checked by the physical consistency rule, for example, if the fuel consumption and engine working time ratio deviate from the historical working condition by ±20%, it is counted as an abnormality classification; if all detectors are not marked as abnormal and pass the physical consistency check, the multi-source heterogeneous preliminary valid data is output, and the confidence score and a number of key feature values used for judgment are recorded. All discrimination results form an anomaly log in time sequence, which is used as the input of the subsequent fuzzy conversion and correction strategy, and the threshold and feature basis for triggering discrimination are marked in the model metadata for auditing.

[0033] Step S24: performing fuzzy quantization feature conversion processing on the multi-source heterogeneous preliminary valid data to obtain multi-source heterogeneous valid fuzzy quantization feature data; In the embodiment of the present application, when the multi-source heterogeneous preliminary effective data is subjected to fuzzy quantization feature conversion, the language membership set is defined for continuous quantity, and the language items are definitely defined as "low", "normal" and "high". The membership function adopts a triangular or trapezoidal function and is constructed by taking historical quantiles as base points: the lower limit Q0 takes the 0.1 percentile, the first limit point Q1 takes the 25 percentile, the median Q2 takes the 50 percentile, the third limit point Q3 takes the 75 percentile, and the upper limit Q4 takes the 99.9 percentile. The "low" membership function is calculated in a linear increasing or decreasing manner in [Q0, Q1, Q2], the "normal" membership function is dominant in [Q1, Q2, Q3], and the "high" membership function is defined in [Q2, Q3, Q4]. The membership degree calculation adopts a linear interpolation formula, and the output is in the interval [0, 1]. The category features are mapped to the membership degree vector by frequency distribution, and the batch identification is obtained by similarity calculation with the historical high-quality batch template. The membership degrees of each variable are concatenated to form a multi-dimensional fuzzy quantization feature vector, and principal component analysis is performed on the vector as necessary to compress the dimension and retain more than 90% of the variance. In order to facilitate subsequent numerical similarity calculation, the barycentric defuzzification method is used to calculate the centroid of part of the fuzzy sets that need to be numerized to obtain the representative numerical value, while the membership degree vector is retained as the fuzzy description.

[0034] Step S25: transmitting the multi-source heterogeneous effective fuzzy quantization feature data to the carbon emission factor heterogeneous type anomaly recognition model to perform anomaly fuzzy feature similarity analysis processing on each heterogeneous type effective data, and outputting multi-source heterogeneous second anomaly data and / or outputting carbon emission factor multi-source heterogeneous verification data; In the embodiment of the present application, when the multi-source heterogeneous effective fuzzy quantization feature data is transmitted to the anomaly recognition model to perform anomaly fuzzy feature similarity analysis, a plurality of prototype clusters are constructed based on the historical verified fuzzy feature library, and the prototype clusters are obtained by fuzzy C-means clustering. Each cluster center is saved in the form of a membership degree vector. A plurality of similarity indexes are calculated for the to-be-tested fuzzy vector: fuzzy cosine similarity, fuzzy Jaccard index and normalized value of fuzzy Euclidean distance. The similarity threshold is obtained by searching the verification set with the maximum F1 as the target. For example, if the similarity is less than 0.65 or the normalized distance is greater than 0.35, it is considered as multi-source heterogeneous second anomaly data; on the contrary, if the similarity is not lower than the threshold and the confidence of the historical prototype distribution is not lower than 0.7, the carbon emission factor multi-source heterogeneous verification data is output. In order to enhance the time sequence consistency, the moving average of the similarity of the last 5 time windows is calculated, and the weighted combination of the moving average and the current similarity is used as the final discrimination basis. The similarity discrimination result records the similarity value, the matched cluster identifier and the confidence, and the verification data is accompanied by distribution explanatory information to support the subsequent correction logic.

[0035] Step S26: Based on the multi-source heterogeneous first abnormal data and the multi-source heterogeneous second abnormal data, each heterogeneous type of abnormal data correction processing is performed to obtain multi-source heterogeneous abnormal correction data, and the multi-source heterogeneous abnormal correction data is fed back to the carbon emission factor multi-source heterogeneous verification data to perform abnormal data repair feedback adjustment processing of the carbon emission factor, and generate carbon emission factor multi-source heterogeneous standard data.

[0036] In the embodiment of the application, when the abnormal correction processing is performed based on the multi-source heterogeneous first abnormal data and the second abnormal data, a hierarchical correction strategy is adopted according to the abnormal type, and multi-source heterogeneous abnormal correction data is output. For transient spikes, a robust smoothing based on median filtering is adopted, a median with a window length of 5 minutes is used to replace the original point, and the residual before and after replacement is recorded; for sensor drift, baseline estimation is realized by comparing the reference sensor, the drift trend is fitted by piecewise linear regression, and the trend is deducted to restore the baseline value; for long-time missing, Gaussian process regression is adopted, the radial basis white noise kernel is selected as the kernel function, the last 7 days of smooth section data are used for training in the training section, and the estimated value and prediction variance are output; for cross-source contradictory items, the inverse variance weighted least squares method is used to solve the consistency estimation, and the weight is determined by the reciprocal of the uncertainty of each observation item. Each correction operation simultaneously performs uncertainty propagation, calculates the confidence interval of the corrected carbon emission element by Monte Carlo sampling (for example: 1000 times sampling), and records the uncertainty decomposition. The correction result forms a multi-source heterogeneous abnormal correction data set, and then the correction result is fed back to the multi-source heterogeneous verification data to trigger model adaptive training: based on the corrected labeled samples, the threshold and weight of the abnormal identification model are periodically re-estimated according to the sliding window, and the model performance change is recorded after each retraining. Finally, the carbon emission factor multi-source heterogeneous standard data is output, each standard record contains original observation, corrected value, correction method code, uncertainty before and after correction, quality mark and complete traceability link, to meet the input requirements of the subsequent accounting model.

[0037] Further, step S22 includes the following steps: According to the historical data of the carbon emission factor multi-source heterogeneous data, the abnormal prior data analysis of the heterogeneous type data is performed to generate the carbon emission factor heterogeneous type abnormal prior data; In the embodiment of the present application, historical observation records of no less than twelve months are used as the prior construction basis, and the historical records are first segmented according to sensor categories and working conditions, the segmentation length is in units of days, and at least 360 daily samples are kept to cover seasonal changes. Quality screening is first performed on each segment; the quality screening rule is: if the missing rate is higher than 5%, it is marked as low quality; if the continuous zero value duration exceeds 30 minutes, it is marked as an abnormal segment; the measurement distribution skewness is tested by skewness and kurtosis and recorded. The time series stability is evaluated by the enhanced Dickey-Fuller test, if it is not stationary, it is processed by difference or seasonal decomposition, the seasonal decomposition uses local weighted regression (STL) to extract the trend and seasonal components, and the residual is used as the modeling object of this stage. The residual distribution is fitted to a parameterized distribution, and the normal distribution and the lognormal distribution are compared according to experience to minimize the Akaike information criterion (AIC) to determine the distribution type and save the mean and variance as the parameter prior. For the frequency of sudden abnormalities, the drift rate and the missing distribution, Bernoulli-Beta process and Gamma distribution are used to model the event rate and variance, the hyperparameters are estimated using the empirical Bayes method, and the hierarchical Bayes structure is used at the group level to reflect the population differences of similar sensors and output the prior parameter set of the population level and the individual level. The shrinkage covariance estimation (Ledoit-Wolf) is used to obtain a robust covariance matrix, which is used as a multivariate prior covariance for subsequent anomaly detector regularization. Simulated observations are generated by sampling from the prior distribution and compared with the key statistics quantile of the actual observation, if the difference is significant, the hyperparameters are iteratively adjusted until the fitting degree meets the set test criterion (example: Kolmogorov-Smirnov test p value greater than 0.05). All prior parameters and test results are recorded as "carbon emission factor isomer type anomaly prior data", and are archived in the form of time stamp and version number for traceability.

[0038] According to the carbon emission factor isomer type characteristic data, the abnormal feature index data of the carbon emission factor isomer type is generated, and a preliminary carbon emission factor isomer type anomaly recognition model is established based on the carbon emission factor isomer type abnormal feature index data; In the embodiment of the present application, the carbon emission factor isomer type characteristic data is taken as input, and feature engineering is first performed to define an abnormality discrimination index set. The time domain features include mean value, standard deviation, maximum and minimum value, maximum slope, start-stop ratio and peak value ratio in a sliding window (window length 30 minutes, step 5 minutes); the frequency domain features are obtained by Welch spectrum estimation, and the window segment length is 256 points and the main frequency energy and bandwidth ratio are extracted; the time-frequency features are extracted by continuous wavelet transform (mother wavelet Morlet) on scales 1 to 128 to extract the transient pulse energy distribution; the complexity features include sample entropy parameters m=2, r=0.2×standard deviation, m represents the embedding dimension for constructing a vector, and r represents the tolerance threshold for matching similar vectors; the multivariate features include Mahalanobis distance, covariance principal component projection and energy conservation residual (for example, the deviation amount of the ratio of fuel consumption to engine working hours from the historical baseline). The class and batch information is constructed to form a statistical frequency and a transition matrix, and the Kullback-Leibler divergence of the transition matrix is used for sequence anomaly measurement. Based on the above features, a preliminary recognition model set is established: single variable uses dynamic z-score combined with historical quantile threshold for rapid screening; density and isolation detection uses isolation forest, and the parameter setting is that the number of base learners is 200, the upper limit of subsample is 1024, and the initial pollution rate is estimated from the history (example range 0.1% to 1%); associated anomaly detection uses LSTM self-encoding structure, and the sequence length is set to 120 steps, the number of encoder layers is 3, the number of hidden units is compressed from 128 to 32 layer by layer, the training target is to minimize the reconstruction error, and L2 regularization coefficient 1e-4 and dropout 0.2 are used to control overfitting. The feature importance is sorted by random forest feature importance score and mutual information evaluation, and the recursive feature elimination method is used to retain the feature subset with the highest explanation degree, and the obtained feature index and its threshold, detector structure and initial hyperparameter jointly constitute the preliminary carbon emission factor isomer type abnormality recognition model, and the performance indicators such as ROC curve, AUC, precision, recall and median detection delay are recorded under the time series segmentation cross-validation framework for verification.

[0039] According to the carbon emission factor isomer type abnormality prior data, the preliminary carbon emission factor isomer type abnormality recognition model is subjected to isomer type abnormality checking model training and parameter weight adjustment processing to generate a carbon emission factor isomer type abnormality recognition model.

[0040] In the embodiment of the present application, the isomorphic type abnormal prior data of carbon emission factors is used as the source of regularization and sample weight, and the preliminary carbon emission factor isomorphic type abnormality identification model is supervised and unsupervised parallel training. The historical labeled abnormal samples and normal samples are used for supervised training, the class imbalance is balanced by self-re-sampling method, and the abnormal samples are given higher weight to reduce the risk of false negatives, and the weight ratio is initially set to normal:abnormal=1:5, and adjusted according to the cost function in the verification process. The unsupervised model parameter optimization adopts Bayesian regularization based on prior distribution: the model weight is set to a priori Gaussian distribution with a mean of zero and a variance determined by the prior noise variance, and the corresponding objective function is the negative log posterior. When optimizing, both the data fitting term and the prior penalty term are considered. The time series blocking cross-validation is used for hyperparameter adjustment, which is divided into five forward rolling folds. The hyperparameter grid includes the isolated forest pollution rate range [0.001, 0.01, 0.05], the LSTM learning rate range [1e-4, 1e-3, 1e-2], the encoding layer width range [64, 128, 256], and the regularization coefficient range [1e-6, 1e-3]. The Gaussian process-based Bayesian optimization is used for efficient search of continuous hyperparameters, and the F1 score and false alarm rate are used for comprehensive evaluation. The threshold is set to keep the false alarm rate below 2 times a day and the false negative rate below 5%. The early stopping strategy is embedded in the training process, and the tolerance is 10 consecutive validation cycles without improvement. The final model requires PR-AUC to exceed 0.75 on the independent posterior test set and meet the minimum recall rate constraint according to the sensor category. The model release outputs the model configuration file, the prior parameters during training, the performance index table, and the version identifier, and records the feedback mechanism for online calibration: when the distribution of continuous observations and the prior Kullback-Leibler divergence exceeds the threshold value 0.2 or the number of newly added manually labeled abnormalities exceeds 200, the model retraining process is triggered to maintain the consistency and robustness of the identification ability.

[0041] Further, step S23 comprises: when the carbon emission factor isomorphic type abnormality identification model identifies the matching abnormal data in the carbon emission factor multi-source isomorphic data, outputting the multi-source isomorphic first abnormal data, and / or when the carbon emission factor isomorphic type abnormality identification model does not identify the matching abnormal data in the carbon emission factor multi-source isomorphic data, outputting the multi-source isomorphic preliminary effective data.

[0042] In this embodiment of the invention, during the online discrimination stage of the carbon emission factor heterogeneity anomaly identification model, a discrimination input vector is constructed using a 30-minute sliding window and a 5-minute step size. The input vector includes standardized time-domain statistics (mean, variance, skewness, kurtosis), instantaneous slope, start / stop ratio, sample entropy, power spectrum frequency and bandwidth, equipment operating condition label, and spatial neighborhood mean. For each continuous signal, univariate threshold detection and anomaly score determination based on isolated forest are performed simultaneously: the upper and lower bounds of the univariate threshold are taken as the 99.5th percentile and 0.5th percentile of the historical data, respectively; any observation exceeding this range is considered a threshold anomaly; an isolated forest anomaly score not less than 0.70 is considered an isolated anomaly. For multivariate association anomalies, a Long Short-Term Memory (LSTM) autoencoder reconstruction error detection is performed; reconstruction errors exceeding the 95th percentile threshold during training are considered reconstruction anomalies. For categorical sequence features, a Hidden Markov Model (HMM) transition probability deviation detection is used; a transition probability deviation exceeding 0.30 is considered a sequence anomaly. The physical consistency rule performs parallel verification of several operating condition constraints: fuel consumption must not be negative; an instantaneous fuel consumption exceeding 1.5 times the equipment's rated consumption is considered a physical anomaly; a single material mass conservation deviation exceeding ±2% is considered a mass inconsistency. The judgment logic adopts a hierarchical consistency rule: when any two types of detectors trigger an anomaly within the same time window, or when the same detector triggers repeatedly within three consecutive sliding windows, the first anomaly data for multi-source heterogeneous systems is output; when all detectors do not trigger, all physical consistency checks pass, and the overall confidence level is not less than 0.85, preliminary valid multi-source heterogeneous data is output. The overall confidence level is calculated using the formula: Confidence = 1 - (w_T) Score_T + w_ IF Score _IF+w_AE× Score _AE), where Score _T represents the normalized threshold deviation value, Score_IF represents the isolated forest anomaly score, and Score_AE represents the normalized autoencoder reconstruction error value; the weights are set to w_T=0.4, w_IF=0.3, and w_AE=0.3. The output includes: record type identifier (first anomaly or preliminary validity), anomaly category code, trigger detector list, trigger timestamp, sensor identifier, adjacent window context vector, physical consistency residual, and confidence score. For data labeled as the first anomaly, a correction priority code is automatically appended according to the anomaly type and queued for use in subsequent repair processes; for data whose output is preliminary validity, a verification token is generated and passed to the next step of fuzzy quantization processing.

[0043] Further, step S24 comprises: when the carbon emission factor heterogeneous type anomaly identification model identifies that the abnormal fuzzy feature similarity score of the multi-source heterogeneous effective fuzzy quantization feature data is greater than a preset abnormal similarity score threshold, outputting multi-source heterogeneous second abnormal data, and / or when the carbon emission factor heterogeneous type anomaly identification model identifies that the abnormal fuzzy feature similarity score of the multi-source heterogeneous effective fuzzy quantization feature data is not greater than a preset abnormal similarity score threshold, outputting carbon emission factor multi-source heterogeneous verification data.

[0044] In the embodiment of the present application, after the multi-source heterogeneous preliminary effective data determined is subjected to fuzzy quantization feature conversion, three types of language items "low", "normal" and "high" are defined for each continuous feature according to historical distribution, the membership function adopts a triangular or trapezoidal function and membership parameters are constructed with historical quantiles Q0=0.1%, Q1=25%, Q2=50%, Q3=75% and Q4=99.9% as nodes. The continuous values are calculated for membership degrees by linear interpolation and membership vectors are generated; the category features are mapped to frequency to generate membership degree vectors. The fuzzy C-means clustering is applied based on the historical verified fuzzy vector sample library to obtain prototype clusters, the number of clusters is set to 8, the fuzzy rate m=2, the maximum iteration is 300, the convergence threshold is 1e-5, and the cluster centers are saved in the prototype library in the format of membership vectors. The current fuzzy vector is calculated for similarity with abnormal prototype clusters and normal prototype clusters: the weighted combination of fuzzy cosine similarity and normalized fuzzy Euclidean distance is adopted, and the similarity score S is calculated according to the formula S=0.7 S_cos+0.3 (1-D_norm). S_cos is the fuzzy cosine similarity, and D_norm is the distance index of the fuzzy Euclidean distance normalized by the historical maximum value. If the maximum value of the similarity of the current vector with any abnormal prototype cluster is not less than 0.60, the multi-source heterogeneous second abnormal data is output, and the trigger prototype ID, the similarity value, the membership dimension with the highest contribution and the time window index are written in the output record. If the similarity of the abnormal prototype is less than 0.60 and the maximum value of the similarity with the normal prototype cluster is not less than 0.75, the carbon emission factor multi-source heterogeneous verification data is output, and the matching normal prototype ID and the matching confidence are simultaneously attached. For the vectors in the fuzzy boundary (abnormal similarity <0.60 and normal similarity <0.75), time consistency expansion is performed: the moving average of the similarity of the last 5 sliding windows is calculated and the similarity threshold is re-evaluated by the moving average; if the normal similarity of the moving average is not less than 0.70, the carbon emission factor multi-source heterogeneous verification data is output, otherwise the second abnormal data is output and marked as a fuzzy boundary abnormal candidate.

[0045] Further, as one embodiment of the present application, referring to Figure 2 , the carbon emission factor multi-source heterogeneous anomaly identification model is trained by the multi-source heterogeneous effective fuzzy quantization feature data and the multi-source heterogeneous second abnormal data and the carbon emission factor multi-source heterogeneous verification data. Figure 1The detailed step flowchart of step S3 is shown in the figure. In this embodiment, step S3 includes the following steps: Step S31: Perform modal division processing on the carbon emission factor multi-source heterogeneous standard data to generate carbon emission factor modal data, wherein the carbon emission factor modal data includes carbon emission factor time series data, carbon emission factor static structured data, and carbon emission factor spatial data. In the embodiment of the present application, when the carbon emission factor multi-source heterogeneous standard data is divided into modes, first, automatic labeling is performed according to observation attributes and metadata rules to clearly define the time granularity, spatial elements, and structured fields of each record. The criteria are defined as follows: if the record contains a continuous timestamp field and the average sampling frequency is higher than once per hour, it is classified as time series data; if the record is mainly based on batch attributes, material attributes, and specification parameters and is in discrete event units, it is classified as static structured data; if the record contains latitude and longitude or local projection coordinates and needs to be mapped, it is classified as spatial data. All records are added with metadata fields during modal labeling, including sampling frequency, coordinate reference system (WGS84 or local UTM projection), measurement unit, measurement point accuracy, and measurement uncertainty estimation. Time stamps are uniformly corrected in time zone and aligned according to UTC reference. Sampling rate labels are applied to time series data records with cross-sampling frequency for subsequent resampling strategy application. Spatial data is simultaneously projected and transformed to obtain metric distance measurement, and event window boundaries are set for discrete events. The modal division result is output as an index table and metadata list of the three types of data sets.

[0046] Step S32: Extract carbon emission time series dynamic characteristics from carbon emission time series data to generate carbon emission time series dynamic characteristic data. In this embodiment of the invention, when extracting time-series dynamic features from carbon emission factor time-series data, a multi-scale sliding window strategy is adopted. The window is divided into a short-term window (5 minutes in length, 1 minute in step), a medium-term window (60 minutes in length, 15 minutes in step), and a long-term window (24 hours in length, 1 hour in step). Within each window, statistical features are calculated, including mean, variance, skewness, kurtosis, interquartile range, maximum value, minimum value, and median. Time-domain dynamic features are calculated, including instantaneous slope, start-stop frequency, duty cycle, and peak frequency. Frequency-domain features are calculated using Welch spectrum estimation to extract the dominant frequency and bandwidth ratio, and continuous wavelet transform (Morlet wavelet) is used to extract scaling coefficients for short-term transients to characterize pulse energy distribution. Seasonal-trend decomposition (STL) is performed on non-stationary trends to separate the trend from the residuals, and sample entropy, approximate entropy, and several lag values ​​of the autocorrelation function are calculated for the residuals. To reduce feature dimensionality, principal component analysis is used to retain 90% of the variance. The output feature vector for each time window includes multi-scale statistics, time-frequency energy spectrum coefficients, complexity measures, and trend residual energy. These features are used to describe the dynamic characteristics of equipment energy consumption, load fluctuations, and transient events, and serve as input for subsequent correlation analysis.

[0047] Step S33: Perform static structured interaction feature analysis on the static structured data of carbon emission factors to generate static interaction feature data of carbon emission factors; In this embodiment of the invention, when performing static interaction feature analysis on static structured data of carbon emission factors, categorical variables are first target-coded using the historical average emission amount at the batch level as the coded value, and the coding uncertainty is recorded; numerical variables are normalized and interaction terms with categorical variables are calculated, such as the product of material quality and supplier reputation score, and material units. The ratio of equivalent to arrival time difference and the difference between the rated power of the equipment and the actual load ratio were used. To identify higher-order interaction effects, a second-order polynomial interaction matrix was constructed, and the effect strength of each interaction term was evaluated. The effect strength was measured by partial regression coefficients and confidence intervals. Statistical significance was assessed using hypothesis testing, and interaction terms with p-values ​​less than 0.05 were retained. For high-cardinality categories, binning and hierarchical aggregation strategies were used to control dimensionality expansion, and the information contribution of categories to the target variable was evaluated using mutual information and chi-square statistics. The output is a static interaction feature matrix, including core interaction terms, standardized regression coefficients, confidence indices, and variable importance ranking, providing structured factors and prior knowledge constraints for cross-modal causal analysis.

[0048] Step S34: Perform spatial distribution characteristic analysis on the spatial data of carbon emission factors to generate spatial distribution characteristic data of carbon emission factors; In the embodiment of the present application, when analyzing the spatial distribution characteristics of the carbon emission factor spatial type data, the map matching algorithm is used to align the transportation trajectory points to the road segment vectors and generate travel statistics (driving distance, stay time, average speed and loading rate) according to the road segments. At the geographic grid level, the transportation intensity and emission estimation value are aggregated in a 200m*200m grid unit, and the unit area emission density of each unit is calculated. The ordinary Kriging interpolation method is used for spatial interpolation of sparse observation points to obtain a continuous emission surface, and the semi-variation function is calculated to estimate the spatial autocorrelation scale. The Moran's I index is used to test the global spatial autocorrelation, and the Getis-Ord Gi The significance level is set to 0.01 when identifying high value aggregation (hot spot). The spatial gradient is calculated at the road segment level to measure the derivative of the emission intensity with respect to the location, and the road segment slope, traffic density and loading frequency are used as explanatory variables for spatial regression. The spatial error model is used to correct the estimation bias caused by spatial autocorrelation. The spatial analysis output includes emission density at the road segment and grid levels, hot spot identification, spatial autocorrelation parameters and local trend vector, which provides geographical constraints and spatial influence factors for cross-modal correlation.

[0049] Step S35: Based on the Bayesian network algorithm, the carbon emission factor cross-modal correlation processing is performed on the carbon emission time series dynamic characteristic data, the carbon emission factor static interaction characteristic data and the carbon emission factor spatial distribution characteristic data, and the carbon emission factor correlation modal data is generated; In the embodiment of the present application, when performing cross-modal correlation processing on the three types of modal characteristics based on the Bayesian network algorithm, a hybrid structure learning process is used to balance the constraint information and data driven evidence. First, the time series dynamic characteristics, static interaction characteristics and spatial distribution characteristics are discretized or kept continuous according to the variable type, and a candidate node set is constructed. The constraint driven conditional independence test (Fisher Z test for continuous variables and G test for discrete variables) is performed to exclude unsupported edges, and then the score driven greedy search algorithm is applied to evaluate the network structure with Bayesian information criterion (BIC) and perform local optimization in the candidate space. To resist the randomness of single learning, the self-help method guided multiple learning is run, and the edge frequency is counted. Only the edges with an appearance frequency not less than 0.7 are retained as stable causal relationships. The maximum likelihood estimation is used for parameter estimation, the conditional Gaussian distribution fitting is applied to continuous parent nodes, and the conditional probability table is established for discrete nodes. The output results include directed acyclic graph structure, conditional probability or regression coefficient of each directed edge and edge weight value. The influence strength matrix between modes is calculated from the edge weight, and the modal mutual influence attention parameter is generated accordingly, which is used to guide the initial attention configuration and cross-modal weighting strategy of the subsequent neural network.

[0050] Step S36: based on the neural network algorithm, the carbon emission factor correlation mode data is analyzed for nonlinear characteristics of the correlation mode, nonlinear characteristic data of the carbon emission factor correlation mode is generated, and global connection processing of the nonlinear characteristics of the correlation mode is performed according to the nonlinear characteristic data of the carbon emission factor correlation mode, and global fusion feature data of the carbon emission factor is generated.

[0051] In the embodiment of the application, when the nonlinear characteristic analysis is performed on the carbon emission factor correlation mode data, first, the data is jointly modeled with the carbon emission cross-modal causal relationship data, and the multi-layer mapping capability of the neural network structure is used to extract the nonlinear correlation mode. In a specific implementation, the input layer is set to a high-dimensional vector composed of weighted combinations of time series dynamic features, static interaction features and spatial distribution features, and a causal relationship matrix is attached before the input layer as a constraint condition. By embedding the causal relationship regular term in the network structure, it is ensured that the model retains the causal dependence logic between modes during feature extraction. The hidden layer part is provided with multiple layers of nonlinear activation units, each layer using different nonlinear functions to capture the nonlinear variation law of complex features, such as non-stationary patterns in time series fluctuations, local aggregation effects in spatial distribution and nonlinear interactions between static features. Through the back propagation training mechanism, the network parameters are iteratively updated, and the parameter convergence direction is adjusted using the causal constraint loss function, thereby generating carbon emission factor correlation mode nonlinear characteristic data that can reflect the deep nonlinear relationship between modes. After the extraction of the correlation mode nonlinear characteristic data is completed, further global connection processing is performed to realize the fusion of multi-modal features at the global level. In the global connection process, first, the nonlinear feature vectors of different modes are standardized to keep the same dimension. Subsequently, a global connection graph is constructed based on the causal relationship weight parameters, which takes each modal feature as a node and the causal weight as an edge weight. Through the convolution operation of the graph structure, the feature vectors of different modes are aggregated across modes. This process can realize the combination of local features and global structure, so that different modal features are not only superimposed at the numerical level, but also form global connections at the topological structure level. Finally, the unified dimension feature representation is output through the full connection layer, and the global fusion feature data of the carbon emission factor is obtained. The data contains nonlinear coupling information of time series, static and spatial modes, and can provide higher precision and higher robustness input support for subsequent optimization of carbon emission accounting relationship models.

[0052] In another embodiment of the present application, in view of the nonlinear feature data processing requirement of the correlation mode of carbon emission factors, a multi-layer neural network structure is constructed. First, in the input layer, the network receives the pre-processed and normalized multi-source feature vector. The dimension of the feature vector is determined according to the number of carbon emission factors and the modal characteristic expansion, such as construction equipment energy consumption factor, building material carbon emission factor, construction site environmental factor, etc. After feature fusion, a 128-dimensional input vector is formed to ensure complete coverage of multi-modal features. In the design of the hidden layer, a multi-layer fully connected structure combined with nonlinear mapping is adopted. The first hidden layer is set to 256 neurons, which increases the dimension to improve the ability to capture complex feature relationships, and uses the ReLU activation function to avoid the problem of gradient disappearance. The second hidden layer is set to 128 neurons, which continues to maintain a high feature abstraction capability, and introduces batch normalization (Batch Normalization) processing in this layer to improve the training stability and alleviate the risk of overfitting. The third hidden layer is set to 64 neurons, which uses the tanh activation function to enhance the expression ability of the nonlinear correlation between carbon emission factors, especially suitable for capturing small fluctuations and asymmetric features. In the design of the deep structure of the network, a double-branch parallel subnetwork is also introduced to process the differences between different modal features. One sub-branch focuses on time series features (such as dynamic fluctuations of energy consumption and emissions), which adopts a double-layer LSTM structure, each layer containing 64 hidden units to capture long-term dependencies. The other sub-branch is for static structured features (such as material carbon emission coefficients, device model parameters), which adopts a two-layer fully connected structure containing 128 and 64 neurons respectively, and uses the ReLU activation function for mapping. Finally, the output results of the two sub-branches are fused through the Concatenate Layer to form a comprehensive high-dimensional feature representation. In the design of the output layer, combined with the target requirements of carbon emission accounting, the fused high-dimensional features are mapped to a single regression node for outputting the interaction feature information of the correlation mode of the correlation mode of carbon emission factor correlation mode data, and a linear activation function is used to ensure the continuity of the numerical value. In addition, to prevent overfitting, a Dropout layer (dropout rate set to 0.3) is added before the output layer to randomly mask some neuron connections, enhancing the model's generalization ability. This neural network structure takes into account the feature abstraction capability and modal difference processing capability, and through the combination of multi-layer fully connected network and double-branch parallel design, it realizes efficient modeling of the nonlinear feature data of the correlation mode of carbon emission factors, making it easier for the interaction features between carbon emission factors in road construction projects to reflect the actual carbon emission accounting status.

[0053] Further, step S35 includes the following steps: The carbon emission cross-modal causal relationship data is generated by performing carbon emission cross-modal causal relationship analysis on the carbon emission time series dynamic characteristic data, the carbon emission factor static interaction characteristic data and the carbon emission factor spatial distribution characteristic data through a Bayesian network algorithm; The carbon emission modal mutual influence correlation characteristic data is generated by performing carbon emission modal mutual influence correlation characteristic analysis according to the carbon emission cross-modal causal relationship data, and the carbon emission modal mutual influence attention weight parameter is designed according to the carbon emission modal mutual influence correlation characteristic data; The carbon emission factor correlation modal data is generated by performing attention weighted cross-modal correlation processing on the carbon emission time series dynamic characteristic data, the carbon emission factor static interaction characteristic data and the carbon emission factor spatial distribution characteristic data based on the carbon emission modal mutual influence attention weight parameter.

[0054] In the embodiment of the present application, when the carbon emission time series dynamic characteristic data, carbon emission factor static interaction characteristic data and carbon emission factor spatial distribution characteristic data are analyzed for cross-modal causal relationship, the three types of modal characteristics are variable standardized and distribution corrected to ensure that data from different sources are comparable in the same probability framework. Based on the Bayesian network structure learning algorithm, a candidate node set is set, wherein the time series dynamic characteristic data node represents the energy consumption rate, fuel consumption rate, instantaneous power fluctuation and the like in the equipment operation process; the static interaction characteristic data node represents the interaction term of material carbon emission factor and supply batch parameter, the coupling term of construction process and equipment parameter; the spatial distribution characteristic data node represents the emission density of transportation route, regional emission hotspot index and spatial autocorrelation statistic. Through conditional independence test (combined with Fisher Z test and chi-square test), the insignificant dependent relationship is gradually removed, and the optimal directed acyclic graph is constructed based on Bayesian information criterion (BIC). The directed edges in the graph structure represent the potential causal relationship, for example, the positive causal effect of the fuel consumption rate of the construction equipment on the emission density of the transportation route. Further, the maximum likelihood estimation is used to fit the conditional probability distribution of the edges, the continuous variable uses the conditional Gaussian model, and the discrete variable uses the conditional probability table, to generate the carbon emission cross-modal causal relationship data, and to clearly define the causal transmission path and influence strength between different modalities. After obtaining the cross-modal causal relationship data, the modal mutual influence relationship is quantitatively analyzed. According to the edge weight and conditional probability output by the Bayesian network, the mutual influence strength value between different modalities is calculated, which is mapped to the [0, 1] interval by normalization processing, and is used to measure the degree of action of different modalities in the overall causal structure. The modal mutual influence matrix is constructed, which records the influence direction and influence strength between any two types of modalities, for example, the causal weight value of the time series dynamic characteristic on the static interaction characteristic is 0.65, and the causal weight value of the spatial distribution characteristic on the time series dynamic characteristic is 0.42. Through row normalization processing of the matrix, the relative influence contribution value between modalities is obtained, and the weighted importance of the modal to the global carbon emission change is calculated combined with the modal characteristic dimension scale. Based on these results, the modal mutual influence attention weight parameter is designed, the attention weight of the high causal strength modality is set in the interval of 0.6-0.8, and the attention weight of the low causal strength modality is set in the interval of 0.2-0.4, thereby establishing a strict weight allocation mechanism. The carbon emission modal mutual influence correlation characteristic data and the corresponding attention weight parameter are generated, ensuring that the contribution of different modalities in the subsequent fusion is consistent with the real causal relationship. In the attention weighted cross-modal association processing, the three types of modal characteristic vectors are aligned according to the time stamp and spatial label, ensuring that the input features have comparability in the time dimension and spatial dimension.The modal interaction attention weight parameter is used to weight each modal feature, for example, in the fusion of time sequence dynamic features and spatial distribution features, if the causal weight shows that the spatial feature has a greater impact on the time sequence feature, then the spatial feature is given a higher weight value in the calculation of the weighted sum. Through this weighting method, not only the core feature information of each type of modal can be retained, but also the real causal contribution between different modes can be highlighted. In the specific calculation process, first, single-modal weighted aggregation is performed to obtain the weighted feature vectors of time sequence, static and space; then cross-modal weighted combination is performed to obtain global associated modal data. The data contains not only the dynamic change law of the time dimension, but also the constraint condition of the static parameter, and reflects the regional difference of the spatial distribution. The finally generated carbon emission factor associated modal data is a set of high-dimensional vectors, which internally contains the weighted coupling results of each modal feature, and can be used as the input of subsequent neural network nonlinear feature analysis to realize the multi-dimensional fusion expression of road engineering carbon emission.

[0055] Further, step S4 includes the following steps: Step S41: performing link division processing on the carbon emission factor global fusion feature data to obtain link division carbon emission factor global fusion feature data; In the embodiment of the present application, when performing link division processing on the carbon emission factor global fusion feature data for the whole life cycle, first, the link categories and their time and space boundaries are clearly defined, and the link categories are strictly divided into material production stage, transportation stage, construction stage, maintenance stage and acceptance stage. The material production stage is bounded by the material factory delivery time to the loading time, and the judgment basis is the RFID batch timestamp and the factory scale delivery record; the transportation stage is bounded by the loading time to the unloading time, and the judgment basis is the vehicle trajectory start and end point and the loading state label, and the transportation segment needs to meet the travel duration greater than 10 minutes and the average speed greater than 5 kilometers / hour to be included; the construction stage is bounded by the first entry of the equipment into the construction area to the engineering acceptance time, and the judgment basis is the cumulative working hours of the equipment positioning in the construction area and the contract acceptance document timestamp; the maintenance stage is bounded by the equipment maintenance record time window, and the acceptance stage is bounded by the project completion acceptance time. The link division slices the global fusion features in time sequence, and each link slice output includes link identifier, time interval, coverage space boundary, global fusion feature matrix in the time period (summarized by minutes or hours), link-level statistics (cumulative energy consumption, material entry quantity, transportation mileage, average load) and uncertainty estimation item. To ensure mapping accuracy, the link division processing also implements consistency verification rules: batch-level material mass conservation verification, transportation mileage and track mileage comparison, and physical consistency verification of equipment working hours and energy consumption ratio.

[0056] Step S42: Perform multi-factor coupled carbon emission conversion feature analysis of link division according to the link division carbon emission factor global fusion feature data, and generate link division coupled carbon emission conversion feature data; In the embodiment of the application, when performing multi-factor coupled carbon emission conversion feature analysis of link division according to the link division carbon emission factor global fusion feature data, a hybrid physical-statistical modeling process is adopted. For each link, an initial conversion equation is first constructed based on the conservation of physical quantities and energy balance relationship, for example, the linear and quadratic relationship between fuel consumption and transportation mileage, loading rate, and rated fuel consumption rate of the vehicle in the transportation stage; a second-order polynomial coupling term is established between the equipment energy consumption and the operation power, start-stop frequency, and working load rate in the construction stage. Subsequently, data-driven parameter fitting is performed on these initial equations: a generalized additive model is used to smooth fit the nonlinear terms, and a polynomial regression with interaction terms is used to evaluate the high-order coupling effect, the regression process introduces L1 and L2 regularization to suppress overfitting and evaluate the generalization error through time series blocking cross-validation. To depict the propagation path of potential variables and observed variables, a structural equation model is constructed to quantify the indirect influence of hidden factors (such as construction intensity or material humidity) on observed emissions, and the maximum likelihood method is used for parameter estimation and the self-help method is used for swing test of parameter confidence interval. Sensitivity analysis is performed on each link, and local elasticity coefficient and global Sobol index are used to measure the contribution rate of each input variable to link emission, and the result generates a link-level coupled feature matrix, the matrix elements are standardized conversion coefficients, interaction term strength and uncertainty decomposition, for subsequent use by mapping model and physical interpretation.

[0057] Step S43: Establish the carbon emission factor fusion feature and carbon emission accounting mapping relationship of each link according to the link division coupled carbon emission conversion feature data, and generate a carbon emission accounting relationship model; In the embodiment of the present application, when establishing the mapping relationship between the carbon emission factor fusion features of each link and the carbon emission accounting based on the coupling carbon emission conversion characteristic data of the links, an integrated mapper architecture is adopted. The mapper of each link is composed of a random forest regressor and a gradient boosting regression tree in parallel. The random forest is set to have 500 base learners, a maximum tree depth of 20, and a minimum leaf node sample size of 10 to ensure robust fitting of non-linear and outlier values. The gradient boosting regression tree is set to have 1000 iterations, a learning rate of 0.01, and an early stopping method on the validation set to control the number of iterations. The mapper takes the link input feature vector (including global fusion features, expanded items of the coupling feature matrix, and uncertainty indicators) as the independent variable, and outputs the link-level carbon dioxide equivalent estimate and prediction quantiles (0.05, 0.5, 0.95). The training uses historical calibration emission samples and uses blocked cross-validation evaluation, and the loss function is weighted mean squared error, with weights set according to the inverse of the observation uncertainty. After training, the mapper is subjected to local interpretive analysis, the local contribution of each input feature to the output is calculated based on SHAP values, and a feature importance ranking report is generated. To support auditability, each mapper also outputs residual distribution, segmented error statistics, and model stability indicators, and the mapper parameters and the prior information version on the training day are archived to facilitate item-by-item comparison when the accounting results need to be traced.

[0058] Step S44: Perform actual working condition dynamic feature optimization processing on the carbon emission conversion relationship model to generate an optimized carbon emission accounting relationship model; In the embodiment of the present application, when performing actual working condition dynamic feature optimization processing on the carbon emission conversion relationship model, an uncertainty learning and online calibration mechanism is introduced. First, the quantile consistency of the mapper output is tested, a quantile regression forest is used to construct a prediction interval, and the prediction interval coverage rate is used to calibrate the confidence level. When the distribution of historical prediction residuals and the Kullback-Leibler divergence of recent residual distribution exceed the threshold value 0.2 or the average relative deviation of continuous 24 hours exceeds the set tolerance (for example, 10%), the dynamic adjustment mechanism is triggered. The dynamic adjustment mechanism includes two paths: path one is incremental retraining, which uses high-quality calibration samples of the last 7 days to update the mapper for a limited number of iterations, and applies decay weights to old samples to retain long-term information and highlight recent working conditions; path two is model weight adaptive redistribution, which weights the parallel models according to recent prediction performance, and the weights of good models are redistributed according to their inverse mean squared error. To ensure stability, the update process implements model rollback points and early stopping strategies, and consistency testing is performed before and after updating in the independent validation window and the performance indicators are recorded. All dynamic optimization operations record the version number, trigger factors and update parameters, and after updating, a new optimized carbon emission accounting relationship model is generated for real-time accounting, and the confidence interval of the model output is continuously monitored to evaluate the optimization effect.

[0059] Step S45: Perform carbon emission intelligent accounting work for the road construction project based on the optimized carbon emission accounting relationship model.

[0060] In the embodiment of the present application, when the carbon emission intelligent accounting work is performed for the road construction project based on the optimized carbon emission accounting relationship model, a hierarchical accounting process and traceable result publishing are implemented. The accounting is performed according to a predetermined time granularity (for example, by hour, by day, or by construction phase), first, the link mapper is called to estimate the value and confidence interval of each link time slice output point, then the engineering level is accumulated and summarized, and the slicing display is realized according to the source, the link and the spatial grid. The accounting result contains fields: time interval, link identifier, source type, carbon dioxide equivalent estimate, uncertainty decomposition and traceability link. In addition, the accounting process outputs trend analysis indicators (comparative, same period, key driving factor change rate) for management decision-making and abnormal alarm records, the alarm conditions include sudden increase outside the prediction confidence interval, abnormal rise of single driving factor influence rate and spatial hotspot mutation, and the alarm and backtracking conclusion form a closed-loop management to ensure the accuracy, interpretability and traceability of intelligent accounting.

[0061] Further, step S44 includes the following steps: According to the link division carbon emission factor global fusion feature data, the uncertainty characteristics of each link actual working condition are analyzed, and link division actual working condition uncertainty feature data are generated; Based on the random forest algorithm, a link actual working condition tree model is established, and the link division actual working condition uncertainty feature data are mapped to the link actual working condition tree model for uncertainty characteristic learning and processing of each link actual working condition, to generate a link division carbon emission actual working condition uncertainty relationship model; Through the link division carbon emission actual working condition uncertainty relationship model, the actual working condition dynamic characteristic optimization processing of the carbon emission accounting of the carbon emission conversion relationship model is performed, to generate an optimized carbon emission accounting relationship model.

[0062] In the embodiment of the present application, when analyzing the uncertain characteristics of the actual working conditions of each link according to the global fusion feature data of the carbon emission factor, a working condition fluctuation detection mechanism is first established. For the material production link, the standard deviation and coefficient of variation of the production batch energy consumption record are used to calculate the working condition stability, and the CUSUM test is introduced to determine the mutation points in the energy consumption curve, and the mutation point interval is taken as the high-risk interval of uncertainty. For the transportation link, the speed standard deviation, instantaneous acceleration root mean square and idling time proportion of the GPS trajectory speed sequence are calculated, which are combined into a transportation working condition fluctuation index. The larger the index, the higher the uncertainty of the transportation link. For the construction link, based on the time series of engine speed, hydraulic pressure and load rate of construction equipment, sample entropy and multi-scale permutation entropy indicators are extracted. The above indicators constitute the actual working condition uncertainty feature vector of each link, and are simultaneously attached with multi-level uncertainty labels. Each link has one-level labels such as "fluctuation type", "stable type" and "mutation type", and two-level labels such as road fluctuation, road stability, or equipment standby running stability, equipment running fluctuation, etc. Finally, the actual working condition uncertainty feature data of the link division is formed, providing an input feature space for subsequent learning. When establishing the actual working condition tree model of each link based on the random forest algorithm, 500 decision trees are used to form the forest, the maximum depth of each tree is limited to 15 layers, the minimum leaf node sample number is set to 5, and the Gini index is used as the splitting standard. The input features are the actual working condition uncertainty feature data of the link division, including fluctuation index, energy consumption standard deviation, speed variation coefficient, sample entropy value, permutation entropy value and other multi-dimensional features, and the output is the link-level uncertainty category and its quantitative score. In the training process, the self-sampling method is introduced to ensure sample diversity, and the out-of-bag error (OOB Error) is used to evaluate the generalization performance of the model. The internal feature importance ranking of the model is calculated by the average impurity reduction, for example, the construction equipment speed entropy indicator ranks first in the construction link uncertainty classification, and the transportation speed fluctuation coefficient has a high weight in the transportation link. The actual working condition uncertainty relationship model of the link division carbon emission generated by the learning process can quantitatively describe the influence degree of different working condition fluctuation modes on the accuracy of carbon emission accounting. When using the actual working condition uncertainty relationship model of the link division carbon emission to dynamically optimize the conversion relationship model of carbon emission, the output of the uncertainty relationship model is first introduced as a dynamic adjustment factor into the conversion relationship model. Specifically: when the uncertainty score of a certain link is higher than 0.7, the system automatically reduces the weight of the conversion coefficient of the link, and relaxes the upper and lower limits of the confidence interval in carbon emission accounting to reflect the higher uncertainty. For example, in the transportation link, when the speed fluctuation is significantly increased, the conversion relationship weight between transportation fuel consumption and mileage is reduced by 10%, and at the same time, the uncertainty adjustment term is increased. If the uncertainty score is less than 0.3, the original accounting relationship is maintained and the narrow confidence interval is maintained.A sliding time window mechanism is introduced in the dynamic optimization process, taking the uncertainty trend of the last 72 hours as a reference to smooth the adjustment of parameters, avoiding the over-sensitivity of the model caused by single-point anomalies. The final output of the optimized carbon emission accounting relationship model contains dynamic weight correction items, link-level uncertainty confidence factors, and updated mapping functions. The optimization model realizes the robust correction of carbon emission accounting results under actual working conditions, ensuring that the accounting can truly reflect the emission level under fluctuating conditions, and has clear uncertainty boundaries and interpretability.

[0063] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, the scope of the application being defined by the appended claims and not by the above description, therefore all variations falling within the meaning and scope of the equivalent requirements of the application file are intended to be included within the present application.

[0064] The above description is merely one specific implementation of the application, which enables those skilled in the art to understand or implement the application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the application. Therefore, the application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A road engineering carbon emission analysis method based on multi-source heterogeneous data fusion, characterized in that, Comprise the following steps: Step S1: Collecting and data preprocessing of multi-source heterogeneous carbon emission influencing factors of road construction project in whole life cycle by using multi-source integrated sensor system, generating multi-source heterogeneous carbon emission factor data; Step S2: Multi-source heterogeneous data anomaly identification and correction processing of carbon emission factor multi-source heterogeneous data, generating carbon emission factor multi-source heterogeneous standard data; Step S3: Carbon emission factor cross-modal global connection fusion processing of carbon emission factor multi-source heterogeneous standard data, generating carbon emission factor global fusion feature data; Step S4: Establishing an optimized relationship model for actual working condition carbon emission accounting based on carbon emission factor global fusion feature data, generating an optimized carbon emission accounting relationship model; Based on the optimized carbon emission accounting relationship model, the carbon emission intelligent accounting work of the road construction project is executed; The establishment of the optimized relationship model for actual working condition carbon emission accounting based on carbon emission factor global fusion feature data comprises: Step S41: Dividing the carbon emission factor global fusion feature data into links in whole life cycle to obtain link-divided carbon emission factor global fusion feature data; Step S42: Multi-factor coupled carbon emission conversion feature analysis of link-divided carbon emission factor global fusion feature data, generating link-divided coupled carbon emission conversion feature data; Step S43: Establishing the carbon emission factor fusion feature and carbon emission accounting mapping relationship of each link according to the link-divided coupled carbon emission conversion feature data, generating a carbon emission accounting relationship model; Step S44: Actual working condition dynamic feature optimization processing of carbon emission accounting of carbon emission conversion relationship model, generating an optimized carbon emission accounting relationship model.

2. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 1, characterized in that, The multi-source integrated sensor system comprises a construction equipment energy consumption sensor module, a building material monitoring sensor module, and a construction site monitoring sensor module, and step S1 comprises the following steps: Step S11: Collecting construction equipment energy consumption data of road construction project in construction period by using construction equipment energy consumption sensor module, to obtain construction equipment energy consumption data; Step S12: Collecting carbon emission data of road construction project in road building material production and transportation stage by using building material monitoring sensor module, to obtain material production and transportation carbon emission data; Step S13: Monitoring and processing site environment data of road construction project in construction period by using construction site monitoring sensor module, to obtain construction environment data; Step S14: Multi-source heterogeneous integration processing of carbon emission factors in whole life cycle of construction equipment energy consumption data, material production and transportation carbon emission data, and construction environment data, to obtain preliminary carbon emission factor multi-source heterogeneous data; Step S15: Data preprocessing of preliminary carbon emission factor multi-source heterogeneous data, generating carbon emission factor multi-source heterogeneous data.

3. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 1, characterized in that, Step S2 comprises the following steps: Step S21: Analyzing the heterogeneous type characteristics of carbon emission factor multi-source heterogeneous data, generating carbon emission factor heterogeneous type characteristic data; Step S22: based on the carbon emission factor heterogeneous type characteristic data, a data anomaly recognition relationship model of the carbon emission factor heterogeneous type is established, and a carbon emission factor heterogeneous type anomaly recognition model is obtained; Step S23: the carbon emission factor multi-source heterogeneous data is transmitted to the carbon emission factor heterogeneous type anomaly recognition model to perform anomaly data recognition processing of each heterogeneous type, and multi-source heterogeneous first anomaly data and / or multi-source heterogeneous preliminary effective data are output; Step S24: fuzzy quantization feature conversion processing is performed on the multi-source heterogeneous preliminary effective data to obtain multi-source heterogeneous effective fuzzy quantization feature data; Step S25: the multi-source heterogeneous effective fuzzy quantization feature data is transmitted to the carbon emission factor heterogeneous type anomaly recognition model to perform anomaly fuzzy feature similarity analysis processing of each heterogeneous type effective data, and multi-source heterogeneous second anomaly data and / or carbon emission factor multi-source heterogeneous verification data are output; Step S26: based on the multi-source heterogeneous first anomaly data and the multi-source heterogeneous second anomaly data, anomaly data correction processing of each heterogeneous type is performed, multi-source heterogeneous anomaly correction data is obtained, and the multi-source heterogeneous anomaly correction data is fed back to the carbon emission factor multi-source heterogeneous verification data for anomaly data repair feedback adjustment processing of the carbon emission factor, and carbon emission factor multi-source heterogeneous standard data is generated.

4. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 3, characterized in that, Step S22 includes the following steps: According to the historical data of the carbon emission factor multi-source heterogeneous data, data anomaly priori data analysis of the heterogeneous type is performed, and carbon emission factor heterogeneous type anomaly priori data is generated; According to the carbon emission factor heterogeneous type characteristic data, anomaly feature index analysis of the carbon emission factor heterogeneous type is performed, carbon emission factor heterogeneous type anomaly feature index data is generated, and a preliminary carbon emission factor heterogeneous type anomaly recognition model is established based on the carbon emission factor heterogeneous type anomaly feature index data; According to the carbon emission factor heterogeneous type anomaly priori data, model training and parameter weight adjustment processing of each heterogeneous type anomaly verification of the preliminary carbon emission factor heterogeneous type anomaly recognition model are performed, and a carbon emission factor heterogeneous type anomaly recognition model is generated.

5. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 3, characterized in that, Step S23 includes: when the carbon emission factor heterogeneous type anomaly recognition model identifies matching anomaly data in the carbon emission factor multi-source heterogeneous data, the multi-source heterogeneous first anomaly data is output, and / or when the carbon emission factor heterogeneous type anomaly recognition model does not identify matching anomaly data in the carbon emission factor multi-source heterogeneous data, the multi-source heterogeneous preliminary effective data is output.

6. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 3, characterized in that, Step S24 includes: when the carbon emission factor heterogeneous type anomaly recognition model identifies that the anomaly fuzzy feature similarity score of the multi-source heterogeneous effective fuzzy quantization feature data is greater than a preset anomaly similarity score threshold, the multi-source heterogeneous second anomaly data is output, and / or when the carbon emission factor heterogeneous type anomaly recognition model identifies that the anomaly fuzzy feature similarity score of the multi-source heterogeneous effective fuzzy quantization feature data is not greater than the preset anomaly similarity score threshold, the carbon emission factor multi-source heterogeneous verification data is output.

7. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 1, characterized in that, Step S3 includes the following steps: Step S31: Carbon emission factor modal division processing is performed on the multi-source heterogeneous standard data of carbon emission factors to generate carbon emission factor modal data, wherein the carbon emission factor modal data includes carbon emission factor time series data, carbon emission factor static structured data, and carbon emission factor spatial data. Step S32: Carbon emission time series dynamic feature extraction is performed on the carbon emission factor time series data to generate carbon emission time series dynamic feature data. Step S33: Static structured interaction feature analysis is performed on the carbon emission factor static structured data to generate carbon emission factor static interaction feature data. Step S34: Carbon emission spatial distribution feature analysis is performed on the carbon emission factor spatial data to generate carbon emission factor spatial distribution feature data. Step S35: Based on the Bayesian network algorithm, carbon emission factor cross-modal correlation processing is performed on the carbon emission time series dynamic feature data, the carbon emission factor static interaction feature data, and the carbon emission factor spatial distribution feature data to generate carbon emission factor correlation modal data. Step S36: Based on the neural network algorithm, correlation modal nonlinear feature analysis is performed on the carbon emission factor correlation modal data to generate carbon emission factor correlation modal nonlinear feature data, and global fusion feature data of carbon emission factors is generated through global connection processing of the correlation modal nonlinear feature data.

8. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 7, characterized in that, Step S35 includes the following steps: Carbon emission cross-modal causal relationship analysis is performed on the carbon emission time series dynamic feature data, the carbon emission factor static interaction feature data, and the carbon emission factor spatial distribution feature data through the Bayesian network algorithm to generate carbon emission cross-modal causal relationship data; Carbon emission modal mutual influence correlation feature analysis is performed according to the carbon emission cross-modal causal relationship data to generate carbon emission modal mutual influence correlation feature data, and carbon emission modal mutual influence attention weight parameters are designed according to the carbon emission modal mutual influence correlation feature data; Based on the carbon emission modal mutual influence attention weight parameters, attention weighted cross-modal correlation processing is performed on the carbon emission time series dynamic feature data, the carbon emission factor static interaction feature data, and the carbon emission factor spatial distribution feature data to generate carbon emission factor correlation modal data.

9. The method for road engineering carbon emission analysis based on multi-source heterogeneous data fusion according to claim 1, characterized in that, Step S44 includes the following steps: Uncertain feature analysis of each link actual working condition is performed on the link divided carbon emission global fusion feature data to generate link divided actual working condition uncertain feature data; Based on the random forest algorithm, each link actual working condition tree model is established, and the link divided actual working condition uncertain feature data is mapped to each link actual working condition tree model for uncertain feature learning processing of each link actual working condition to generate a link divided carbon emission actual working condition uncertain relationship model; Through the link divided carbon emission actual working condition uncertain relationship model, carbon emission accounting actual working condition dynamic feature optimization processing is performed on the carbon emission conversion relationship model to generate an optimized carbon emission accounting relationship model.

Citation Information

Patent Citations

  • Highway construction project carbon emission accounting evaluation method and system

    CN119624168A

  • Method and system for dynamically monitoring carbon emission of Yellow River basin by using big data

    CN119850229A

  • Method and system for optimizing complete-cycle carbon emission of AI-driven highway engineering

    CN120851308A

Cited By

  • Carbon emission fast report generation method and prediction system based on energy structure transformation

    CN121365654A

  • Industrial equipment control method based on multi-protocol fusion and related device

    CN121433172A

  • Industrial device control method based on multi-protocol fusion and related apparatus

    CN121433172B

  • Carbon emission monitoring method and system based on multi-source monitoring data

    CN122087772A

  • Traffic flow sensing method and system

    CN122176927A