Building environment and traffic accident causal effect analysis method based on bi-robust learning
Through a dual-stable learning method, combined with federated learning and generative adversarial network, the shortcomings of mixed bias, privacy protection and counterfactual data generation in causal inference of urban traffic accidents are solved, and efficient cross-regional data collaborative analysis and high-fidelity counterfactual scenario generation are achieved, which significantly improves the reliability of causal effect verification.
Patent Information
- Application Number
- CN202510464793.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing technology has mixed bias, privacy protection and counterfactual data generation in causal inference of urban traffic accidents, making it difficult to effectively distinguish between confounding variables in high-dimensional data from real causal effects, and the fidelity and diversity of cross-regional data collaboration and counterfactual scenario generation.
A method of causal effect analysis based on double-stable learning is adopted to achieve cross-regional data collaborative analysis through federated learning, combined with differential privacy encryption to ensure data compliance, a spatiotemporal causal dynamic weighted GBDT prediction model is built, and a high-fidelity counterfactual scenario is introduced to generate a generative adversarial network.
It realizes distributed collaborative analysis of cross-regional traffic safety data without sharing original data, generates counterfactual data that is highly consistent with the real built environment, significantly improves the reliability and scenario coverage of causal effect verification, and reduces bias caused by missed variables or model misconfiguration.
Smart Images

Figure CN119990785A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent traffic safety and traffic accident risk assessment, and specifically relates to a built environment and traffic accident causal effect analysis method based on double robust learning. Background Art
[0002] With the acceleration of global urbanization, the complexity of urban transportation systems and the incidence of traffic accidents continue to rise, becoming a core challenge threatening public safety. Traditional research methods are mostly based on statistics and regression models, analyzing the causes of accidents from multiple dimensions such as people, vehicles, roads, and the environment. Although combined with machine learning models, their limitations are still quite significant, as shown in the following: (1) Confounding bias and model misspecification problems in traditional causal inference methods; Existing technologies (such as correlation analysis and instrumental variable method) rely on strong exogeneity assumptions and cannot effectively distinguish between confounding variables and true causal effects in high-dimensional data (such as the correlation between road network density and traffic accidents). Traditional regression models (generalized linear models, random forests) are susceptible to noise interference, resulting in biased conditional average treatment effect (CATE) estimates; (2) The conflict between privacy and efficiency in cross-regional data collaboration; Multi-regional traffic safety data is subject to privacy regulations (such as GDPR). Traditional centralized processing methods require sharing of raw data, which poses a risk of leakage. Existing encryption technologies (homomorphic encryption, differential privacy) significantly increase computational complexity while ensuring privacy, and lead to a decrease in model accuracy (AUC-ROC decreases by 8%-12%), restricting the feasibility of global causal analysis. (3) The fidelity and diversity of counterfactual data generation are insufficient; Existing methods (interpolation method, basic generative adversarial network) are difficult to simulate the complex spatial heterogeneity of the built environment (such as road topology, POI distribution). The distribution of generated data is very different from that of real scenes (KL divergence ≥ 0.5), resulting in low confidence in causal effect verification and inability to support virtual experiments with dynamic environmental changes. Summary of the invention
[0003] In response to the complex needs of urban traffic safety, the present invention proposes a built environment and traffic accident causal effect analysis method based on dual robust learning, which effectively breaks through the bottlenecks of existing technologies in causal inference robustness, privacy collaboration capabilities and counterfactual generation accuracy, and provides technical support for urban traffic planning and safety management.
[0004] The present invention is implemented by adopting the following technical solution: a method for analyzing the causal effect of built environment and traffic accidents based on double robust learning, comprising the following steps: Step A: Collect multi-source data, perform federated preprocessing and feature alignment on it, implement cross-regional data collaborative analysis through federated learning, and combine differential privacy encryption to ensure data compliance; Step B: Build a spatiotemporal causal dynamic weighted GBDT prediction model and train and optimize it, deeply integrating the spatiotemporal attention mechanism, dual robust causal constraints, and regional adaptive splitting strategy; Step C: Based on step B, a generative adversarial network is introduced to generate high-fidelity counterfactual scenarios, and a dual robust learning framework is coupled to quantify the conditional average treatment effect CATE of the built environment on traffic accident risk; Step D: Finally, the causal effect between the built environment and traffic accidents is tested, and a comprehensive analysis report and a causal effect quantification table are generated.
[0005] Furthermore, in step A, the multi-source data includes traffic accident data, satellite remote sensing data and built environment data; the traffic accident data includes the time, location and severity of the accident; after preprocessing the satellite remote sensing data, a multidimensional data set including spectral characteristics, texture parameters, night light intensity and population density gradient is obtained to support the association modeling of the built environment and traffic safety; the built environment data includes road networks and multi-category points of interest; and a multi-scale built environment quantitative indicator system is formed through three-level verification.
[0006] Furthermore, in step A, when performing federated preprocessing and feature alignment, a multi-head self-attention mechanism is introduced in the fusion layer of the federated model, and feature weights are dynamically allocated through attention scores to capture the synergistic impact of cross-modal interaction "RS (remote sensing) × road network density" on accident risk, specifically including: (1) Horizontal federated learning: By sharing encrypted model parameters among multiple participants and dynamically adjusting node weights based on KL divergence, collaborative analysis and adaptive fusion of cross-regional data can be achieved; (2) Encrypted feature alignment: Use the Z-score standardization method to standardize the feature values to a distribution with a mean of 0 and a standard deviation of 1; (3) Differential privacy encryption: Calculate the global sensitivity of the query function and add Laplace noise to the query results.
[0007] Furthermore, in step B, the spatiotemporal causal dynamic weighted GBDT prediction model includes: Spatiotemporal attention gating module: dynamically captures regional specificity and temporal evolution, transforming geographic grids and time series Encoded as a spatiotemporal feature vector, the regional weights are dynamically calculated through a multi-head attention mechanism, and dynamic regional weights are generated through a dual-channel gating mechanism; Regional adaptive splitting module: proposes a regional adaptive splitting strategy, dynamically adjusts the splitting gain, and introduces smoothing constraints; Double robust causal constraint module: introduces double robust causal constraints to ensure the robustness of the conditional average treatment effect CATE estimate; Then, the robustness and generalization ability of the spatiotemporal causal dynamic weighted GBDT prediction model are evaluated through feature screening, hyperparameter tuning and federated cross-validation, and the model is optimized through iteration.
[0008] Furthermore, step C is specifically implemented in the following manner: Step C1, start double robust learning for causal inference: by combining inverse probability weighting and regression model, double correction is achieved, and the conditional average treatment effect CATE is estimated using the spatiotemporal causal dynamic weighted GBDT prediction model; Step C2: Generate high-fidelity counterfactual data through generative adversarial networks: Input real traffic accident data and built environment features, and the generator generates counterfactual data based on the spatial attention mechanism; Step C3: Compare the generated counterfactual data with the real data to verify the authenticity and diversity of the counterfactual data, and ensure that the generated counterfactual data can reflect the impact of different built environment characteristics on traffic accidents; and further realize the quantitative analysis of the conditional average treatment effect CATE.
[0009] Furthermore, in step D, the significance of the causal effect is verified by estimating the conditional average treatment effect CATE and combining cross-validation: (1) When significance is established, integrate the data of the entire experimental cycle, generate a comprehensive analysis report and a causal effect quantitative table, and propose targeted planning strategies; (2) When the significance is not established: adjust the DRL parameters, start the double robust estimation, and re-optimize the spatiotemporal causal dynamic weighted GBDT prediction model.
[0010] Compared with the prior art, the advantages and positive effects of the present invention are: This solution integrates a multi-dimensional causal analysis framework of federated learning, generative adversarial networks, and dual robust learning to accurately infer the causal effect between the built environment and traffic accidents, and supports cross-regional data collaboration and privacy protection: 1. Privacy protection and cross-regional collaboration capabilities: This solution uses a federated learning framework to achieve distributed collaborative analysis of multi-regional traffic safety data for the first time. Compared with the limitations of traditional centralized data processing, this solution uses encrypted model parameter aggregation technology without sharing original data to break through data island barriers and meet the compliance requirements of the Data Security Law and the Personal Information Protection Law. 2. High-fidelity counterfactual scenario generation capability: Through the deep combination of generative adversarial networks (GANs) and attention mechanisms, this solution can generate counterfactual data that is highly consistent with the real built environment (such as virtual scenarios such as road network structure adjustment and land use change), significantly improving the reliability and scenario coverage of causal effect verification, and overcoming the defects of high distortion and insufficient diversity of data generated by traditional interpolation methods; 3. Robust causal inference and control of confounding factors: Traditional methods (such as instrumental variable method and ordinary least squares method) are sensitive to confounding factors and rely on strong assumptions. The dual robust learning framework of this scheme integrates inverse probability weighting (IPW) and nonlinear regression model (GBDT) to realize a dual correction mechanism in causal effect estimation, effectively reducing the bias caused by omitted variables or model misspecification and improving the generalizability of the conclusions. 4. Advantages of multi-source heterogeneous data fusion: Integrate multi-dimensional information such as satellite remote sensing data (such as nighttime light dynamic monitoring), POI distribution, and traffic accident records to build a built environment indicator system with enhanced spatiotemporal features. Compared with the limitations of a single data source in existing technologies, this solution achieves collaborative analysis of high-dimensional data through federated feature extraction and cross-modal alignment technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 This is a schematic diagram of a method flow chart of an embodiment of the present invention; Figure 2 This is a schematic diagram of feature extraction of satellite remote sensing data and built environment data according to an embodiment of the present invention; Figure 3 This is a diagram of the federated learning architecture of an embodiment of the present invention; Figure 4 Schematic diagram of GAN counterfactual generation according to an embodiment of the present invention. DETAILED DESCRIPTION
[0012] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described below in conjunction with the accompanying drawings and embodiments. In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention can also be implemented in other ways different from those described herein, and therefore, the present invention is not limited to the specific embodiments disclosed below.
[0013] Examples, such as Figure 1 As described above, this embodiment proposes a method for analyzing the causal effect of built environment and traffic accidents based on dual robust learning, comprising the following steps: Step A: Collect multi-source data and perform federated preprocessing and feature alignment on it; use federated learning to achieve cross-regional data collaborative analysis, and combine differential privacy encryption to ensure data compliance; Step B: Construct a spatiotemporal causal dynamic weighted GBDT prediction model and train and optimize it; deeply integrate the spatiotemporal attention mechanism, dual robust causal constraints and regional adaptive splitting strategy; Step C: Based on step B, a generative adversarial network is introduced to generate high-fidelity counterfactual scenarios, and a dual-robust learning framework is coupled to quantify the conditional average treatment effect of the built environment on traffic accident risks; Step D: Finally, the causal effect between the built environment and traffic accidents is verified, and a comprehensive analysis report and a causal effect quantification table are generated.
[0014] Specifically, in order to more clearly understand the solution of the present invention, the specific implementation steps are described in detail below: Step A: Collect multi-source data and perform federated preprocessing and feature alignment on them; Step 1: Collect data from multiple sources Collecting multi-source data is the basis for verifying the causal effect of built environment and traffic accidents, where the multi-source data includes traffic accident data, satellite remote sensing data, and built environment data.
[0015] 1. Traffic accident data: Obtain traffic accident records from 2021 to 2024 from the traffic police department of Q City, including information such as the time, location, and severity of the accident. The sample size is expected to be no less than 22,400. For traffic accident data, de-identification, spatiotemporal coordinate system conversion (WGS84), and outlier filtering are performed in sequence. The processed data is organized according to a standardized data structure to generate structured tabular data so that subsequent work can be carried out based on the tabular data.
[0016] 2. Satellite remote sensing data: Combination Figure 1 In this embodiment, Sentinel-2 L1C, Landsat-9 L1 data, and NOAA VIIRS night light monthly products (excluding the period when the moon phase is >80%) are obtained from Google Earth Engine and USGS Earth Explorer. The acquired data is preprocessed by atmospheric correction, radiation calibration, calibration, etc., and the spatial resolution is unified and spatial and temporal fusion is performed. The VIIRS night light data is downscaled to 10m resolution by calculating the spectral index and extracting texture features. The study area is divided into 5×5km grids (WGS84 UTM projection), and 200 sample units are extracted in layers. The accident frequency is counted with a 500m buffer zone at the traffic accident point, and finally a multidimensional data set is integrated for modeling the association between the built environment and traffic safety.
[0017] 3. Built environment data: Based on the open source data of OpenStreetMap (OSM), the road network and multi-category points of interest (POI) are extracted to construct a built environment database. The extracted data is routinely processed, including correcting road geometry errors, classifying road grades, quantifying road density, calculating road network connectivity and accessibility index, and analyzing POI spatial characteristics. The 100-1000m gradient buffer is used to analyze the POI density attenuation law, and the Sentinel-2 image boundary is integrated to construct a "road-POI-land use" collaborative matrix. After three levels of verification to ensure data quality, a multi-scale built environment quantitative indicator system tabular data is generated to provide structured data support for subsequent research on the causal effect of built environment and traffic accidents.
[0018] Step 2: Federated preprocessing and feature alignment Federated preprocessing and feature alignment are key steps to ensure effective collaboration of multi-source data in the federated learning framework. Feature alignment can eliminate heterogeneity between different data sources and ensure data consistency and comparability. By integrating multi-regional data (such as traffic accident data from the traffic police department in Q City and satellite remote sensing data) through the federated learning framework, the original data is retained locally and only the encrypted model parameters are shared.
[0019] Satellite data with high-dimensional spatial features and structured indicators of traffic accidents and built environment tabular data have huge differences in feature dimensions and semantics. Direct splicing or simple weighting will lead to information redundancy or loss. The data distribution of federated nodes varies significantly. The traditional FedAvg algorithm aggregates model parameters by sample size, which easily leads to the global model being biased towards nodes with high sample sizes. Federated learning requires that local data is not shared, and attention weight calculation depends on global feature distribution, but traditional methods require centralized access to data at each node, which violates the privacy principle. Therefore, how to dynamically capture the cross-regional feature correlation of multimodal data and adaptively allocate fusion weights in a federated learning environment where the original data is not shared is an urgent problem that needs to be solved.
[0020] There is a nonlinear synergistic effect between satellite data and tabular data. The attention mechanism explicitly models this interactive relationship through similarity calculation, improving the model's ability to capture complex causal relationships. Traditional centralized attention calculation requires global access to data, which violates the privacy principle of federated learning. This embodiment achieves dynamic fusion through encrypted weight transmission and localized calculation without exposing the original data.
[0021] Specific: A multi-head self-attention mechanism is introduced in the fusion layer of the federated model to extract query, key, and value matrices from the satellite data branch and the table data branch respectively.
[0022] Dynamically assign feature weights through attention scores, capture the synergistic impact of cross-modal interaction "RS (remote sensing) × road network density" on accident risk, and construct a spatiotemporal feature fusion matrix (Q, K, V). Calculate the original attention weights at the local node, and use Paillier homomorphic encryption technology to transmit the gradient; implement Laplace noise injection (ε=0.5) at the central server to achieve differential privacy protection. Introduce a regularized compensation model: ; in, To optimize the objective function as a whole; is the loss function for traffic accident risk prediction; is a hyperparameter that controls the strength of the regularization constraint. It was determined by cross-validation in the experiment (0.001-0.1); is the attention weight of the i-th sample in the j-th federation area; is the global average weight of the j-th federal region.
[0023] The model uses the regional variance penalty term (Var(w j )) Suppresses noise disturbances, breaks through the linear fusion limitations of traditional federated learning frameworks, and realizes nonlinear collaborative modeling of multimodal features in the encrypted domain for the first time, providing a new paradigm for complex causal inference in privacy protection scenarios.
[0024] 1. Horizontal Federated Learning Horizontal federated learning enables collaborative analysis of cross-regional data by sharing encrypted model parameters among multiple participants. Each participant trains the model locally and aggregates the model parameters through the federated learning framework to ensure data privacy.
[0025] In the design of the federated learning architecture, City Q is divided into 7 federated nodes (Area A-Area G). Each node trains the model independently, and the central server is responsible for aggregating and encrypting model parameters. Communication uses gRPC encrypted channels for data transmission to ensure data privacy. The model is designed with a dual-branch neural network structure: the table data branch uses a three-layer perceptron (MultilayerPerceptron, MLP) to process accident features, and the satellite data branch extracts spatial features through a lightweight MobileNetV3-Small (a lightweight convolutional neural network), and introduces an attention gating mechanism to dynamically fuse multimodal data. KL divergence is used to quantify the similarity of node feature distribution and build an adaptive aggregation model: ; The model adapts to the data characteristics of different regions (for example, District A of City Q focuses on POI distribution, District B focuses on terrain texture, etc.).
[0026] The improved FedAvg algorithm is used to update the global model, and global model distillation is performed after every 10 rounds of aggregation to compress the model size. The formula is: ; in, Output for the current global large model, Output of the lightweight model. After distillation, the subsequent training uses the lightweight model as the new global model.
[0027] Node similarity evaluation: Calculate the data distribution of each client With global distribution KL divergence of: ; Dynamic Weight Allocation: ; Weighted aggregation formula: ; In terms of privacy protection, the client uses Paillier to encrypt the gradient and injects Gaussian noise ( = 1e-3) and Laplace noise ( =0.5) to achieve hybrid differential privacy (( , )-DP, =1e-5). The security test simulates member inference attacks (accuracy ≤ 55%) and parameter inversion attacks (reconstructed image PSNR < 20dB), and uses Rényi differential privacy to quantify the risk of privacy leakage, ensuring the privacy security and computational efficiency of the model in cross-regional collaborative modeling.
[0028] 2. Encrypted feature alignment A two-layer processing architecture is adopted: first, the multi-source data is scaled to a standard normal distribution with a mean of 0 and a standard deviation of 1 through Z-score standardization to eliminate dimensional differences; then, principal component analysis (PCA) is used to reduce the feature dimension, compressing the feature dimension to 1 / 3 of the original space while retaining 95% of the information. The alignment effect is verified by constructing a feature covariance matrix to ensure that the mean deviation of the standardized data is less than 0.05 and the standard deviation dispersion is controlled within 0.1. This method realizes the nonlinear alignment of multimodal data and lays the data foundation for subsequent spatiotemporal causal modeling.
[0029] 3. Differential Privacy Encryption The research process of this scheme involves statistical analysis of the average severity of traffic accidents in a specific area and calculation of the total amount of a certain type of built environment characteristics. In the data processing system, the function that implements a specific query operation on a data set and realizes the corresponding calculation function is called a query function. This function plays a key role in the differential privacy encryption link. Its main function is to extract valuable information from the original data and provide a data basis for subsequent data analysis, model training and other work.
[0030] Global Sensitivity Calculation: Calculate the global sensitivity of the query function.
[0031] ; Among them, x and x′ are adjacent data sets, f is a query function.
[0032] Laplace Noise Addition: Add Laplace noise to the query results.
[0033] ; in, is the privacy budget, set to 0.5. is the Laplace noise.
[0034] The noise injection intensity is determined by calculating the global sensitivity of the query function, and then Laplace noise is added to the query result to achieve privacy protection. The differential privacy effect is ensured by verifying that the difference between the query result and the original value after the noise is added is less than the privacy budget ε. On this basis, the impact of privacy protection measures on model performance is verified through model accuracy comparison experiments. Finally, the standardized feature set after federation preprocessing is input into the spatiotemporal causal dynamic weighted GBDT model to complete the multimodal data fusion modeling.
[0035] Step B: Construct a spatiotemporal causal dynamic weighted GBDT prediction model and train and optimize it, which specifically includes the following steps: Step 3: Construct spatiotemporal causal dynamic weighted GBDT prediction model In causal inference, the Gradient Boosting Decision Tree (GBDT) can be used to estimate the Conditional Average Treatment Effect (CATE), which is the average effect of the treatment variable on the outcome variable given the confounding factors. Spatiotemporal causal dynamic weighting refers to dynamically adjusting the weights of the model to capture the temporal and spatial heterogeneity of built environment characteristics, thereby improving the accuracy of causal effect estimation.
[0036] Traditional GBDT uses a unified splitting strategy and cannot dynamically capture regional specificity or heterogeneity of subgroups, resulting in large prediction errors in highly heterogeneous areas (such as urban-rural fringe areas). Moreover, the CATE estimation of traditional GBDT relies on the local fitting of a single tree, and although the dynamic splitting strategy can improve regional adaptability, it does not explicitly constrain the structural relationship between the treatment variable T and the outcome Y. Although traditional GBDT can handle high-dimensional data, its ability to handle confounding factors in causal inference is limited, especially when there is a complex interaction between the confounding factors and the treatment variables.
[0037] How to capture this regional specificity and improve the accuracy of CATE estimation; how to ensure that GBDT gives priority to confounding factors related to treatment effects when splitting while avoiding introducing bias; and how to extract dynamic weights from regional features and use them as weighting coefficients for node splitting to better capture regional specificity are all key research contents of this solution. To this end, this embodiment proposes a spatiotemporal causal dynamic weighted GBDT prediction model (ST-DW-GBDT), which is the first ensemble learning strategy that deeply integrates spatiotemporal attention mechanism, dual robust causal constraints and regional adaptive splitting strategy, which largely solves the problems of traditional GBDT in highly heterogeneous regions. Specifically, the spatiotemporal causal dynamic weighted GBDT prediction model includes: 1. Spatiotemporal Attention Gating Module (ST-AGM): A spatiotemporal attention gating module is introduced to dynamically capture regional specificity and temporal evolution.
[0038] Geographic Raster and time series Encoded as a spatiotemporal feature vector: ; in, is the spatial-temporal feature dimension, Contains 8 categories of socio-economic indicators (such as population density, night light index), For time series (seasonal migration, etc.).
[0039] Dynamically calculate regional weights through a multi-head attention mechanism: ; ; in, is the input feature, is a learnable parameter, is the attention head dimension.
[0040] Generate dynamic regional weights through a dual-channel gating mechanism: ; in, is the weight matrix, is the bias term, is the Sigmoid activation function.
[0041] 2. Regional Adaptive Split Strategy (RASS) The traditional GBDT splitting criterion does not take regional heterogeneity into consideration, while ST-DW-GBDT proposes a regional adaptive splitting strategy to dynamically adjust the splitting gain calculation.
[0042] Weighted split gain formula: ; in, For the The spatiotemporal weight of the region, For Region The entropy of is the number of splitting directions.
[0043] Regional smoothing constraints: In order to prevent weight fluctuations from causing tree structure oscillation, smoothness constraints are introduced: ; Among them, β is the smoothing coefficient, and the most effective solution is obtained through experiments.
[0044] 3. Dual Robust Causal Constraints (DRCC) Traditional GBDT does not explicitly constrain causal effect estimation, while ST-DW-GBDT introduces dual robust causal constraints to ensure the robustness of CATE estimation.
[0045] Dual robust loss function: ; in, To process variables, is the outcome variable, and is the counterfactual predicted value, is the propensity score.
[0046] Custom total loss: ; in, and is the regularization coefficient, which is determined by cross-validation.
[0047] The training and optimization process of the spatiotemporal causal dynamic weighted GBDT prediction model is as follows: 1. Feature screening (1) Data preparation The dataset that has undergone federated preprocessing, feature alignment, and privacy encryption in step A is used as the model input.
[0048] The data set is divided into: training set: 70% (spatial stratified sampling), validation set: 15%, test set: 15%; (2) Initial feature screening Pre-screen based on Pearson correlation coefficient and retain Secondly, the variance is eliminated by variance threshold analysis. The low-variance characteristics of the data were used to eliminate the interference of data noise. Combined with the expert knowledge in the field of transportation engineering, population density gradient, night light intensity, and land mixed use index were listed as core features. Finally, through redundancy analysis, the VIF test (variance inflation factor>10) and the correlation coefficient matrix (r>0.8) were combined to identify and eliminate the impact of multicollinearity. Redundant feature combinations such as GDP and night light intensity, traffic flow and number of public transportation stations, commercial land area and commercial facility density were eliminated.
[0049] (3) Training the initial GBDT model First, the model is configured and the parameters are initialized. MSE is used as the loss function for regression problems. At the same time, the optimal number of trees, maximum depth, and learning rate combination are determined through cross-validation combined with grid search. Then, the model is trained. To prevent overfitting, the subsampling rate is set. The early stopping mechanism is used during training. When the error of the validation set does not decrease for 5 consecutive rounds, the training stops. Finally, the feature importance is calculated based on the split gain to evaluate the role of each feature in the model.
[0050] Feature importance calculation, based on split gain: ; : No. The set of split nodes of a tree : indicator function (characteristic Used for splitting Time = 1) (4) Secondary screening based on feature importance After completing the feature importance calculation, perform secondary feature screening.
[0051] Set importance threshold, keep standard: importance , and sort them in descending order of importance, selecting the top 20 features. Next, based on the correlation matrix, redundant features are eliminated and the correlation coefficients between features are calculated. ,like > 0.7, the features with lower importance are removed. Through these steps, the final feature subset is generated as the feature set for subsequent modeling.
[0052] (5) Model verification and tuning The iterative optimization process of the spatiotemporal causal dynamic weighted GBDT model (ST-DW-GBDT) is as follows: based on the secondary screened feature subset (Top20), nested cross-validation (5-fold outer layer + 3-fold inner layer) is used for model training, where the outer layer validation is used for performance evaluation and the inner layer loop implements hyperparameter optimization (Bayesian search, number of trees 100-500, learning rate 0.01-0.3, maximum depth 3-8).
[0053] The model uses AUC-ROC to evaluate classification performance, F1-Score to measure the ability to balance positive and negative samples, and CATE estimation deviation to verify the accuracy of causal inference. When the AUC-ROC of the validation set is lower than the preset threshold (such as 0.85) or the CATE deviation exceeds 15%, the dynamic adjustment mechanism of feature screening parameters is triggered: by adjusting the feature importance threshold k and the redundant feature removal threshold r, a new feature subset is generated and re-modeled. This closed-loop optimization strategy optimizes and improves the robustness of the model in the task of traffic accident risk prediction while maintaining the stability of the feature space dimension (≤20 dimensions).
[0054] 2. Hyperparameter Tuning Hyperparameter tuning is a key step to improve model performance. The combination space of hyperparameters (such as tree depth and learning rate) of the spatiotemporal causal dynamic weighted GBDT prediction model is huge. Traditional hyperparameter tuning methods (such as grid search and SHA) need to traverse all possible parameter combinations, which is computationally expensive and time-consuming. In the federated learning framework, multi-node collaborative training requires rapid iteration of model parameters, and traditional methods cannot meet real-time requirements.
[0055] This embodiment applies ASHA to model optimization, and the optimized model is called "Hyper-GBDT". ASHA is a hyperparameter optimization algorithm proposed on the basis of the Successive Halving Algorithm (SHA). The SHA algorithm randomly initializes multiple sets of hyperparameter combinations, evenly distributes the budget and evaluates them, screens them according to the loss value of the validation set, eliminates half of the combinations with poor performance in each round, and iterates until the optimal combination is found. ASHA uses a continuous halving strategy (eliminating half of the inefficient combinations in each round) combined with asynchronous parallelism to quickly focus on the high-performance parameter area and accelerate convergence to the optimal solution.
[0056] Compared with traditional methods, ASHA greatly improves computing resource utilization through asynchronous parallel execution and shortens debugging time by one third. It can adapt to complex scenarios such as cluster learning, dynamically allocate budgets to optimize key parameters, and improve model generalization capabilities. It expands the parameter search space, overcomes the limitations of traditional methods, and ensures that the global optimal solution is found through an efficient elimination mechanism.
[0057] In order to achieve the ideal accuracy value (AUC-ROC ≥ 0.85, MAE ≤ 0.1), the spatiotemporal causal dynamic weighted GBDT prediction model in this embodiment needs to be quickly optimized within the federated learning framework, and the asynchronous ASHA model is naturally adapted to the distributed federated learning architecture, which can efficiently complete the setting of hyperparameters and ensure the accuracy of subsequent causal reasoning (step 5). ASHA has significant advantages in performance, resource utilization and result quality, and is the main optimization tool supporting the technical solution of the present invention.
[0058] The specific operation is as follows: divide the data into training set, validation set and test set in the ratio of 70%, 15% and 15%. Set the initial hyperparameters: number of neurons, learning rate, batch size and drop rate. Set the ASHA algorithm parameters: number of hyperparameter combinations 32, minimum budget 10, maximum budget 100, decay factor 3, minimum early stopping rate 0.2.
[0059] The ASHA algorithm is used to optimize the hyperparameters of the spatiotemporal causal dynamic weighted GBDT prediction model. Multiple sets of hyperparameter combinations are randomly initialized, the budget is evenly distributed and evaluated. By comparing the validation set loss values, the hyperparameter combinations with good performance are screened out and the combinations with poor performance are eliminated. Repeat the iteration until the optimal hyperparameter combination is found. Finally, the spatiotemporal causal dynamic weighted GBDT prediction model is retrained using the optimized hyperparameter combination.
[0060] 3. Federated Cross-Validation Federated cross-validation evaluates the robustness and generalization ability of the model by cross-validating between multiple participants. Each participant trains and validates the model locally, and the central server aggregates the validation results. A 5-fold cross-validation (K=5) is used to evaluate the model performance. The training set is divided into K subsets, and K-1 subsets are used as the training set each time, and the remaining 1 subset is used as the validation set. Repeat K times and calculate the performance indicators of the model on each validation set. Indicators such as AUC-ROC, F1-Score, mean square error (MSE), MAE (mean absolute error), and R² score are calculated.
[0061] Step 4: Is the accuracy up to standard? Evaluation Metrics: AUC-ROC: used to evaluate the model's ability to distinguish between positive and negative samples, especially to distinguish between high-risk and low-risk areas in traffic accident risk classification, with a target value of ≥0.85. Each federated node calculates the area under the ROC curve based on the local test set data. The central server aggregates the AUC-ROC values of each node through weighted averaging (the weight is the proportion of the sample size of each node).
[0062] If the global AUC-ROC is less than 0.85: re-execute feature screening (step 3.1), increase the information gain threshold, and screen the top 20 features. Adjust the GBDT hyperparameters and use the ASHA algorithm to search for the optimal combination again.
[0063] F1-Score is the harmonic mean of precision and recall, with a target value of ≥ 0.8. Each node calculates precision and recall based on the confusion matrix, and then calculates F1-Score. Global aggregation uses sample size weighted average, with the same formula as AUC-ROC.
[0064] If the global F1-Score is less than 0.8: Introduce category weights in GBDT training (high-risk sample weight = 3, low-risk = 1) to balance category influences. Dynamically adjust the classification threshold (from 0.5 to 0.3) to optimize the recall rate. Through the federated learning framework, migrate some low-risk samples from high-sample-volume nodes to low-sample-volume nodes.
[0065] MAE: measures the average absolute deviation between the model's predicted traffic accident risk value and the true value, with a target value of ≤0.1. Each node calculates the MAE of the local test set. The global aggregation weighted average aggregates the results of each node, with the weight being the proportion of the node's test set samples.
[0066] If the global MAE>0.1: add spatial interaction features (such as road density × POI density) to improve the ability to capture nonlinear relationships. Through horizontal federated learning, use the model parameters of high-precision nodes (MAE≤0.08) to guide the training of low-precision nodes.
[0067] The R² score (coefficient of determination) is used to measure the goodness of fit of the model to the data. In this study, the target value R² was set to ≥ 0.7. This means that the expected model can explain 70% or more of the variation in traffic accident risk data, thereby ensuring that the model has a good fit to the data and has a high explanatory power. Each federated node calculates the R² score based on the local test set data, and the central server aggregates the R² scores of each node through weighted average to obtain a global R² score.
[0068] Only when AUC-ROC≥0.85, F1-Score≥0.8, MAE≤0.1, and R²≥0.7 are satisfied at the same time, the accuracy is judged to be up to standard and the process goes to step 5 (doubly robust causal inference). If any indicator fails to meet the standard, the adjustment mechanism is triggered.
[0069] Adjust the priority: Prioritize optimization of MAE: MAE directly reflects the prediction error and affects the accuracy of causal effect estimation.
[0070] Sub-optimal F1-Score: Ensures the recognition reliability of high-risk areas.
[0071] Finally, optimize AUC-ROC: improve the overall classification performance.
[0072] Perform up to three rounds of adjustments. If the target is still not met, recheck the data quality (step 1) or the federated learning framework (step 2).
[0073] After the accuracy reaches the target, the risk weight and feature importance score output by the GBDT model are passed to step 5 (double robust causal inference) for inverse probability weighting (IPW) and CATE estimation. If the target is not reached, it is fed back to step 3 (GBDT model construction) to form a closed-loop optimization. The risk prediction value output by the GBDT model is used as the conditional variable of the generative adversarial network to guide the generator to generate counterfactual scenarios consistent with the real data distribution.
[0074] Step C: Generate high-fidelity counterfactual scenarios using a generative adversarial network, and couple the dual robust learning framework to quantify the conditional average treatment effect of the built environment on traffic accident risk. Combine the generated counterfactual data with the real data to construct a causal verification dataset for causal effect estimation and significance verification of the dual robust learning framework. Specifically: Step 5: Enable doubly robust learning for causal inference In traditional causal analysis methods, methods such as instrumental variables and discontinuity regression rely on strong exogeneity assumptions and often fail in practical applications due to variable selection bias. On the other hand, standard regression models (such as GLM and random forests) can handle high-dimensional data, but cannot distinguish between confounding variables and true causal effects, which may lead to serious biases in estimating the conditional average effect of treatment.
[0075] Doubly Robust Learning (DRL) combines inverse probability weighting (IPW) and regression models to achieve double correction, reduce confounding bias, and improve the robustness of causal effect estimation. The core idea of DRL is to adjust sample weights through inverse probability weighting and estimate causal effects through regression models.
[0076] Phase 1: Inverse Probability Weighting (IPW) is a commonly used causal inference method that estimates causal effects by adjusting sample weights to balance the impact of confounding factors. The core idea of IPW is to calculate the propensity score for each sample and use these scores to weight the samples so that the distribution of the treatment group and the control group on confounding factors tends to be balanced. Based on the above data on treatment variables, outcome variables (traffic accident risk) and confounding factors (such as road density, land use changes, night light intensity, etc.), the data is divided into a training set (70%) and a test set (30%).
[0077] Use the propensity score model to estimate the propensity score for the treatment variable: ; in, is the processing variable, is a covariate.
[0078] The logistic regression model has the form: ; in, is the weight vector, is the bias term.
[0079] For the treatment group samples (Right now =1), inverse probability weight for: ; For the control group samples (Right now =0), inverse probability weight for: ; Use the calculated weights We weight the samples to get a weighted dataset. We use the assign method in the Pandas library to add weights to the dataset.
[0080] Phase 2: Causal effect estimation GBDT regression modeling: Using the weighted dataset in Phase 1, the data is divided into a training set (70%) and a test set (30%).
[0081] Use the spatiotemporal causal dynamic weighted GBDT prediction model to train the causal effect estimation model. Set the model parameters, consistent with the above, using inverse probability weights as sample weights.
[0082] The conditional average treatment effect (CATE) is estimated using the GBDT regression model: ; ; in, is the outcome variable (traffic accident risk), is the propensity score.
[0083] The loss function of the GBDT model is the mean square error (MSE): ; in, is the inverse probability weight, is the true value, is the predicted value.
[0084] Use the test set data to verify the performance of the model, calculate the mean square error (MSE) and R² score of the model, compare the performance of the model on the training set and the test set, and ensure the generalization ability of the model. Calculate the propensity score of each sample through the logistic regression model , check the distribution of the propensity score to ensure it is between 0 and 1. Calculate the inverse probability weight for each sample . Check the distribution of weights to ensure that they make sense. Estimate the conditional mean treatment effect , analyze the impact of different built environment characteristics on traffic accidents.
[0085] Step 6: Generative Adversarial Network (GAN) In this scheme, the spatiotemporal causal dynamic weighted GBDT model in step B and the generative adversarial network (GAN) in step 6 are deeply coordinated through feature space mapping, causal constraint coupling and closed-loop verification mechanism. The risk weight matrix and feature importance map output by GBDT provide causal constraints for GAN's counterfactual generation, and the high-fidelity counterfactual data generated by GAN serves as the key data set for GBDT federated verification. The two use a two-way feedback mechanism to iteratively optimize causal effect estimation, forming a closed-loop system of "causal modeling-counterfactual generation-verification optimization", which significantly improves the robustness of CATE estimation.
[0086] FedGAN data augmentation uses GAN to generate high-fidelity counterfactual data to support causal effect verification. The generator and discriminator are trained adversarially to make the generated counterfactual data close to the real data distribution. The introduction of the spatial attention mechanism can improve the quality of the counterfactual data generated by the generator. This mechanism adjusts the generator output by calculating the attention weight, making the generated counterfactual data closer to the real data in terms of spatial features.
[0087] 1. Generator design and training The Generator is one of the core components of GAN. The Generator maps random noise and conditional variables to the data space by learning the distribution of real data to generate counterfactual scenarios. In the specific implementation, real traffic accident data and built environment characteristics are input, and the Generator generates counterfactual scenarios (such as virtual data after the road network structure of a certain area is adjusted) based on the spatial attention mechanism.
[0088] Network structure design: Input layer: random noise vector The dimension is 100, the conditional variable The dimension is 10.
[0089] Hidden layer: Use two layers of MLP, the number of neurons in each layer is 256 and 128 respectively, and the activation function is ReLU.
[0090] Attention layer: A multi-head self-attention module is introduced to enhance the generator’s ability to extract spatial features.
[0091] Output layer: The output dimension is consistent with the real data, and the activation function is Tanh.
[0092] The generator parameters are initialized using the He initialization method, and its loss function goal is to maximize the probability of misjudgment of the generated data by the discriminator. In practical applications, real traffic accident data and built environment features are input, and the generator can generate counterfactual scenarios based on the spatial attention mechanism, such as virtual data after the road network structure of a certain area is adjusted. Then the generator is trained using the training set data, and the number of training rounds is set to 1000 rounds and the batch size is 64, so as to ensure that the generator can generate counterfactual data that is close to the distribution of real data. During the training process, the generator will be continuously optimized using the feedback from the discriminator. In order to facilitate tracking and evaluating the performance of the model, the model will be saved every 100 rounds and verified using the test set data. Finally, the trained generator is used to input a random noise vector and condition variables , we can generate counterfactual data in a format consistent with the real data.
[0093] 2. Generate counterfactual data Counterfactual data generation is one of the core tasks of GAN. Through the generator, counterfactual samples that are highly similar to the real data distribution can be generated to verify the causal effect. The key to generating counterfactual data is to ensure the authenticity and diversity of the generated data. In simple terms, the generator generates high-quality counterfactual samples by learning the distribution characteristics of real data to support the accuracy and reliability of causal inference.
[0094] 3. Discriminator design and optimization The discriminator is another core component of GAN, which is mainly responsible for distinguishing real data from generated data. During adversarial training, the generator continues to generate data that is closer to real data, and the discriminator also continuously improves its ability to distinguish. To optimize the performance of the discriminator, its loss function is used. Minimize the probability of misjudging real data as generated data and generated data as real data. The quality of generated data is evaluated by KL divergence to ensure that the distribution difference between generated data and real data is minimized. This iterative optimization mechanism enables the generator to generate more realistic and diverse counterfactual data, thereby improving the performance and robustness of the overall model.
[0095] Step 7: Data Verification The generated counterfactual scenarios are used to simulate the impact of changes in the built environment on traffic accidents, and a causal verification data set is constructed in combination with real data. The consistency between the generated data and the real data is evaluated by calculating the KL divergence between the generated data and the real data, and the target value is limited to 0.3 or below; statistical indicators such as mean and standard deviation are used to compare to ensure that the generated data is close to the real data. At the same time, the performance of the discriminator is evaluated by the classification accuracy of the real and generated data to ensure that the classification accuracy of the real data is higher than that of the generated data. After completing the above evaluation, the generated counterfactual data is compared with the real data to verify the authenticity and diversity of the counterfactual data, and to ensure that the generated counterfactual data can reflect the impact of different built environment characteristics on traffic accidents.
[0096] After verifying the consistency and reliability of the data set, the quantitative analysis phase of the conditional average treatment effect (CATE) is entered. If the data verification fails to meet the preset standards, the attention mechanism in the generative adversarial network (GAN) needs to be adjusted to enhance the quality and authenticity of the counterfactual samples produced by the generator. Specifically, by integrating multi-head self-attention units, the generator's ability to identify and simulate spatial features is enhanced, thereby generating counterfactual samples that are closer to the actual data distribution.
[0097] Step D: Finally, test the causal effect of built environment and traffic accidents, generate a comprehensive analysis report and causal effect quantitative table, specifically: Step 8: Estimate CATE and cross-validate robustness The conditional average treatment effect (CATE) was estimated using a dual robust learning framework combined with counterfactual data generated by GAN, and the causal effect significance was verified through cross-validation (p-value < 0.05).
[0098] Cross-validation is an effective method to evaluate the robustness and generalization ability of a model. By dividing the data set into multiple subsets and using different subsets as training sets and validation sets in turn, the performance of the model on different data subsets can be evaluated, thereby verifying the robustness of the model. This example still evaluates the stability of the model through 5-fold cross-validation to ensure that the error of the causal effect estimation is controlled within 15%.
[0099] Step 9: Is the causal relationship significant? To observe whether the causal effect is significant, statistical tests and confidence interval estimates can be used to evaluate it. The significance level is set to 0.05. If the p-value is less than 0.05, the causal effect is considered significant. Use analysis of variance (ANOVA) to test the significance of the causal effect, and use the F distribution to calculate the p-value of the causal effect. If the confidence interval of the causal effect does not contain 0, the causal effect is also considered significant. Use the bootstrap method to estimate the confidence interval of the causal effect. The bootstrap method is a resampling technique that estimates the distribution of a statistic by randomly drawing samples (with replacement) from the original data.
[0100] Step 10: End signal; (1) Output report: When significance is established, the system integrates the data of the entire experimental cycle to generate a comprehensive analysis report and a causal effect quantification table, accurately quantifying the causal effect of each physical feature in the urban built environment on the risk of traffic accidents, clarifying the direction and intensity of the impact of each feature on the risk of accidents, and proposing targeted planning strategies. Through the optimization of the built environment, it is expected to reduce the accident rate in high-risk areas by 10%-15%; providing a data-driven decision-making basis for urban traffic planning and avoiding blind infrastructure investment.
[0101] (2) Adjusting DRL parameters: If the significance threshold is not reached, the system will start the parameter optimization procedure of the doubly robust estimation method (DRL). Specifically, the system will optimize the covariate selection strategy of the inverse probability weighted estimator (IPW) and the functional parameter configuration combination of the outcome model, focusing on reducing potential confounding bias. In the inverse probability weighted (IPW) module, the complexity of the propensity score estimation model (such as the regularization coefficient λ) needs to be adjusted to ensure that the overlap assumption is met; in the regression model part, the selection of the base learner and the feature engineering strategy should be optimized to reduce confounding bias and improve the prediction accuracy of the model. By iteratively optimizing the DRL parameters, the robustness and statistical power of the causal effect estimation can be effectively improved, providing a more reliable quantitative basis for subsequent decision-making.
[0102] The method used in this embodiment constructs an adaptive closed-loop optimization system. The system uses distributed cross-validation (5-fold federated validation) and dynamic parameter adjustment (asynchronous continuous halving algorithm ASHA) to iteratively optimize the model accuracy (required to reach AUC-ROC ≥ 0.85, F1-Score ≥ 0.8) and the significance of causal effects. During operation, the system can automatically trigger operations such as feature re-screening, weight balancing, and GAN attention optimization through a feedback mechanism, thereby forming a closed-loop system of "modeling-verification-tuning". This closed-loop system ensures that the system still has good stability and generalization capabilities in complex scenarios.
[0103] The above description is only a preferred embodiment of the present invention and does not limit the present invention in other forms. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.
Claims
1. A method for analyzing the causal effects of built environment and traffic accidents based on dual robust learning, characterized in that: The following steps are involved: Step A: Collect multi-source data, perform federated preprocessing and feature alignment on it, implement cross-regional data collaborative analysis through federated learning, and combine differential privacy encryption to ensure data compliance; Step B: Build a spatiotemporal causal dynamic weighted GBDT prediction model and train and optimize it, deeply integrating the spatiotemporal attention mechanism, dual robust causal constraints, and regional adaptive splitting strategy; Step C: Based on step B, a generative adversarial network is introduced to generate high-fidelity counterfactual scenarios, and a dual robust learning framework is coupled to quantify the conditional average treatment effect CATE of the built environment on traffic accident risk; Step D: Finally, the causal effect between the built environment and traffic accidents is tested, and a comprehensive analysis report and a causal effect quantification table are generated.
2. The method for analyzing the causal effect of built environment and traffic accidents based on dual robust learning according to claim 1 is characterized by: In step A, the multi-source data includes traffic accident data, satellite remote sensing data and built environment data; the traffic accident data includes the time, location and severity of the accident; after preprocessing the satellite remote sensing data, a multidimensional data set including spectral characteristics, texture parameters, night light intensity and population density gradient is obtained to support the association modeling between the built environment and traffic safety; The built environment data includes road networks and multi-category points of interest; a multi-scale built environment quantitative indicator system is formed through three-level verification.
3. The method for analyzing the causal effect of built environment and traffic accidents based on dual robust learning according to claim 1, characterized in that: In step A, when performing federated preprocessing and feature alignment, a multi-head self-attention mechanism is introduced in the fusion layer of the federated model to dynamically allocate feature weights through attention scores to capture the synergistic impact of cross-modal interaction "remote sensing × road network density" on accident risk, specifically including: (1) Horizontal federated learning: By sharing encrypted model parameters among multiple participants and dynamically adjusting node weights based on KL divergence, collaborative analysis and adaptive fusion of cross-regional data can be achieved; (2) Encrypted feature alignment: Use the Z-score standardization method to standardize the feature values to a distribution with a mean of 0 and a standard deviation of 1; (3) Differential privacy encryption: Calculate the global sensitivity of the query function and add Laplace noise to the query results.
4. The method for analyzing the causal effect of built environment and traffic accidents based on dual robust learning according to claim 1, characterized in that: In step B, the spatiotemporal causal dynamic weighted GBDT prediction model includes: Spatiotemporal attention gating module: dynamically captures regional specificity and temporal evolution, transforming geographic grids and time series Encoded as a spatiotemporal feature vector, the regional weights are dynamically calculated through a multi-head attention mechanism, and dynamic regional weights are generated through a dual-channel gating mechanism; Regional adaptive splitting module: proposes a regional adaptive splitting strategy, dynamically adjusts the splitting gain, and introduces smoothing constraints; Double robust causal constraint module: introduces double robust causal constraints to ensure the robustness of the conditional average treatment effect CATE estimate; Then, the robustness and generalization ability of the spatiotemporal causal dynamic weighted GBDT prediction model are evaluated through feature screening, hyperparameter tuning and federated cross-validation, and the model is optimized through iteration.
5. The method for analyzing the causal effect of built environment and traffic accidents based on dual robust learning according to claim 1 is characterized by: The step C is specifically implemented in the following manner: Step C1, start double robust learning for causal inference: by combining inverse probability weighting and regression model, double correction is achieved, and the conditional average treatment effect CATE is estimated using the spatiotemporal causal dynamic weighted GBDT prediction model; Step C2: Generate high-fidelity counterfactual data through generative adversarial networks: Input real traffic accident data and built environment features, and the generator generates counterfactual data based on the spatial attention mechanism; Step C3: Compare the generated counterfactual data with the real data to verify the authenticity and diversity of the counterfactual data, and ensure that the generated counterfactual data can reflect the impact of different built environment characteristics on traffic accidents; and further realize the quantitative analysis of the conditional average treatment effect CATE.
6. The method for analyzing the causal effect of built environment and traffic accidents based on dual robust learning according to claim 1, characterized in that: In step D, the significance of the causal effect is verified by estimating the conditional average treatment effect CATE and combining cross-validation: (1) When significance is established, integrate the data of the entire experimental cycle, generate a comprehensive analysis report and a causal effect quantitative table, and propose targeted planning strategies; (2) When the significance is not established: adjust the DRL parameters, start the double robust estimation, and re-optimize the spatiotemporal causal dynamic weighted GBDT prediction model.
Citation Information
Patent Citations
Machine learning architecture for quantifying and monitoring event-based risk
CA3172010A1
Private data protection method and system based on homomorphic encryption and federated learning
CN119513919A
Causal inference via neuroevolutionary selection
US20230376776A1
Ai-enhanced simulation and modeling experimentation and control
US20240348663A1
Cited By
Causal inference method and device for traffic accident data processing and prediction model optimization
CN120781977A
Causal inference method and device for traffic accident data processing and prediction model optimization
CN120781977B
E-commerce platform multi-source data security fusion method and system based on federated learning
CN120930169A
Self-evolution federal element learning method for cross-domain heterogeneous space-time intelligence
CN121279362A
A self-evolution federated meta-learning method for cross-domain heterogeneous spatio-temporal intelligence
CN121279362B