A Multi-Source Data-Driven Road Performance Prediction Method
By employing a multi-source data-driven approach, spatiotemporal alignment and uncertainty quantification techniques, combined with a Bayesian neural network model and dynamic feedback mechanism, the challenge of multi-source data fusion was solved, enabling highly reliable prediction of pavement performance and precise maintenance decisions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA HIGHWAY ENG CONSULTING GRP CO LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-26
AI Technical Summary
Existing pavement performance prediction methods fail to effectively integrate multi-source data, resulting in insufficient robustness of prediction results, making it difficult to support accurate maintenance decisions, and often leading to resource waste or deterioration of road conditions.
A multi-source data-driven approach is adopted, which processes multi-source data through spatiotemporal alignment and uncertainty quantification, uses a Bayesian neural network model to predict pavement performance, outputs probabilistic prediction results, and combines Monte Carlo simulation and dynamic feedback mechanism to optimize maintenance timing.
It improves the reliability and accuracy of pavement performance prediction, reduces the risk of maintenance misjudgment, realizes refined simulation and prediction of the future evolution of pavement performance, and supports differentiated precision maintenance strategies.
Smart Images

Figure CN122087568A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of road engineering and traffic infrastructure management technology, specifically relating to a multi-source data-driven method for predicting pavement performance. Background Technology
[0002] In the field of road maintenance and management, accurately predicting the trend of pavement performance degradation and formulating scientific maintenance plans accordingly is key to improving the service level of the road network and optimizing the allocation of maintenance funds. Most existing pavement performance prediction methods rely on a single or limited number of data sources. For example, they may use only historical field inspection data for regression analysis or use simple time series models for extrapolation prediction.
[0003] With the development of detection technology, the data sources available for pavement performance assessment are becoming increasingly abundant. These include time-series remote sensing data capable of acquiring pavement surface information over a wide area and periodically, GIS spatial data containing rich geographic and environmental contextual information, and high-precision fixed-point field detection data. Theoretically, fusing these multi-source data can more comprehensively depict the evolution of pavement conditions. However, in practice, due to significant differences in spatiotemporal resolution, accuracy, format, and physical meaning (i.e., spatiotemporal heterogeneity) among these data, effective fusion faces enormous challenges. Existing technologies often only perform simple data stacking or shallow fusion, failing to address the spatiotemporal alignment issues between data points and neglecting the propagation and amplification effects of uncertainties inherent in each data source in the fusion and prediction chain. This results in insufficient robustness of the constructed prediction models in complex real-world environments, large fluctuations in prediction results, low reliability, and difficulty in supporting accurate maintenance decisions. This often leads to maintenance being performed too early or too late, resulting in resource waste or road condition deterioration. Summary of the Invention
[0004] This application provides a multi-source data-driven method for predicting road performance to address one of the aforementioned technical problems.
[0005] The technical solution adopted in this application is as follows:
[0006] This application provides a multi-source data-driven method for predicting road performance, including:
[0007] Collect multi-source data for the target road segment, including time-series remote sensing data, GIS spatial data, and on-site detection data;
[0008] The multi-source data is preprocessed and fused to generate a spatiotemporally aligned fused data sequence, and the uncertainty of the fused data sequence is quantified.
[0009] The fused data sequence is input into the pavement performance degradation prediction model to obtain the pavement performance prediction sequence output by the model.
[0010] Based on the pavement performance prediction sequence, the timing of pavement maintenance for the target road section is calculated;
[0011] The pavement performance degradation prediction model is trained based on historical multi-source data. It is used to characterize the correlation between pavement performance and multi-source data and outputs prediction results in probabilistic form.
[0012] According to one embodiment of this application, preprocessing and fusing the multi-source data includes:
[0013] The time-series remote sensing data, GIS spatial data and field detection data are unified into a common spatiotemporal framework using a spatiotemporal registration algorithm.
[0014] Kriging interpolation is applied to spatially standardize the GIS spatial data, and dynamic time warping algorithm is applied to time-align the time-series remote sensing data and field detection data.
[0015] A joint uncertainty map is generated by fusing the uncertainties of various data sources using the Bayesian method. The uncertainties include the signal-to-noise ratio of remote sensing data, the interpolation variance of GIS data, and the measurement error distribution of field detection data.
[0016] According to one embodiment of this application, the pavement performance degradation prediction model is a Bayesian neural network model, and its training process includes:
[0017] Historical multi-source data was used as input, and pavement performance index was used as label;
[0018] The optimization objective is to minimize the negative log-likelihood loss, and incorporates a physical constraint equation for pavement performance degradation. This physical constraint equation is based on fracture mechanics and takes the following form:
[0019]
[0020] Where P is the pavement performance index, k and α are attenuation parameters, and ϵ is random noise;
[0021] We employ a combination of stochastic gradient descent and Monte Carlo Dropout for model training to handle large-scale data.
[0022] According to one embodiment of this application, calculating the timing of pavement maintenance based on the pavement performance prediction sequence includes:
[0023] Confidence intervals for pavement performance prediction sequences are generated by sampling multiple predictions from the output distribution of the pavement performance degradation prediction model using Monte Carlo simulation.
[0024] When the probability of the pavement performance index falling below a preset threshold exceeds a preset probability threshold, maintenance recommendations are triggered.
[0025] The preset probability threshold is optimized based on historical maintenance data.
[0026] According to one embodiment of this application, a dynamic feedback step is also included:
[0027] Newly acquired multi-source data is periodically received and compared with the pavement performance prediction sequence to calculate residuals.
[0028] If the residual exceeds the preset tolerance, the retraining of the pavement performance degradation prediction model will be automatically triggered to update the model parameters.
[0029] The dynamic feedback step forms a closed-loop technical logic, ensuring that the model adapts to data changes.
[0030] A second aspect of this application provides a multi-source data-driven road performance prediction device, comprising:
[0031] The data acquisition module is used to collect multi-source data of the target road segment, including time-series remote sensing data, GIS spatial data and on-site detection data;
[0032] The data preprocessing and fusion module is used to preprocess and fuse the multi-source data to generate a spatiotemporally aligned fused data sequence, and to perform uncertainty quantification on the fused data sequence.
[0033] The model prediction module is used to input the fused data sequence into the pavement performance degradation prediction model to obtain the pavement performance prediction sequence output by the model.
[0034] The maintenance timing calculation module is used to calculate the pavement maintenance timing of the target road section based on the pavement performance prediction sequence.
[0035] The pavement performance degradation prediction model is trained based on historical multi-source data. It is used to characterize the correlation between pavement performance and multi-source data and outputs prediction results in probabilistic form.
[0036] A third aspect of this application provides a computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the steps described in the method.
[0037] A fourth aspect of this application provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method as described.
[0038] Due to the adoption of the above technical solution, the beneficial effects achieved by this application are as follows:
[0039] This application's solution performs rigorous spatiotemporal alignment and uncertainty quantification on multi-source data during the data preprocessing stage, and employs a pavement performance degradation prediction model capable of outputting probabilistic prediction results (such as a Bayesian neural network) at the model level. The final output pavement performance prediction sequence contains a probability distribution or confidence interval. This allows maintenance management departments to intuitively assess prediction risks; for example, they can not only know when performance might fall below a threshold, but also the probability of this event occurring. This achieves a leap from "deterministic decision-making" to "risk-aware decision-making," significantly reducing the risk of maintenance misjudgments (false alarms or missed alarms) caused by prediction uncertainty.
[0040] By quantifying the uncertainty of the fused data sequence and incorporating this information into the model training and prediction process, this approach enables the model to recognize and adapt to the noise and uncertainty inherent in the data itself. This makes the final prediction model more tolerant of various disturbances during the data acquisition process (such as cloud cover in remote sensing images, sensor errors, and sampling errors in on-site detection), ensuring that the prediction system maintains stable output performance even under conditions of fluctuating multi-source data quality.
[0041] By integrating time-series remote sensing data, GIS spatial data, and field detection data, this solution combines the advantages of macroscopic temporal trends, microscopic spatial variations, and high-precision location information. This enables the constructed model to capture complex degradation patterns that are difficult to detect from a single data source (e.g., accelerated local performance degradation caused by specific geographical environments or uneven spatial distribution of traffic loads). This allows for a more refined simulation and prediction of the future evolution trajectory of pavement performance, more closely approximating reality, and laying a solid foundation for developing differentiated and precise maintenance strategies. Attached Figure Description
[0042] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0043] Figure 1 A flowchart illustrating a multi-source data-driven road performance prediction method provided in this application embodiment;
[0044] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0045] Figure label:
[0046] 810, Processor; 820, Communication interface; 830, Memory; 840, Communication bus. Detailed Implementation
[0047] To more clearly illustrate the overall concept of this application, a detailed explanation is provided below with reference to the accompanying drawings.
[0048] Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below. It should be noted that, unless otherwise specified, the embodiments of this application and the features thereof can be combined with each other.
[0049] In this application, unless otherwise expressly specified and limited, the "above" or "below" of the second feature can mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. In the description of this specification, references to terms such as "an embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples.
[0050] Example 1
[0051] like Figure 1 As shown, a multi-source data-driven method for predicting road performance includes:
[0052] Collect multi-source data for the target road segment, including time-series remote sensing data, GIS spatial data, and on-site detection data.
[0053] As mentioned above, this step involves systematically collecting various types of data sources from the target road segment to construct a comprehensive, multi-dimensional pavement condition information database. Among these, time-series remote sensing data refers to pavement image sequences acquired periodically through remote sensing platforms such as satellites or drones, reflecting the dynamic trends of pavement surface attributes (such as reflectivity and texture) over time; GIS spatial data refers to static or quasi-static spatial attributes such as road topology, elevation models, and traffic network distribution integrated based on geographic information systems, used to provide the environmental context and spatial correlation information of the pavement; and field inspection data refers to pavement physical parameters directly obtained through field measurement equipment (such as sensors or manual inspections), such as deflection values, damage rates, or friction coefficients. These data have high accuracy and reliability and can serve as benchmarks for model calibration. The fundamental purpose of collecting this multi-source data is to overcome the limitations of a single data source by integrating macroscopic temporal trends, microscopic spatial variations, and high-precision location information, providing rich and complementary input features for subsequent pavement performance prediction models, thereby more accurately capturing the complex mechanisms of pavement degradation.
[0054] For example, time-series remote sensing data can be represented by monthly satellite multispectral image sequences. By analyzing changes in infrared reflectance, the evolution of road surface aging or water accumulation can be indirectly monitored. GIS spatial data can be specifically represented by road slope distribution, traffic flow heat maps, or surrounding land use types obtained from public databases or dedicated sensors. For instance, a digital elevation model (DEM) can be used to calculate road longitudinal slope to assess the impact of load stress distribution on road performance. Field inspection data may include road deflection value sequences measured periodically using automatic deflectometers, or road crack density data collected by damage inspection vehicles. These data are linked to spatiotemporal coordinates via GPS positioning to ensure spatial consistency with remote sensing and GIS data. For example, in a highway maintenance project, a multi-source dataset can be constructed by combining Landsat satellite time-series imagery (reflecting seasonal changes in road surface), road alignment data from GIS (identifying vulnerable areas such as curves), and field smoothness data collected by a mobile inspection vehicle. This provides realistic and diverse samples for model training.
[0055] The multi-source data is preprocessed and fused to generate a spatiotemporally aligned fused data sequence, and the uncertainty of the fused data sequence is quantified.
[0056] As mentioned above, this step is the core data processing stage for building a highly reliable prediction model. Its purpose is to transform multi-source raw data of varying origins and formats into a standardized data sequence that is strictly aligned in time and space and includes its own reliability assessment, for subsequent model consumption. The so-called "preprocessing and fusion" first includes spatiotemporal alignment, which involves using a series of algorithms to unify time-series remote sensing data (such as satellite image sequences), GIS spatial data (such as road slope maps), and field detection data (such as discrete point deflection values) onto the same spatiotemporal reference. For example, geographic coordinate registration is used to map all data to the same map projection, and temporal interpolation or resampling techniques are employed to normalize all data sequences to the same and equally spaced timestamps, thereby eliminating misalignments caused by different collection times and frequencies. Based on this, data fusion is performed, integrating the aligned data at the feature level. For example, for each spatial unit (such as a road segment) at each time point, its remote sensing reflectance features, GIS slope features, and field deflection features are combined to form a multi-dimensional feature vector. Furthermore, uncertainty quantification measures and labels the credibility of each data point or feature in the fused data sequence. Its core lies in identifying and quantifying the inherent errors of each data source (such as the systematic errors of remote sensing sensors, the uncertainty of GIS interpolation, and the random errors of field detection), and may use probabilistic methods (such as Bayesian modeling) to merge these uncertainties, giving the fused data sequence an overall or point-by-point reliability index, thereby clearly informing subsequent models which information is relatively reliable or questionable.
[0057] For example, the process begins with preprocessing and fusion: quarterly Sentinel satellite remote sensing imagery (reflecting changes in the macroscopic condition of the road surface) is fused with high-precision LiDAR scans obtained from GIS elevation data (providing accurate road slope), and semi-annual International Roughness Index (IRI) data collected by automated inspection vehicles. During processing, spatial registration algorithms are used to precisely map the pixels of the satellite imagery to the elevation points in the GIS data, mapping them to specific road station locations. For time alignment, since the satellite data is quarterly and the field inspections are semi-annual, time series interpolation algorithms (such as linear interpolation or physical model-based interpolation) can be used to generate monthly sequences with values available from all data sources. After fusion, the data feature vector for each road segment unit in each month includes the average remote sensing reflectance for that month, the average slope value for that road segment, and the estimated smoothness value for that month obtained through interpolation. Uncertainty quantification is then performed: For satellite reflectivity data, the uncertainty may stem from cloud cover, which can be addressed by providing a signal-to-noise ratio as an uncertainty indicator through an image quality assessment report; for GIS slope data, the uncertainty arises from the accuracy of lidar measurements, which can be defined by providing an elevation error range and calculating the slope uncertainty using the propagation law; for interpolated flatness data, the uncertainty can be estimated using the residuals of the interpolation model. Finally, a comprehensive uncertainty score is assigned to the fused feature vector for each month. This score can be a scalar or a covariance matrix, representing the overall reliability of the monthly data.
[0058] The fused data sequence is input into the pavement performance degradation prediction model to obtain the pavement performance prediction sequence output by the model.
[0059] As described above, this step is the core reasoning and prediction stage of the method. Its core lies in using a specially designed and trained "pavement performance degradation prediction model" to perform deep analysis on the fused data sequence generated in the preceding steps, which has undergone spatiotemporal alignment and uncertainty quantification, thereby outputting a quantitative prediction of pavement performance evolution over a future period. This model is not a simple regression or classifier, but a hybrid architecture capable of simultaneously capturing physical mechanisms and data-driven patterns. It guides the learning process by embedding physical constraints (e.g., performance degradation differential equations based on fracture mechanics or material fatigue theory as regularization terms or structural priors), ensuring that the prediction results conform to basic engineering principles. Simultaneously, it utilizes data-driven methods (such as deep neural networks) to learn complex nonlinear mapping relationships and influencing factors not fully covered by the physical model (such as spatiotemporal variations in traffic flow and micro-environment erosion) from the rich fused data sequence. Crucially, the model's output is a probabilistic prediction sequence. For example, instead of directly outputting "the Road Condition Index (PCI) will be 80 one year from now," it outputs "the probability distribution of the PCI one year from now, with a mean of 80, a standard deviation of 2, and a 15% probability that the PCI will be below the maintenance threshold of 75." This output format, by explicitly expressing the uncertainty of the prediction, elevates a single point estimate to a probabilistic forecast that includes risk information, providing a fundamental basis for subsequent risk-based decision-making.
[0060] For example, a 24-month (two-year historical) fusion data sequence, updated monthly (containing the monthly average remote sensing reflectance, GIS slope, traffic volume estimates, and interpolated field deflection values for each road segment), can be input into a prediction model built using a Bayesian neural network. During training, the network's loss function includes not only the mean square error between predicted and actual values but also a decay equation based on pavement mechanics as a physical constraint. After receiving this 24-month sequence, the model encodes it and infers the evolution trajectory of the pavement performance index (such as PCI) for the next 36 months (three-year prediction). The output is not a single curve but rather a probability distribution of the PCI value for each future month, obtained through multiple forward propagation sampling techniques such as Monte Carlo Dropout or variational inference. Ultimately, the model's output "pavement performance prediction sequence" can be represented as a set of quantile curves (such as the 10%, 50%, and 90% quantiles), clearly depicting the optimistic, most likely, and pessimistic scenarios of performance degradation. Maintenance managers can use this information to determine the optimal timing for maintenance, provided that sufficient reliability is ensured (e.g., when the probability of performance falling below a threshold exceeds 90%).
[0061] Based on the pavement performance prediction sequence, the timing of pavement maintenance for the target road section is calculated;
[0062] The pavement performance degradation prediction model is trained based on historical multi-source data. It is used to characterize the correlation between pavement performance and multi-source data and outputs prediction results in probabilistic form.
[0063] As mentioned above, this step is a crucial leap from "prediction" to "decision-making." Its core lies in transforming the probabilistic pavement performance prediction sequence output by the aforementioned model into a specific, executable, and cost-effective maintenance action time point. This calculation process is not a simple threshold comparison, but a decision-making process that integrates reliability analysis, economic trade-offs, and risk management. Specifically, it utilizes the probability distribution information in the prediction sequence (e.g., the probability that the pavement performance index will fall below a critical value at various future time points) combined with predefined maintenance strategy rules for comprehensive judgment. A typical decision logic is: when the probability that the pavement performance index will fall below a preset maintenance threshold at a future time point exceeds a pre-set reliability confidence level (e.g., 90% or 95%), that time point is determined as the recommended maintenance time. This method is essentially a risk warning mechanism, transforming maintenance decision-making from a passive mode of "acting only after performance inevitably deteriorates" to a proactive prevention mode of "intervening before performance is at extremely high risk of deterioration." Among them, the training features of the pavement performance degradation prediction model are crucial, as they form the basis for the model to achieve high-precision probabilistic prediction. The model is trained based on historical multi-source data (i.e., past time-series remote sensing, GIS spatial, and field detection data). Through algorithms such as deep learning, it automatically learns and quantifies the complex nonlinear correlation and uncertainty between pavement performance and various influencing factors from these historical data. Thus, when a new fused data sequence is input, it can reproduce this correlation and output probabilistic prediction results that reflect future uncertainties.
[0064] For example, the model outputs a predicted sequence of Road Surface Condition Index (PCI) for the next five years, presented as a probability distribution for each month. The decision-making system analyzes each future month sequentially: for instance, by month 36, the PCI probability distribution shows that 85% of the sampled results are below the maintenance threshold of 70; by month 37, this probability rises to 93%. At this point, if the system's preset confidence level for triggering maintenance action is 90%, then month 37 will be officially calculated and output as the recommended maintenance time. This means that the system's judgment that maintenance should be carried out in month 37 can prevent the actual road surface performance from deteriorating below the threshold with over 90% reliability, thereby maximizing the extension of the maintenance cycle and optimizing the efficiency of fund utilization while ensuring the road service level. To achieve this accurate probability output, the prediction model must be fully trained, for example, using historical satellite imagery of the road segment every month over the past ten years, continuous GIS traffic flow data, and annual manual inspection reports as training samples, allowing the model to learn how to identify the patterns and uncertainties of performance degradation from these multi-source historical data.
[0065] According to one embodiment of this application, preprocessing and fusing the multi-source data includes:
[0066] The time-series remote sensing data, GIS spatial data and field detection data are unified into a common spatiotemporal framework using a spatiotemporal registration algorithm.
[0067] Kriging interpolation is applied to spatially standardize the GIS spatial data, and dynamic time warping algorithm is applied to time-align the time-series remote sensing data and field detection data.
[0068] A joint uncertainty map is generated by fusing the uncertainties of various data sources using the Bayesian method. The uncertainties include the signal-to-noise ratio of remote sensing data, the interpolation variance of GIS data, and the measurement error distribution of field detection data.
[0069] As described above, a spatiotemporal registration algorithm is used to unify the time-series remote sensing data, GIS spatial data, and field detection data into a common spatiotemporal framework. This step aims to address the inconsistencies in spatial coordinate systems and time bases between data from different sources. Spatial registration maps various types of data to a unified geographic coordinate system; time base unification aligns the timestamps of all data to a standard time series, establishing a consistent spatiotemporal foundation for subsequent fusion.
[0070] The GIS spatial data are spatially standardized using Kriging interpolation. Based on geostatistical principles, this method performs optimal unbiased interpolation on discrete spatial observation points according to the variogram characteristics of the spatial data, generating continuous spatial distribution data. It also provides an estimated variance for each interpolation point, providing a basis for subsequent uncertainty quantification.
[0071] A dynamic time warping algorithm is applied to align the time-series remote sensing data and field detection data. This algorithm can effectively process time-series data with different sampling frequencies and observation time points. By finding the optimal time warping path, it non-linearly aligns different time series on the time axis, eliminating time phase differences caused by inconsistent sampling times.
[0072] This method uses a Bayesian approach to fuse uncertainties from various data sources, generating a joint uncertainty map. It takes the signal-to-noise ratio of remote sensing data, the interpolation variance of GIS data, and the measurement error distribution of field detection data as prior information. Through a probabilistic inference model, it integrates these sources of uncertainty, quantifies the overall uncertainty level of the fused data at each spatiotemporal location, and forms a joint uncertainty map reflecting the reliability of the data.
[0073] According to one embodiment of this application, the pavement performance degradation prediction model is a Bayesian neural network model, and its training process includes:
[0074] Historical multi-source data was used as input, and pavement performance index was used as label;
[0075] The optimization objective is to minimize the negative log-likelihood loss, and incorporates a physical constraint equation for pavement performance degradation. This physical constraint equation is based on fracture mechanics and takes the following form:
[0076]
[0077] Where P is the pavement performance index, k and α are attenuation parameters, and ϵ is random noise;
[0078] We employ a combination of stochastic gradient descent and Monte Carlo Dropout for model training to handle large-scale data.
[0079] As described above, historical multi-source data is used as the model input, including time-series remote sensing data, GIS spatial data, and field detection data; at the same time, the measured values of the pavement performance index for the corresponding period are used as training labels to establish the data foundation for supervised learning.
[0080] During training, the optimization objective of the model is set to minimize the negative log-likelihood loss. This loss function can effectively measure the difference between the probability distribution of the model output and the true observations, enabling the trained model to not only provide point predictions but also quantify the uncertainty of the predictions.
[0081] Meanwhile, the training process incorporates a physical constraint equation based on fracture mechanics. This equation describes the decay of the pavement performance index P over time t, where k is the decay coefficient, α is a material property parameter, and ε represents a random noise term. In implementation, this physical constraint is introduced into the loss function through a regularization term to ensure that the model predictions conform to the fundamental laws of materials mechanics.
[0082] The model training employs a combination of stochastic gradient descent and Monte Carlo Dropout. Monte Carlo Dropout maintains the model in an active state during both training and prediction phases, approximating Bayesian inference through multiple forward propagation random sampling, thereby estimating the posterior distribution of the model parameters. This training method effectively handles large-scale, multi-source data while preserving the model's ability to quantify uncertainty.
[0083] According to one embodiment of this application, calculating the timing of pavement maintenance based on the pavement performance prediction sequence includes:
[0084] Confidence intervals for pavement performance prediction sequences are generated by sampling multiple predictions from the output distribution of the pavement performance degradation prediction model using Monte Carlo simulation.
[0085] When the probability of the pavement performance index falling below a preset threshold exceeds a preset probability threshold, maintenance recommendations are triggered.
[0086] The preset probability threshold is optimized based on historical maintenance data.
[0087] As described above, multiple samples are taken from the output distribution of the pavement performance degradation prediction model using Monte Carlo simulation. Specifically, a trained Bayesian neural network is used for multiple forward propagations, with each propagation generating different prediction results by randomly dropping some network connections (Monte Carlo Dropout), thus obtaining multiple sets of pavement performance prediction sequences. These prediction sequences are then statistically analyzed to generate confidence intervals for pavement performance prediction sequences with a certain confidence level.
[0088] Maintenance decisions are made based on the generated confidence intervals. The system continuously monitors the predicted distribution of pavement performance index at various future time points. When it detects that the probability of the pavement performance index falling below a preset threshold at a certain future time point exceeds a preset probability threshold, maintenance recommendations are automatically triggered.
[0089] The preset probability threshold is optimized by analyzing historical maintenance data. Specifically, it is determined by statistically analyzing road surface condition data from historical maintenance records when maintenance was actually required, and combining this with maintenance effectiveness evaluation results, to establish an optimal probability threshold that ensures both road surface performance and optimizes maintenance resource allocation.
[0090] According to one embodiment of this application, a dynamic feedback step is also included:
[0091] Newly acquired multi-source data is periodically received and compared with the pavement performance prediction sequence to calculate residuals.
[0092] If the residual exceeds the preset tolerance, the retraining of the pavement performance degradation prediction model will be automatically triggered to update the model parameters.
[0093] The dynamic feedback step forms a closed-loop technical logic, ensuring that the model adapts to data changes.
[0094] As described above, the system periodically receives newly acquired multi-source data, including updated time-series remote sensing data, GIS spatial data, and field detection data. After undergoing the same preprocessing and fusion process as the training data, this new data is compared with the pavement performance prediction sequences generated in previous prediction periods.
[0095] The system calculates the residual between the actual observed values and the predicted values. This residual is obtained by comparing the measured values of pavement performance in the newly collected data with the model predictions at the corresponding time points, reflecting the degree of deviation between the model predictions and the actual conditions.
[0096] When the calculated residual exceeds the preset tolerance value, the system automatically triggers the retraining mechanism of the pavement performance degradation prediction model. This retraining process uses a complete historical dataset including new data to update the weight parameters of the Bayesian neural network, enabling the model to adapt to changes in data distribution.
[0097] This dynamic feedback mechanism constitutes a complete closed-loop technical logic. Through the cyclical execution of three stages—continuous monitoring, residual evaluation, and model updating—the predictive model is ensured to adapt to changes in pavement performance degradation patterns over time, maintaining predictive accuracy.
[0098] A second aspect of this application provides a multi-source data-driven road performance prediction device, comprising:
[0099] The data acquisition module is used to collect multi-source data of the target road segment, including time-series remote sensing data, GIS spatial data and on-site detection data;
[0100] The data preprocessing and fusion module is used to preprocess and fuse the multi-source data to generate a spatiotemporally aligned fused data sequence, and to perform uncertainty quantification on the fused data sequence.
[0101] The model prediction module is used to input the fused data sequence into the pavement performance degradation prediction model to obtain the pavement performance prediction sequence output by the model.
[0102] The maintenance timing calculation module is used to calculate the pavement maintenance timing of the target road section based on the pavement performance prediction sequence.
[0103] The pavement performance degradation prediction model is trained based on historical multi-source data. It is used to characterize the correlation between pavement performance and multi-source data and outputs prediction results in probabilistic form.
[0104] A second aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any of the embodiments of the first aspect above.
[0105] Figure 2 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 2 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call logical instructions in the memory 830 to execute the method in any of the embodiments of the first aspect described above, the method including:
[0106] Collect multi-source data for the target road segment, including time-series remote sensing data, GIS spatial data, and on-site detection data;
[0107] The multi-source data is preprocessed and fused to generate a spatiotemporally aligned fused data sequence, and the uncertainty of the fused data sequence is quantified.
[0108] The fused data sequence is input into the pavement performance degradation prediction model to obtain the pavement performance prediction sequence output by the model.
[0109] Based on the pavement performance prediction sequence, the timing of pavement maintenance for the target road section is calculated;
[0110] The pavement performance degradation prediction model is trained based on historical multi-source data. It is used to characterize the correlation between pavement performance and multi-source data and outputs prediction results in probabilistic form.
[0111] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0112] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer being able to perform the methods provided by the above methods, the method comprising:
[0113] Collect multi-source data for the target road segment, including time-series remote sensing data, GIS spatial data, and on-site detection data;
[0114] The multi-source data is preprocessed and fused to generate a spatiotemporally aligned fused data sequence, and the uncertainty of the fused data sequence is quantified.
[0115] The fused data sequence is input into the pavement performance degradation prediction model to obtain the pavement performance prediction sequence output by the model.
[0116] Based on the pavement performance prediction sequence, the timing of pavement maintenance for the target road section is calculated;
[0117] The pavement performance degradation prediction model is trained based on historical multi-source data. It is used to characterize the correlation between pavement performance and multi-source data and outputs prediction results in probabilistic form.
[0118] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods provided by the above methods, the method comprising:
[0119] Collect multi-source data for the target road segment, including time-series remote sensing data, GIS spatial data, and on-site detection data;
[0120] The multi-source data is preprocessed and fused to generate a spatiotemporally aligned fused data sequence, and the uncertainty of the fused data sequence is quantified.
[0121] The fused data sequence is input into the pavement performance degradation prediction model to obtain the pavement performance prediction sequence output by the model.
[0122] Based on the pavement performance prediction sequence, the timing of pavement maintenance for the target road section is calculated;
[0123] The pavement performance degradation prediction model is trained based on historical multi-source data. It is used to characterize the correlation between pavement performance and multi-source data and outputs prediction results in probabilistic form.
[0124] Example 2
[0125] This technical solution constructs a closed-loop, multi-source data-driven prediction system. The core of the solution lies in introducing an uncertainty quantification module and a dynamic feedback mechanism to ensure adaptive optimization of the data fusion and prediction processes. The technical logic is complete and closed-loop: from data acquisition, preprocessing, and feature extraction to model training, prediction output, and feedback updates, a self-validating loop is formed.
[0126] 1. Data Preprocessing and Fusion Module
[0127] Data Acquisition: Integrating multi-source data, including:
[0128] Temporal remote sensing data: such as road surface reflectance sequences collected by satellites or drones, with a time resolution of weekly or monthly, covering infrared and visible light bands, used to monitor changes in road surface.
[0129] GIS spatial data: such as road network topology, elevation models, and traffic flow distribution, with spatial resolution at the meter level, providing road structure and environmental context.
[0130] On-site testing data, such as deflection testing and breakage rate measurement, are obtained through mobile sensors or manual inspections. The data points are sparse but the accuracy is high.
[0131] Data alignment and cleaning: Spatiotemporal registration algorithms are used to unify data from different sources into a common spatiotemporal framework. For example, Kriging interpolation is used to spatially standardize GIS spatial data, with the following formula:
[0132]
[0133] in, These are predicted point values. The weights are calculated based on a semi-variogram to minimize spatial errors. For time-series data, the Dynamic Time Warping (DTW) algorithm is applied to align the time series and address the time offset between remote sensing and detection data.
[0134] Uncertainty quantification: An uncertainty index is introduced for each data source. For example, the uncertainty of remote sensing data is represented by the signal-to-noise ratio (SNR), the uncertainty of GIS data is measured by interpolation variance, and the uncertainty of field data is described by the distribution of measurement error. These uncertainties are fused using Bayesian methods to form a joint uncertainty map.
[0135] 2. Feature Engineering and Dimensionality Reduction Module
[0136] Feature extraction: Extracting key features from fused data, including:
[0137] Temporal characteristics, such as the attenuation slope of road surface reflectivity and the amplitude of seasonal fluctuations, are fitted using an autoregressive integral moving average (ARIMA) model, with the following formula:
[0138]
[0139] Where ϕ(B) and θ(B) are polynomials, B is the lag operator, and d is the difference order, used to capture time trends.
[0140] Spatial features, such as road curvature and slope gradient, are calculated based on GIS data using geometric differential algorithms.
[0141] Field characteristics: such as the statistical distribution of deflection values (mean, variance).
[0142] Dimensionality Reduction and Selection: Principal Component Analysis (PCA) or Variational Autoencoder (VAE) is applied to reduce the dimensionality of high-dimensional features, retaining more than 95% of the variance to reduce overfitting and highlight key influencing factors. Simultaneously, the mutual information criterion is used to select features with the highest correlation to pavement performance.
[0143] 3. Road surface performance degradation prediction model construction and training module
[0144] Model Architecture: A hybrid model is employed, combining physics-driven and data-driven approaches. The core is a Bayesian Neural Network (BNN) to quantify prediction uncertainty. The model input consists of the aforementioned features, and the output is the future sequence of pavement performance indices (such as PCI, pavement condition index).
[0145] BNN structure: consists of an input layer, hidden layers, and an output layer. The hidden layers use random variables to represent weights and learn the posterior distribution through variational inference. During prediction, the output is a probability distribution, not a point estimate.
[0146] Attenuation dynamics modeling: Introducing physical constraints, such as performance degradation equations based on fracture mechanics:
[0147]
[0148] Where P is the performance index, k and α are decay parameters, and ϵ is random noise, which is combined with BNN through maximum likelihood estimation.
[0149] Training process: The model is trained using historical data, and the optimization objective is to minimize the negative log-likelihood loss, as shown in the formula:
[0150]
[0151] in, This is the output distribution of the BNN, and the KL term ensures consistency between the prior and posterior weights. Training employs a combination of stochastic gradient descent and Monte Carlo Dropout to efficiently handle large-scale data.
[0152] 4. Uncertainty Propagation and Dynamic Feedback Module
[0153] Uncertainty propagation: In the prediction chain, the accumulation of uncertainty from data input to model output is quantified. Monte Carlo simulations are used to sample multiple predictions from the posterior distribution of the BNN and calculate the confidence interval of the performance curve. For example, the prediction distribution of the future performance index P(t) is:
[0154]
[0155] Where, μ(t) and The mean and variance obtained through sampling are used to evaluate the reliability of predictions.
[0156] Maintenance timing decision: Based on the predicted distribution, a maintenance trigger threshold is defined. For example, when the probability of P(t) falling below a certain quantile (e.g., 5%) exceeds 90%, a maintenance recommendation is triggered. This avoids false alarms caused by a single threshold.
[0157] Closed-loop feedback: The system periodically receives new data (such as the latest remote sensing or detection data), compares it with the prediction results, and calculates the residuals. If the residuals exceed a preset tolerance, model retraining is automatically triggered, updating the BNN parameters. This feedback loop ensures the model adapts to environmental changes, forming a closed-loop logic: data → prediction → validation → update.
[0158] 5. Output and Visualization Module
[0159] Output results: Generate future performance curves (such as PCI over time), with confidence intervals; also output maintenance timing recommendations, including recommended time and uncertainty range.
[0160] Visual interface: Through integration with the GIS platform, it displays spatial distribution predictions, helping users intuitively understand the hot spots of road surface attenuation.
[0161] For any parts not mentioned in this application, existing technologies may be used or referenced.
[0162] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0163] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A multi-source data-driven method for predicting road performance, characterized in that, include: Collect multi-source data for the target road segment, including time-series remote sensing data, GIS spatial data, and on-site detection data; The multi-source data is preprocessed and fused to generate a spatiotemporally aligned fused data sequence, and the uncertainty of the fused data sequence is quantified. The fused data sequence is input into the pavement performance degradation prediction model to obtain the pavement performance prediction sequence output by the model. Based on the pavement performance prediction sequence, the timing of pavement maintenance for the target road section is calculated; The pavement performance degradation prediction model is trained based on historical multi-source data. It is used to characterize the correlation between pavement performance and multi-source data and outputs prediction results in probabilistic form.
2. The method according to claim 1, characterized in that, Preprocessing and fusing the multi-source data includes: The time-series remote sensing data, GIS spatial data and field detection data are unified into a common spatiotemporal framework using a spatiotemporal registration algorithm. Kriging interpolation is applied to spatially standardize the GIS spatial data, and dynamic time warping algorithm is applied to time-align the time-series remote sensing data and field detection data. A joint uncertainty map is generated by fusing the uncertainties of various data sources using the Bayesian method. The uncertainties include the signal-to-noise ratio of remote sensing data, the interpolation variance of GIS data, and the measurement error distribution of field detection data.
3. The method according to claim 1, characterized in that, The pavement performance degradation prediction model is a Bayesian neural network model, and its training process includes: Historical multi-source data was used as input, and pavement performance index was used as label; The optimization objective is to minimize the negative log-likelihood loss, and incorporates a physical constraint equation for pavement performance degradation. This physical constraint equation is based on fracture mechanics and takes the following form: Where P is the pavement performance index, k and α are attenuation parameters, and ϵ is random noise; We employ a combination of stochastic gradient descent and Monte Carlo Dropout for model training to handle large-scale data.
4. The method according to claim 1, characterized in that, Calculating the timing of pavement maintenance based on the pavement performance prediction sequence includes: Confidence intervals for pavement performance prediction sequences are generated by sampling multiple predictions from the output distribution of the pavement performance degradation prediction model using Monte Carlo simulation. When the probability of the pavement performance index falling below a preset threshold exceeds a preset probability threshold, maintenance recommendations are triggered. The preset probability threshold is optimized based on historical maintenance data.
5. The method according to claim 1, characterized in that, It also includes a dynamic feedback step: Newly acquired multi-source data is periodically received and compared with the pavement performance prediction sequence to calculate residuals. If the residual exceeds the preset tolerance, the retraining of the pavement performance degradation prediction model will be automatically triggered to update the model parameters. The dynamic feedback step forms a closed-loop technical logic, ensuring that the model adapts to data changes.
6. A multi-source data-driven pavement performance prediction device, characterized in that, include: The data acquisition module is used to collect multi-source data of the target road segment, including time-series remote sensing data, GIS spatial data and on-site detection data; The data preprocessing and fusion module is used to preprocess and fuse the multi-source data to generate a spatiotemporally aligned fused data sequence, and to perform uncertainty quantification on the fused data sequence. The model prediction module is used to input the fused data sequence into the pavement performance degradation prediction model to obtain the pavement performance prediction sequence output by the model. The maintenance timing calculation module is used to calculate the pavement maintenance timing of the target road section based on the pavement performance prediction sequence. The pavement performance degradation prediction model is trained based on historical multi-source data. It is used to characterize the correlation between pavement performance and multi-source data and outputs prediction results in probabilistic form.
7. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1-5.
8. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1-5.