Roadbed soil moisture content nondestructive determination method based on radar signal characteristic value

By combining ground-penetrating radar with signal feature analysis and machine learning algorithms, the problem of non-destructive, continuous, and efficient measurement of roadbed soil moisture content has been solved, enabling real-time quality control and accurate condition assessment of road engineering and avoiding structural damage.

CN121633129APending Publication Date: 2026-03-10GUANGXI TRANSPORTATION SCI & TECH GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies are insufficient for large-scale, continuous, efficient, and non-destructive determination of subgrade soil moisture content. Furthermore, traditional methods are prone to damaging the subgrade structure or resulting in discontinuous measurements, failing to meet the demands of road engineering for real-time quality control and accurate condition assessment.

Method used

By combining ground-penetrating radar with signal feature analysis and machine learning algorithms, radar data is preprocessed and multi-domain features are extracted. Feature values ​​with high correlation coefficients are selected, and a regression model is used to predict water content, avoiding structural damage and enabling continuous measurement.

Benefits of technology

It enables non-destructive, continuous, and efficient determination of subgrade soil moisture content, meeting the needs of road engineering for real-time quality control and accurate condition assessment, and providing intelligent and precise technical support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121633129A_ABST
    Figure CN121633129A_ABST
Patent Text Reader

Abstract

The invention discloses a roadbed soil moisture content nondestructive determination method based on a radar signal characteristic value, which comprises the following steps: collecting original data of a ground penetrating radar, and preprocessing to obtain radar profile data; extracting a time domain, a frequency domain and an instantaneous characteristic value; screening characteristic values based on the correlation coefficient matrix to obtain a dimension reduction characteristic set; and inputting the dimension reduction feature set into a pre-trained moisture content prediction model, and carrying out inversion to obtain the moisture content of the roadbed soil. According to the invention, continuous, efficient and lossless measurement of the water content of the roadbed soil is realized, and the problems of structure damage and low efficiency of traditional point type measurement are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of non-destructive testing of roadbed moisture content using ground-penetrating radar, and particularly relates to a non-destructive testing method for roadbed soil moisture content based on radar signal characteristic values. Background Technology

[0002] Accurate determination of the moisture content of subgrade soil is a core aspect of road engineering quality control and condition assessment, as its value directly determines the compaction efficiency, mechanical strength, and long-term stability of the soil. Currently, traditional methods for moisture content determination include in-situ detection techniques such as the drying method and time-domain reflectometry (TDRS). While the drying method is widely recognized for its accuracy, it is a point-based measurement with a small and discontinuous measurement range. Furthermore, its sampling method disrupts the integrity of the subgrade structure, and the excessively long drying cycle cannot meet the urgent need for real-time feedback during construction. Although the TDRS improves detection efficiency to some extent, the probe inserted during measurement still falls into the category of minimal-destruction measurement, essentially remaining a point-based measurement and unable to efficiently and continuously obtain moisture content distribution information across the entire subgrade cross-section.

[0003] In recent years, ground-penetrating radar (GPR), as a highly efficient and non-destructive geophysical exploration technology, has been increasingly applied in the field of road engineering inspection. It detects underground conditions by emitting high-frequency electromagnetic waves and receiving reflected signals from the subsurface medium. Existing GPR-based water content estimation methods largely rely on the accurate identification of specific reflective interfaces. This involves first calculating the wave velocity using the depth and one-way travel time of the interface, then calculating the dielectric constant using the relationship between the radar wave propagation speed and the dielectric constant, and finally using empirical formulas such as the Topp formula to calculate the water content. However, the accurate identification of reflective interfaces is easily affected by the inhomogeneity and complex internal structure of the subgrade materials, posing challenges to the accuracy and engineering practicality of the inversion results. Therefore, developing a method for determining the water content of subgrade soil that is wide-ranging, continuous, efficient, non-destructive, and reliable has become an urgent technical need to promote the intelligent and precise development of road engineering construction and maintenance. Summary of the Invention

[0004] This invention proposes a non-destructive method for determining the moisture content of subgrade soil based on radar signal feature values, in order to solve the problems existing in the prior art.

[0005] To achieve the above objectives, this invention provides a non-destructive method for determining the moisture content of subgrade soil based on radar signal characteristic values, comprising the following steps:

[0006] The raw radar data collected by ground-penetrating radar on the target roadbed is acquired, and the raw radar data is preprocessed to obtain radar profile data characterizing the spatiotemporal response features.

[0007] Multi-domain feature extraction is performed on the radar profile data to obtain several radar signal feature values;

[0008] Based on the correlation coefficient matrix, the radar signal feature values ​​are filtered to obtain the target feature set after dimensionality reduction.

[0009] The target feature set is input into a pre-trained moisture content prediction model to obtain the predicted moisture content of the target roadbed.

[0010] The moisture content prediction model is obtained by regression training on a sample dataset, which includes several sample radar signal feature values ​​and their corresponding true moisture content.

[0011] Optionally, preprocessing of the raw radar data includes:

[0012] The original radar data is reorganized into a two-dimensional sampling matrix in spatial order;

[0013] The two-dimensional sampling matrix is ​​sequentially subjected to DC removal, gain adjustment, filtering, and background denoising to obtain the radar profile data.

[0014] Optionally, the multi-domain feature extraction includes:

[0015] Instantaneous feature values ​​of the radar profile data are extracted using Hilbert transform analysis;

[0016] The temporal characteristic values ​​of the radar profile data are extracted through temporal curve analysis.

[0017] Frequency domain feature values ​​of the radar profile data are extracted by power spectral density analysis.

[0018] Optionally, the instantaneous characteristic values ​​include the average instantaneous frequency and the average instantaneous phase;

[0019] The time-domain characteristic values ​​include total absolute amplitude, mean absolute amplitude, maximum peak amplitude, amplitude variance, amplitude kurtosis, total energy, and energy half-life.

[0020] The frequency domain characteristic values ​​include peak frequency, main frequency band energy, low frequency band energy, average signal frequency, and 100 Mbps bandwidth energy percentage.

[0021] Optional filtering operations include:

[0022] From the aforementioned radar signal feature values, feature values ​​with a correlation coefficient higher than the first threshold for the true moisture content are selected.

[0023] The target feature set is obtained by removing redundant feature values ​​whose correlation coefficients with each other are higher than the second threshold from the selected feature values.

[0024] Optionally, the moisture content prediction model is trained through the following steps:

[0025] Acquire sample radar data and their corresponding true moisture content to construct a sample dataset;

[0026] The sample radar data in the sample dataset are preprocessed and multi-domain feature extracted to obtain several sample radar signal feature values;

[0027] Feature filtering is performed on the feature values ​​of the aforementioned sample radar signals to obtain a dimension-reduced sample feature set;

[0028] The sample feature set is divided into a training set and a validation set;

[0029] The regression model is trained using the training set, and the prediction accuracy of each regression model is evaluated using the validation set. The model with the highest prediction accuracy is selected as the moisture content prediction model.

[0030] Optionally, the regression model includes an artificial neural network model, a random forest model, and a lightweight gradient booster model.

[0031] Optionally, the raw radar data is obtained by continuously collecting data along the survey line using a vehicle-mounted or hand-push ground-penetrating radar.

[0032] Optionally, the method further includes:

[0033] Before conducting actual measurements, a roadbed structure model was constructed for forward modeling, and simulated radar data was generated using electromagnetic simulation software to expand the sample dataset.

[0034] Compared with the prior art, the present invention has the following advantages and technical effects:

[0035] This invention represents a significant advancement in subgrade soil moisture content determination technology. Compared to traditional point-based detection methods such as drying and time-domain reflectometry, this invention employs ground-penetrating radar combined with signal feature analysis and machine learning algorithms. This enables large-scale, continuous, and efficient non-destructive moisture content determination without compromising the structural integrity of the subgrade. This method not only completely avoids structural damage and measurement interruptions caused by sampling or probe insertion, but also rapidly acquires moisture content distribution information across the entire cross-section of the survey line. It effectively meets the urgent needs of road engineering for real-time quality control and accurate condition assessment, providing reliable technical support for intelligent and precise road construction and maintenance. Attached Figure Description

[0036] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0037] Figure 1 This is a flowchart of a method according to an embodiment of the present invention;

[0038] Figure 2 This is a schematic diagram of the roadbed structure model established according to an embodiment of the present invention;

[0039] Figure 3 This is a diagram of the eigenvalue Pearson correlation coefficient matrix of an embodiment of the present invention;

[0040] Figure 4 This is a diagram of the ANN algorithm structure according to an embodiment of the present invention;

[0041] Figure 5 This is a graph showing the regression model analysis results of an embodiment of the present invention;

[0042] Figure 6 This is a comparison chart of the MAE of the three regression models in this embodiment of the invention;

[0043] Figure 7 This is a comparison chart of the MSE of three regression models in this embodiment of the invention;

[0044] Figure 8 R values ​​for the three regression models in this embodiment of the invention. 2 Comparison chart;

[0045] Figure 9 This is a graph showing the relationship between the predicted and actual moisture content of the roadbed in a field test according to an embodiment of the present invention. Detailed Implementation

[0046] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0047] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0048] Example 1

[0049] like Figure 1 As shown, this embodiment provides a non-destructive method for determining the moisture content of subgrade soil based on radar signal feature values, including the following steps:

[0050] The raw radar data collected by ground-penetrating radar on the target roadbed is acquired, and the raw radar data is preprocessed to obtain radar profile data characterizing the spatiotemporal response features.

[0051] Multi-domain feature extraction is performed on the radar profile data to obtain several radar signal feature values;

[0052] Based on a pre-built feature filtering model, several radar signal feature values ​​are filtered to obtain a dimensionality-reduced target feature set.

[0053] The target feature set is input into a pre-trained moisture content prediction model to obtain the predicted moisture content of the target subgrade.

[0054] The moisture content prediction model is obtained by regression training on a sample dataset, which includes several sample radar signal feature values ​​and their corresponding true moisture content.

[0055] Specifically, it can be set as follows:

[0056] (1) Construct a forward model for measuring water content using ground penetrating radar and obtain A-scan data through forward simulation using gprMax. Reconstruct the ground penetrating radar A-Scan data (amplitude-time series) into a two-dimensional sampling matrix (time-range series) in spatial order, and then process the two-dimensional sampling matrix by DC removal, gain adjustment, filtering and background denoising to obtain B-Scan data that characterizes the spatiotemporal response features.

[0057] (2) Extract time-domain feature values, frequency-domain feature values ​​and instantaneous feature values ​​from the radar signal B-Scan data through Hilbert transform analysis, time-domain curve analysis and power spectral density analysis.

[0058] (3) To optimize the model's generalization ability and reduce the level of explanation in the regression algorithm, the correlation coefficient matrix is ​​used to screen the water content of radar signal features and target features to obtain a dimension-reduced feature matrix. The feature screening method is to first select radar signal features with high correlation coefficients to the water content of target features, and then retain radar signal features with low correlation coefficients to eliminate redundant features.

[0059] (4) The dimensionality reduction feature matrix dataset is randomly divided into 80% training set and 20% validation set. Then, regression analysis is performed on the dataset using regression models such as Artificial Neural Network (ANN), Random Forest (RF), and Lightweight Gradient Boosting Machine (LGBM). Finally, the model with the highest prediction accuracy is selected by using the model error evaluation index.

[0060] (5) Preprocess the ground-penetrating radar measured data. The preprocessing of the original radar data includes: reorganizing the original radar data into a two-dimensional sampling matrix in spatial order; performing DC removal, gain adjustment, filtering and background denoising on the two-dimensional sampling matrix in sequence to obtain the radar profile data, converting it into B-Scan data, and then extracting the radar signal feature values ​​selected in step (3) from the B-scan data through Hilbert transform analysis, time-domain curve analysis and power spectral density analysis, and substituting them into the trained model with the highest prediction accuracy, so that the subgrade soil moisture content can be inferred from the radar signal feature values.

[0061] (6) Randomly select points on the field survey line, take soil samples and use the drying method to determine the actual moisture content, compare it with the regression model inversion results, and verify the accuracy of the inversion results.

[0062] The following experiment was conducted in this embodiment:

[0063] (1) Construct a forward model for measuring water content using ground penetrating radar and obtain A-scan data through forward simulation using gprMax. Reconstruct the ground penetrating radar A-Scan data (amplitude-time series) into a two-dimensional sampling matrix (time-range series) in spatial order, and then process the two-dimensional sampling matrix by DC removal, gain adjustment, filtering and background denoising to obtain B-Scan data that characterizes the spatiotemporal response features.

[0064] ① Model building and generation of forward simulation A-Scan dataset.

[0065] Based on the actual roadbed conditions, the roadbed structure is defined as three layers: an air layer, an upper embankment, and a lower embankment. The dielectric constant and conductivity of the medium in each layer of the roadbed are shown in Table 1.

[0066] Table 1

[0067]

[0068] The first layer is the air layer on the roadbed. The dielectric constant of air is 1, and the conductivity is 0 S / m.

[0069] The second layer is the upper embankment of the roadbed, which consists of 40cm thick test soil. The dielectric constant of the test soil is taken as 2~20 according to the fitting formula newly established in Chapter 5, and the electrical conductivity is taken as 0.001S / m.

[0070] The third layer is the lower embankment of the roadbed, which is 40cm thick Laibin soil with a dielectric constant of 9.5 and an electrical conductivity of 0.001S / m.

[0071] In the forward modeling, the test soil type for the second layer of the upper embankment was Laibin soil from Guangxi Zhuang Autonomous Region. The set variable was the dielectric constant of the second layer of the upper embankment. The dielectric constant value of this layer is related to the soil type, water content, and compaction degree. That is, the dielectric constant is taken according to the formula of the upper embankment in Table 2 (this formula is obtained by fitting the test data of dielectric constant ~ water content ~ compaction degree of Laibin soil from Guangxi Zhuang Autonomous Region).

[0072] Table 2

[0073]

[0074] This embodiment studies the application of ground-penetrating radar in forward modeling of highway subgrades. Figure 2 The roadbed structure model established for this embodiment.

[0075] Based on the stability conditions and numerical dispersion conditions of the solution of the FDTD algorithm, the parameters for the forward modeling of ground penetrating radar are set as shown in Table 3.

[0076] Table 3

[0077]

[0078] By following the steps and setting the model parameters as described above, ground-penetrating radar A-Scan (amplitude-time series) datasets under different compaction degrees and moisture contents can be generated through gprMax forward modeling.

[0079] ② The forward modeling process converts the A-scan dataset into a B-scan dataset.

[0080] Spatial reorganization of A-Scan data;

[0081] The forward modeling obtained through gprMax is a series of one-dimensional A-Scan signals acquired at spatial locations, with each A-Scan representing an amplitude-time series at a specific location. The first step in constructing B-Scan data is to reassemble these isolated A-Scan signals into a two-dimensional sampling matrix in spatial order.

[0082] Suppose there are M measurement points in total, and each A-Scan has N time sampling points.

[0083] Construct a two-dimensional matrix B(m,n), where m = 1, 2, ..., M represents the measurement point number; n = 1, 2, ..., N represents the time sampling point number. Each row of the matrix corresponds to an A-Scan signal: B(m,n) = [Am(t1), Am(t2), ..., Am(tN)]. This two-dimensional matrix reflects spatial position changes in the horizontal direction and temporal changes in the vertical direction.

[0084] *B-Scan data processing;

[0085] The two-dimensional sampling matrix is ​​processed by DC removal, gain adjustment, filtering, and background denoising to obtain B-Scan data that characterizes the spatiotemporal response features.

[0086] DC drift removal: Data on the radar profile may be entirely positive, entirely negative, or have an asymmetry between the positive and negative half-cycles. In this case, the data contains DC drift. Before further processing, the DC component needs to be eliminated or suppressed. The processing method is shown in formula (1).

[0087] (1)

[0088] In the formula: These are the matrix elements after DC removal; N is the number of sampling points; It is a two-dimensional sampling matrix.

[0089] #Gain: Automatic gain control (AGC gain) is used. The gain automatically decreases when the signal is strong and automatically increases when the signal is weak, thus ensuring the uniformity of strong and weak signals and facilitating the tracking of effective waves. Automatic gain is achieved by multiplying the radar record by a time-varying factor, and the algorithm is shown in formula (2).

[0090] (2)

[0091] In the formula: For radar recording after automatic gain control; For radar recording before automatic gain control; For an automatic gain weighting function, the weighting function value is the gain weighting factor for a fixed time.

[0092] The recording time of the ground penetrating radar channel is divided into 6 control points, namely K1, K2, K3, K4, K5 and K6. The average amplitude in each time period is calculated according to formula (3).

[0093] (3)

[0094] In the formula, pi represents a specific time period, and B represents the number of sample points within the time period pi. The average increase over the pi time period is used to calculate the control point gain weight according to formula (4).

[0095] (4)

[0096] In the formula: This represents the maximum average amplitude over various time periods. This represents the control point gain weights; the meanings of the other symbols are the same as before.

[0097] Filtering: Filtering can suppress interference signals outside the effective signal band, so as to highlight the effective signal and improve the profile signal-to-noise ratio. If the unit impulse response is a finite-length sequence, such a system is called a "finite-length unit impulse response system", i.e., an FIR filter, and is filtered according to formula (5).

[0098] (5)

[0099] In the formula: It is a discrete signal; This is the filtered signal; It is a finite impulse response sequence.

[0100] # Background Denoising: Background denoising uses a weighted moving average background denoising algorithm. When calculating the background value, first find the channel with strong signal and assign it a smaller weight when averaging. For the channel with weak signal, assign it a larger weight when averaging. In this way, the calculated background value contains as few abnormal signals as possible. The algorithm is shown in formula (6).

[0101] (6)

[0102] In the formula: For average background lane; Weights for each point; The amplitude is the average processed amplitude; N is the total number of channels in the large window.

[0103] (2) Extract time-domain characteristic values ​​(total absolute amplitude, average absolute amplitude, maximum peak amplitude, amplitude variance, amplitude kurtosis, total energy, energy half-life), frequency-domain characteristic values ​​(peak frequency, main frequency band energy, low frequency band energy, average signal frequency, 100 MHz bandwidth energy percentage) and instantaneous characteristic values ​​(average instantaneous frequency, average instantaneous phase) from the radar signal B-Scan data through Hilbert transform analysis, time-domain curve analysis and power spectral density analysis.

[0104] The formula for calculating the characteristic value of radar signals is as follows.

[0105] Total absolute amplitude:

[0106] (7)

[0107] In the formula: The total absolute amplitude; The absolute amplitude of each channel.

[0108] Mean absolute amplitude:

[0109] (8)

[0110] In the formula: n is the number of sampling points in the time window; This represents the average absolute amplitude; the other symbols have the same meaning as before.

[0111] Maximum peak amplitude:

[0112] (9)

[0113] In the formula: The maximum peak amplitude; It is a time series of finite length.

[0114] Amplitude variance:

[0115] (10)

[0116] In the formula: This represents the amplitude variance; the other symbols have the same meaning as before.

[0117] Amplitude kurtosis:

[0118] (11)

[0119] In the formula: This represents the amplitude variance; the other symbols have the same meaning as before.

[0120] Total Energy:

[0121] (12)

[0122] In the formula: This represents the amplitude variance; the other symbols have the same meaning as before.

[0123] Energy half-life:

[0124] (13)

[0125] In the formula: When the energy is half-decayed; Let be the energy of the signal at time t. This is the initial energy.

[0126] Peak frequency:

[0127] (14)

[0128] In the formula: Peak frequency; This is the amplitude spectrum curve.

[0129] Main frequency band energy:

[0130] (15)

[0131] In the formula: The main frequency band energy is represented by the line sx(f) = 0.707, which is used to cut the power spectrum curve at the corresponding two frequencies f1 and f2. The main frequency band is represented by Δf = f2 - f1. The meanings of the other symbols are the same as before.

[0132] Low-frequency energy:

[0133] (16)

[0134] In the formula: This represents low-frequency energy; the other symbols have the same meaning as before.

[0135] Average signal frequency:

[0136] (17)

[0137] The frequency at which the area under the power spectrum curve is divided into two equal parts is the average frequency FA(i). The lower the center frequency, the stronger the low-frequency energy.

[0138] Percentage of energy used in 100 Mbps bandwidth:

[0139] (18)

[0140] In the formula: Percentage of bandwidth energy used in a 100 Mbps connection; The power spectral density of the signal; For bandwidth; MHz.

[0141] Instantaneous phase:

[0142] (19)

[0143] In the formula: It is the instantaneous phase; This is a real signal; The result is the Hilbert transform of a real signal.

[0144] Instantaneous frequency:

[0145] (20)

[0146] In the formula: This represents the instantaneous frequency; the other symbols have the same meaning as before.

[0147] The "average instantaneous frequency" and "average instantaneous phase" can be obtained by averaging the instantaneous frequency and instantaneous phase along the time axis (or space axis).

[0148] (3) In order to optimize the model generalization ability and explanation level reduction problem in the algorithm regression, the radar signal feature values ​​are screened through the correlation coefficient matrix: from the radar signal feature values, feature values ​​with a correlation coefficient higher than the first threshold are screened; redundant feature values ​​with a correlation coefficient higher than the second threshold are removed from the screened feature values ​​to obtain the target feature set.

[0149] Specifically, the Pearson correlation coefficient is calculated for the water content of radar signal characteristic values ​​and target characteristic values, and the correlation coefficient matrix is ​​shown below. Figure 3 The eigenvalues ​​are selected based on the correlation coefficient matrix. The eigenvalue selection method is to first select radar signal eigenvalues ​​with high correlation coefficients to the target eigenvalue moisture content, and then retain radar signal eigenvalues ​​with low correlation coefficients to eliminate redundant eigenvalues.

[0150] The formula for calculating the Pearson correlation coefficient is shown in formula (21).

[0151] (21)

[0152] In the formula: and These are sample values ​​of two variables. and , where are the average values ​​corresponding to the invariable variables, and n is the sample size.

[0153] To ensure that the gprMax forward modeling results can fully characterize the complex interaction between electromagnetic waves and the water-bearing medium, it is necessary to retain as many radar eigenvalues ​​as possible. Figure 3 It is known that the correlation coefficients between the existing 14 radar signal feature values ​​and water content all exceed 0.43, and all of them are worth retaining. To address the high correlation between feature values, a domain-specific (time domain / frequency domain / instantaneous attribute) strategy is adopted for redundancy removal, with a correlation coefficient threshold set at 0.9.

[0154] Among the time-domain metrics, amplitude kurtosis has a relatively weak correlation with the other six highly correlated features (such as total energy and amplitude variance) (correlation coefficient of 0.72). Theoretically, to eliminate feature redundancy, only amplitude kurtosis and any collinear feature need to be retained. However, to prevent underfitting or decreased model accuracy due to too few features, based on the principle of "lowest correlation with amplitude kurtosis," amplitude variance and total energy (both with a correlation coefficient of 0.62 with kurtosis) were ultimately selected as retained terms.

[0155] In the frequency domain metrics, the correlation between peak frequency, low-band energy, and average signal frequency with other features did not exceed the threshold, thus qualifying for retention. However, due to the limitations of the narrowband antenna, the peak frequency degenerates into a constant across multiple samples, failing to provide effective information for water content discrimination, and is therefore excluded. Furthermore, the main band energy is highly correlated with the percentage of energy in the 100 MHz bandwidth. Given that the latter has a stronger correlation with water content, the percentage of energy in the 100 MHz bandwidth is preferred for retention.

[0156] In the instantaneous attributes, the average instantaneous phase is highly correlated with the frequency. Given that the absolute value of the correlation coefficient between the average instantaneous phase and the moisture content is larger (-0.90 > -0.84), it is more sensitive to changes in moisture content and is therefore retained.

[0157] In summary, seven radar signal characteristic values—amplitude variance, amplitude kurtosis, total energy, low-frequency band energy, average signal frequency, percentage of bandwidth (100 MHz), and average instantaneous phase—were selected as indicators for subsequent regression analysis. Table 4 shows some of the characteristic values ​​and target values ​​obtained after time-domain analysis, frequency-interval analysis, Hilbert transform analysis, and centering, which will serve as the dataset for subsequent model training. Due to space limitations, only a portion of the data is listed.

[0158] Table 4

[0159]

[0160] (4) The dimensionality reduction feature matrix dataset is randomly divided into 80% training set and 20% validation set. Then, regression analysis is performed on the dataset using regression models such as Artificial Neural Network (ANN), Random Forest (RF), and Lightweight Gradient Boosting Machine (LGBM). Finally, the model with the highest prediction accuracy is selected as the water content prediction model through the model error evaluation index.

[0161] The moisture content prediction model is trained through the following steps: acquiring sample radar data and their corresponding true moisture content to construct a sample dataset; preprocessing and extracting multi-domain features from the sample radar data in the sample dataset to obtain several sample radar signal feature values; performing feature filtering on the several sample radar signal feature values ​​to obtain a dimensionality-reduced sample feature set; dividing the sample feature set into a training set and a validation set; using the training set to train the regression model, and evaluating the prediction accuracy of each regression model through the validation set, selecting the model with the highest prediction accuracy as the moisture content prediction model.

[0162] ① An Artificial Neural Network (ANN) is a computational model inspired by the neural structure of the human brain. It consists of multiple neurons, each connected to other neurons via weights, and these connections transmit and process information. An ANN model typically includes an input layer, hidden layers, and an output layer, with multiple hidden layers. Each neuron receives input signals from the neurons in the previous layer, performs a weighted sum of these signals, processes them through an activation function, and then passes the result to the next layer of neurons. By continuously adjusting the connection weights, the ANN model can learn complex nonlinear relationships from the input data and perform tasks such as classification, regression, and clustering. The specific structure of the ANN in this invention patent is as follows... Figure 4 As shown.

[0163] ② Random Forest (RF) is an ensemble learning method based on decision trees. It generates multiple decision trees by random sampling and randomly selects features at each node to reduce the risk of overfitting. For regression problems, the output of each decision tree is a continuous value, and the final output of the model is the average of all decision trees, which is expressed mathematically as shown in formula (22).

[0164] (22)

[0165] In the formula Let be the regression result of the t-th decision tree; is the number of decision trees in the random forest.

[0166] This invention patent utilizes GridSearchCV to systematically fine-tune key parameters when constructing a random forest regression model: the number of trees is tested within the range of [150, 200, 250, 300], and the maximum depth is explored within the range of [5, 8, 10, 13, 15]. The model employs Friedman_MSE as the node splitting criterion, uses bootstrapping, and comprehensively evaluates the performance of different parameter combinations through 5-fold cross-validation, ultimately automatically selecting the optimal parameter configuration.

[0167] ③ The Lightweight Gradient Boosting Machine (LGBM) is an efficient machine learning framework based on Gradient Boosting Decision Tree (GBDT). Its core innovation lies in optimizing the efficiency and accuracy of the traditional GBDT algorithm. LGBM significantly reduces computational complexity through two key techniques: Gradient One-Sided Sampling (GOSS) and Mutually Exclusive Feature Bundling (EFB). GOSS reduces the data size while maintaining model accuracy by retaining samples with large gradients and randomly sampling samples with low gradients. EFB effectively reduces feature dimensionality and further accelerates the training process by merging mutually exclusive features (i.e., features that do not simultaneously take non-zero values) into a single feature. In addition, LGBM uses a histogram algorithm to bin continuous features, discretizing feature values ​​into histogram intervals, which greatly reduces computational overhead.

[0168] In constructing the LightGBM regression model, this invention systematically fine-tuned the three core structural parameters using GridSearchCV. The number of trees was selected within the range of [150, 200, 250], the maximum tree depth explored three levels [4, 5, 6], and the number of leaf nodes was tested within the range of [30, 60, 90]. The model uses Gradient Boosting Decision Tree (GBDT) as the basic algorithm, with regression as the target task, a learning rate of 0.1, and finally determines the optimal parameter combination through 5-fold cross-validation.

[0169] ④ Artificial neural networks (ANN), random forests (RF), and lightweight gradient boosters (LGBM) were used to perform regression training on 2965 sets of data with 8 features (7 radar signal features and 1 moisture content). (20% of the total data was analyzed, and the remaining 80% was used for model training). The results are shown in […]. Figure 5 As shown.

[0170] The mean absolute error (MAE), mean square error (MSE), and coefficient of determination (R²) were used. 2 The predictive capabilities of ANN, RF, and LGBM models were evaluated, and the results are shown in [link to evaluation]. Figures 6-8 .

[0171] Depend on Figures 6-8 The following conclusions can be drawn:

[0172] First, the MAE values ​​of RF, LGBM, and ANN are 0.76716, 0.82538, and 1.40641, respectively. RF has the smallest MAE. If MAE is used as the standard to evaluate the predictive performance of the three regression models, RF has the best predictive performance.

[0173] Secondly, the MSE values ​​of RF, LGBM, and ANN are 1.40853, 1.48711, and 3.59003, respectively. RF has the smallest MSE. If the predictive performance of the three regression models is evaluated based on MSE, RF has the best predictive performance.

[0174] Thirdly, the R of RF, LGBM, and ANN. 2 The values ​​are 0.98586, 0.98507, and 0.96395 respectively, representing the RF of RF. 2 Maximum, if R 2 To evaluate the predictive performance of the three regression models, RF has the best predictive performance.

[0175] Fourth, it integrates MAE, MSE, and R. 2 The results show that RF has the best predictive performance for the moisture content of typical soils in Guangxi Zhuang Autonomous Region.

[0176] (5) The ground-penetrating radar measured data is processed by DC removal, gain, filtering and background denoising, etc., and converted into B-Scan data. Then, through Hilbert transform analysis, time-domain curve analysis and power spectral density analysis, the seven radar signal feature values ​​selected in step (3) are extracted: amplitude variance, amplitude kurtosis, total energy, low frequency band energy, average signal frequency, 100 MHz bandwidth percentage and average instantaneous phase. These are substituted into the trained RF model with the highest prediction accuracy, and the subgrade soil moisture content can be inverted based on the seven radar signal feature values.

[0177] The methods for preprocessing the ground-penetrating radar (GPR) measured data, such as removing DC signals, and extracting radar signal feature values ​​in this step are the same as described above and will not be repeated here. The steps for inverting the subgrade soil moisture content are as follows: ① Data calibration: Ensure that the radar signal feature values ​​of the new data are consistent with the input format during model training. ② Model invocation: Invoke the already trained RF model. ③ Model prediction: Input the calibrated radar signal feature values ​​into the trained model to obtain the predicted moisture content value.

[0178] (6) Randomly select points on the field survey line, take soil samples and use the drying method to determine the actual moisture content, compare it with the regression model inversion results, and verify the accuracy of the inversion results.

[0179] The accuracy of the RF model in predicting the moisture content of the subgrade soil was verified at a highway subgrade construction site. Three subgrade moisture contents (15%, 18%, and 21%) were constructed at 3% mass moisture content intervals, following the prescribed test methods and procedures. The ground-penetrating radar used in the field test was an LTD-2600, with an antenna frequency of 400MHz, a repetition frequency of 16, a scan rate of 256 scans / s, a time window of 40ns, a sampling point count of 1024 points / channel, a signal position of 1, and automatic gain control and filtering functions enabled.

[0180] The data files generated after ground-penetrating radar scanning were converted into B-Scan data. After centering, seven radar signal characteristic values ​​were extracted: amplitude variance, amplitude kurtosis, total energy, low-frequency band energy, average signal frequency, percentage of 100 MHz bandwidth, and average instantaneous phase. These values ​​were then substituted into the RF model, and the predicted values ​​and true values ​​are shown in Table 5. Figure 9 As shown.

[0181] Table 5

[0182]

[0183] Through Table 5 and Figure 9 It can be seen that the maximum relative error between the predicted and actual mass moisture content from the roadbed field test is 1.7% (within 2%), which meets the error requirements in practical engineering. MAE, MSE, and R... 2 The predictive ability of the typical soil RF model in Guangxi Zhuang Autonomous Region for the construction site of the Hezhou-Fuping Expressway subgrade was quantitatively evaluated. The model error analysis is shown in Table 6.

[0184] Table 6

[0185]

[0186] Table 6 shows the MAE, MSE, and R of the inverted moisture content at a certain highway subgrade construction site. 2 The values ​​were 1.06, 1.30, and 0.88, respectively, indicating that the RF model had a high accuracy in predicting the moisture content at a highway subgrade construction site.

[0187] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for non-destructive determination of moisture content of subgrade soil based on eigenvalues of radar signal characteristics, characterized in that, The method comprises the following steps: Obtaining original radar data collected by ground penetrating radar on a target roadbed, and preprocessing the original radar data to obtain radar profile data representing spatiotemporal response characteristics; Extracting multi-domain features from the radar profile data to obtain a plurality of radar signal feature values; Based on a correlation coefficient matrix, screening the plurality of radar signal feature values to obtain a target feature set after dimension reduction; Inputting the target feature set into a pre-trained water content prediction model to obtain a water content prediction value of the target roadbed; The water content prediction model is obtained by regression training on a sample data set, and the sample data set includes a plurality of sample radar signal feature values and their corresponding true water content.

2. The method of claim 1, wherein, The preprocessing of the original radar data includes: Reorganizing the original radar data into a two-dimensional sampling matrix in spatial order; Processing the two-dimensional sampling matrix in turn to remove direct current, gain adjustment, filtering and background noise to obtain the radar profile data.

3. The method of claim 1, wherein, The multi-domain feature extraction includes: Extracting the instantaneous feature values of the radar profile data by Hilbert transform analysis; Extracting the time domain feature values of the radar profile data by time domain curve analysis; Extracting the frequency domain feature values of the radar profile data by power spectral density analysis.

4. The method of claim 3, wherein, The instantaneous feature values include average instantaneous frequency and average instantaneous phase; The time domain feature values include total absolute amplitude, average absolute value amplitude, maximum peak value amplitude, amplitude variance, amplitude kurtosis, total energy and energy half-life; The frequency domain feature values include peak frequency, main frequency band energy, low frequency band energy, average signal frequency and megahertz bandwidth energy percentage.

5. The method of claim 1, wherein, The screening of the plurality of radar signal feature values includes: From the plurality of radar signal feature values, screening out feature values with a correlation coefficient higher than a first threshold value with the true water content; From the screened feature values, eliminating redundant feature values with a correlation coefficient higher than a second threshold value with each other to obtain the target feature set.

6. The method of claim 1, wherein, The water content prediction model is trained by the following steps: Obtaining sample radar data and their corresponding true water content to construct a sample data set; Preprocessing and multi-domain feature extraction of the sample radar data in the sample data set to obtain a plurality of sample radar signal feature values; Feature screening of the plurality of sample radar signal feature values to obtain a sample feature set after dimension reduction; Dividing the sample feature set into a training set and a validation set; Training a regression model using the training set, and evaluating the prediction accuracy of each regression model through the validation set, and selecting the model with the highest prediction accuracy as the water content prediction model.

7. The method of claim 6, wherein, The regression model includes an artificial neural network model, a random forest model and a light gradient boosting machine model.

8. The method of claim 1, wherein, The original radar data is obtained by a vehicle-mounted or hand-held ground penetrating radar along the survey line.

9. The method of claim 1, wherein, The method further comprises: Before actual measurement, constructing a roadbed structure model for forward simulation, and generating simulated radar data by electromagnetic simulation software for expanding the sample data set.