Machine learning-based satellite-borne GNSS-R soil moisture fusion inversion method and system

By integrating spaceborne GNSS-R and SMAP data through machine learning methods, optimizing the inversion model and weight calculation, the problem of low accuracy in spaceborne GNSS-R soil moisture inversion was solved, achieving high-precision soil moisture inversion, which is applicable to agricultural irrigation zoning and hydrological models.

CN120950899BActive Publication Date: 2025-12-30SHANDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511475709.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2025-12-30
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

Existing soil moisture retrieval technology based on spaceborne GNSS-R suffers from low retrieval accuracy and poor adaptability, making it difficult to meet the requirements of high precision and high reliability. Furthermore, it does not fully integrate external auxiliary data and multi-source observation data from different GNSS systems, failing to meet the continuous grid data requirements of agricultural irrigation zones and hydrological models.

Method used

Using machine learning methods, a multi-source dataset is constructed by acquiring internal feature value data from spaceborne GNSS-R and external auxiliary data from SMAP. Quality control and spatiotemporal matching are performed, and the inversion model is optimized using the XGBoost library and Bayesian algorithm. Combined with the improved adaptive CRITIC weighting method, grid point weighting and fusion are carried out to achieve high-precision soil moisture inversion.

Benefits of technology

It achieves high-precision soil moisture inversion, can stably output results under different surface environments, meets the needs of agricultural irrigation zoning and hydrological models, and improves inversion accuracy and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950899B_ABST
    Figure CN120950899B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of GNSS-R remote sensing, in particular to a kind of machine learning-based spaceborne GNSS-R soil moisture fusion inversion method and system, the method comprising obtaining the multi-source data consisting of spaceborne GNSS-R internal eigenvalue data and SMAP external auxiliary data;Based on the multi-source data obtained, preprocessing is carried out, the matching data set is used as input to construct machine learning inversion model, orbit point soil moisture inversion is carried out using optimized inversion model, grid point weight fusion is carried out according to orbit point soil moisture inversion value, the weight of each GNSS system is calculated using improved adaptive CRITIC weight method, based on the soil moisture inversion value of grid point, result verification and application are carried out, the present application weights each GNSS system orbit point eigenvalue in each grid, effectively resist abnormal value interference, and the output continuous grid data can be directly adapted to agricultural irrigation zoning, hydrological model input, climate research and other business scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of GNSS-R remote sensing technology, and in particular to a machine learning-based method and system for fusion and inversion of soil moisture in spaceborne GNSS-R. Background Technology

[0002] Soil moisture, as a key parameter characterizing the hydrological state of the Earth's surface, plays an irreplaceable core role in precision agricultural production management, regional hydrological cycle simulation, global climate change response, and ecological environment monitoring. However, traditional methods of obtaining soil moisture have significant limitations. Against this backdrop, the spaceborne Global Navigation Satellite System Reflection Signal (GNSS-R) technology utilizes a bistatic radar observation mechanism to indirectly retrieve soil moisture by receiving echo signals from navigation satellites reflected off the Earth's surface. It has significant advantages such as all-weather, low-cost, and wide-coverage capabilities, and shows great application potential, especially in cloudy and rainy areas and large-scale regional monitoring. It has become a research hotspot in the field of soil moisture remote sensing monitoring.

[0003] However, existing soil moisture retrieval technologies based on spaceborne GNSS-R still face many technical challenges, making it difficult to meet the demands of practical applications for high-precision and high-reliability retrieval results. First, current technologies mostly rely on the limited feature values ​​of a single data source, spaceborne GNSS-R, for retrieval, failing to fully integrate complementary information from external auxiliary data. This results in insufficient input dimensions for the retrieval model, making it difficult to characterize the influencing factors of soil moisture in complex surface environments. Second, different GNSS systems exhibit varying sensitivity to soil moisture due to differences in signal frequency, coverage, and polarization. Retrieval from a single data source struggles to leverage the synergistic advantages of multi-source observation data, and customized retrieval models are not constructed for data from different GNSS systems and polarization methods, leading to low accuracy in orbital point soil moisture retrieval. Finally, existing technologies, when converting discrete orbital point soil moisture data into continuous grid data, do not consider the differences in information contribution from different GNSS systems within the grid, failing to meet the requirements for continuous grid data in scenarios such as agricultural irrigation zoning and hydrological model input. Currently, a machine learning-based spaceborne GNSS-R soil moisture fusion retrieval method and system is needed. Summary of the Invention

[0004] To address the issues of low monitoring accuracy and poor adaptability in traditional soil moisture retrieval methods, this invention provides a machine learning-based spaceborne GNSS-R soil moisture fusion retrieval method and system.

[0005] This invention provides a machine learning-based method for fusion and inversion of soil moisture in spaceborne GNSS-R, employing the following technical solution:

[0006] A machine learning-based method for fusion and inversion of soil moisture in spaceborne GNSS-R, comprising:

[0007] S1. Acquire multi-source data consisting of internal eigenvalue data from spaceborne GNSS-R and external auxiliary data from SMAP;

[0008] S2. Based on the acquired multi-source data, preprocessing is performed, including quality control of the internal feature value data of the spaceborne GNSS-R, and then spatiotemporal matching of the quality-controlled internal feature value data of the spaceborne GNSS-R with the external auxiliary data of SMAP to obtain a matching dataset.

[0009] S3. Construct an inversion model using the matching dataset as input, including building the inversion model using the XGBoost library as a framework, and optimizing the model parameters by combining Bayesian algorithm, grid search and random search.

[0010] S4. Use the optimized inversion model to invert soil moisture at orbital points, including classifying the preprocessed GNSS-R internal feature data and using the inversion model to invert each type of data.

[0011] S5. Based on the soil moisture inversion values ​​of the orbit points, perform grid point weighting and fusion, including dividing the soil moisture inversion values ​​of the orbit points into grids according to the region and assigning them to the corresponding grids and constructing a decision matrix. The improved adaptive CRITIC weighting method is used to calculate the weights of each GNSS system to obtain the soil moisture inversion values ​​of the grid points.

[0012] S6. Verify and apply the results based on the soil moisture inversion values ​​of the grid points.

[0013] Furthermore, the preprocessing based on the acquired multi-source data includes constructing a satellite GNSS-R data quality assessment index system. The index system includes data validity index, surface adaptability index, signal morphology index, and parameter rationality index. Based on the index system, quality thresholds are set to control the quality of the internal feature value data of the satellite GNSS-R, resulting in optimized satellite GNSS-R data. The nearest neighbor matching method is used to perform spatiotemporal matching between the satellite GNSS-R data and external SMAP auxiliary data to form a matching dataset.

[0014] Furthermore, the method of using the nearest neighbor matching to perform spatiotemporal matching between spaceborne GNSS-R data and SMAP external auxiliary data includes determining the coordinates of the specular reflection point corresponding to the spaceborne GNSS-R data based on the Fibonacci search algorithm, constructing a distance-to-sigma function of the signal propagation path between the spaceborne receiver and the GNSS satellite, taking the position corresponding to the minimum value of the distance-to-sigma function as the specular reflection point, extracting the observation timestamps of the spaceborne GNSS-R data and the SMAP external auxiliary data using the day of the year as the time base, finally calculating the great circle distance between the specular reflection point and the center point of the SMAP external auxiliary data grid, setting a spatial distance threshold, and completing the spatiotemporal matching based on the great circle distance. The formula for calculating the great circle distance is:

[0015] ,

[0016] in, The average radius of the Earth The latitude of the point of mirror reflection. The latitude of the center point of the SMAP grid. This represents the difference in longitude between the specular reflection point and the center point of the SMAP grid.

[0017] Furthermore, the matching dataset is used as input to construct an inversion model. This includes dividing the matching dataset into training and test sets, performing Z-score standardization on the training set, determining core features based on the feature importance ranking function of the XGBoost library as model input, using CART decision trees as base learners, and constructing the inversion model through forward step-by-step iterations. Mean squared error is used as the base loss function, and L2 regularization and tree complexity regularization terms are introduced to construct a comprehensive objective function.

[0018] ,

[0019] Where N represents the total number of samples in the training set. This represents the mean squared error loss function. This represents the true soil moisture value of the i-th training sample. This represents the soil moisture model prediction value for the i-th training sample. This represents the total number of base learners in the model. This represents the leaf node weights of the t-th CART base learner. This represents the L2 regularization parameter. The regularization parameter represents the tree complexity. This represents the total number of leaf nodes in the t-th CART base learner.

[0020] Furthermore, the optimization of model parameters includes dividing the value range of each core parameter into a discrete grid, traversing all parameter combinations, training the inversion model and calculating the root mean square error of the test set, selecting the parameter combination with the smallest RMSE, then randomly sampling within the value range of each core parameter, selecting the parameter combination corresponding to the smallest average RMSE after multiple runs, constructing a mapping relationship between parameters and RMSE based on a Gaussian process model, using the desired improved EI function as the acquisition function, and iteratively searching for the optimal parameter combination multiple times. The EI function calculation formula is:

[0021] ,

[0022] in, For the new parameter combination to be evaluated, This is the current optimal parameter combination. The root mean square error of the model on the test set, It is a non-negative truncation function. It represents the mathematical expectation.

[0023] Furthermore, the classification of the preprocessed GNSS-R internal feature data includes constructing a hardware identifier dimension based on the inherent hardware parameters of the spaceborne GNSS-R constellation, setting several discrete classification intervals for the channel identifier parameters, the hardware identifier dimension including channel identifier parameters for distinguishing signal receiving channels and satellite identifier parameters for distinguishing individual satellites, constructing a signal characteristic dimension based on the physical properties of the received signal, the signal characteristic dimension including GNSS system type parameters for distinguishing signal source systems and polarization mode parameters for distinguishing signal polarization states, constructing an association mapping relationship between the hardware identifier dimension parameters and the signal characteristic dimension parameters, and grouping all GNSS-R internal feature data of orbit observation points belonging to the same combination of GNSS system type parameters and polarization mode parameters into one category according to the association mapping relationship, resulting in multiple independent subsets.

[0024] Furthermore, the inversion model is used to invert various types of data, including standardizing each subset, calling the optimized inversion model framework, adaptively adjusting the model parameters based on the subsets to obtain inversion sub-models, inputting the orbital observation data that were not included in the sub-datasets into the corresponding inversion sub-models, and calculating and outputting the soil moisture inversion value for each orbital point through the inversion sub-models. Using the true SMAP soil moisture value as a benchmark, the root mean square error, mean absolute percentage error, and coefficient of determination of the inversion sub-models are calculated to verify the results. The formula for calculating the soil moisture inversion value is:

[0025] ,

[0026] in, Let K be the soil moisture inversion value at the t-th orbital point, and K be the number of CART base learners in the customized sub-model. For the k-th CART base learner, the feature vector of the t-th orbital point is... The predicted value.

[0027] Furthermore, the step of dividing the orbital point soil moisture inversion values ​​into grids by region and assigning them to corresponding grids to construct a decision matrix includes dividing the study area into regular grids according to a preset resolution, determining the spatial range of each grid by the latitude and longitude of its center point, assigning the orbital point inversion values ​​to corresponding grids based on the latitude and longitude coordinates corresponding to the orbital point soil moisture inversion values, obtaining the correspondence between the grids and the set of orbital point inversion values, counting the number of orbital points corresponding to each GNSS system in each grid, taking the minimum number of orbital points for each system as m, and constructing a decision matrix for the soil moisture inversion values ​​within the grid using the selected m samples as rows and the GNSS systems as columns.

[0028] Furthermore, the improved adaptive CRITIC weighting method for calculating the weights of each GNSS system includes normalizing the decision matrix data using a standardization algorithm that includes standardization of benefit-type and cost-type indicators; quantifying the information content by comparing the dispersion and information carrying capacity of the data from each GNSS system; constructing a Pearson correlation coefficient matrix between GNSS systems; calculating conflict based on the correlation coefficients; obtaining the core metric of each GNSS system by multiplying the information content and conflict; obtaining the GNSS system weights after normalization; calculating the median of the orbital point inversion values ​​for each GNSS system within each grid; and weighted summing of the medians to obtain the grid point soil moisture inversion value. The formula for calculating the grid point soil moisture inversion value is as follows:

[0029] ,

[0030] in, , , and These represent the median values ​​of orbit point inversion for GNSS systems from different signal sources. , , and These are the system weights corresponding to the inversion medians.

[0031] Secondly, a machine learning-based spaceborne GNSS-R soil moisture fusion and inversion system includes:

[0032] The data acquisition module is configured to acquire multi-source data consisting of internal eigenvalue data from the spaceborne GNSS-R and external auxiliary data from SMAP.

[0033] The preprocessing module is configured to preprocess the acquired multi-source data, including quality control of the onboard GNSS-R internal feature data, and spatiotemporal matching of the quality-controlled onboard GNSS-R internal feature data with SMAP external auxiliary data to obtain a matching dataset.

[0034] The model module is configured to: construct an inversion model using the matching dataset as input, including constructing the inversion model using the XGBoost library as a framework, and optimizing the model parameters by combining Bayesian algorithms, grid search and random search;

[0035] The inversion module is configured to: perform soil moisture inversion at orbital points using the optimized inversion model, including classifying the preprocessed GNSS-R internal feature value data and performing inversion on each type of data using the inversion model;

[0036] The planning module is configured to perform grid point weighting and fusion based on the soil moisture inversion values ​​of the orbit points, including dividing the soil moisture inversion values ​​of the orbit points into grids according to the region and assigning them to the corresponding grids and constructing a decision matrix, and using the improved adaptive CRITIC weighting method to calculate the weights of each GNSS system to obtain the soil moisture inversion values ​​of the grid points.

[0037] The verification module is configured to verify and apply the results based on the soil moisture inversion values ​​of the grid points.

[0038] In summary, the present invention has the following beneficial technical effects:

[0039] 1. This invention achieves comprehensive quality control of spaceborne GNSS-R data by constructing a four-dimensional quality assessment index system that considers data validity, surface adaptability, signal morphology, and parameter rationality, rather than using a single threshold for screening. This system can accurately eliminate low-quality data such as equipment malfunctions, atmospheric interference, and non-target surfaces, ensuring that the retained data fits the soil moisture monitoring scenario and provides high-quality basic input for subsequent inversion, avoiding model training bias caused by data impurities.

[0040] 2. This invention uses CART decision trees as base learners and constructs an integrated model through forward step-by-step iteration. Each round of base learner training is based on historical residuals, gradually optimizing the fitting ability to GNSS-R signal features and the complex nonlinear relationship of soil moisture. At the same time, L2 regularization and tree complexity regularization are introduced to construct a comprehensive objective function. L2 regularization suppresses abnormal fluctuations in leaf node weights, while tree complexity regularization controls the complexity of the base learner. This dual constraint avoids model overfitting and ensures that high-precision inversion results can be stably output under different land surface types.

[0041] 3. This invention trains a separate inversion sub-model for each type of dataset, rather than using a uniform model to process all data, ensuring that the model parameters are highly matched with the features of the subset datasets. At the same time, it eliminates dimensional differences through standardization and ensures the reliability of inversion by combining SMAP ground truth verification. Ultimately, it achieves high-precision calculation of soil moisture at discrete orbital points, providing high-quality basic data for regional grid fusion.

[0042] 4. The improved adaptive CRITIC weighting method of this invention dynamically determines the weights of each GNSS system through standardization to eliminate dimensions, standard deviation to quantify information content, correlation coefficient calculation to resolve conflicts, and core metric normalization. This method abandons the traditional extensive equal-weight fusion model, accurately identifies GNSS systems with rich information content and low information overlap with other systems, assigns them higher weights, fully leverages the unique advantages of each system, and avoids interference from systems with low information value in the fusion results.

[0043] 5. This invention weights and fuses the feature values ​​of each GNSS system orbit point within each grid, effectively resisting outlier interference. Then, it performs weighted summation based on adaptive weights to transform discrete orbit point data into continuous grid point data. This not only solves the problem that traditional orbit point data is spatially discontinuous and cannot meet the needs of regional applications, but also further improves the inversion accuracy through the integration of information from multiple systems. The output continuous grid data can be directly adapted to business scenarios such as agricultural irrigation zoning, hydrological model input, and climate research. Attached Figure Description

[0044] Figure 1 This is a schematic diagram of the overall process of a machine learning-based spaceborne GNSS-R soil moisture fusion and inversion method according to an embodiment of the present invention.

[0045] Figure 2 This is a scatter density map of the soil moisture inversion values ​​and true values ​​at the orbital points in an embodiment of the present invention.

[0046] Figure 3 This is a scatter density map of the soil moisture inversion values ​​and true values ​​of the grid points in an embodiment of the present invention.

[0047] Figure 4 This is a scatter density plot of the inverted values ​​and ground truth values ​​of the test set of the Bayesian-optimized XGBoost machine learning algorithm for multi-GNSS fusion grid points according to an embodiment of the present invention. Detailed Implementation

[0048] The present invention will be further described in detail below with reference to the accompanying drawings.

[0049] Example 1

[0050] Reference Figure 1This embodiment of a machine learning-based method for fusion and inversion of soil moisture in spaceborne GNSS-R includes:

[0051] S1. Acquire multi-source data consisting of internal eigenvalue data from spaceborne GNSS-R and external auxiliary data from SMAP;

[0052] S2. Based on the acquired multi-source data, preprocessing is performed, including quality control of the internal feature value data of the spaceborne GNSS-R, and then spatiotemporal matching of the quality-controlled internal feature value data of the spaceborne GNSS-R with the external auxiliary data of SMAP to obtain a matching dataset.

[0053] S3. Construct an inversion model using the matching dataset as input, including building the inversion model using the XGBoost library as a framework, and optimizing the model parameters by combining Bayesian algorithm, grid search and random search.

[0054] S4. Use the optimized inversion model to invert soil moisture at orbital points, including classifying the preprocessed GNSS-R internal feature data and using the inversion model to invert each type of data.

[0055] S5. Based on the soil moisture inversion values ​​of the orbit points, perform grid point weighting and fusion, including dividing the soil moisture inversion values ​​of the orbit points into grids according to the region and assigning them to the corresponding grids and constructing a decision matrix. The improved adaptive CRITIC weighting method is used to calculate the weights of each GNSS system to obtain the soil moisture inversion values ​​of the grid points.

[0056] S6. Verify and apply the results based on the soil moisture inversion values ​​of the grid points.

[0057] Specifically, a machine learning-based method for fusion and inversion of soil moisture in spaceborne GNSS-R includes the following steps:

[0058] like Figure 1 As shown, S1, acquire multi-source data consisting of internal eigenvalue data from the spaceborne GNSS-R and external auxiliary data from SMAP;

[0059] To acquire multi-source data consisting of internal eigenvalue data from spaceborne GNSS-R and external SMAP auxiliary data, the data source for the spaceborne GNSS-R data was first determined to be the Tianmu-1 (TM-1) constellation. This constellation contains 22 satellites (TM01 to TM22) and has 8 observation channels (C1 to C8). It can receive signals from four major GNSS systems: BDS, GPS, Galileo, and GLONASS, and supports four polarization modes: left-hand circular polarization, right-hand circular polarization, horizontal polarization, and vertical polarization. Using the Python programming language, a data parsing library was invoked to read the raw observation files of the TM-1 constellation (containing raw information such as signal propagation delay, Doppler shift, and signal power). Eight core internal eigenvalues ​​reflecting the interaction between the signal and the Earth's surface were extracted from the raw observation data. Among these, geographical location features were obtained by analyzing satellite orbital parameters. The latitude and longitude of the specular reflection point are calculated based on the signal propagation path to determine the spatial coordinates of the observation point. Signal morphology characteristics are extracted by analyzing the waveform data of the time-delay Doppler (DDM) image to obtain the leading edge slope and normalized backscattering cross section. The former reflects the steepness of the DDM signal leading edge, while the latter characterizes the surface's ability to scatter GNSS-R signals. Signal strength characteristics are obtained by calculating the ratio of effective signal energy to noise energy to obtain the signal-to-noise ratio. Simultaneously, reflectivity is extracted based on the energy attenuation after signal reflection, directly correlated with surface moisture content. Observation angle characteristics are obtained by calculating the incident angle through the spatial geometric relationship between the satellite, reflection point, and receiver, affecting the depth of interaction between the signal and the surface. Signal distribution characteristics are obtained by statistically analyzing the probability distribution of the DDM signal to extract kurtosis and skewness, respectively describing the steepness and asymmetry of the signal distribution, and assisting in correcting for interference from surface roughness.

[0060] Secondly, the source of external auxiliary data for SMAP was determined to be NASA's official data platform. SMAP satellite L3 level data products from the same period as the TM-1 constellation observation data were downloaded. These products have undergone preprocessing to eliminate atmospheric and radiation interference, resulting in high accuracy. The SMAP data files were loaded using Python's h5py library, and five types of external auxiliary data closely related to soil moisture inversion were extracted. Soil moisture data served as label data for model training and a ground truth benchmark for validation, used to measure the accuracy of the inversion results. Roughness coefficient data quantified the degree of surface undulation, used to correct for the interference of surface roughness on GNSS-R signal reflection. Surface temperature data reflected the surface... Thermodynamic state, correlated with soil moisture evaporation rate, can optimize the model's inversion performance in temperature-sensitive areas. Vegetation optical thickness and vegetation water content data quantify the occlusion and attenuation effects of vegetation on GNSS-R signals from two dimensions: the vegetation's attenuation ability on microwave signals and the vegetation's own moisture status, thereby improving the inversion accuracy in high-vegetation-coverage areas. Finally, the extracted spaceborne GNSS-R internal feature value data and SMAP external auxiliary data are converted and preliminarily processed to unify the data's timestamp format (e.g., converting to year-day + hour-minute-second format) and coordinate reference system (e.g., using the WGS84 coordinate system), forming an initial multi-source dataset, laying the foundation for subsequent preprocessing steps.

[0061] S2. Based on the acquired multi-source data, preprocessing is performed, including quality control of the internal feature value data of the spaceborne GNSS-R, and then spatiotemporal matching of the quality-controlled internal feature value data of the spaceborne GNSS-R with the external auxiliary data of SMAP to obtain a matching dataset.

[0062] Based on the acquired multi-source data, preprocessing is performed, and a multi-dimensional quality assessment system is constructed to complete the quality control of spaceborne GNSS-R data. High-precision spatiotemporal matching technology is used to associate multi-source data, and finally a standardized matching dataset is formed. First, a spaceborne GNSS-R data quality assessment index system is constructed. This system covers four core dimensions: data validity index, surface adaptability index, signal morphology index, and parameter rationality index. Each dimension is designed around the actual needs of soil moisture monitoring scenarios to ensure that high-quality data is selected from different levels. Among them, the data validity index uses the quality status indicators inherent in the spaceborne GNSS-R data as the basis for judgment. These indicators include twenty different status indicators, each corresponding to different data quality conditions (such as normal equipment observation, signal interference, invalid data, etc.). Only records where the indicators show no abnormalities in data quality and stable observation status are retained. The surface adaptability index focuses on the selection of target observation areas. Based on the surface cover type parameters in the data, only land type data is selected, excluding data from non-soil moisture monitoring target areas such as ocean, sea ice, and coastal areas. The signal morphology index uses the time-delay Doppler image (i.e., the image reflecting the energy distribution of the GNSS-R signal after reflection from the ground) as the analysis object. By extracting the location information of the signal power peak in the time-delay Doppler image, it is determined whether there is a time delay or frequency shift deviation in the signal. The parameter rationality index sets threshold ranges that conform to physical meaning and observation laws for key parameters in the internal characteristic values ​​of the spaceborne GNSS-R, ensuring that the parameter values ​​are reasonable and reliable.

[0063] Based on the aforementioned quality assessment index system, specific quality thresholds were set to conduct comprehensive quality control on the internal characteristic value data of the spaceborne GNSS-R. In data validity screening, quality status indicators were read using programming tools, and records with abnormal indicator displays were directly removed. In surface adaptability screening, records with a surface cover type parameter value of "1" (representing land) were retained, while other values ​​(such as "0" representing ocean and "2" representing sea ice) were removed. In signal morphology screening, the column and row numbers of the power peak in the time delay Doppler image were extracted, and only records with 6 to 15 columns and 11 to 51 rows were retained. If the number of columns exceeded this range, it indicated that the signal was abnormal. Severe time delay deviations, if the number of rows exceeds this range, indicate a severe Doppler frequency shift deviation in the signal, and such data must be discarded. In the parameter rationality screening, thresholds are set for incident angle, signal-to-noise ratio, and reflectivity as "less than 65 degrees", "between -50 and 50", and "between 0 and 0.2", respectively. At the same time, it is ensured that the leading edge slope, normalized backscattering cross section, and signal-to-noise ratio are not invalid values ​​(such as -9999). Through multi-dimensional threshold collaborative screening, low-quality and abnormal data are eliminated to obtain optimized spaceborne GNSS-R data.

[0064] Next, we construct the distance sum function for the signal propagation path between the spaceborne receiver and the GNSS satellite. This function is defined as the sum of the distance from the GNSS satellite to a point on the Earth's surface and the distance from that point to the spaceborne receiver. Let any point on the Earth's surface be P, the GNSS satellite be S, and the spaceborne receiver be R. Then the expression for the distance sum function is: ,in, This represents the distance from GNSS satellite S to point P on the Earth's surface. This represents the distance from a point P on the Earth's surface to the satellite receiver R. The distance calculation is based on an Earth ellipsoid model (such as the WGS84 ellipsoid), and the ellipsoidal distance formula is used to ensure accuracy. Then, the interval shortening ratio is determined based on the Fibonacci sequence. The Fibonacci sequence is generated according to the rule that "the next term equals the sum of the previous two terms," ​​and its recursive formula is: The initial conditions are set as follows: Subsequent terms are calculated sequentially using a recursive formula; simultaneously, the initial search interval for the specular reflection point is determined. This section is the surface arc segment formed by the line connecting the nadir point of the spaceborne receiver and the nadir point of the GNSS satellite. The starting coordinates of the GNSS satellite's nadir point on the arc segment. The coordinates of the satellite receiver's nadir point on the arc segment are given, and the coordinate units are converted to arc length parameters in kilometers. Next, the mirror reflection point is determined by iteratively calculating the distance and function values. In the k-th iteration, two trial points are determined based on the Fibonacci sequence ratio. The calculation formulas are as follows: , ,in, Let the left trial point be within the search interval of the k-th iteration. Let be the right trial point within the search interval of the k-th iteration. , and All numbers are specific terms in the Fibonacci sequence, without units. The sequence is generated using a recursive formula. 'n' is the preset total number of iterations, determined based on the initial interval length and positioning accuracy requirements. For example, if the initial interval length is 20 kilometers and the accuracy requirement is 10 meters, then n = 17. and Let these be the start and end points of the interval in the k-th iteration; calculate the distance and value corresponding to the two trial points. ,in, As a test point The distance and, As a test point The distance and sum, if This indicates that the distance and minimum value are located in the left subinterval. Let the interval of the (k+1)th iteration be... ;like This indicates that the distance and minimum value are located in the right subinterval. Let the interval of the (k+1)th iteration be... ;like Then, arbitrarily select either the left or right sub-interval as the next iteration interval; repeat the above iteration process until the number of iterations reaches the preset total number of iterations n, or the length of the current iteration interval. , If a preset positioning accuracy threshold is set, such as 10 meters, then the midpoint of the current interval, or the ground position corresponding to the test point with the smallest distance and value, is taken as the coordinates of the mirror reflection point of the spaceborne GNSS-R data, so as to achieve high-precision positioning of the reflection point.

[0065] Using the cumulative day of the year as the time base, time matching was completed between the spaceborne GNSS-R data and the SMAP external auxiliary data. The cumulative day of the year is the number of days calculated from January 1st of the current year (e.g., February 4th, 2024 corresponds to 35 cumulative days of the year). First, the observation timestamps of the spaceborne GNSS-R data and the SMAP external auxiliary data were extracted and converted into the cumulative day of the year format. Then, records with timestamps belonging to the same observation day were selected from both types of data to ensure that the soil moisture status corresponding to the spaceborne GNSS-R data and SMAP data participating in the matching were in the same time period. This avoids matching deviations caused by time differences (e.g., significant changes in soil moisture in a short period of time due to single-day precipitation), laying a foundation for time consistency for subsequent spatial matching.

[0066] Calculate the great circle distance between the specular reflection point and the center point of the SMAP external auxiliary data grid, complete spatial matching and form a matching dataset, obtain the coordinates (latitude, longitude) of the grid center point of the SMAP external auxiliary data, and for each GNSS-R specular reflection point coordinate ( , ), calculate its relationship with all SMAP grid center points ( , The great circle distance, combined with the coordinates of the already located mirror reflection point, is used to calculate the spatial distance between the two points using the great circle distance formula, which is:

[0067] ,

[0068] in, The average radius of the Earth The latitude of the point of mirror reflection. The latitude of the SMAP grid center point. The longitude difference between the specular reflection point and the SMAP grid center point is used. The Earth's average radius is taken as 6371 km. Both the latitude and longitude of the specular reflection point and the SMAP grid center point are converted to radians. The longitude difference between the two points is the absolute value of the difference in their longitude values ​​(also converted to radians). A maximum spatial distance threshold of 10 km is set. For each spaceborne GNSS-R observation point, a search is conducted in the SMAP data for grid center points with a great circle distance less than 10 km. If multiple matching grid points exist, the one with the smallest distance is selected for matching; otherwise, the spaceborne GNSS-R observation point is discarded. Finally, the spatially matched spaceborne GNSS-R data is correlated with the corresponding SMAP external auxiliary data. Each correlated data set contains optimized spaceborne GNSS-R internal feature values ​​and SMAP external auxiliary parameters, forming a spatiotemporally consistent and structurally unified matching dataset, providing high-quality input for subsequent inversion model construction.

[0069] S3. Construct an inversion model using the matching dataset as input, including building the inversion model using the XGBoost library as a framework, and optimizing the model parameters by combining Bayesian algorithm, grid search and random search.

[0070] like Figure 2 , Figure 3 , Figure 4 As shown, the spatiotemporal matching dataset generated in the second step is first divided and its features standardized to provide standardized input for model training. A stratified random partitioning method is used, splitting the matching dataset into training and testing sets in a 75%:25% ratio. The stratification is based on land cover type (e.g., farmland, woodland, grassland) and soil moisture range (e.g., 0~0.1cm³ / cm³, 0.1~0.2cm³ / cm³, 0.2~0.3cm³ / cm³), ensuring consistency in sample distribution characteristics between the training and testing sets and avoiding bias in model generalization ability due to uneven sample distribution. Subsequently, Z-Score standardization is performed on the feature data of the training set to eliminate the interference of differences in feature dimensions on model training. The standardization formula is:

[0071] ,in, These are the standardized eigenvalues. To train a set of original feature values ​​(such as the reflectivity of a spaceborne GNSS-R or the vegetation optical thickness of an SMAP). This is the mean of the feature in the training set (calculated by iterating through all samples in the training set). The standard deviation of this feature in the training set is calculated by iterating through the training set samples. At the same time, the SMAP soil moisture data in the training set and the test set are used as the "label ground truth" for model training and the "benchmark ground truth" for subsequent model accuracy verification, respectively, to complete the preprocessing of the model input data.

[0072] The basic model framework of the XGBoost library is invoked. Using all features of the standardized training set from step 1 (including internal features of the spaceborne GNSS-R and external auxiliary data from SMAP) as input and the true SMAP soil moisture value as output, preliminary model training is performed (without setting complex regularization parameters, only for feature importance assessment). After training, the importance score of each feature is extracted using the "feature_importances_" attribute provided by the XGBoost library. The calculation logic for this score is: the sum of the "split gain" generated by the feature during the splitting process of all base learners (CART decision trees). The larger the split gain, the higher the contribution of the feature to the soil moisture inversion result. A feature importance threshold of 0.01 is set, and redundant features with scores below this threshold are removed (such as kurtosis features under a certain polarization mode; if their total split gain is extremely small, their contribution to inversion accuracy is negligible, so they are removed). Core features with scores above the threshold (such as reflectivity, incident angle, vegetation optical thickness, roughness coefficient, etc.) are retained as the final input features of the subsequent inversion model, thereby reducing the model's computational load and improving the effectiveness of the input features.

[0073] Using CART (Classification and Regression Tree) decision trees as base learners, an XGBoost ensemble model structure is constructed using a forward step-by-step iterative strategy. Specifically, the initial model predictions are set as the mean of the true soil moisture values ​​in the training set. Where N is the total number of samples in the training set. This represents the true soil moisture value for the i-th sample; in the t-th iteration, the prediction residuals of the model from the previous t-1 iterations are first calculated. Given the predicted value of the i-th sample from the previous t-1 iterations of the model, train the t-th CART base learner using this residual as the target, so that the new base learner focuses on correcting the prediction errors of the historical model; after training, the predicted value of the t-th base learner is... , The core feature vector of the i-th sample is accumulated into the prediction result of the preceding model to obtain the model output of the t-th iteration. Repeat the above iterative process until the number of base learners reaches a preset initial value (e.g., 100 trees), completing the construction of the basic model structure. This structure can fit the complex nonlinear relationship between core features and soil moisture through the integration of multiple weak learners (CART trees).

[0074] To balance model fitting accuracy and generalization ability, a comprehensive objective function is designed, which uses mean squared error (MSE) as the basic loss function and integrates L2 regularization and tree complexity regularization terms. The expression is as follows:

[0075] ,

[0076] Where N represents the total number of samples in the training set. The mean squared error loss function is expressed as follows: , This represents the true soil moisture value of the i-th training sample. This represents the soil moisture model prediction value for the i-th training sample. This represents the total number of base learners in the model. This represents the leaf node weights of the t-th CART base learner. This represents the L2 regularization parameter. The regularization parameter represents the tree complexity. This represents the total number of leaf nodes in the t-th CART base learner. Through this objective function, the model will simultaneously minimize the fitting error and model complexity during training, effectively suppressing overfitting and ensuring inversion stability under different surface environments (such as high vegetation cover, arid / humid areas).

[0077] For the scenario of fusion and inversion of soil moisture from spaceborne GNSS-R and SMAP multi-source data, five core optimization parameters and their corresponding value ranges are determined based on the learning ability, complexity control and generalization ability of the XGBoost model.

[0078] Number of base learners ( ): The value range is ([50,300]), which represents the total number of CART decision trees in the model. Too few trees can easily lead to underfitting of the model (unable to capture complex nonlinear relationships), while too many trees will increase the computational cost and may cause overfitting.

[0079] Sample downsampling parameter: value range [0.5, 1.0], used to control the sample sampling ratio when training each base learner (e.g., a value of 0.8 means that each tree only uses 80% of the training samples). By randomly downsampling, the problem of uneven sample distribution is balanced, and the risk of overfitting is reduced.

[0080] Maximum depth of decision tree: ranges from [3,10], which controls the structural complexity of a single CART decision tree. Too deep a depth will cause the model to memorize training set noise (such as abnormal orbit point data), while too shallow a depth will fail to depict the fine relationship between features and soil moisture.

[0081] Parameter update step size: The value range is [0.01, 0.3], also known as the learning rate. It is used to adjust the weight contribution of the prediction results of each base learner in the model ensemble. If the step size is too small, it will prolong the model training time. If it is too large, it may cause the model to converge to a local optimum.

[0082] L2 regularization parameter The value range is [0.01, 1.0]. By penalizing the squared weights of the leaf nodes of the base learner, the model's oversensitivity to extreme samples is suppressed, and the model's generalization ability under different land surface types (such as farmland and forest land) is enhanced.

[0083] The range of each parameter value is divided into a discrete grid. In this embodiment, the number of base learners is divided into 50, 100, ..., 300 with a step size of 50, and the maximum depth of the decision tree is divided into 3, 4, ..., 10 with a step size of 1. All parameter combinations are traversed (a total of 3×6×8×3×3=1296 possibilities). After training the model for each combination, the root mean square error (RMSE) of the test set is calculated. The RMSE formula is:

[0084] ,

[0085] Where M is the number of samples in the test set. To determine the true value of the j-th sample in the test set, (For the predicted value), select the parameter combination with the smallest RMSE, then perform 100 random samplings within the range of each parameter value to generate 100 sets of parameter combinations. Calculate the RMSE after training the model for each set, repeat the process 10 times and take the average RMSE, then select the parameter combination with the smallest average.

[0086] Based on the Gaussian process model (using the radial basis function RBF as the kernel function), a mapping relationship between parameters and RMSE is constructed, with the expected improvement (EI) function used as the acquisition function. The EI function formula is as follows:

[0087] ,

[0088] in, For the new parameter combination to be evaluated, This is the current optimal parameter combination. The root mean square error of the model on the test set, It is a non-negative truncation function. To represent the mathematical expectation, the number of iterations is set to 50. In each iteration, the parameter combination with the largest EI value is selected to train the model and update the Gaussian process model.

[0089] Finally, the RMSE and mean absolute percentage error (MAPE) of the optimal parameter combinations obtained by the three algorithms on the test set are compared. (Formula) With the coefficient of determination ,formula , M represents the mean of the true values ​​in the test set, and M represents the number of samples in the test set. The parameter combination with the best overall performance is selected as the final parameters of the XGBoost inversion model, thus completing the construction and optimization of the entire inversion model.

[0090] S4. Use the optimized inversion model to invert soil moisture at orbital points, including classifying the preprocessed GNSS-R internal feature data and using the inversion model to invert each type of data.

[0091] Based on the optimized XGBoost inversion model framework from the third step, and targeting the preprocessed internal eigenvalue data of the spaceborne GNSS-R from the second step, high-precision inversion of soil moisture at each orbital point is achieved through two-dimensional data classification, subset standardization, customized sub-model training, orbital point inversion and validation. The specific implementation steps are as follows:

[0092] First, based on the inherent hardware parameters of the spaceborne GNSS-R constellation and the physical properties of the received signals, a two-dimensional classification system of hardware identification and signal characteristics is constructed to complete data classification and generate independent subsets. Taking the hardware parameters of the Tianmu-1 (TM-1) constellation as the basis, this dimension includes two types of parameters: one is the channel identification parameter used to distinguish the signal receiving channels (corresponding to the constellation's C1~C8, a total of 8 observation channels), which is divided into 8 discrete classification intervals according to the channel number (i.e., C1, C2, ..., C8 are independent intervals); the other is the satellite identification parameter used to distinguish individual satellites (corresponding to the constellation's TM01~TM22, a total of 22 satellites), which is set with 22 independent classification identifiers according to the satellite number (i.e., TM01, TM02, ..., TM22 are independent identifiers).

[0093] Next, the signal characteristic dimension is constructed: based on the physical properties of the signals received by the spaceborne GNSS-R, this dimension also includes two types of parameters: one is the GNSS system type parameter used to distinguish the signal source system, which is divided into four categories according to the constellation receiving capability: BDS, GPS, Galileo, and GLONASS; the other is the polarization mode parameter used to distinguish the signal polarization state, which is divided into four categories according to the signal polarization characteristics: left-hand circular polarization, right-hand circular polarization, horizontal polarization, and vertical polarization.

[0094] By using the hardware parameter configuration documents for the TM-1 constellation, a mapping relationship between the hardware identification dimension and the signal characteristic dimension is established. For example, it is specified that the C1-C4 channels of satellites TM01-TM10 only receive left-hand circularly polarized signals from the GPS system, while the C5-C8 channels of satellites TM11-TM22 can receive vertically polarized signals from the BDS system. Based on this mapping relationship, the GNSS-R internal characteristic value data (valid data that has passed quality control) after the second step of preprocessing is filtered and grouped: all orbital observation point data corresponding to the same combination of GNSS system type parameters and polarization mode parameters are grouped into one category, such as GPS system plus left-hand circularly polarized, "BDS system and vertical polarization," etc. Each category of data forms an independent subset, and each subset contains the GNSS-R internal characteristic values ​​(such as latitude and longitude of specular reflection points, leading edge slope, reflectivity, etc.) of all orbital observation points in that category, ensuring that the data characteristics of each subset are highly matched with the corresponding "GNSS system-polarization mode" combination.

[0095] For each independent subset generated in step 1, the Z-Score normalization algorithm, consistent with that in step 3, is used to eliminate dimensional differences, ensuring that the feature format of the subset is compatible with the optimized inversion model framework. The above normalization operation is performed on all GNSS-R internal feature values ​​(such as front slope, normalized backscattering cross section, signal-to-noise ratio, etc.) in each subset to obtain the normalized subset, providing input data in a unified format for subsequent customized sub-model training.

[0096] The S3-optimized XGBoost inversion model framework (including the basic structure of optimal parameter combinations, such as base learner type, loss function, regularization term form, etc.) is invoked to adaptively adjust the model parameters and train a customized inversion sub-model for each standardized subset of data.

[0097] Each standardized subset is divided into a training set and a test set in a 75%:25% ratio. The division is done using a stratified random partitioning method, based on the land cover type (e.g., farmland, woodland) and incident angle range (e.g., 0°~30°, 30°~60°) of the orbital points in the subset. This ensures that the sample distribution of the training set and the test set is consistent and avoids deviations in the generalization ability of the sub-model.

[0098] Using the standardized features of the sub-training set as input and the SMAP soil moisture ground truth of the corresponding orbital points as labels, a customized inversion sub-model is trained based on the XGBoost framework. Based on the core parameters optimized in the third step (such as the number of base learners and the maximum depth of the decision tree), adaptive fine-tuning of the parameters is performed within ±20% to address the feature differences in the sub-datasets (such as different signal-to-noise ratio distributions of different GNSS system signals and different reflectivity ranges of different polarization modes). (For example, if a sub-dataset corresponds to GPS system signals and its data dispersion is high, then...) The regularization value was fine-tuned from 0.5 in the third step to 0.6 to enhance the regularization effect. During training, the comprehensive objective function designed in the third step (including the mean squared error loss term and the double regularization term) was still used to ensure that the sub-model fits the sub-training set data while avoiding overfitting.

[0099] Repeat the above steps to train a corresponding customized inversion sub-model for each independent subset of data, such as the GPS-left circular polarization sub-model and the BDS-vertical polarization sub-model.

[0100] For each subset of data, input the orbital observation data that did not participate in the sub-model training (i.e., all standardized orbital point data outside the sub-test set) into the corresponding customized inversion sub-model. The sub-model will then calculate and output the soil moisture inversion value for each orbital point. The formula for calculating the inversion value is as follows:

[0101] ,

[0102] in, Let be the soil moisture inversion value at the t-th orbital point, and K be the number of CART base learners in the customized sub-model, with a value consistent with the number of base learners optimized in the third step. For the k-th CART base learner, the feature vector of the t-th orbital point is... The predicted value is calculated by the base learner based on the input feature vector. It traverses the split nodes of its own decision tree and finally falls into the corresponding leaf node. The weight value of that leaf node is... .

[0103] Following the above method, inversion calculations are performed on the orbital observation data corresponding to all subsets one by one to obtain soil moisture inversion results covering all orbital points, forming a correlated dataset.

[0104] Using the true SMAP soil moisture value as a benchmark, the inversion accuracy of the corresponding customized inversion sub-model was validated using subtest sets of each subset dataset. Validation metrics included root mean square error. Mean absolute percentage error With the coefficient of determination Set accuracy standards: , Only the orbital point inversion results corresponding to the sub-models that meet the standard are retained, and abnormal orbital point data with unacceptable accuracy are removed, thus forming a high-quality orbital point soil moisture inversion dataset.

[0105] S5. Based on the soil moisture inversion values ​​of the orbit points, perform grid point weighting and fusion, including dividing the soil moisture inversion values ​​of the orbit points into grids according to the region and assigning them to the corresponding grids and constructing a decision matrix. The improved adaptive CRITIC weighting method is used to calculate the weights of each GNSS system to obtain the soil moisture inversion values ​​of the grid points.

[0106] The grid resolution is set to 0.25°×0.25° (this resolution balances inversion accuracy and data continuity, adapts to the grid scale of SMAP data, and facilitates subsequent result comparison and verification). The latitude and longitude range of the study area is used as the boundary (e.g., the study area is 110°~120°E, 45°~35°S). Regular grids are divided according to the rule of "0.25° intervals for longitude and 0.25° intervals for latitude." Each grid cell has a unique spatial range determined by the latitude and longitude of its center point. For example, if the center point of a grid cell is (110.125°E, 44.875°S), then the spatial range of that grid cell is 110°~110.25°E, 45°~44.75°S, ensuring that all grid cells do not overlap and completely cover the study area.

[0107] Extract the latitude and longitude coordinates (i.e., the latitude and longitude of the specular reflection point) of each orbital point in the soil moisture inversion dataset obtained in step four. Determine the grid to which the orbital point belongs through spatial coordinate matching: for each orbital point, calculate whether its latitude and longitude fall within the spatial range of a certain grid (e.g., an orbital point with longitude 110.1° and latitude -44.9° falls within the grid with center points of 110.125°E and 44.875°S). If it falls within a grid range, assign the orbital point's soil moisture inversion value and its GNSS system type (BDS, GPS, Galileo, GLONASS) to that grid. If the orbital point's latitude and longitude exceed the grid range of the study area, it is removed. Finally, establish the association between each grid and the corresponding set of orbital point inversion values. For example, grid A corresponds to a set containing inversion values ​​of 20 GPS orbital points, 15 BDS orbital points, 18 Galileo orbital points, and 12 GLONASS orbital points.

[0108] For each grid, based on the set of inverted orbit points associated with it, samples are selected and a decision matrix is ​​constructed to provide structured input data for subsequent weight calculation. The number of valid orbit points corresponding to each GNSS system (BDS, GPS, Galileo, GLONASS) within each grid is counted. For example, in grid B, the GPS system has 25 orbit points, the BDS system has 22, the Galileo system has 20, and the GLONASS system has 18. The minimum number of orbit points for each system is taken as m (m=18 in this example). For each GNSS system, m samples are selected from its corresponding orbit point inversion values. The selection rule is to prioritize orbit points with "high signal-to-noise ratio (SNR > -10) and small incident angle (incident angle < 50°)". This is because high SNR signals are less affected by noise interference, and small incident angle signals can better reflect the surface soil moisture characteristics, thus improving the representativeness of the samples. If all orbit points of a certain system meet the selection criteria, m samples are selected by random sampling to ensure that the number of samples participating in matrix construction is consistent for each system, and to avoid affecting the fairness of weight calculation due to differences in sample size.

[0109] Using the selected m samples as rows and the GNSS system as columns, construct the decision matrix for soil moisture inversion values ​​within the grid. The matrix expression is as follows:

[0110] ,

[0111] in, This represents the soil moisture inversion value of the i-th sample at the orbital point in the GPS system. This represents the inversion value of the i-th sample in the BDS system. This represents the inversion value of the i-th sample in the Galileo system. This represents the inversion value of the i-th sample in the GLONASS system (i=1,2,…,m), with a matrix dimension of m rows × 4 columns, fully covering the representative inversion samples of each GNSS system within the grid.

[0112] To eliminate the dimensional differences in inversion values ​​from different GNSS systems (although feature standardization has been performed previously, the inversion values ​​may still have numerical range differences due to system characteristics), a standardization algorithm is used to normalize the decision matrix X, providing two standardization methods: benefit-based and cost-based. In this embodiment, the soil moisture inversion value is a benefit-based index, where a larger value is better; therefore, the benefit-based index standardization formula is selected:

[0113] ,

[0114] in, For the standardized matrix element in the i-th row and j-th column, These are the original inversion values ​​in the decision matrix. The minimum value of the original inversion value in column j (corresponding to a certain GNSS system) The maximum value of the original inversion value in column j is used to normalize all matrix elements to the interval [0,1], ensuring the fairness of subsequent information content and conflict calculations.

[0115] The dispersion of data from each GNSS system is quantified by standard deviation. Higher dispersion indicates richer spatial variation information about soil moisture and a greater amount of information contained in the system's data. The calculation formula is as follows:

[0116] ,

[0117] in, Let be the standard deviation of the standardized data of the j-th GNSS system. The mean of the standardized data of the j-th GNSS system is calculated by traversing the m standardized elements in the j-th column, where m is the number of samples.

[0118] Next, a Pearson correlation coefficient matrix is ​​constructed among the various GNSS systems to quantify the information overlap between them. A correlation coefficient closer to 1 indicates a higher degree of information overlap and lower conflict between the two systems; conversely, a lower correlation coefficient indicates a higher degree of conflict. The formula for calculating the Pearson correlation coefficient is as follows:

[0119] ,

[0120] in, Let be the correlation coefficient between the j-th GNSS system and the k-th GNSS system. For the standardized element in the i-th row and k-th column, The mean of the standardized data for the k-th GNSS system is used as the basis for calculating the conflict rate of each GNSS system based on the correlation coefficient. The conflict rate is the sum of the absolute values ​​of the correlation coefficients (1 - absolute values) of the system's data with all other systems. The calculation formula is as follows:

[0121] ,

[0122] in, The absolute value of the correlation coefficient. To reflect the independence between the two systems (the larger the value, the stronger the independence and the lower the information overlap). The larger the value, the higher the conflict between the system and other systems, and the more significant the contribution of unique information.

[0123] The core metric for each GNSS system is obtained by multiplying the information content by the conflict rate. The higher the core metric, the higher the information value of the system, and the higher its weight should be assigned. The formula for calculating the core metric is as follows:

[0124] ,

[0125] in, For the j-th GNSS system, For information content, To mitigate conflicts, the core metrics of all GNSS systems are normalized to obtain the weights of each system. The normalization formula is as follows:

[0126] ,

[0127] in, Let j be the weight of the j-th GNSS system. It is the sum of the four core metrics of the GNSS system, and satisfies For example, if GPS system BDS system Galileo system GLONASS system The sum is 1.35, and the GPS weight is... BDS weights Galileo weights GLONASS weights .

[0128] For all valid orbit point inversion values ​​of each GNSS system within each grid (not just the selected m samples, but all orbit points of the system within the grid), the median is calculated as the representative inversion value of the system within the current grid. The median is chosen over the mean because the median has a stronger resistance to outliers (such as extreme inversion values ​​caused by signal interference) and can more accurately reflect the inversion level of the system.

[0129] The specific calculation method is as follows: For all inversion values ​​of a certain GNSS system within the current grid, sort them from smallest to largest. If the number of orbit points is odd, take the middle value after sorting as the median; if it is even, take the average of the two middle values ​​as the median. Finally, the inversion medians of the four GNSS systems are obtained: GPS system inversion median. BDS system inversion median Galileo system inversion median GLONASS system inversion median .

[0130] Based on the GNSS system weights obtained in step 3 and the inversion median obtained in step 4, the final soil moisture inversion value of the current grid point is calculated by weighted summation. The fusion formula is as follows:

[0131] ,

[0132] in, , , and The inversion medians are for GPS, BDS, Galileo, and GLONASS systems, respectively. , , and These are the system weights corresponding to the inversion medians.

[0133] S6. Verify and apply the results based on the soil moisture inversion values ​​of the grid points.

[0134] The SMAP satellite L3 soil moisture product (spatial resolution 36km, calibrated by global stations, accuracy 0.04cm³ / cm³) from the same period as the study area was selected as the true benchmark. The SMAP data resolution was matched to 0.25°×0.25° through spatial resampling. For all resampled networks in the study area, the statistical indicators of the inversion values ​​in this embodiment and the true SMAP values ​​were calculated: root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²).

[0135] This study introduces soil moisture observation data from ground meteorological stations within the research area. For each ground station, grid point inversion values ​​within a 1km radius of the station are extracted, and the deviation between the inversion values ​​and the ground observation values ​​is calculated. Simultaneously, a spatial distribution map of the ground observation data is generated using a spatial interpolation algorithm (such as Kriging interpolation). This map is then visually and statistically compared with the spatial distribution map of the grid point inversion values. If the spatial distribution trends of the two data are consistent in high-humidity areas (such as farmland irrigation areas) and low-humidity areas (such as arid grassland areas), and the deviation between the inversion values ​​and the observation values ​​around the ground station is ≤0.03cm³ / cm³, then the spatial consistency of the inversion results is considered acceptable. The grid point soil moisture inversion dataset, verified through multiple dimensions, can be adapted to three core business scenarios: agriculture, hydrology, and ecology. This provides high-precision, high-spatiotemporal resolution soil moisture data support for agriculture, hydrology, and ecology. For example, based on the grid point soil moisture inversion values, combined with crop type distribution maps within the research area (such as wheat, corn, and rice planting areas), and according to the matching relationship between soil moisture and crop water requirements, dynamic irrigation suggestions can be output by utilizing changes in soil moisture.

[0136] Example 2

[0137] The difference between this embodiment and Embodiment 1 is that this embodiment provides a machine learning-based spaceborne GNSS-R soil moisture fusion and inversion system, including:

[0138] The data acquisition module is configured to acquire multi-source data consisting of internal eigenvalue data from the spaceborne GNSS-R and external auxiliary data from SMAP.

[0139] The preprocessing module is configured to preprocess the acquired multi-source data, including quality control of the onboard GNSS-R internal feature data, and spatiotemporal matching of the quality-controlled onboard GNSS-R internal feature data with SMAP external auxiliary data to obtain a matching dataset.

[0140] The model module is configured to: construct an inversion model using the matching dataset as input, including constructing the inversion model using the XGBoost library as a framework, and optimizing the model parameters by combining Bayesian algorithms, grid search and random search;

[0141] The inversion module is configured to: perform soil moisture inversion at orbital points using the optimized inversion model, including classifying the preprocessed GNSS-R internal feature value data and performing inversion on each type of data using the inversion model;

[0142] The planning module is configured to perform grid point weighting and fusion based on the soil moisture inversion values ​​of the orbit points, including dividing the soil moisture inversion values ​​of the orbit points into grids according to the region and assigning them to the corresponding grids and constructing a decision matrix, and using the improved adaptive CRITIC weighting method to calculate the weights of each GNSS system to obtain the soil moisture inversion values ​​of the grid points.

[0143] The verification module is configured to verify and apply the results based on the soil moisture inversion values ​​of the grid points.

[0144] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A machine learning based space-borne GNSS-R soil moisture fusion inversion method, characterized in that, The application relates to a soil moisture inversion method based on multi-source data, and belongs to the technical field of remote sensing. The application comprises the following steps: acquiring multi-source data composed of satellite-borne GNSS-R internal characteristic value data and SMAP external auxiliary data; performing preprocessing based on the acquired multi-source data, including quality control on the satellite-borne GNSS-R internal characteristic value data, and then performing space-time matching on the satellite-borne GNSS-R internal characteristic value data after quality control and the SMAP external auxiliary data to obtain a matching data set; constructing an inversion model by taking the matching data set as input, including constructing the inversion model by taking the XGBoost library as a framework, and optimizing model parameters by combining a Bayesian algorithm, a grid search and a random search; performing orbit point soil moisture inversion by using the optimized inversion model, including classifying the GNSS-R internal characteristic value data after preprocessing, and respectively inverting each type of data by using the inversion model; performing grid point weighted fusion according to the orbit point soil moisture inversion value, including dividing the orbit point soil moisture inversion value into grids according to regions, distributing the orbit point soil moisture inversion value to corresponding grids and constructing a decision matrix, and calculating the weight of each GNSS system by using an improved adaptive CRITIC weight method to obtain the soil moisture inversion value of the grid point; the step of dividing the orbit point soil moisture inversion value into grids according to regions, distributing the orbit point soil moisture inversion value to corresponding grids and constructing a decision matrix comprises the following steps: dividing the research region into regular grids according to a preset resolution, determining the spatial range of each grid through the central point longitude and latitude, distributing the orbit point inversion value to the corresponding grid based on the longitude and latitude coordinates corresponding to the orbit point soil moisture inversion value, obtaining the corresponding relationship between the grid and the orbit point inversion value set, taking the minimum value of the number of orbit points corresponding to each GNSS system in each grid as m, and taking the m samples after screening as rows and the GNSS system as columns to construct the decision matrix of the soil moisture inversion value in the grid; ; wherein, , , and are the GNSS system orbit point inversion medians of different signal sources, respectively, , , and are the system weights corresponding to each inversion median, respectively. the step of calculating the weight of each GNSS system by using the improved adaptive CRITIC weight method comprises the following steps: normalizing the decision matrix data by using a standardization algorithm including benefit-type index standardization and cost-type index standardization, quantitatively calculating the information amount by comparing the discrete degree and information carrying capacity of the data of each GNSS system, then constructing a Pearson correlation coefficient matrix between the GNSS systems, calculating the conflictiveness based on the correlation coefficient, obtaining the core measurement of each GNSS system through the product of the information amount and the conflictiveness, and then obtaining the weight of the GNSS system through normalization processing, calculating the median of the orbit point inversion value of each GNSS system in each grid, and obtaining the grid point soil moisture inversion value by weighted summation according to the medians, wherein the grid point soil moisture inversion value calculation formula is: performing result verification and application based on the soil moisture inversion value of the grid point. The application has the advantages that the soil moisture inversion method based on multi-source data can effectively improve the accuracy of soil moisture inversion, and can be applied to the fields of meteorology, hydrology, agriculture, ecology and the like.

2. The machine learning based spaceborne GNSS-R soil moisture fusion inversion method according to claim 1, characterized in that, The preprocessing based on the obtained multi-source data includes constructing a satellite-borne GNSS-R data quality evaluation index system, the index system including data validity indexes, ground surface adaptability indexes, signal form indexes and parameter rationality indexes, setting quality thresholds based on the index system to perform quality control on satellite-borne GNSS-R internal characteristic value data, obtaining optimized satellite-borne GNSS-R data, and using a nearest neighbor matching method to perform time-space matching of satellite-borne GNSS-R data and SMAP external auxiliary data to form a matched data set.

3. The machine learning based spaceborne GNSS-R soil moisture fusion inversion method according to claim 2, characterized in that, The time-space matching of satellite-borne GNSS-R data and SMAP external auxiliary data using the nearest neighbor matching method includes determining the coordinates of the specular reflection point corresponding to the satellite-borne GNSS-R data based on a Fibonacci search algorithm, taking the position corresponding to the minimum value of the range sum function as the specular reflection point by constructing the range sum function of the signal propagation path between the satellite-borne receiver and the GNSS satellite, extracting the observation time stamps of the satellite-borne GNSS-R data and the SMAP external auxiliary data based on the year accumulation day as the time reference, and finally calculating the great circle distance between the specular reflection point and the grid center point of the SMAP external auxiliary data, setting a spatial distance threshold and completing the time-space matching according to the great circle distance, and the great circle distance calculation formula is: ; where, is the Earth's mean radius, is the latitude of the specular point, is the latitude of the SMAP grid center point, is the difference in longitude between the specular point and the SMAP grid center point.

4. The machine learning based spaceborne GNSS-R soil moisture fusion inversion method according to claim 1, characterized in that, The matched data set is used as input to construct an inversion model, including dividing the matched data set into a training set and a test set, performing Z-Score standardization processing on the training set, determining core features as model inputs based on the feature importance ranking function of the XGBoost library, taking the CART decision tree as the base learner, constructing the inversion model through forward step-by-step iteration, taking the mean square error as the basic loss function, introducing the L2 regularization term and the tree complexity regularization term to construct a comprehensive objective function, ; where N denotes the total number of training samples, denotes the mean square error loss function, denotes the real value of soil moisture of the i-th training sample, denotes the predicted value of soil moisture of the i-th training sample by the model, denotes the total number of base learners in the model, denotes the leaf node weight of the t-th CART base learner, denotes the L2 regularization parameter, denotes the tree complexity regularization parameter, denotes the total number of leaf nodes of the t-th CART base learner.

5. The machine learning based spaceborne GNSS-R soil moisture fusion inversion method according to claim 1, characterized in that, The optimization of the model parameters includes dividing the value range of each core parameter into a discrete grid, traversing all parameter combinations, training the inversion model and calculating the root mean square error of the test set, selecting the parameter combination with the minimum RMSE, then randomly sampling within the value range of each core parameter, taking the parameter combination corresponding to the minimum average value of the RMSE after multiple runs, constructing the mapping relationship between the parameters and the RMSE based on the Gaussian process model, taking the expected improvement EI function as the acquisition function, and iteratively searching for the optimal parameter combination multiple times, and the EI function calculation formula is: ; wherein, is the new parameter combination to be evaluated, is the current best parameter combination, is the root mean squared error of the model on the test set, is a non-negative truncation function, denotes the mathematical expectation.

6. The machine learning based spaceborne GNSS-R soil moisture fusion inversion method according to claim 1, characterized in that, The classification of the pre-processed GNSS-R internal eigenvalue data comprises constructing a hardware identification dimension based on inherent hardware parameters of a spaceborne GNSS-R constellation, setting a channel identification parameter to several discrete classification intervals, the hardware identification dimension comprising a channel identification parameter for distinguishing signal receiving channels and a satellite identification parameter for distinguishing satellite individuals, constructing a signal characteristic dimension based on physical properties of received signals, the signal characteristic dimension comprising a GNSS system type parameter for distinguishing signal source systems and a polarization mode parameter for distinguishing signal polarization states, constructing an associated mapping relationship between the hardware identification dimension parameter and the signal characteristic dimension parameter, and grouping all orbit observation point GNSS-R internal eigenvalue data corresponding to the same GNSS system type parameter and the same polarization mode parameter combination into a class according to the associated mapping relationship, to obtain a plurality of independent sub-data sets.

7. The machine learning based spaceborne GNSS-R soil moisture fusion inversion method according to claim 1, characterized in that, The inversion of each type of data by the inversion model comprises standardizing each sub-data set, calling an optimized inversion model framework, adaptively adjusting model parameters based on the sub-data set, obtaining an inversion sub-model, inputting orbit observation data not involved in the division in the sub-data set into the corresponding inversion sub-model, and calculating and outputting soil moisture inversion values of each orbit point by the inversion sub-model. The root mean square error, the mean absolute percentage error and the determination coefficient of the inversion sub-model are calculated based on the SMAP soil moisture true value to complete result verification. The soil moisture inversion value calculation formula is: ; wherein, is the soil moisture inversion value for the tth orbit point, K is the number of CART-based learners in the customized sub-model, is the prediction value of the kth CART-based learner for the tth orbit point feature vector . 8.A machine learning based space-borne GNSS-R soil moisture fusion inversion system, performing the method of claim 1, characterized in that, It comprises: A data acquisition module configured to acquire multi-source data composed of spaceborne GNSS-R internal eigenvalue data and SMAP external auxiliary data; A preprocessing module configured to preprocess the acquired multi-source data, including quality control of the spaceborne GNSS-R internal eigenvalue data, and spatio-temporal matching of the quality-controlled spaceborne GNSS-R internal eigenvalue data and the SMAP external auxiliary data to obtain a matched data set; A model module configured to construct an inversion model by taking the matched data set as input, including constructing an inversion model based on the XGBoost library, and optimizing model parameters by combining Bayesian algorithms, grid search and random search; An inversion module configured to perform orbit point soil moisture inversion by using the optimized inversion model, including classifying the pre-processed GNSS-R internal eigenvalue data and inverting each type of data by the inversion model; A planning module configured to perform grid point weight fusion based on the orbit point soil moisture inversion value, including dividing the orbit point soil moisture inversion value into grids according to regions, assigning the orbit point soil moisture inversion value to corresponding grids, and constructing a decision matrix, calculating the weight of each GNSS system by using an improved adaptive CRITIC weight method, and obtaining the soil moisture inversion value of the grid point; A verification module configured to perform result verification and application based on the soil moisture inversion value of the grid point.

Citation Information

Patent Citations

  • Satellite-borne GNSS-R soil humidity downscaling inversion method

    CN117405704A

  • GNSS-IR real-time soil humidity inversion method considering inter-system deviation correction

    CN118706865A