Remote sensing-observation relation machine learning space extrapolation method and device based on geographic feature constraint
Through the remote sensing-observation relationship machine learning spatial extrapolation method with geographical feature constraints, a linear mapping model is established using the gradient enhancement regression algorithm, which solves the problem of insufficient prediction accuracy of water quality models in areas without measured data, and realizes high-precision total nitrogen concentration prediction and water quality monitoring supplement.
Patent Information
- Application Number
- CN202510585309.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-01
AI Technical Summary
The remote sensing inversion data in the existing technology in areas without actual measured data cannot accurately reflect the water quality conditions. The traditional method is insufficient in spatial heterogeneous areas, resulting in low prediction accuracy of water quality model.
Through the remote sensing-observation relationship machine learning spatial extrapolation method based on geographical feature constraints, a linear mapping model is established using the gradient enhancement regression algorithm, and equivalent measured values are generated in combination with remote sensing data to predict the total nitrogen concentration in the sub-basin without measured results.
Accurate prediction in spatial heterogeneity areas is achieved, the accuracy and reliability of total nitrogen concentration simulation is improved, high-precision supplementary information for water quality monitoring is provided, and scientific pollution prevention and control strategies are supported.
Smart Images

Figure CN120409271A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hydrological monitoring, and particularly to a machine learning spatial extrapolation method and device for remote sensing-observation relationship based on geographical feature constraints. Background Art
[0002] Hydrological and water quality models can simulate the dynamic changes of runoff, erosion, sediment transport, and nutrients such as total nitrogen and total phosphorus, and play an important role in the fields of watershed environmental management, agricultural planning, and environmental protection.
[0003] However, in practical applications, there are often problems of uneven distribution of monitoring stations and insufficient measured data. Especially in some sub-watersheds, due to the difficulty of fully covering the monitoring network of the watershed, the measured data of key water quality indicators are seriously lacking. This phenomenon of insufficient data makes the traditional method of establishing the mapping relationship between remote sensing inversion data and observation data unable to accurately reflect the real situation of the unmonitored area, thus affecting the overall inversion accuracy and reliability of the model.
[0004] To make up for the deficiency of measured data, in recent years, by applying remote sensing inversion technology to water quality monitoring models, it is possible to provide water quality information such as the distribution information of total nitrogen concentration in a large range continuously to a certain extent by taking advantage of the wide spatial coverage and high acquisition frequency of remote sensing data. However, due to the limitation of the spatial resolution of remote sensing technology itself, it is difficult to capture local subtle changes, resulting in uncertainties in calibration and data consistency of remote sensing data. When remote sensing data is used as the only data source, it often fails to meet the requirements of high-precision simulation; traditional methods often use linear or nonlinear regression techniques to establish a mapping model between remote sensing data and measured data in the sub-watersheds with measured data. However, this method often has the limitation of low detection accuracy in the areas without measured data and is difficult to accurately deduce the "equivalent measured" data of these areas. In addition, geographical conditions between regions, such as differences in watershed area, river length, width, depth, and land use, all have an impact on the spatial distribution of water quality indicators, and traditional methods have not fully considered these factors, resulting in insufficient applicability and low prediction accuracy when the existing models are popularized and applied. Summary of the Invention
[0005] The object of the present invention is to provide a machine learning spatial extrapolation method and device for remote sensing-observation relationship based on geographical feature constraints. By establishing linear regression parameters between the total nitrogen concentration observation values and the remote sensing inversion values based on the sub-watersheds with measured data, and extracting the river morphology features, land morphology features and spatial proximity features of each sub-watershed, a linear mapping model between geographical features and regression parameters is established by using the gradient boosting regression algorithm to predict the equivalent regression parameters of the sub-watersheds without measured data; finally, equivalent measured values are generated by combining remote sensing data, which can break through the generalization bottleneck of traditional methods in spatially heterogeneous regions, achieve accurate prediction of regional regression parameters, and construct an error model with the "equivalent measured" data obtained by using the spatial extrapolation method and the measured data, providing a theoretical basis for subsequent data fusion, thereby improving the accuracy and reliability of the simulation of the total nitrogen concentration, a key water quality parameter.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] In the first aspect, a machine learning spatial extrapolation method and device for remote sensing-observation relationship based on geographical feature constraints are provided, which collect sub-watershed data with actual observation values of target observation factors and sub-watershed data with remote sensing inversion values of target observation factors from the remote sensing historical database;
[0008] The sub-watershed data with actual observation values of target observation factors is taken as the first target watershed, and the sub-watershed data with remote sensing inversion values of target observation factors is taken as the second target watershed;
[0009] Based on a distributed model with the watershed as the geographical unit, the geographical feature data of the first target watershed and the second target watershed are sequentially extracted, and the first geographical feature data and the second geographical feature data are correspondingly obtained;
[0010] A linear relationship model between the first target watershed and the second target watershed is established by using the spatial extrapolation method to obtain the slope β1 and intercept β0 of the linear relationship model;
[0011] Taking the slope β1 and intercept β0 as the target variables and the first geographical feature data or the second geographical feature data as the input variables, regression training is carried out by using the gradient boosting algorithm to construct a sub-watershed parameter prediction extrapolation model;
[0012] The extrapolation result is output by using the sub-watershed parameter prediction extrapolation model, and the prediction accuracy and robustness of the sub-watershed parameter prediction model are verified based on the extrapolation result.
[0013] As a further solution of the present invention: the area of the first target watershed is smaller than the area of the second target watershed.
[0014] As a further aspect of the present invention: both the first geographical feature data and the second geographical feature include river morphology features, land morphology features, and spatial proximity features; wherein,
[0015] The river morphology features include basin area, river length, average width, and average depth; the river morphology features are used to characterize the hydrological transmission capacity within the sub-basin;
[0016] The land morphology features include the proportion of cultivated land area; the land morphology features are used to reflect the contribution degree of agricultural activities to the total nitrogen concentration;
[0017] The spatial proximity features include the great circle distance value from the sub-basin data with observed values.
[0018] As a further aspect of the present invention: the distributed model with the basin as the geographical unit is the SWAT model.
[0019] As a further aspect of the present invention: the actual observed values with target observation factors are obtained from the actual observation database of the observation stations in the first target basin;
[0020] The remotely sensed inversion values with target observation factors are obtained from the remotely sensed inversion database of the second target basin. The remotely sensed inversion database of the second target basin is a visualization platform constructed by applying space. The visualization platform is a visualization platform constructed by applying the AI-Earth space. The visualization platform constructed by applying the AI-Earth space includes the total nitrogen database of the Pearl River Basin with a resolution of 5km×1d from 2015 to 2024. The average determination coefficient R of the overall inversion result of the remotely sensed data in the total nitrogen database of the Pearl River Basin 2 is not less than 0.79, which is conducive to obtaining rich and highly accurate remotely sensed inversion data.
[0021] As a further aspect of the present invention: the spatial extrapolation method is used to establish a linear relationship model between the first target basin and the second target basin, and the slope β1 and intercept β0 of the linear relationship model are obtained, including:
[0022] The formula:
[0023] U = f(Z) = β0 + β1·Z,
[0024] is used to establish a linear relationship model between the first target basin and the second target basin, where U is the remotely sensed inversion value with target observation factors, Z is the actual observed value of the target observation factor, and β0 and β1 are the regression coefficient intercept and slope respectively.
[0025] As a further solution of the present invention: using the slope β1 and the intercept β0 as the target variables, and the first geographical feature data or the second geographical feature data as the input variables, a regression training is performed using the gradient boosting algorithm to construct a sub-watershed parameter prediction extrapolation model, including:
[0026] Determine the target sub-watershed for extrapolation and judge the type of the target sub-watershed for extrapolation;
[0027] Among them, judging the type of the target sub-watershed for extrapolation includes:
[0028] When the target sub-watershed only has the remote sensing inversion value U of the target observation factor, the following formula is used:
[0029] And
[0030]
[0031] Perform spatial extrapolation. In the formula, β0 (i) and β1 (i) are respectively the intercept and slope of the target sub-watershed i for extrapolation; and are respectively the characteristic coefficients of the target sub-watershed i for extrapolation; Area i is the watershed area of the target sub-watershed i for extrapolation; Len i is the length of the river in the target sub-watershed i for extrapolation; Wid i is the average width of the river in the target sub-watershed i for extrapolation; Dep i is the average depth of the river in the target sub-watershed i for extrapolation; CropF i is the proportion of cultivated land area of the target sub-watershed i for extrapolation; S is the set of target sub-watersheds with actual observed values of the target observation factor; ω j and ω' j are distance weight coefficients, and their values reflect the influence degree of the j-th target sub-watershed on the parameter extrapolation of the i-th target sub-watershed; Dist i,j is the distance from the i-th target sub-watershed to the j-th target sub-watershed.
[0032] As a further solution of the present invention: verifying the prediction accuracy and robustness of the sub-watershed parameter prediction model based on the extrapolation results, including:
[0033] Compare the extrapolation results with the measured data to obtain the difference between the extrapolation results and the measured data. The measured data is actually observed from the observation stations of the target sub-watershed and is used to verify the accuracy of the extrapolation results of the target sub-watershed;
[0034] [[ID=5J]]Judge whether the difference is within the threshold range of the target difference. If the difference is within the threshold range of the target difference, it is judged that the sub-watershed parameter prediction model is qualified;
[0035] If the difference exceeds the threshold range of the target difference, it is determined that the sub-basin parameter prediction model is sub-qualified. By comparing the difference between the extrapolation result and the measured data, the accuracy and robustness of the sub-basin parameter prediction model can be judged, which is conducive to accurately evaluating the credibility of the prediction model.
[0036] As a further solution of the present invention: the target observation factor is the total nitrogen concentration.
[0037] In a second aspect, a device is provided, including a data acquisition module, a data extraction module, a spatial extrapolation module, a model construction module, and a model verification module. The data acquisition module is configured to collect sub-basin data with actual observation values of the target observation factor and sub-basin data with remotely sensed inversion values of the target observation factor from the remote sensing historical database;
[0038] The data extraction module is configured to sequentially extract the geographical feature data of the first target basin and the second target basin based on a distributed model with the basin as the geographical unit, and correspondingly obtain the first geographical feature data and the second geographical feature data;
[0039] The spatial extrapolation module is configured to establish a linear relationship model between the first target basin and the second target basin by using the spatial extrapolation method, and obtain the slope β1 and intercept β0 of the linear relationship model;
[0040] The model construction module is configured to use the slope β1 and intercept β0 as target variables, and the first geographical feature data or the second geographical feature data as input variables, and perform regression training by using the gradient boosting algorithm to construct a sub-basin parameter prediction extrapolation model;
[0041] The model verification module uses the sub-basin parameter prediction extrapolation model to output the extrapolation result, and verifies the prediction accuracy and robustness of the sub-basin parameter prediction model based on the extrapolation result.
[0042] Compared with the prior art, the beneficial effects of the present invention are:
[0043] 1. In the present invention, by establishing a linear regression parameter between the total nitrogen concentration observation value and the remotely sensed inversion value based on the sub-basin with measured data, and extracting the river morphology characteristics, land morphology characteristics, and spatial proximity characteristics of each sub-basin, a linear mapping model between the geographical characteristics and the regression parameter is established by using the gradient boosting regression algorithm to predict the equivalent regression parameter of the sub-basin without measured data; finally, combining the remote sensing data to generate an equivalent measured value, the generalization bottleneck of the traditional method in the spatially heterogeneous area can be broken through, the accurate prediction of the regional regression parameter is realized, and an error model is constructed with the "equivalent measured" data obtained by using the spatial extrapolation method and the measured data, providing a theoretical basis for subsequent data fusion, thereby improving the accuracy and reliability of the simulation of the total nitrogen concentration in the key water quality parameters.
[0044] 2. In the present invention, by constructing a linear relationship model between the first target basin and the second target basin, the mapping relationship between the observation data and the remote sensing inversion data in the sub-basin with observation data is extrapolated to the sub-basin without monitoring stations, thereby deriving "equivalent measured" data. Through this method, the existing measured total nitrogen concentration data and the remote sensing inversion total nitrogen concentration data obtained over a large area can be used in some sub-basins to establish a mapping relationship between the two. Further, combined with the river geometric characteristics and geographical characteristics of the sub-basin, a machine learning algorithm is used to perform regression fitting on the slope and intercept of the linear model, and spatial extrapolation of total nitrogen concentration data is realized for areas without monitoring stations.
[0045] 3. The present invention has obvious advantages in dealing with data uncertainty and spatial heterogeneity, and the verification results show that the model has an R 2 The values are all higher than 0.8, which has high application value and applicability, and can solve the problems of inaccurate mapping and large inversion errors caused by insufficient observation data coverage in existing technologies.
[0046] 4. The extrapolation method of the present invention can derive "equivalent measured" data in areas where monitoring stations are unevenly distributed and data is scarce within the basin, providing supplementary information for real-time monitoring of water quality in the basin. In areas that are not covered by monitoring stations, high-precision total nitrogen concentration prediction values can be obtained by utilizing remote sensing inversion data and geographic feature information, which helps water quality departments to fully understand the pollution situation in the basin, adjust management measures in a timely manner, and formulate scientific and reasonable pollution prevention and environmental protection strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flow chart of the method of the present invention;
[0048] Figure 2 This is a diagram showing the simulation results of the first sub-basin of the present invention;
[0049] Figure 3 This is a diagram showing the simulation results of the second sub-basin of the present invention;
[0050] Figure 4 This is a diagram showing simulation results of the third sub-basin of the present invention;
[0051] Figure 5 It is a module structure diagram of the present invention.
[0052] In the figure: 1. Data acquisition module; 2. Data extraction module; 3. Spatial extrapolation module; 4. Model construction module; 5. Model verification module. DETAILED DESCRIPTION
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] Embodiment:
[0055] Please refer to Figure 1 , the embodiment of the present invention provides a machine learning spatial extrapolation method for remote sensing-observation relationship based on geographical feature constraints, including the following steps:
[0056] S1: Collect sub-basin data with actual observation values of target observation factors and sub-basin data with remote sensing inversion values of target observation factors from the remote sensing historical database;
[0057] S2: Take the sub-basin data with actual observation values of target observation factors as the first target basin, and take the sub-basin data with remote sensing inversion values of target observation factors as the second target basin;
[0058] S3: Based on a distributed model with the basin as the geographical unit, sequentially extract the geographical feature data of the first target basin and the second target basin, and correspondingly obtain the first geographical feature data and the second geographical feature data;
[0059] S4: Use the spatial extrapolation method to establish a linear relationship model between the first target basin and the second target basin, and obtain the slope β1 and intercept β0 of the linear relationship model;
[0060] S5: Take the slope β1 and intercept β0 as the target variables, and the first geographical feature data or the second geographical feature data as the input variables, and use the gradient boosting algorithm for regression training to construct a sub-basin parameter prediction extrapolation model;
[0061] S6: Use the sub-basin parameter prediction extrapolation model to output the extrapolation result, and verify the prediction accuracy and robustness of the sub-basin parameter prediction model based on the extrapolation result.
[0062] Preferably, the area of the first target basin is smaller than the area of the second target basin.
[0063] Preferably, both the first geographical feature data and the second geographical feature include river morphological features, land morphological features, and spatial proximity features; among them,
[0064] The river morphological features include basin area, river length, average width, and average depth; the river morphological features are used to characterize the hydrological transmission capacity within the sub-basin;
[0065] The land morphological characteristics include the proportion of cultivated land area; the land morphological characteristics are used to reflect the contribution degree of agricultural activities to the total nitrogen concentration;
[0066] The spatial proximity characteristics include the great circle distance value from the sub-basin data with observed values.
[0067] Preferably, the distributed model with the basin as the geographical unit is the SWAT model.
[0068] Preferably, a decision tree learner is used to collect the sub-basin data with the actual observed values of the target observation factors and the sub-basin data with the remotely sensed inversion values of the target observation factors from the historical remote sensing database.
[0069] Preferably, the spatial extrapolation method is used to establish the linear relationship model between the first target basin and the second target basin, and the slope β1 and intercept β0 of the linear relationship model are obtained, including:
[0070] Using the formula:
[0071] U = f(Z) = β0 + β1·Z(1),
[0072] The linear relationship model between the first target basin and the second target basin is established, where U is the remotely sensed inversion value of the target observation factor, Z is the actual observed value of the target observation factor, and β0 and β1 are the regression coefficient intercept and slope respectively.
[0073] Preferably, with the slope β1 and intercept β0 as the target variables and the first geographical feature data or the second geographical feature data as the input variables, the gradient boosting algorithm is used for regression training to construct the sub-basin parameter prediction extrapolation model, including:
[0074] Determine the target sub-basin for extrapolation and judge the type of the target sub-basin for extrapolation;
[0075] Among them, judging the type of the target sub-basin for extrapolation includes:
[0076] When the target sub-basin only has the remotely sensed inversion value U of the target observation factor, the following formula is used:
[0077]
[0078] And
[0079]
[0080] Perform spatial extrapolation, where β0 (i) 、β1 (i) are the intercept and slope of the target sub-basin i for extrapolation respectively; and are the characteristic coefficients of the target sub-basin i for extrapolation respectively; Areai is the catchment area of the target sub-catchment i for extrapolation; Len i is the length of the river in the target sub-catchment i for extrapolation; Wid i is the average width of the river in the target sub-catchment i for extrapolation; Dep i is the average depth of the river in the target sub-catchment i for extrapolation; CropF i is the proportion of arable land area in the target sub-catchment i for extrapolation; S is the set of target sub-catchments for extrapolation with actual observed values of the target observation factor; ω j and ω' j are distance weight coefficients, and their values reflect the influence degree of the j-th target sub-catchment on the parameter extrapolation of the i-th target sub-catchment; Dist i,j is the distance from the i-th target sub-catchment to the j-th target sub-catchment, ∑ j∈sωj ·Dist i,j is the inverse distance weighting term; by using the inverse distance weighting term to assign higher weights to the measured data of adjacent target sub-catchments, the constraint of spatial autocorrelation on parameter extrapolation can be strengthened. For sub-catchments without measured data, input their river morphological characteristics into the trained sub-catchment parameter prediction and extrapolation model to predict the corresponding extrapolated regression parameters β0 (i) and β1 (i) , and then combine with the remote sensing inversion value U of the target observation factor in this area to invert the actual observed value Z of the target observation factor.
[0081] Preferably, based on the extrapolation results, verify the prediction accuracy and robustness of the sub-catchment parameter prediction model, including:
[0082] Compare the extrapolation result with the measured data to obtain the difference between the extrapolation result and the measured data. The measured data is actually observed from the observation stations of the target sub-catchment, and is used to verify the accuracy of the extrapolation result of the target sub-catchment;
[0083] Judge whether the difference is within the threshold range of the target difference. If the difference is within the threshold range of the target difference, then judge that the sub-catchment parameter prediction model is qualified;
[0084] If the difference exceeds the threshold range of the target difference, then judge that the sub-catchment parameter prediction model is sub-qualified. By comparing the difference between the extrapolation result and the measured data, the accuracy and robustness of the sub-catchment parameter prediction model can be judged, which is beneficial to accurately evaluate the credibility of the prediction model.
[0085] Preferably, based on the extrapolation results, verify the prediction accuracy and robustness of the sub-catchment parameter prediction model, and also include: comparing the prediction results on the validation set using the sub-catchment parameter prediction and extrapolation model;
[0086] Cross-validate the sub-basin samples with measured data and calculate the determination coefficient R of the predicted regression parameters 2 and the root mean square error. When the prediction accuracy of the sub-basin parameter prediction extrapolation model on the validation set is within the threshold range of the original accuracy, it indicates that the derived "equivalent measured" data is reliable in reflecting the regional total nitrogen concentration information.
[0087] Preferably, the target observation factor is the total nitrogen concentration.
[0088] The first target basin is the Dongjiang Basin. The sub-basins of the Dongjiang Basin are the Longchuan City Railway Bridge, Boluo City Lower, and Dadunzi Basins. The Longchuan City Railway Bridge is the first sub-basin, Boluo City Lower is the second sub-basin, and Dadunzi Basin is the third sub-basin;
[0089] The second target basin is the Pearl River Basin;
[0090] Collect the measured daily-scale total nitrogen concentration data for two years from 2021 to 2022 at multiple monitoring stations in the Dongjiang Basin; and obtain the remotely sensed inversion total nitrogen concentration data of the Pearl River Basin on the visualization platform constructed by the AI-Earth application space. The resolution of the remotely sensed inversion total nitrogen concentration data is 5km×1d;
[0091] The visualization platform is the visualization platform constructed by the AI-Earth application space. The visualization platform constructed by the AI-Earth application space contains the total nitrogen database of the Pearl River Basin with a resolution of 5km×1d from 2015 to 2024. The average determination coefficient R of the overall inversion results of the remotely sensed data in the total nitrogen database of the Pearl River Basin 2 is not less than 0.79, which is conducive to obtaining rich and high-precision remotely sensed inversion data;
[0092] Select the remotely sensed inversion data point closest to the Pearl River Basin station as the remotely sensed data of the Pearl River Basin; use the established SWAT model to extract river morphological characteristics, land morphological characteristics, and spatial proximity characteristics:
[0093] Obtain the areas of the Longchuan City Railway Bridge, Boluo City Lower, and Dadunzi Basins, as well as the total lengths, average widths, average depths, proportion of cultivated land area, and great circle distance between sub-basins of the rivers within the Longchuan City Railway Bridge, Boluo City Lower, and Dadunzi Basins through the output of GIS data and the SWAT model.
[0094] Among them, the proportion of cultivated land area is the proportion of cultivated land occupied by agricultural land within the sub-basin;
[0095] Adopt the great circle distance calculation formula to obtain the great circle distance between sub-basins through the distance between the data points of the Longchuan City Railway Bridge, Boluo City Lower, and Dadunzi Basins and the data points of the nearest sub-basin with measured data;
[0096] As described above, the formula (1) is constructed by using the least squares method for fitting to obtain the regression coefficients β0 and β1 corresponding to the Longchuan City Railway Bridge, under Boluo City, and the Dadunzi Basin. The extracted features are modeled through formulas (2) and (3) to establish the relationships between β0, β1 and the river length, average width, average depth, proportion of cultivated land area, and great circle distance between sub-basins. and are the characteristic coefficients of sub-basin i, which can be obtained by fitting historical data; S is the set of partially measured sub-basins; ω j and ω' j are the distance weight coefficients, and their values reflect the influence degree of the j-th sub-basin on the extrapolation of the parameters of the i-th sub-basin. ∑ j∈S ω j ·Dist i,j The overall part can be understood as the inverse distance weighting from the sub-basin to be predicted to the measured sub-basins. This formula is a general formula, and according to the actual situation, the coefficients in front of individual features can be 0.
[0097] The sub-basin samples with measured data are divided into a training set and a validation set according to a ratio of 7:3.
[0098] The Bayesian optimization algorithm is used to determine the optimal hyperparameters of the gradient boosting model, and the model is trained on the training set to obtain the prediction model.
[0099] The coefficient of determination R 2 and the root mean square error are calculated on the validation set. For the sub-basins without measured data, their geographical feature data are input into the trained gradient boosting model to predict the corresponding extrapolated regression parameters β0 (i) and β1 (i) , and the spatial extrapolation prediction result table is obtained as shown in Table 1 below.
[0100]
[0101] The sub-basins with measured data are used as the test set to perform cross-validation on the extrapolation model, calculate the prediction errors of the regression parameters of each sub-basin, and statistically analyze the R 2 and root mean square error indicators on the validation set; compare the performance of the traditional linear regression and the extrapolation method of the present invention on the same data set to verify the prediction accuracy of this method, so as to apply it to the sub-basins without measured data.
[0102] It can be seen that the coefficients of determination R 2 of the Longchuan City Railway Bridge, under Boluo City, and the Dadunzi Basin are all greater than 0.8.
[0103] As Figure 4As shown, a device is provided, including a data acquisition module 1, a data extraction module 2, a spatial extrapolation module 3, a model construction module 4, and a model verification module 5. The data acquisition module 1 is configured to collect sub-basin data with actual observed values of target observation factors and sub-basin data with remotely sensed inversion values of target observation factors from a remote sensing historical database;
[0104] The data extraction module 2 is configured to sequentially extract geographical feature data of a first target basin and a second target basin based on a distributed model with the basin as the geographical unit, and correspondingly obtain first geographical feature data and second geographical feature data;
[0105] The spatial extrapolation module 3 is configured to establish a linear relationship model between the first target basin and the second target basin using the spatial extrapolation method, and obtain the slope β1 and intercept β0 of the linear relationship model;
[0106] The model construction module 4 is configured to use the slope β1 and intercept β0 as target variables, and the first geographical feature data or the second geographical feature data as input variables, and perform regression training using the gradient boosting algorithm to construct a sub-basin parameter prediction extrapolation model;
[0107] The model verification module 5 outputs an extrapolation result using the sub-basin parameter prediction extrapolation model, and verifies the prediction accuracy and robustness of the sub-basin parameter prediction model based on the extrapolation result.
[0108] The present invention realizes the global extrapolation of the local observation-remote sensing relationship by constructing a feature-driven model with spatial correlation, and establishes a linear regression parameter between the total nitrogen concentration observed value and the remotely sensed inversion value based on the sub-basin with measured data; extracts the river morphology characteristics, land morphology characteristics, and spatial proximity characteristics of each sub-basin in the hydrological soil evaluation model; establishes a linear mapping model between geographical features and regression parameters by using the gradient boosting regression algorithm to predict the equivalent regression parameters of the sub-basin without measured data; finally generates equivalent measured values in combination with remote sensing data, and is verified by the Dongjiang River Basin. The extrapolation parameter prediction R 2 reaches more than 0.8, breaking through the generalization bottleneck of traditional methods in spatially heterogeneous regions, providing an interpretable machine learning solution for parameterization of unobserved areas in hydrological models, and can be extended and applied to the fields of watershed non-point source pollution assessment and water quality supervision.
[0109] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A machine learning spatial extrapolation method for remote sensing-observation relationship based on geographical feature constraints, characterized in that Including: Collecting sub - basin data with actual observed values of target observation factors and sub - basin data with remotely sensed inversion values of target observation factors from a remote sensing historical database; Taking the sub - basin data with actual observed values of target observation factors as the first target basin, and taking the sub - basin data with remotely sensed inversion values of target observation factors as the second target basin; Based on a distributed model with the basin as the geographical unit, successively extracting the geographical feature data of the first target basin and the second target basin, and correspondingly obtaining the first geographical feature data and the second geographical feature data; Using spatial extrapolation to establish a linear relationship model between the first target basin and the second target basin, and obtaining the slope β1 and intercept β0 of the linear relationship model; Using the slope β1 and intercept β0 as target variables, and the first geographical feature data or the second geographical feature data as input variables, and performing regression training using the gradient boosting algorithm to construct a sub - basin parameter prediction extrapolation model; Using the sub - basin parameter prediction extrapolation model to output the extrapolation result, and verifying the prediction accuracy and robustness of the sub - basin parameter prediction model based on the extrapolation result.
2. A machine learning spatial extrapolation method for remote sensing-observation relationship based on geographical feature constraints according to claim 1, characterized in that: The area of the first target basin is smaller than the area of the second target basin.
3. A machine learning spatial extrapolation method for remote sensing-observation relationship based on geographical feature constraints according to claim 1, characterized in that: Both the first geographical feature data and the second geographical feature data include river morphological features, land morphological features, and spatial proximity features; among them, River morphological features include basin area, river length, average width, and average depth; Land morphological features include the proportion of cultivated land area; Spatial proximity features include the great - circle distance value from the sub - basin data with observed values.
4. A machine learning spatial extrapolation method for remote sensing-observation relationship based on geographical feature constraints according to claim 1, characterized in that: The distributed model with the basin as the geographical unit is the SWAT model.
5. A machine learning spatial extrapolation method for remote sensing-observation relationship based on geographical feature constraints according to claim 1, characterized in that: The actual observed values of target observation factors are obtained from the actual observation database of the observation stations in the first target basin; The remotely sensed inversion values of target observation factors are obtained from the remotely sensed inversion database of the second target basin, and the remotely sensed inversion database of the second target basin is a visualization platform constructed using the application space.
6. A machine learning spatial extrapolation method for remote sensing-observation relationship based on geographical feature constraints according to claim 1, characterized in that: Using spatial extrapolation to establish a linear relationship model between the first target basin and the second target basin, and obtaining the slope β1 and intercept β0 of the linear relationship model, including: Using the formula: U = f(Z)=β0 + β1·Z, To establish a linear relationship model between the first target basin and the second target basin, where U is the remotely sensed inversion value of the target observation factor, Z is the actual observed value of the target observation factor, and β0 and β1 are the regression coefficient intercept and slope respectively.
7. A machine learning spatial extrapolation method for remote sensing-observation relationship based on geographical feature constraints according to claim 6, characterized in that: Using the slope β1 and intercept β0 as target variables, and the first geographical feature data or the second geographical feature data as input variables, and performing regression training using the gradient boosting algorithm to construct a sub - basin parameter prediction extrapolation model, including: Determining the target sub - basin for extrapolation and judging the type of the target sub - basin for extrapolation; Among them, judging the type of the target sub - basin for extrapolation includes: When the target sub - basin only has the remotely sensed inversion value U of the target observation factor, using the following formula: and Perform spatial extrapolation, where β0 (i) and β1 (i) are the intercept and slope of the target sub - watershed i for extrapolation, respectively; and are the characteristic coefficients of the target sub - watershed i for extrapolation; Area i is the basin area of the target sub - watershed i for extrapolation; Len i is the length of the river in the target sub - watershed i for extrapolation; Wid i is the average width of the river in the target sub - watershed i for extrapolation; Dep i is the average depth of the river in the target sub - watershed i for extrapolation; CropF i is the proportion of cultivated land area of the target sub - watershed i for extrapolation; S is the set of target sub - watersheds for extrapolation of the actual observed values of the target observation factors; ω j and ω' j are the distance weight coefficients, and their values reflect the influence degree of the j - th target sub - watershed on the parameter extrapolation of the i - th target sub - watershed; Dist i,j is the distance from the i - th target sub - watershed to the j - th target sub - watershed.
8. A machine learning spatial extrapolation method for remote sensing-observation relationship based on geographical feature constraints according to claim 1, characterized in that: Verifying the prediction accuracy and robustness of the sub - basin parameter prediction model based on the extrapolation result, including: Comparing the extrapolation result with the measured data to obtain the difference between the extrapolation result and the measured data; Determine whether the difference is within the threshold range of the target difference. If the difference is within the threshold range of the target difference, determine that the sub-basin parameter prediction model is qualified; If the difference exceeds the threshold range of the target difference, determine that the sub-basin parameter prediction model is sub-qualified.
9. A machine learning spatial extrapolation method for remote sensing-observation relationship based on geographical feature constraints according to claim 1, characterized in that: The target observation factor is the total nitrogen concentration.
10. A device, comprising: A data acquisition module configured to acquire sub-basin data with actual observation values of a target observation factor and sub-basin data with remotely sensed inversion values of the target observation factor from a remote sensing historical database; A data extraction module configured to sequentially extract the geographical feature data of a first target basin and a second target basin based on a distributed model with the basin as the geographical unit, and correspondingly obtain first geographical feature data and second geographical feature data; A spatial extrapolation module configured to establish a linear relationship model between the first target basin and the second target basin by using the spatial extrapolation method, and obtain the slope β1 and intercept β0 of the linear relationship model; A model construction module configured to use the slope β1 and intercept β0 as target variables and the first geographical feature data or the second geographical feature data as input variables, and perform regression training by using the gradient boosting algorithm to construct a sub-basin parameter prediction extrapolation model; A model verification module configured to output an extrapolation result by using the sub-basin parameter prediction extrapolation model, and verify the prediction accuracy and robustness of the sub-basin parameter prediction model based on the extrapolation result.