InVEST area water yield multi-source data fusion method and system
By performing quality assessment and processing on multi-source data, high-precision fused data is generated, which solves the problem of inconsistent data quality in the water production assessment of the InVEST area and improves the accuracy and reliability of water production assessment.
Patent Information
- Application Number
- CN202511688496.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-17
AI Technical Summary
Existing technologies lack a systematic quality assessment mechanism for multi-source data in InVEST regional water yield assessment, resulting in inconsistent data quality. This affects the reliability of hydrological data and soil type composite data, as well as the accuracy of water yield assessment. In particular, there are problems of poor spatial continuity and insufficient accuracy in heterogeneous data assimilation and spatial overlay analysis.
By conducting quality assessments on multi-source data, data quality levels are generated. Based on these levels, heterogeneous data assimilation, spatial interpolation processing, soil and land use composite data generation, and eco-hydrological data aggregation are performed. Finally, high-precision fused data is generated and input into the InVEST model for evaluation.
It improved the accuracy and spatial continuity of hydrological data, enhanced the reliability of soil type composite data, achieved high-precision regional water yield assessment, and ensured the relevance and effectiveness of data processing.
Smart Images

Figure CN121542998A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for fusing multi-source water production data in the InVEST region. Background Technology
[0002] In the InVEST regional water yield assessment process, multi-source data is the core input foundation, but existing technologies lack a systematic quality assessment mechanism for multi-source data. The failure to comprehensively verify the characteristics of different data types results in data of varying quality being directly fed into subsequent processing stages. This leads to deviations in results from heterogeneous data assimilation and spatial overlay analysis, severely impacting the reliability of composite hydrological and soil type data and posing a threat to the accuracy of water yield assessment.
[0003] Existing technologies have significant shortcomings in the multi-source data fusion process: On the one hand, when assimilating heterogeneous precipitation and evapotranspiration data, reasonable weights are not assigned based on data quality levels, and effective spatial consistency optimization methods are lacking, resulting in poor spatial continuity and insufficient accuracy of the fused hydrological data; on the other hand, in the integration of soil land data and hydrological grid data, spatial registration accuracy is low, spatial-spectral information is not closely correlated, and scientific standardization transformation is not performed on eco-hydrological data, resulting in inconsistent dimensions and mismatched spatial scales in the final input data to the InVEST model, making it difficult to support high-precision regional water yield assessment. Therefore, how to improve the effectiveness and accuracy of multi-source data fusion has become an urgent problem to be solved in the field of InVEST water yield assessment. Summary of the Invention
[0004] This invention provides a method and system for fusing multi-source water production data in the InVEST region to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, this invention provides a method for fusing multi-source water production data in an InVEST region, comprising: S1. Perform quality assessment on the multi-source data of the target area to obtain the data quality level of the multi-source data; S2. Based on the data quality level, perform heterogeneous data assimilation on precipitation data and evapotranspiration data in the multi-source data to obtain hydrological data of the target area; S3. Perform spatial interpolation on the hydrological data to obtain the hydrological grid data of the hydrological data; S4. Based on the data quality level, perform spatial overlay analysis on the land use data and soil data in the multi-source data to obtain the composite soil land use data of the target area; S5. Aggregate the spatial spectral information of the soil type composite data and the hydrological grid data to obtain the eco-hydrological data of the target area; S6. Based on the data quality level, the eco-hydrological data is standardized to obtain high-precision fused data of the target area, and the high-precision fused data is input into the InVEST model to complete the assessment of the water yield of the target area.
[0006] In a preferred embodiment, the step of performing quality assessment on multi-source data of the target area to obtain the data quality level of the multi-source data includes: Obtain the data source description of precipitation data from multi-source data in the target area, and perform precipitation assessment on the precipitation data to obtain the precipitation index of the precipitation data; Spatiotemporal alignment is performed on the evapotranspiration data in the multi-source data to obtain a consistency identifier for the evapotranspiration data; The land use data in the multi-source data is reviewed for land category coding to obtain the coding standard compliance of the land use data. The soil data in the multi-source data is subjected to attribute completeness checks to obtain the attribute evaluation results of the soil data; Based on the data source description, the precipitation index, the consistency identifier, the coding standard compliance, and the attribute evaluation results, the data quality level of the multi-source data is generated.
[0007] In a preferred embodiment, the step of heterogeneous data assimilation of precipitation and evapotranspiration data from the multi-source data based on the data quality level to obtain hydrological data for the target area includes: Based on the data quality level, the quality weight ratio of the precipitation data and the evapotranspiration data is determined; The precipitation data and the evapotranspiration data are gridded to unify them into the same spatial grid. Within the spatial grid, based on the mass weight ratio, the gridded precipitation data and the evapotranspiration data are weighted and fused to generate initial hydrological parameters for the precipitation data and the evapotranspiration data. Spatial consistency optimization is performed on the initial hydrological parameters to generate hydrological data for the target area.
[0008] In a preferred embodiment, the step of performing spatial interpolation processing on the hydrological data to obtain hydrological grid data includes: Based on the spatial distribution characteristics of the hydrological data, a distribution framework for the hydrological data is constructed. Within the distribution framework, interpolation parameters are determined based on the numerical gradient changes of the hydrological data; Based on the interpolation parameters, the missing areas in the hydrological data are spatially filled to generate the initial interpolation of the hydrological data; The initial interpolation is smoothed at the boundary to obtain the hydrological grid data of the hydrological data.
[0009] In a preferred embodiment, the step of performing spatial overlay analysis on land use data and soil data from the multi-source data based on the data quality level to obtain composite soil land use data for the target area includes: Based on the data quality level, assign superimposed weights to the land use data and the soil data; Spatial registration is performed on the land use data and the soil data respectively to obtain the registered land use spatial data and the registered soil spatial data. Spatial integration of the land use spatial data and the soil spatial data yields spatial framework data for the land use data and the soil data. Based on the superposition weight, the land use type code and soil type code in the spatial frame data are weighted and combined to obtain the composite land category code of the spatial frame data. The composite land cover code is updated to the spatial framework data to generate composite soil land cover data for the target area.
[0010] In a preferred embodiment, the spatial integration of the land use spatial data and the soil spatial data to obtain spatial framework data for the land use data and the soil data includes: Establish a spatial index between the land use spatial data and the soil spatial data; Based on the spatial index, the geometric boundaries of the land use spatial data and the soil spatial data are merged, and the attribute structures of the land use spatial data and the soil spatial data are constructed simultaneously. Based on the geometric boundaries and the attribute structure, spatial framework data for the land use data and the soil data are generated.
[0011] In a preferred embodiment, the step of aggregating the soil type composite data and the hydrological grid data using spatial spectral information to obtain the eco-hydrological data of the target area includes: Based on the spatial reference of the soil land use composite data, the hydrological grid data is spatially resampled to obtain hydrological grid data with the same spatial resolution as the soil land use composite data. Spatial registration is performed on the resampled hydrological grid data and the soil land use composite data to obtain the correspondence between the spatial locations of the hydrological grid data and the soil land use composite data; Based on the correspondence, the soil land use composite attributes and hydrological grid attributes of the spatial location are associated and integrated to generate a composite parameter set of the hydrological grid data and the soil land use composite data; Multi-source data normalization is performed on the composite parameter set to generate eco-hydrological data for the target area.
[0012] In a preferred embodiment, the step of associating and integrating the soil land use composite attributes and hydrological grid attributes based on the correspondence to generate a composite parameter set of the hydrological grid data and the soil land use composite data includes: Based on the correspondence, feature pairs of soil land use composite attributes and hydrological grid attributes in the spatial location are obtained; Calculate the fusion parameters of the feature pairs, wherein the formula for calculating the fusion parameters is as follows: ; In the formula, For the first Fusion parameters for each spatial location, For the first Numerical representation of the composite attributes of soil land use at a spatial location For the first Numerical representation of hydrological grid attributes for each spatial location The preset weighting coefficients, The preset weighting coefficients, The preset weighting coefficients, To obtain the maximum value, To obtain the minimum value; Spatial consistency verification is performed on the fusion parameters, and a composite parameter set of the hydrological grid data and the soil land use composite data is generated based on the verification results.
[0013] In a preferred embodiment, the step of standardizing the eco-hydrological data based on the data quality level to obtain high-precision fused data of the target area, and inputting the high-precision fused data into the InVEST model to complete the assessment of water yield in the target area, includes: Based on the data quality level, the transformation weights of the parameters in the eco-hydrological data are determined; Based on the transformation weights, the soil parameters, vegetation parameters and hydrological parameters in the eco-hydrological data are standardized to obtain regularized eco-hydrological data. Spatial scale adaptation is performed on the regularized data to obtain high-precision fused data of the target region; The high-precision fused data is input into the InVEST model to complete the water production assessment results for the target area.
[0014] To address the aforementioned problems, this invention also provides an InVEST regional permeability multi-source data fusion system, the system comprising: The data quality assessment module is used to assess the quality of multi-source data in the target area and obtain the data quality level of the multi-source data. The heterogeneous data assimilation module is used to perform heterogeneous data assimilation on precipitation data and evapotranspiration data in the multi-source data based on the data quality level, so as to obtain the hydrological data of the target area. The spatial interpolation processing module is used to perform spatial interpolation processing on the hydrological data to obtain the hydrological grid data of the hydrological data; The spatial overlay analysis module is used to perform spatial overlay analysis on land use data and soil data in the multi-source data based on the data quality level, so as to obtain composite soil land use data of the target area; The spatial spectral information aggregation module is used to aggregate the spatial spectral information of the soil land type composite data and the hydrological grid data to obtain the eco-hydrological data of the target area. The normalization conversion module is used to normalize the eco-hydrological data based on the data quality level to obtain high-precision fused data of the target area, and input the high-precision fused data into the InVEST model to complete the assessment of the water yield of the target area.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention, by conducting quality assessments on multi-source data from a target area, can accurately control the quality status of precipitation, evapotranspiration, land use, and soil data, providing a clear basis for subsequent data processing. Assimilation of heterogeneous data based on this quality level ensures the rationality of precipitation and evapotranspiration data fusion. Combined with spatial interpolation processing to fill in missing areas and smooth boundaries in hydrological data, it significantly improves the accuracy and spatial continuity of hydrological data. Furthermore, when performing spatial overlay analysis on land use and soil data, incorporating quality level-based weighting enhances the reliability of composite soil and land use data, laying a high-quality data foundation for subsequent fusion processes.
[0016] 2. This invention, through spatial spectral information aggregation, enables precise correlation and integration of soil land use composite data and hydrological grid data. The resulting eco-hydrological data more accurately reflects regional ecological and hydrological characteristics. Furthermore, by standardizing the eco-hydrological data according to its quality level, the data dimensions are unified and spatial scales are adapted, resulting in high-precision fused data. Inputting this data into the InVEST model effectively improves the accuracy of water yield assessment in the target area. Simultaneously, the entire data fusion process proceeds systematically around the data quality level, ensuring the relevance and effectiveness of each processing step, further enhancing the overall efficiency and quality of InVEST regional water yield multi-source data fusion. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a method for fusing multi-source water production data in an InVEST region according to an embodiment of the present invention. Figure 2 This is a functional block diagram of an InVEST regional water production multi-source data fusion system provided in an embodiment of the present invention; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] This application provides a method for fusing multi-source water production data in an InVEST region. The executing entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0020] Reference Figure 1 The diagram shown is a flowchart illustrating a method for fusing multi-source water production data in an InVEST region according to an embodiment of the present invention. In this embodiment, the method includes: S1. Perform quality assessment on the multi-source data of the target area to obtain the data quality level of the multi-source data; In this embodiment of the invention, the step of performing quality assessment on multi-source data of the target area to obtain the data quality level of the multi-source data includes: Obtain the data source description of precipitation data from multi-source data in the target area, and perform precipitation assessment on the precipitation data to obtain the precipitation index of the precipitation data; Spatiotemporal alignment is performed on the evapotranspiration data in the multi-source data to obtain a consistency identifier for the evapotranspiration data; The land use data in the multi-source data is reviewed for land category coding to obtain the coding standard compliance of the land use data. The soil data in the multi-source data is subjected to attribute completeness checks to obtain the attribute evaluation results of the soil data; Based on the data source description, the precipitation index, the consistency identifier, the coding standard compliance, and the attribute evaluation results, the data quality level of the multi-source data is generated.
[0021] When obtaining the data source description of precipitation data from multi-source data in a target area, it is necessary to obtain metadata documents or data description files from the precipitation data provider. From these documents, information such as the data source type, data collection time range, collection frequency, observation instrument model, or remote sensing sensor type should be extracted to form a complete data source description. When assessing precipitation data, it is necessary to compare it with the historical average precipitation data for the same period in the region to check whether there are extreme outliers in the precipitation data that exceed the climate characteristics of the region. At the same time, the percentage of missing days within the collection time range should be statistically analyzed. Based on the percentage of outliers and the data completeness, a precipitation index should be determined. This index is characterized by a grade of excellent, good, medium, or poor or a specific numerical range. A data completeness of 95% or higher is considered excellent. Finally, both the data source description and the precipitation index of the precipitation data are obtained.
[0022] When performing spatiotemporal alignment of evapotranspiration data from multiple sources, firstly, in the time dimension, if evapotranspiration data from different sources have different time resolutions (some are hourly data and some are daily data), all evapotranspiration data are uniformly converted to the same time unit, becoming daily-scale data. For missing data in the time dimension, evapotranspiration data from adjacent time points are linearly extended and supplemented. In the spatial dimension, using a pre-defined standard spatial coordinate system for the target area as a reference, coordinate transformation tools are used to transform the spatial coordinates of all evapotranspiration data from different sources to this standard coordinate system, ensuring that all evapotranspiration data correspond precisely in spatial location. After spatiotemporal alignment, the numerical differences of evapotranspiration data from different sources at the same spatial location and time point are compared. If the difference is within the normal fluctuation range of evapotranspiration in the region, it is marked as consistent; if the difference exceeds this range, it is marked as inconsistent, thus obtaining a consistency identifier for the evapotranspiration data.
[0023] When reviewing land use data from multi-source data for land category coding, first determine the national or industry coding standards that the land use data should follow, clarify the scope of land category coding stipulated in the standards and the land category name corresponding to each code; then check the coding of each land category patch in the land use data one by one, check whether the code is within the valid range stipulated in the standards, whether the code and the land category name labeled on the patch are completely matched, and whether there are any duplicate or missing coding labels; count the proportion of land category patches that conform to the coding standards to the total number of land category patches in the land use data, and this proportion is the coding standard compliance degree of the land use data. When checking the completeness of soil data from multi-source datasets, the core attribute items that the soil data must include are first identified. These attribute items are key information supporting the subsequent generation of composite soil land use data. Then, the attribute data corresponding to each soil sampling point or soil patch is checked one by one, and the presence of missing data for each core attribute item is recorded. The proportion of soil samples with no missing data for each core attribute item is counted to the total number of soil samples. Finally, the overall completeness of the soil data is evaluated by combining the completeness of all core attribute items. If the completeness of all core attribute items exceeds 90%, it is evaluated as high completeness; if the completeness of some core attribute items is below 60%, it is evaluated as low completeness. This evaluation result is the attribute assessment result of the soil data.
[0024] When generating data quality levels for multi-source data based on data source descriptions, precipitation indicators, consistency identifiers, coding standard compliance, and attribute assessment results, quantitative scoring standards are first set for each of these five information items. For the data source description, 100 points are awarded if it includes complete data source type, collection time range, collection frequency, and instrument or sensor information; 25 points are deducted for each missing key piece of information. For precipitation indicators, 100 points are awarded for excellent, 80 points for good, 60 points for average, and 40 points for poor. In the consistency identifier, 100 points are awarded for consistent spatial locations and time points exceeding 90%, and 80%-89% for... 80 points, 70%-79% get 60 points, and below 70% get 40 points; coding standard compliance is converted into points according to the actual proportion, with 95% compliance getting 95 points and 80% compliance getting 80 points; attribute evaluation results are 100 points for high completeness, 70 points for medium completeness, and 40 points for low completeness; then the scores of the five information items are calculated into a total score according to each accounting for 20%, and the quality level is divided according to the total score range: 90-100 points is level one, 80-89 points is level two, 70-79 points is level three, and below 70 points is level four, finally generating the data quality level of multi-source data.
[0025] The beneficial effects are that by conducting targeted quality checks on precipitation, evapotranspiration, land use, and soil data from multiple sources in the target area, specific quality information for each type of data can be obtained comprehensively and accurately. The generated data source descriptions, precipitation indicators, consistency identifiers, coding standard compliance, and attribute evaluation results can clearly reflect the quality details of each data. Based on this information, the multi-source data quality level generated can provide a clear quality basis for subsequent data processing steps such as heterogeneous data assimilation and spatial overlay analysis, avoiding subsequent processing deviations caused by unclear data quality. This ensures the reliability of basic data in the process of fusion of multi-source water yield data in the InVEST area from the source, laying a solid foundation for the subsequent generation of high-precision fused data.
[0026] S2. Based on the data quality level, perform heterogeneous data assimilation on precipitation data and evapotranspiration data in the multi-source data to obtain hydrological data of the target area; In this embodiment of the invention, the step of assimilating heterogeneous data of precipitation and evapotranspiration data from the multi-source data based on the data quality level to obtain hydrological data of the target area includes: Based on the data quality level, the quality weight ratio of the precipitation data and the evapotranspiration data is determined; The precipitation data and the evapotranspiration data are gridded to unify them into the same spatial grid. Within the spatial grid, based on the mass weight ratio, the gridded precipitation data and the evapotranspiration data are weighted and fused to generate initial hydrological parameters for the precipitation data and the evapotranspiration data. Spatial consistency optimization is performed on the initial hydrological parameters to generate hydrological data for the target area.
[0027] When determining the quality weight ratio of precipitation and evapotranspiration data based on data quality levels, the previously established rules for corresponding quality levels and weight coefficients are strictly followed: Level 1 corresponds to a weight coefficient of 1.0, Level 2 to 0.8, Level 3 to 0.6, and Level 4 to 0.4, ensuring consistency in weight allocation standards. Complete quality level information for precipitation and evapotranspiration data is extracted from the previously generated multi-source data quality level reports. This level information has been determined through multi-dimensional evaluation of data source descriptions, precipitation indicators, and consistency markers, providing a clear basis for reliability. Based on the extracted quality levels, the corresponding weight coefficients are precisely matched, clearly distinguishing between the weight coefficients for precipitation and evapotranspiration data. Using the weight coefficient for precipitation as the numerator and the weight coefficient for evapotranspiration as the denominator, a standard division operation yields the specific numerical result, which is the quality weight ratio of precipitation and evapotranspiration data, providing a quantitative basis for subsequent weighted fusion.
[0028] When gridding precipitation and evapotranspiration data, the complete spatial extent of the target area is first comprehensively defined, including the region's latitude and longitude boundaries, administrative boundaries, or natural geographical boundaries, ensuring complete grid coverage. Considering the topographical complexity of the target area, the original data resolution, and the spatial scale requirements for subsequent hydrological data applications, commonly used grid scale standards in the field of hydrological assessment for this region are investigated, and a unified and fixed spatial grid size is set to ensure both data accuracy and computational efficiency. Using the established spatial grid as a standardized benchmark framework, precipitation and evapotranspiration data are projected one by one into the corresponding grid cells according to their original spatial coordinates, ensuring that each original data point is accurately matched to its assigned grid cell. For each grid cell, the specific values of all original data points within its coverage area are statistically analyzed. By calculating the arithmetic mean of these values, a representative value for that grid cell is obtained, and this mean is used as the final value record for that grid cell. Through the above standardization process, precipitation and evapotranspiration data from different sources and in different formats are ultimately unified into the same spatial grid, resulting in gridded precipitation and evapotranspiration data with a regular structure and consistent scale.
[0029] When performing weighted fusion of gridded precipitation and evapotranspiration data based on mass weight ratios within a spatial grid, precise calculations are performed for each independent spatial grid cell. First, the gridded precipitation and evapotranspiration data values corresponding to that grid cell are quickly retrieved using a data index to ensure accurate data extraction. Following the numerator and denominator allocation rules of the mass weight ratio, the gridded precipitation data value is multiplied by the numerator of the mass weight ratio to obtain the weighted contribution value of precipitation data in that grid cell; the gridded evapotranspiration data value is multiplied by the denominator of the mass weight ratio to obtain the weighted contribution value of evapotranspiration data in that grid cell. The sum of these two weighted contribution values is calculated, and then divided by the sum of the numerator and denominator of the mass weight ratio. This weighted average calculation yields a result that comprehensively reflects the contributions of both types of data; this result is the initial hydrological parameter corresponding to that spatial grid cell. Following this unified process, all spatial grid cells within the target area are traversed sequentially, and the weighted fusion calculation for each cell is completed, ultimately generating a dataset covering all grid cells in the target area and containing complete initial hydrological parameters.
[0030] When optimizing the spatial consistency of initial hydrological parameters, the spatial fluctuation patterns of historical hydrological parameters in the target area are first analyzed based on long-term hydrological observation data, climate characteristics, and topographic conditions. A reasonable spatial fluctuation range for the initial hydrological parameters is then determined, which must cover the parameter variation range under normal hydrological scenarios while excluding interference from extreme outliers. The initial hydrological parameters of each spatial grid cell are then numerically compared with those of the eight adjacent grid cells. The numerical difference between the grid cell and each adjacent cell is calculated, and the distribution of all differences is statistically analyzed. If the initial hydrological parameter of a certain grid cell differs from the numerical difference of most of its adjacent grid cells beyond the preset reasonable fluctuation range, it indicates a spatial anomaly that may affect the continuity of the overall data. Using the arithmetic mean of the initial hydrological parameters of the eight adjacent grid cells as a reference, and considering the original data characteristics of the anomalous grid cell, its initial hydrological parameters are slightly adjusted. The adjustment range does not exceed 10% of the mean, ensuring that the adjusted values conform to the trends of the surrounding data without deviating from the reasonableness of the original data. After the initial hydrological parameters of all grid units have been checked and adjusted, the optimized parameters of all grid units are integrated to form a set of hydrological data that fully covers the target area, is numerically continuous and coordinated, and conforms to the spatial distribution law of hydrology, i.e., the hydrological data of the target area.
[0031] The beneficial effects are as follows: The quality weight ratio determined by the data quality level ensures that the fusion process of precipitation and evapotranspiration data fully reflects their quality differences, with higher-quality data receiving higher weights. This guarantees the scientific validity and rationality of the fusion logic and improves the reliability of the fusion results. Data gridding, through a unified spatial grid standard, effectively eliminates the spatial scale differences between the two types of heterogeneous data, establishing a consistent basic framework for subsequent fusion calculations and avoiding errors caused by spatial mismatch. The weighted fusion method based on weight ratios accurately quantifies the contribution of each type of data in each grid cell, ensuring the targeted and accurate calculation of initial hydrological parameters. Spatial consistency optimization, through neighborhood comparison and anomaly adjustment, successfully eliminates abrupt spatial fluctuations in initial parameters, significantly improving the spatial continuity and overall coordination of the data. The final generated hydrological data possesses good accuracy, spatial integrity, and logical rationality, providing high-quality and highly adaptable hydrological foundational data support for subsequent multi-source data fusion processes such as soil and land use composite data integration and eco-hydrological data aggregation in the target area, effectively reducing the error risk in subsequent data processing.
[0032] S3. Perform spatial interpolation on the hydrological data to obtain the hydrological grid data of the hydrological data; In this embodiment of the invention, the step of performing spatial interpolation processing on the hydrological data to obtain the hydrological grid data includes: Based on the spatial distribution characteristics of the hydrological data, a distribution framework for the hydrological data is constructed. Within the distribution framework, interpolation parameters are determined based on the numerical gradient changes of the hydrological data; Based on the interpolation parameters, the missing areas in the hydrological data are spatially filled to generate the initial interpolation of the hydrological data; The initial interpolation is smoothed at the boundary to obtain the hydrological grid data of the hydrological data.
[0033] When constructing a hydrological data distribution framework based on the spatial distribution characteristics of hydrological data, the spatial density distribution of hydrological data within the target area is first comprehensively analyzed to clarify the specific ranges of concentrated and sparse areas of the data. Simultaneously, the high-to-low numerical clustering patterns of the data are analyzed to determine the spatial distribution patterns of high, medium, and low numerical regions, as well as the overall spatial spread trend of the data. Based on the complete spatial boundary of the target area, the unified grid scale previously set for data gridding is used to determine the spatial coverage of the distribution framework, ensuring that the framework can completely encompass all valid hydrological data points within the target area. At the same time, a grid density consistent with the existing spatial grid precision is set, so that the distribution framework not only adapts to the spatial distribution characteristics of the hydrological data but also provides a unified and stable spatial benchmark for subsequent interpolation operations, ultimately forming a structurally well-structured hydrological data distribution framework.
[0034] When determining interpolation parameters based on the numerical gradient changes of hydrological data within a distributed framework, all adjacent valid hydrological data points within the framework are first selected. The numerical difference between each pair of adjacent data points is then calculated. The magnitude of the difference indicates the rate of change of the numerical gradient; a larger difference indicates a faster gradient change, and a smaller difference indicates a slower gradient change. Corresponding interpolation influence ranges are set for different gradient change scenarios. For regions with rapid gradient changes, a smaller interpolation influence range is set to ensure accurate preservation of local numerical details; for regions with slow gradient changes, a larger interpolation influence range is set to ensure data continuity over a large spatial area. Simultaneously, the weighting rules for interpolation calculation are defined: valid hydrological data points closer to the location to be interpolated have a greater weight in the interpolation calculation, while those farther away have a smaller weight. Through gradient change analysis and weighting rule setting, all parameters used for subsequent interpolation are fully determined.
[0035] When generating initial interpolation for spatial filling of missing regions in hydrological data based on interpolation parameters, the process begins by traversing all grid cells within the distribution frame, identifying each missing grid cell that does not contain valid hydrological data, and recording the spatial location information of each missing cell. For each missing grid cell, based on the determined interpolation influence range, all valid hydrological data points within the cell's surrounding area are selected. Then, according to the set weighting rules and the spatial distance between each valid data point and the missing cell, a corresponding weight value is assigned to each valid data point. The sum of the products of these valid data point values and their corresponding weight values is calculated and divided by the sum of all weight values to obtain the weighted average value of the missing grid cell. This value is then assigned to the corresponding missing grid cell, completing the numerical filling of a single missing cell. This filling operation is repeated for all missing grid cells within the distribution frame. After all cells are filled, the values of all grid cells within the distribution frame are integrated, including the values of the original valid data cells and the values of the filled missing cells, to generate an initial interpolation covering the entire distribution frame.
[0036] When obtaining hydrological grid data by smoothing the boundaries of the initial interpolation, the focus is first on the boundaries between different numerical regions in the initial interpolation, identifying the boundary grid cells within these regions. The numerical difference between each boundary grid cell and its multiple neighboring grid cells is calculated, and this difference is compared to a preset reasonable continuity range. This reasonable range is determined based on the hydrological characteristics of the target area and by analyzing the spatial fluctuation range of historical hydrological data. If the numerical difference between a boundary cell and its neighboring cells exceeds the reasonable range, resulting in abrupt numerical discontinuities at the boundaries of different regions, the average value of the values of the multiple neighboring cells surrounding that boundary cell is used as the adjustment benchmark. The value of that boundary cell is then slightly adjusted to ensure a smooth transition with the surrounding cell values, eliminating the numerical discontinuity. After all boundary cells have been checked and adjusted, the final values of all grid cells within the distribution framework are integrated to form hydrological grid data that is numerically continuous, has consistent boundaries, complete spatial coverage, and conforms to hydrological characteristics.
[0037] The beneficial effects are as follows: the distribution framework constructed based on the spatial distribution characteristics of hydrological data ensures the spatial coverage integrity and scale consistency of interpolation processing, providing a reliable foundation for subsequent operations; the interpolation parameters determined based on numerical gradient changes make spatial filling more targeted, accurately preserving local data details while ensuring the overall spatial distribution pattern; the missing area filling operation completes the spatial coverage of hydrological data, solving the problem of data sparsity or missing data; the boundary smoothing process effectively eliminates numerical discontinuities and improves the spatial continuity of data. The final hydrological grid data has complete coverage, reasonable numerical gradients, and good boundary coordination, which can provide high-quality and highly adaptable hydrological basic data for subsequent integrated analysis with soil and land use composite data.
[0038] S4. Based on the data quality level, perform spatial overlay analysis on the land use data and soil data in the multi-source data to obtain the composite soil land use data of the target area; In this embodiment of the invention, the step of performing spatial overlay analysis on land use data and soil data from the multi-source data based on the data quality level to obtain composite soil land use data for the target area includes: Based on the data quality level, assign superimposed weights to the land use data and the soil data; Spatial registration is performed on the land use data and the soil data respectively to obtain the registered land use spatial data and the registered soil spatial data. Spatial integration of the land use spatial data and the soil spatial data yields spatial framework data for the land use data and the soil data. Based on the superposition weight, the land use type code and soil type code in the spatial frame data are weighted and combined to obtain the composite land category code of the spatial frame data. The composite land cover code is updated to the spatial framework data to generate composite soil land cover data for the target area.
[0039] The spatial integration of the land use spatial data and the soil spatial data to obtain spatial framework data for the land use data and the soil data includes: Establish a spatial index between the land use spatial data and the soil spatial data; Based on the spatial index, the geometric boundaries of the land use spatial data and the soil spatial data are merged, and the attribute structures of the land use spatial data and the soil spatial data are constructed simultaneously. Based on the geometric boundaries and the attribute structure, spatial framework data for the land use data and the soil data are generated.
[0040] When assigning weights to land use data and soil data based on data quality levels, the correspondence between data quality levels and weight coefficients is continued: Level 1 corresponds to 1.0, Level 2 to 0.8, Level 3 to 0.6, and Level 4 to 0.4. The data quality levels of land use data and soil data are extracted respectively, and the weight coefficients of the two are determined according to the corresponding levels. This coefficient is the weight of the land use data and soil data, clearly distinguishing the weight of the land use data and the weight of the soil data.
[0041] When spatially registering land use data and soil data separately, the standard spatial coordinate system preset for the target area is used as a unified reference. First, the original spatial coordinate information of the land use data is analyzed, and its spatial position is adjusted to the standard coordinate system through coordinate transformation methods to ensure the accuracy of the data's spatial position, thus obtaining the registered land use spatial data. The same standard spatial coordinate system and coordinate transformation method are used to adjust the original spatial coordinates of the soil data so that the spatial position of the soil data is accurately matched with the standard coordinate system, thus obtaining the registered soil spatial data.
[0042] When spatially integrating land use spatial data and soil spatial data, a spatial index is first established for both. The spatial index is used to quickly locate the spatially related areas in the two types of data. Then, the geometric boundaries of the land use spatial data and soil spatial data are merged to ensure that the boundaries are completely coincident and without overlap or omission. At the same time, a unified attribute structure containing land use-related attributes and soil-related attributes is constructed. The attribute information of the two types of data is organized into this structure. Based on the merged geometric boundaries and the unified attribute structure, spatial framework data of land use data and soil data are generated.
[0043] When weighting land use type codes and soil type codes in spatial frame data based on superposition weights, the land use type code and soil type code corresponding to each grid cell in the spatial frame data are first extracted. The land use type code is converted into its corresponding numerical form and multiplied by the superposition weight of the land use data. Then, the soil type code is converted into its corresponding numerical form and multiplied by the superposition weight of the soil data. The sum of the two products is calculated, and this sum is the composite land use code of the spatial frame data corresponding to that grid cell. This calculation is performed by traversing all grid cells to obtain the composite land use code for the entire region. When updating the composite land use code to the spatial frame data, the composite land use code corresponding to each grid cell is filled into the corresponding attribute field of the spatial frame data one by one, replacing the original single code information in that field. This ensures that the attribute field of each grid cell is accurately associated with the corresponding composite land use code. All updated grid cell data are integrated to form a dataset that completely covers the target area and contains composite land use codes, i.e., the composite soil land use data of the target area.
[0044] When establishing spatial indexes for land use spatial data and soil spatial data, the standard spatial coordinate system and grid precision of the two types of data after registration are first defined. Taking the spatial range of the target area as the boundary, the entire area is divided into several non-overlapping spatial units using a spatial division method. A unique spatial identifier code is assigned to each spatial unit. Then, each patch or grid unit of the land use spatial data and each patch or grid unit of the soil spatial data are matched one by one, and the spatial unit identifier code to which it belongs is recorded, forming an index table with spatial identifier codes as the link. Through this index table, the land use spatial data and soil spatial data units corresponding to any spatial location can be quickly located, thus completing the establishment of the spatial index for both.
[0045] When fusing the geometric boundaries of land use spatial data and soil spatial data based on spatial indexing and simultaneously constructing the attribute structure, the following steps are taken: First, the corresponding land use spatial data patches / grids and soil spatial data patches / grids under the same spatial identifier code are extracted through the spatial index table. Topological checks are then performed on the geometric boundaries of the two types of data to eliminate boundary overlaps, gaps, or topological errors. The boundaries of overlapping areas are uniformly fitted to form a seamless and fully covered unified geometric boundary. At the same time, the core attribute items of land use spatial data and soil spatial data are sorted out, and duplicate and redundant attribute fields are eliminated. Following the logical order of "land use-related attributes first, soil-related attributes second," a unified attribute structure containing key information of both types of data is formed, ensuring that each attribute field has a clear definition and data type.
[0046] When generating spatial framework data for land use and soil data based on geometric boundaries and attribute structures, the fused unified geometric boundary is used as a spatial carrier. The constructed unified attribute structure is then associated and bound to this spatial carrier. Corresponding attribute fields are assigned to each geometric unit to ensure that each geometric location accurately corresponds to the relevant attribute information of land use and soil. At the same time, the integrity of the associated overall data is verified to confirm that there are no missing geometric boundaries and no omissions of attribute fields. Finally, a dataset with both unified spatial form and complete attribute information is formed, namely, spatial framework data for land use and soil data.
[0047] The beneficial effects are as follows: by assigning weights based on data quality levels, the overlay analysis of land use data and soil data fully matches their respective quality levels, ensuring the rationality of the analysis logic; spatial registration operation realizes the unification of spatial benchmarks for the two types of data, avoiding the impact of spatial location deviations on the analysis results; spatial integration constructs a unified spatial framework and attribute structure, providing a stable foundation for coding combinations; the composite land use code generated by weighted combination accurately integrates the core information of land use and soil, and the final generated composite soil land use data has the key characteristics of both types of data, providing high-quality composite basic data for subsequent integration with hydrological grid data.
[0048] S5. Aggregate the spatial spectral information of the soil type composite data and the hydrological grid data to obtain the eco-hydrological data of the target area; In this embodiment of the invention, the step of aggregating the spatial-spectral information of the soil type composite data and the hydrological grid data to obtain the eco-hydrological data of the target area includes: Based on the spatial reference of the soil land use composite data, the hydrological grid data is spatially resampled to obtain hydrological grid data with the same spatial resolution as the soil land use composite data. Spatial registration is performed on the resampled hydrological grid data and the soil land use composite data to obtain the correspondence between the spatial locations of the hydrological grid data and the soil land use composite data; Based on the correspondence, the soil land use composite attributes and hydrological grid attributes of the spatial location are associated and integrated to generate a composite parameter set of the hydrological grid data and the soil land use composite data; Multi-source data normalization is performed on the composite parameter set to generate eco-hydrological data for the target area.
[0049] Based on the aforementioned correspondence, the soil land use composite attributes and hydrological grid attributes at the spatial location are correlated and integrated to generate a composite parameter set of the hydrological grid data and the soil land use composite data, including: Based on the correspondence, feature pairs of soil land use composite attributes and hydrological grid attributes in the spatial location are obtained; Calculate the fusion parameters of the feature pairs, wherein the formula for calculating the fusion parameters is as follows: ; In the formula, For the first Fusion parameters for each spatial location, For the first Numerical representation of the composite attributes of soil land use at a spatial location For the first Numerical representation of hydrological grid attributes for each spatial location The preset weighting coefficients, The preset weighting coefficients, The preset weighting coefficients, To obtain the maximum value, To obtain the minimum value; Spatial consistency verification is performed on the fusion parameters, and a composite parameter set of the hydrological grid data and the soil land use composite data is generated based on the verification results.
[0050] When spatially resampling hydrological grid data based on the spatial benchmark of soil and land use composite data, the standard spatial coordinate system, the number of grid rows and columns, and the side length of individual grids used in the soil and land use composite data should be fully defined first. These parameters should be used together as the target benchmark for resampling to ensure spatial consistency in subsequent processing. Neighborhood mean sampling is used as the sole sampling method, with the grid size of the target benchmark as a fixed standard. The original hydrological grid data is processed as follows: if the side length of the original grid is greater than that of the target grid, the original grid is split into multiple smaller grids with the same size as the target benchmark, and the value of each smaller grid is the average value of the original grids; if the side length of the original grid is less than that of the target grid, multiple adjacent original grids are merged into a larger grid with the same size as the target benchmark, and the value of the larger grid is determined by the arithmetic mean of the values of all merged original grids. Through this splitting or merging operation, hydrological grid data with a spatial resolution completely consistent with the soil and land use composite data is finally obtained.
[0051] When spatially registering the resampled hydrological grid data with the soil and land use composite data, the spatial location of the soil and land use composite data is always used as the core reference standard. The planar coordinates of each grid cell in the resampled hydrological grid data are checked one by one. If a slight deviation is found between the coordinates of a certain grid cell and the corresponding grid cell in the soil and land use composite data, a coordinate translation correction method is used. Based on the specific value of the deviation, the coordinates of the grid cell are finely adjusted to the precise position in the horizontal or vertical direction, ensuring that each grid cell in the hydrological grid data is precisely aligned point-to-point with its corresponding grid cell in the soil and land use composite data. After alignment, a detailed location association table is established, clearly recording the one-to-one correspondence between the unique identifier of each hydrological grid cell and the unique identifier of each grid cell in the soil and land use composite data. This clearly reveals the correspondence between the spatial locations of the hydrological grid data and the soil and land use composite data.
[0052] When obtaining feature pairs of soil land use composite attributes and hydrological grid attributes based on the correspondence relationship, attribute extraction is performed on each spatial location one by one, strictly following the established spatial location correspondence relationship. For soil land use composite data, the core composite attributes corresponding to that spatial location are extracted, including key information such as composite land use code, soil texture, land use type, soil thickness, and soil organic matter content. For resampled hydrological grid data, the core hydrological grid attributes corresponding to that spatial location are extracted, including key information such as precipitation, evapotranspiration, surface runoff, and soil moisture content. The set of soil land use composite attributes extracted from the same spatial location is bound to the set of hydrological grid attributes to form a feature pair specific to that spatial location, ensuring that each spatial location has one and only one corresponding feature pair, and that the attribute information in the feature pair is complete and without omission.
[0053] When calculating the fusion parameters of feature pairs, the system first prepares the necessary basic data. The numerical representation of the composite soil land use attributes of a spatial location is obtained by converting each composite soil land use attribute of that spatial location into specific values according to a unified numerical conversion rule, and then integrating them. For example, the composite land use code is converted into a numerical code according to a preset correspondence table, the soil texture is converted according to the rule of sandy 1.0, loam 2.0, clay 3.0, and the land use type is converted according to the rule of cultivated land 1.0, forest land 2.0, etc., and then the values are combined to form a unified numerical representation; The numerical representation of hydrological grid attributes at each spatial location adopts the same logic, directly using the actual physical quantity values of each hydrological grid attribute as the basis. If non-numerical attributes exist, they are converted into values according to corresponding rules, and then integrated to form a unified numerical representation. The first, second, and third preset weight coefficients are pre-set to fixed specific values based on the eco-hydrological characteristics of the target area, the reliability level corresponding to the data quality level, and the application requirements for actual water yield assessment. The sum of the three weight coefficients is 1.0 to ensure the rationality of weight allocation.
[0054] Next, a comprehensive calculation of various basic values is performed to lay the foundation for subsequent normalization operations. The numerical representations of soil land use composite attributes at all spatial locations within the target area are traversed, and all values are compared one by one to select the maximum and minimum values, which are then used as the maximum and minimum values in the numerical representations of soil land use composite attributes at all spatial locations. The same traversal and comparison method is used to process the numerical representations of hydrological grid attributes at all spatial locations, selecting the corresponding maximum and minimum values. For the correlation characteristics between the two types of attributes, the product of the numerical representation of soil land use composite attributes and the numerical representation of hydrological grid attributes at each spatial location is first calculated. Then, the product results at all spatial locations are traversed, and the maximum value is selected by comparing them one by one, serving as the maximum value in the product of the numerical representations of the two types of attributes at all spatial locations.
[0055] Then, a three-step normalization calculation was performed to ensure that the values of different attributes were on a uniform scale. The first step was to calculate the normalized values of the composite attributes of soil land use, using the... The minimum value among the numerical representations of soil land use composite attributes at all spatial locations is subtracted from the numerical representation of soil land use composite attributes at all spatial locations to obtain the attribute difference. This difference is then divided by the difference between the maximum and minimum values among the numerical representations of soil land use composite attributes at all spatial locations. This calculation maps the values of soil land use composite attributes uniformly to the 0-1 interval, ensuring they are within a uniform scale range. The second step calculates the normalized values of hydrological grid attributes, using the exact same calculation logic as the first step. The numerical representation of hydrological grid attributes at each spatial location is subtracted from the corresponding minimum value, and the difference is divided by the difference between the corresponding maximum and minimum values. This maps the hydrological grid attribute values to the 0-1 interval, maintaining scale consistency with the normalized values of soil land use composite attributes. The third step calculates the normalized value of the product of the two types of attributes, using the... The product of the numerical representation of the composite soil land cover attribute at each spatial location and the numerical representation of the hydrological grid attribute is divided by the maximum value among the products of the numerical representations of the two types of attributes corresponding to all spatial locations. The product result is also mapped to the 0-1 interval to further explore the correlation features between the two types of attributes.
[0056] Finally, a weighted summation calculation is performed. The normalized value of the soil land use composite attribute obtained in the first step is multiplied by the first preset weight coefficient to obtain the contribution value of the soil land use composite attribute to the fusion parameter; the normalized value of the hydrological grid attribute obtained in the second step is multiplied by the second preset weight coefficient to obtain the contribution value of the hydrological grid attribute to the fusion parameter; the normalized value of the product of the two types of attributes obtained in the third step is multiplied by the third preset weight coefficient to obtain the contribution value of the correlation feature between the two types of attributes to the fusion parameter; these three contribution values are added together, and the final sum is the first... The fusion parameters of each spatial location enable deep integration of soil land use composite attributes and hydrological grid attributes.
[0057] When performing spatial consistency verification on the fusion parameters and generating a composite parameter set, the fusion parameters at each spatial location are compared and checked against their surroundings. The fusion parameters of the eight adjacent spatial locations are selected as references. The numerical differences between the fusion parameters at this spatial location and those at each adjacent location are calculated. The average and maximum values of all differences are statistically analyzed to determine whether the maximum value exceeds the reasonable range allowed by the eco-hydrological characteristics of the target area. This reasonable range is pre-defined by analyzing the spatial fluctuation patterns of historical eco-hydrological data in the target area. If the difference between the fusion parameters at a certain spatial location and those at most adjacent locations exceeds the reasonable range, it indicates a spatial anomaly. The anomaly parameter is adjusted based on the arithmetic mean of the fusion parameters at the eight adjacent locations to ensure continuity and consistency with surrounding parameters. If the difference does not exceed the reasonable range, the original value of the parameter is retained. After the fusion parameters at all spatial locations have been verified and adjusted, all valid fusion parameters are organized in spatial order to form a complete dataset covering all spatial locations in the target area—a composite parameter set of hydrological grid data and soil type composite data.
[0058] When generating eco-hydrological data for the target area by multi-source data normalization of composite parameter sets, a comprehensive structural review of the composite parameter sets is first performed. The naming definitions of all parameters are standardized; for example, "the first..." The spatial location fusion parameters are standardized and named "Ecological-Hydrological Integrated Parameters," with clear definitions for each parameter. A unified data recording format is adopted, using a fixed format of "unique spatial location identifier – ecological-hydrological integrated parameter value" to ensure data consistency. Each parameter entry is checked individually, eliminating duplicate entries and invalid values. The parameter set is then categorized and organized according to the administrative or natural divisions of the target area, establishing an independent parameter subset for each division to ensure that each spatial division has corresponding complete parameter records for easy subsequent data application. Finally, a comprehensive integrity check is performed on the standardized parameter set to confirm that all spatial locations have corresponding parameter records, no missing parameters, standardized formats, and logical coherence. The checked parameter sets are then integrated to form ecological-hydrological data covering the target area, containing core soil type and hydrological information, and with a standardized structure.
[0059] The beneficial effects are as follows: Spatial resampling, by clarifying the target benchmark and standardizing the sampling method, achieved unified spatial resolution between soil and land use composite data and hydrological grid data, providing a consistent spatial foundation for subsequent fusion operations; Spatial registration, through precise coordinate verification and correction, ensured accurate point-to-point correspondence between the two types of data spatial locations, avoiding the impact of spatial deviations on the fusion results; Feature pair extraction, by comprehensively extracting and binding core attributes, ensured the integrity of the two types of attribute information at the same spatial location; Fusion parameter calculation, through the system's basic data preparation, multi-step normalization processing, and scientific weighted summation, achieved deep integration of the two types of attributes, fully exploring the correlation features between attributes; Spatial consistency verification, through comparison with surrounding parameters and anomaly adjustment, eliminated spatial anomalies in the fusion parameters, ensuring the spatial continuity and rationality of the parameters; Multi-source data normalization, through unified format, removal of invalid data, partitioning and organization, and integrity verification, optimized the data structure and integrity. The final generated eco-hydrological data integrates core information on soil types and hydrological characteristics. The data is highly consistent, complete in information dimensions, and highly accurate, providing high-quality core foundational data for subsequent standardization conversion and InVEST model input, effectively supporting the accuracy of water yield assessment in the target area.
[0060] S6. Based on the data quality level, the eco-hydrological data is standardized to obtain high-precision fused data of the target area, and the high-precision fused data is input into the InVEST model to complete the assessment of the water yield of the target area.
[0061] In this embodiment of the invention, the step of standardizing and transforming the eco-hydrological data based on the data quality level to obtain high-precision fused data of the target area, and inputting the high-precision fused data into the InVEST model to complete the assessment of water yield in the target area, includes: Based on the data quality level, the transformation weights of the parameters in the eco-hydrological data are determined; Based on the transformation weights, the soil parameters, vegetation parameters and hydrological parameters in the eco-hydrological data are standardized to obtain regularized eco-hydrological data. Spatial scale adaptation is performed on the regularized data to obtain high-precision fused data of the target region; The high-precision fused data is input into the InVEST model to complete the water production assessment results for the target area.
[0062] When determining the transformation weights of parameters in eco-hydrological data based on data quality levels, the correspondence between data quality levels and weight coefficients is continued: Level 1 corresponds to 1.0, Level 2 to 0.8, Level 3 to 0.6, and Level 4 to 0.4. First, the three core parameters contained in the eco-hydrological data—soil parameters, vegetation parameters, and hydrological parameters—are identified. The data quality level corresponding to each type of parameter is extracted, and the corresponding weight coefficient is matched according to its respective level. This coefficient is the transformation weight of each type of parameter. The transformation weights of soil parameters, vegetation parameters, and hydrological parameters are clearly distinguished to ensure that the weight of each type of parameter is accurately matched with its own data quality level.
[0063] When unifying the dimensions of soil, vegetation, and hydrological parameters in eco-hydrological data based on transformation weights, standardization is used as the sole method for unifying dimensions. First, the maximum and minimum values of all data in the soil parameter set, vegetation parameter set, and hydrological parameter set are statistically analyzed separately. Then, for each data point in each parameter category, the minimum value of the corresponding parameter set is subtracted from the data point, and the difference is divided by the difference between the maximum and minimum values of the corresponding parameter set to complete the dimension normalization of a single parameter category. Subsequently, the normalized soil parameters are multiplied by the soil parameter transformation weight, the normalized vegetation parameters are multiplied by the vegetation parameter transformation weight, and the normalized hydrological parameters are multiplied by the hydrological parameter transformation weight. All weighted parameters are then integrated to obtain the regularized eco-hydrological data.
[0064] When adapting regularized data to a spatial scale to obtain high-precision fused data for the target region, the spatial scale requirements of the InVEST model for the input data are first clarified, including the model's preset grid size, spatial resolution, and coordinate system. Then, based on these requirements, the spatial grid of the regularized data is adjusted: if the regularized data grid is larger than the model's required grid, a spatial segmentation method is used to split a single regularized data grid into multiple smaller grids with the same size as the model's required grid. The value of each smaller grid uses the weighted value of the original regularized data grid. If the regularized data grid is smaller than the model's required grid, a neighborhood mean method is used to merge multiple adjacent regularized data grids into a larger grid with the same size as the model's required grid. The value of the larger grid is determined by the arithmetic mean of the weighted values of all smaller grids within the merged range. After adjustment, the spatial position of each grid is checked one by one to verify its compatibility with the model's coordinate system, ensuring no positional deviation. Finally, high-precision fused data with a spatial scale that perfectly matches the InVEST model is formed.
[0065] When inputting high-precision fused data into the InVEST model to complete the water yield assessment of the target area, the soil parameters, vegetation parameters, and hydrological parameters in the high-precision fused data are first formatted according to the data input format specified by the InVEST model, clarifying the correspondence between the parameter fields and the model input interface. Then, the formatted high-precision fused data is completely imported into the designated input module of the model through the model's data import function. The model will automatically read various parameter information in the data, combine it with the built-in water yield assessment algorithm, and comprehensively analyze the impact of soil, vegetation, hydrology and other factors on the regional water yield. By calculating and outputting assessment results including information such as the overall water yield of the target area and the distribution of water yield at different spatial locations, the water yield assessment of the target area is completed.
[0066] The beneficial effects include: determining transformation weights based on data quality levels, ensuring that the standardized transformation of various parameters in eco-hydrological data fully matches their own quality levels, and guaranteeing the scientific nature of the transformation logic; unified dimensional operations eliminate scale differences between different types of parameters, avoiding interference from inconsistent dimensions in subsequent assessments; spatial scale adaptation ensures high compatibility between the data and the InVEST model, reducing data adaptation errors during model operation; after inputting high-precision fused data into the model, the model can perform calculations based on complete, standardized, and adapted basic data, effectively improving the accuracy and efficiency of water yield assessment in the target area, and providing reliable data support for regional water resource management and ecological protection.
[0067] like Figure 2 The diagram shown is a functional block diagram of an InVEST regional water production multi-source data fusion system provided in an embodiment of the present invention.
[0068] The InVEST regional water production multi-source data fusion system 100 described in this invention can be installed in an electronic device. Depending on the functions implemented, the InVEST regional water production multi-source data fusion system 100 may include a data quality assessment module 101, a heterogeneous data assimilation module 102, a spatial interpolation processing module 103, a spatial overlay analysis module 104, a spatial spectrum information aggregation module 105, and a normalization conversion module 106. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, stored in the memory of the electronic device.
[0069] In this embodiment, the functions of each module / unit are as follows: The data quality assessment module 101 is used to assess the quality of multi-source data in the target area and obtain the data quality level of the multi-source data. The heterogeneous data assimilation module 102 is used to perform heterogeneous data assimilation on precipitation data and evapotranspiration data in the multi-source data based on the data quality level, so as to obtain the hydrological data of the target area. The spatial interpolation processing module 103 is used to perform spatial interpolation processing on the hydrological data to obtain hydrological grid data of the hydrological data; The spatial overlay analysis module 104 is used to perform spatial overlay analysis on land use data and soil data in the multi-source data based on the data quality level, so as to obtain composite soil land use data of the target area. The spatial spectrum information aggregation module 105 is used to perform spatial spectrum information aggregation on the soil land type composite data and the hydrological grid data to obtain the eco-hydrological data of the target area. The normalization conversion module 106 is used to normalize the eco-hydrological data based on the data quality level to obtain high-precision fused data of the target area, and input the high-precision fused data into the InVEST model to complete the assessment of the water yield of the target area.
[0070] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0071] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0072] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0073] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0074] This application embodiment can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. An InVEST regional water yield multi-source data fusion method, characterized in that, The method comprises: S1, quality evaluation of multi-source data of a target region, to obtain data quality level of the multi-source data; S2, based on the data quality level, heterogeneous data assimilation of precipitation data and evapotranspiration data in the multi-source data, to obtain hydrological data of the target region; S3, spatial interpolation processing of the hydrological data, to obtain hydrological grid data of the hydrological data; S4, based on the data quality level, spatial overlay analysis of land use data and soil data in the multi-source data, to obtain soil class composite data of the target region; S5, spatial information aggregation of the soil class composite data and the hydrological grid data, to obtain eco-hydrological data of the target region; S6, based on the data quality level, normalization conversion of the eco-hydrological data, to obtain high-precision fusion data of the target region, and inputting the high-precision fusion data into an InVEST model to complete the evaluation of water yield of the target region. 2.The InVEST regional water yield multi-source data fusion method according to claim 1, characterized in that, The quality evaluation of multi-source data of a target region, to obtain data quality level of the multi-source data, comprises: obtaining data source description of precipitation data in multi-source data of a target region, and performing precipitation evaluation on the precipitation data to obtain precipitation quantity index of the precipitation data; spatial and temporal alignment of evapotranspiration data in the multi-source data, to obtain consistency identification of the evapotranspiration data; land class coding review of land use data in the multi-source data, to obtain coding specification compliance of the land use data; attribute completeness check of soil data in the multi-source data, to obtain attribute evaluation result of the soil data; based on the data source description, the precipitation quantity index, the consistency identification, the coding specification compliance and the attribute evaluation result, generating data quality level of the multi-source data.
3. The InVEST regional water yield multi-source data fusion method according to claim 1, characterized in that, The heterogeneous data assimilation of precipitation data and evapotranspiration data in the multi-source data based on the data quality level, to obtain hydrological data of the target region, comprises: based on the data quality level, confirming quality weight ratio of the precipitation data and the evapotranspiration data; data gridding of the precipitation data and the evapotranspiration data, to unify the precipitation data and the evapotranspiration data to the same spatial grid; based on the quality weight ratio, weighted fusion of the gridded precipitation data and the evapotranspiration data in the spatial grid, to generate initial hydrological parameters of the precipitation data and the evapotranspiration data; spatial consistency optimization of the initial hydrological parameters, to generate hydrological data of the target region.
4. The InVEST regional water yield multi-source data fusion method according to claim 1, characterized in that, The spatial interpolation processing of the hydrological data, to obtain hydrological grid data of the hydrological data, comprises: based on spatial distribution characteristics of the hydrological data, constructing distribution framework of the hydrological data; determining interpolation parameters according to numerical gradient change of the hydrological data in the distribution framework; based on the interpolation parameters, spatial filling of missing areas in the hydrological data, to generate initial interpolation of the hydrological data; Boundary smoothing is performed on the initial interpolation to obtain hydrological grid data of the hydrological data.
5. The InVEST regional water yield multi-source data fusion method according to claim 1, characterized in that, Based on the data quality level, spatial overlay analysis is performed on land use data and soil data in the multi-source data to obtain soil class composite data of the target region, including: Based on the data quality level, the land use data and the soil data are assigned an overlay weight; The land use data and the soil data are respectively spatially registered, and land use spatial data after registration of the land use data and soil spatial data after registration of the soil data are obtained; The land use spatial data and the soil spatial data are spatially integrated to obtain spatial framework data of the land use data and the soil data; Based on the overlay weight, land use type coding and soil type coding in the spatial framework data are combined by weighting to obtain composite land class coding of the spatial framework data; The composite land class coding is updated to the spatial framework data to generate soil class composite data of the target region.
6. The InVEST regional water yield multi-source data fusion method according to claim 5, characterized in that, The spatial integration of the land use spatial data and the soil spatial data to obtain the spatial framework data of the land use data and the soil data includes: A spatial index of the land use spatial data and the soil spatial data is established; Based on the spatial index, the geometric boundaries of the land use spatial data and the soil spatial data are fused, and the attribute structure of the land use spatial data and the soil spatial data is synchronously constructed; Based on the geometric boundaries and the attribute structure, the spatial framework data of the land use data and the soil data is generated.
7. The InVEST regional water yield multi-source data fusion method according to claim 1, characterized in that, The spatial-spectral information aggregation of the soil class composite data and the hydrological grid data to obtain the ecological hydrological data of the target region includes: Based on the spatial reference of the soil class composite data, the hydrological grid data is spatially resampled to obtain hydrological grid data consistent with the spatial resolution of the soil class composite data; The resampled hydrological grid data and the soil class composite data are spatially registered to obtain a corresponding relationship between the spatial positions of the hydrological grid data and the soil class composite data; Based on the corresponding relationship, the soil class composite attributes and the hydrological grid attributes of the spatial positions are associated and integrated to generate a composite parameter set of the hydrological grid data and the soil class composite data; The multi-source data is regularized to generate ecological hydrological data of the target region.
8. The InVEST regional water yield multi-source data fusion method according to claim 7, characterized in that, Based on the corresponding relationship, the soil class composite attributes and the hydrological grid attributes of the spatial positions are associated and integrated to generate a composite parameter set of the hydrological grid data and the soil class composite data, including: Based on the corresponding relationship, a feature pair of the soil class composite attributes and the hydrological grid attributes in the spatial position is obtained; A fusion parameter of the feature pair is calculated, and the calculation formula of the fusion parameter is as follows: ; wherein, is a fusion parameter for the th spatial location, is a numerical representation of a soil class composite attribute for the th spatial location, is a numerical representation of a hydrological grid attribute for the th spatial location, is a preset weight coefficient, is a preset weight coefficient, is a preset weight coefficient, is a maximum value, is a minimum value. The fusion parameters are checked for spatial consistency, and a composite parameter set of the hydrological grid data and the soil and land class composite data is generated according to a checking result.
9. The InVEST regional water yield multi-source data fusion method according to claim 1, characterized in that, The ecological hydrological data is normalized and converted based on the data quality level, high-precision fusion data of the target region is obtained, and the high-precision fusion data is input into the InVEST model to complete the evaluation of the water yield of the target region. The conversion weight of the parameters in the ecological hydrological data is determined based on the data quality level. The soil parameters, vegetation parameters and hydrological parameters in the ecological hydrological data are unified in dimension based on the conversion weight, and the regularized data of the ecological hydrological data is obtained. The regularized data is subjected to spatial scale adaptation to obtain high-precision fusion data of the target region. The high-precision fusion data is input into the InVEST model to complete the evaluation result of the water yield of the target region.
10. An InVEST regional water yield multi-source data fusion system for implementing the InVEST regional water yield multi-source data fusion method of claim 1, the system comprising: a data quality evaluation module for evaluating the quality of multi-source data of a target region to obtain a data quality level of the multi-source data; a heterogeneous data assimilation module for assimilating precipitation data and evapotranspiration data in the multi-source data based on the data quality level to obtain hydrological data of the target region; a spatial interpolation processing module for performing spatial interpolation processing on the hydrological data to obtain hydrological grid data of the hydrological data; a spatial superposition analysis module for performing spatial superposition analysis on land use data and soil data in the multi-source data based on the data quality level to obtain soil and land class composite data of the target region; an air-spectrum information aggregation module for aggregating air-spectrum information of the soil and land class composite data and the hydrological grid data to obtain ecological hydrological data of the target region; a normalization conversion module for normalizing and converting the ecological hydrological data based on the data quality level to obtain high-precision fusion data of the target region, and inputting the high-precision fusion data into the InVEST model to complete the evaluation of the water yield of the target region.