Geographic information multi-source data fusion processing method oriented to ecological influence evaluation
By performing format conversion, coordinate correction and data fusion on multi-source geographical data, combined with vegetation coverage calculation, the problem of insufficient accuracy in multi-source geographical data fusion is solved, efficient data processing and intelligent mapping are realized, and ecological environment evaluation and land planning are supported.
Patent Information
- Application Number
- CN202510945303.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-07-09
AI Technical Summary
The existing technology is difficult to effectively integrate multi-source geographical data with different resolutions and different phases, resulting in insufficient data accuracy of vegetation type and ecosystem type. At the same time, the national second-definition data is inconsistent with the latest remote sensing image data, which affects the data fusion and analysis accuracy. The mapping template lacks intelligence and cannot meet the needs of diversified analysis.
Through format conversion, coordinate correction, data fusion and administrative division unity, combined with vegetation coverage calculation, decision tree classification algorithm and mapping template are used to realize the in-depth mining and efficient utilization of multi-source geographical data, and generate land use status maps, vegetation type maps and ecosystem maps.
It realizes efficient processing and in-depth mining of multi-source geographical data, improves data fusion accuracy and analysis accuracy, and supports the application of ecological environment evaluation and land planning.
Smart Images

Figure CN120451339A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of geographic information science and data processing technology, and specifically relates to a geographic information multi-source data fusion processing method for ecological impact assessment. Background Art
[0002] With the booming development of geographic information science, access to multi-source geographic data is becoming increasingly convenient, but data processing and analysis face significant challenges. While authoritative, the Second National Land Use National Land Survey data, as crucial foundational land use data, is stored in the MDB format, making it incompatible with current mainstream geographic information system software. Furthermore, its early acquisition time deviates from actual land use conditions. Furthermore, its coordinate system is inconsistent with the latest remote sensing imagery, making data fusion difficult. Furthermore, the administrative divisions in the Second National Land Survey data differ from the latest administrative divisions, impacting statistical and analytical accuracy. Traditional methods struggle to effectively integrate land use raster and vector data of varying resolutions and temporal phases, resulting in insufficiently accurate data on vegetation and ecosystem types. Furthermore, existing mapping templates lack intelligence during the output phase, failing to automatically adjust mapping elements based on the data range, making them incapable of meeting diverse analytical needs. Therefore, a multi-source geographic information data fusion processing method is needed to address these challenges. Summary of the Invention
[0003] In order to solve the above technical problems, the present invention proposes a geographic information multi-source data fusion processing method for ecological impact assessment, which aims to solve the complex and difficult problems of multi-source data processing in the geographic information field and realize the deep mining and efficient utilization of geographic data.
[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is: a method for fusion processing of multi-source geographic information data for ecological impact assessment, comprising the following steps: Step 1: Obtain basic land use data and the latest multi-source geographic information data, including road network data, water system data, remote sensing satellite image data, and administrative division data; Step 2: Perform format conversion and coordinate correction on the basic land use data to obtain the first basic land use data; Step 3: Fuse and overlay the road network data and water system data with the first basic land use data to obtain the second basic land use data; construct vegetation type raster data based on the spectral characteristics and texture characteristics of the remote sensing satellite image data; obtain ecosystem type data based on the vegetation type raster data and the second basic land use data, and improve the attribute table of the second basic land use data based on the vegetation type raster data and the ecosystem type data to obtain the third basic land use data; Step 4: Update the administrative division information in the basic land use data according to the administrative division data to obtain the fourth basic land use data; Step 5: Calculate the vegetation coverage of each pixel in the area of interest based on the remote sensing image, and determine the vegetation coverage level attribute of each pixel based on the set threshold. The calculation formula for vegetation coverage is: ; in, FVC Indicates the vegetation coverage of the pixel, NDVI represents the normalized vegetation index of the pixel, NDVI soil and NDVI veg They represent the normalized vegetation index of pure soil pixels and pure vegetation pixels in the area of interest respectively; Step 6: Call the land use data, vegetation type data, ecosystem type data and vegetation coverage level data of the focus area in the fourth land use basic data to generate a land use status map, a vegetation type map, an ecosystem map and a vegetation coverage map.
[0005] In the step 2, the specific method for coordinate correction of the basic land use data is: selecting multiple feature points in the basic land use data, obtaining the coordinates of the corresponding points in the remote sensing satellite image data, using the least squares method to fit and solve the affine transformation matrix, and performing coordinate correction on the basic land use data according to the affine transformation matrix.
[0006] In the step three, vegetation type data is constructed using a decision tree classification algorithm.
[0007] In step 3, the specific method of constructing vegetation type data using the decision tree classification algorithm is: First, spectral features and texture features are extracted from remote sensing satellite image data as input features; Then, a decision tree is generated by recursively partitioning the data set; classification rules are generated through the decision tree to obtain the vegetation type of each area.
[0008] In step 3, the decision tree classification algorithm selects the best split feature based on maximizing information gain, where the information gain is calculated as: ; in, IG(D,A) Represents split features A The information gain of D is the number of samples in the data set. Indicates the value is v Subset of D v The number of samples in the medium; H(D) represents the entropy of the dataset, H(Dv ) Representation subset D v The entropy of the data set is calculated as follows: ; in, p i Indicates the first i The proportion of class samples; In step 3, the decision tree classification algorithm selects the best split feature based on minimizing the Gini impurity, where the calculation formula of the Gini impurity is: ; in, represents the Gini impurity of the dataset, p i Indicates the first i The ratio of class samples, n represents the number of types of samples in the dataset.
[0009] In step five, five different levels of thresholds are set, namely, high vegetation coverage, medium-high vegetation coverage, medium vegetation coverage, medium-low vegetation coverage, and low vegetation coverage. According to the set thresholds, the vegetation coverage level attributes of each pixel are obtained, namely, the vegetation coverage level data.
[0010] The step five also includes a process of calculating 10-meter NDVI data based on the remote sensing image. When calculating vegetation coverage, the 10-meter NDVI data is used as the normalized vegetation index value of the pixel.
[0011] The step six also includes the step of making mapping templates for the current land use map, vegetation type map, ecosystem type map, and vegetation coverage map. In the mapping templates, different colors are mapped to different types of data.
[0012] The step six also includes the following steps: calculating the area of various types of land use, vegetation, ecosystem, and vegetation coverage level in the area of interest, and generating a statistical table.
[0013] Compared with the prior art, the present invention has the following beneficial effects: the present invention proposes a method for fusion processing of multi-source geographic information data for ecological impact assessment, firstly, losslessly converts the basic land use data, then updates and supplements the basic land use data by using steps such as coordinate correction, data fusion, and administrative division unification, and combines the calculation of vegetation coverage to realize intelligent mapping of various land information such as vegetation coverage through mapping templates. The present invention realizes deep mining and efficient utilization of multi-source geographic data through efficient processing of multi-source geographic information data and organic integration of it with land use basic data, providing a systematic and integrated solution for applications in the field of geographic information, which can be widely used in multiple fields such as ecological environment assessment, land planning, and environmental monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A flowchart of a method for fusion processing of multi-source geographic information data for ecological impact assessment provided by an embodiment of the present invention; Figure 2 Schematic diagram of data flow in an embodiment of the present invention; Figure 3 A land use status map generated in an embodiment of the present invention; Figure 4 A vegetation type map generated in an embodiment of the present invention; Figure 5 An ecological type map generated in an embodiment of the present invention; Figure 6 This is a vegetation coverage map generated in an embodiment of the present invention. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are part of the embodiments of the present invention, not all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0016] like Figures 1 and 2 As shown, the embodiment of the present invention provides a method for fusion processing of multi-source geographic information data for ecological impact assessment, comprising the following steps: Step 1: Data preparation.
[0017] Obtain basic land use data and the latest multi-source geographic information data, including road network data, water system data, remote sensing satellite image data and administrative division data.
[0018] Specifically, in this embodiment, the basic land use data is the National Second National Land Survey data stored in MDB format. The road network data, water system data, remote sensing satellite imagery data, and administrative division data are the latest data collected in 2024. During the collection process, data with high spatial and temporal resolution was selected to ensure data accuracy and applicability.
[0019] Step 2: Data processing.
[0020] The land use data table in the land use basic data is formatted; then, a plurality of feature points are selected from the land use basic data, the coordinates of the corresponding points are obtained from the remote sensing satellite image data, the affine transformation matrix is solved by fitting using the least squares method, and the coordinates of the land use basic data are corrected according to the affine transformation matrix to obtain the first land use basic data.
[0021] In this example, we leveraged the powerful data conversion capabilities of the GDAL (Geospatial Data Abstraction Library) to write a Python script to convert the land use data table format within the land use basic data. During the conversion process, we strictly adhered to the spatial data interoperability specifications established by the OGC (Open Geospatial Consortium) to ensure accurate conversion of the data geometry while preserving the integrity and correctness of the attribute fields. Through logging and error capture during the conversion process, we promptly identified and resolved potential issues during the conversion process, achieving lossless conversion of the land use data table from MDB format to SHP format, ensuring both geometric accuracy and attribute integrity.
[0022] Specifically, the specific formula for coordinate correction is: ; (1) Among them, (x, y) is the coordinate in the original land use basic data, (x', y') is the corrected coordinate, a,b,c,d, e,f are the affine transformation matrix parameters. The optimal parameters are determined by minimizing the sum of squares of the residuals between the actual feature point coordinates and the transformed coordinates.
[0023] Therefore, in this embodiment, based on the principle of affine transformation, multiple evenly distributed feature points are selected from the land use basic data, and the coordinates of the corresponding feature points are obtained in the remote sensing image through ground control points (GCPs). The coordinate correspondence between the feature points in the land use basic data and the remote sensing image is fitted using the least squares method, and the affine transformation matrix parameters are solved. Then, the parameters are substituted into formula (1) to perform coordinate correction on the original land use basic data to obtain the first land use basic data, ensuring the consistency of the position of the land use basic data with the remote sensing image.
[0024] Specifically, in this embodiment, Python's GDAL and Fiona libraries are used in combination with the NumPy numerical calculation library to achieve full process automation of feature point selection, coordinate matching, parameter calculation and data correction.
[0025] Step 3: Data update.
[0026] The road network data and water system data are integrated and superimposed with the first land use basic data to update and supplement the first land use basic data and obtain the second land use basic data; the spectral characteristics and texture characteristics of the remote sensing satellite image data are used to construct vegetation type raster data, and the ecosystem type data is obtained based on the vegetation type raster data and the second land use basic data; and the attribute table of the second land use basic data is improved based on the vegetation type raster data and the ecosystem type data to obtain the third land use basic data.
[0027] Specifically, in this embodiment, vegetation types are divided into evergreen coniferous forests, evergreen broad-leaved forests, grass, cultivated vegetation, no vegetation, etc. By using the spectral characteristics and texture characteristics of remote sensing satellite image data, the vegetation types can be identified to obtain vegetation type raster data. In addition, ecosystem types include broad-leaved forest ecosystems, coniferous forest ecosystems, grass ecosystems, lake ecosystems, cultivated land ecosystems, garden ecosystems, residential ecosystems, bare land ecosystems, etc. After the vegetation type raster data is obtained by identification, it is compared with the second land use basic data to obtain the ecosystem type of each region, and then the ecosystem type data is obtained.
[0028] Specifically, the first basic land use data can be supplemented and updated by fusing and superimposing the first basic land use data with the latest road network and water system data.
[0029] Specifically, in step three, vegetation type data is constructed using a decision tree classification algorithm.
[0030] In step 3, the specific method for constructing vegetation type data using the decision tree classification algorithm is as follows: First, spectral features (such as NDVI, NDWI) and texture features are extracted from remote sensing satellite image data as input features; Next, a decision tree is generated by recursively partitioning the dataset. In this tree, each node represents a feature, each branch represents a decision rule, and the leaf nodes represent the final classification result. The recursive process continues until a stopping condition is met (such as reaching a maximum depth or the number of node samples falling below a threshold). To prevent overfitting, the generated decision tree can be pruned to remove unimportant branches. Classification rules are generated from the decision tree, and the vegetation type of each area is obtained, resulting in vegetation type raster data.
[0031] After obtaining the vegetation type raster data, combined with the second land use basic data, the ecosystem type data can be obtained. Then, using the attribute table editing function of ArcGIS software, the vegetation type raster data and ecosystem type data are added to the attribute table of the second land use basic data to improve the data attribute table, which can ensure the integrity and accuracy of the data.
[0032] Specifically, in step 3, the decision tree classification algorithm selects the best split feature based on maximizing information gain, where the information gain is calculated as follows: ; (2) in, IG(D,A) Represents split features A The information gain of D is the number of samples in the data set. H(D) represents the entropy of the dataset, Indicates the value is v Subset of D v The number of samples in H(D v ) Representation subset D v The entropy of the data set is calculated as follows: ; (3) in, p i Indicates the first i The proportion of class samples; In addition, in step 3, the decision tree classification algorithm can also select the best split feature based on minimizing the Gini impurity, where the calculation formula of the Gini impurity is: ; (4) in, represents the Gini impurity of the dataset, p i Indicates the first i The ratio of class samples, n represents the number of types of samples in the dataset.
[0033] Information gain measures the degree to which a feature reduces the uncertainty of a dataset, while Gini impurity reflects the impurity of a dataset. In this embodiment, selecting the splitting feature with the highest information gain or lowest Gini impurity as the optimal splitting feature for the decision tree classification algorithm can effectively improve the accuracy of vegetation type classification.
[0034] Step 4: Update the administrative division information in the third basic land use data according to the administrative division data to obtain the fourth basic land use data.
[0035] In this example, spatial topology analysis and attribute matching algorithms were used in ArcGIS software using the Spatial Join tool to link the district and county administrative divisions in the land use basic data with the latest administrative divisions. Based on attributes such as the administrative division code and name, and incorporating topological relationships (such as intersection and inclusion), the administrative division information in the third land use basic data was updated to generate the fourth land use basic data, ensuring data consistency and timeliness.
[0036] Step 5: Calculate the vegetation coverage of each pixel in the area of interest based on the remote sensing image, and determine the vegetation coverage level attribute of each pixel based on the set threshold. The calculation formula of the vegetation coverage is: ; (5) in, FVC represents vegetation coverage, NDVI represents the normalized vegetation index of the pixel, NDVI soil and NDVI veg In this embodiment, the normalized vegetation index of each pixel is calculated by the normalized vegetation index of the pure soil pixel and the pure vegetation pixel in the area of interest. NDVI Normalization to obtain vegetation coverage can more significantly reflect the differences in vegetation coverage of each pixel.
[0037] Specifically, in this embodiment, the 10-meter NDVI data of the remote sensing image is first calculated, and then the vegetation cover is calculated using the 10-meter NDVI data. The effects of factors such as topography and atmosphere on the NDVI data are considered, and the remote sensing image is processed using preprocessing methods such as atmospheric correction and topographic correction. The normalized difference vegetation index of each pixel is then calculated to improve the accuracy of the NDVI data. The calculation formula for NDVI is as follows: ; (6) in, NIR represents the reflectivity in the near-infrared band, Red Indicates the reflectivity of the red light band.
[0038] Specifically, in step five, thresholds of five different levels are set, namely, high vegetation coverage, medium-high vegetation coverage, medium vegetation coverage, medium-low vegetation coverage, and low vegetation coverage. According to the set thresholds, the vegetation coverage level attributes of each pixel are obtained, namely, the vegetation coverage level data.
[0039] Step 6: Call the land use data, vegetation type data, ecosystem type data and vegetation coverage level data in the fourth land use basic data of the area of interest to generate land use status map, vegetation type map, ecosystem map and vegetation coverage map, such as Figures 3 to 6 shown.
[0040] Step 6 also includes creating cartographic templates for land use, vegetation type, ecosystem type, and vegetation coverage maps. These templates use different colors for different data types and include Chinese text to enhance map readability. These templates also feature a dynamic update function, automatically adjusting the display style and position of the title, north arrow, scale, and longitude / latitude grid based on the regional data range.
[0041] Specifically, in this example, the project extent and buffer zone are plotted using Python libraries such as Matplotlib and Basemap. Coordinate conversion tools are used to verify and convert coordinates to ensure consistency with the land use data. Regional data is overlaid with administrative division data to obtain specific district and county names, and corresponding land use, vegetation type, and ecosystem type data are retrieved based on these names.
[0042] Step 6 further includes calculating the area of each type of land use, vegetation, ecosystem, and vegetation coverage within the region of interest, and generating a statistical table. Specifically, the area calculation tool in ArcGIS software can be used in conjunction with the Python Pandas library for data processing to calculate the area of each type of land use, vegetation type, ecosystem type, and vegetation coverage within the region, and generate a statistical table to provide data support for subsequent analysis.
[0043] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for fusion processing of multi-source geographic information data for ecological impact assessment, characterized in that: The following steps are involved: Step 1: Obtain basic land use data and the latest multi-source geographic information data, including road network data, water system data, remote sensing satellite image data, and administrative division data; Step 2: Perform format conversion and coordinate correction on the basic land use data to obtain the first basic land use data; Step 3: Fuse and overlay the road network data and water system data with the first basic land use data to obtain the second basic land use data; construct vegetation type raster data based on the spectral characteristics and texture characteristics of the remote sensing satellite image data; obtain ecosystem type data based on the vegetation type raster data and the second basic land use data, and improve the attribute table of the second basic land use data based on the vegetation type raster data and the ecosystem type data to obtain the third basic land use data; Step 4: Update the administrative division information in the third basic land use data according to the administrative division data to obtain the fourth basic land use data; Step 5: Calculate the vegetation coverage of each pixel in the area of interest based on the remote sensing image, and determine the vegetation coverage level attribute of each pixel based on the set threshold. The calculation formula for vegetation coverage is: ; in, FVC Indicates the vegetation coverage of the pixel, NDVI represents the normalized vegetation index of the pixel, NDVI soil and NDVI veg They represent the normalized vegetation index of pure soil pixels and pure vegetation pixels in the area of interest respectively; Step 6: Call the land use data, vegetation type data, ecosystem type data and vegetation coverage level data of the focus area in the fourth land use basic data to generate a land use status map, a vegetation type map, an ecosystem map and a vegetation coverage map.
2. The method for fusion processing of multi-source geographic information data for ecological impact assessment according to claim 1, characterized in that: In the step 2, the specific method for coordinate correction of the basic land use data is: selecting multiple feature points in the basic land use data, obtaining the coordinates of the corresponding points in the remote sensing satellite image data, using the least squares method to fit and solve the affine transformation matrix, and performing coordinate correction on the basic land use data according to the affine transformation matrix.
3. The method for fusion processing of geographical information multi-source data for ecological impact assessment according to claim 1, characterized in that: In the step three, vegetation type data is constructed using a decision tree classification algorithm.
4. The method for fusion processing of multi-source geographic information data for ecological impact assessment according to claim 3, characterized in that: In step 3, the specific method of constructing vegetation type data using the decision tree classification algorithm is: First, spectral features and texture features are extracted from remote sensing satellite image data as input features; Then, a decision tree is generated by recursively partitioning the data set; classification rules are generated through the decision tree to obtain the vegetation type of each area.
5. The method for fusion processing of multi-source geographic information data for ecological impact assessment according to claim 3 is characterized in that: In step 3, the decision tree classification algorithm selects the best split feature based on maximizing information gain, where the information gain is calculated as: ; in, IG(D,A) Represents split features A The information gain of D is the number of samples in the data set. H(D) represents the entropy of the dataset, Indicates the value is v Subset of D v The number of samples in H(D v ) Representation subset D v The entropy of the data set is calculated as follows: ; in, p i Indicates the first i The ratio of class samples, n represents the number of types of samples in the dataset.
6. The method for fusion processing of multi-source geographic information data for ecological impact assessment according to claim 3, characterized in that: In step 3, the decision tree classification algorithm selects the best split feature based on minimizing the Gini impurity, where the calculation formula of the Gini impurity is: ; in, represents the Gini impurity of the dataset, p i Indicates the first i The ratio of class samples, n represents the number of types of samples in the dataset.
7. The method for fusion processing of geographical information multi-source data for ecological impact assessment according to claim 1, characterized in that: In step five, five different levels of thresholds are set, namely, high vegetation coverage, medium-high vegetation coverage, medium vegetation coverage, medium-low vegetation coverage, and low vegetation coverage. According to the set thresholds, the vegetation coverage level attributes of each pixel are obtained, namely, the vegetation coverage level data.
8. The method for fusion processing of multi-source geographic information data for ecological impact assessment according to claim 1, characterized in that: The step five also includes the step of calculating 10-meter NDVI data based on the remote sensing image. When calculating vegetation coverage, the 10-meter NDVI data is used as the normalized vegetation index value of the pixel.
9. The method for fusion processing of multi-source geographic information data for ecological impact assessment according to claim 1, characterized in that: The step six also includes the step of making mapping templates for the current land use map, vegetation type map, ecosystem type map, and vegetation coverage map. In the mapping templates, different colors are mapped to different types of data.
10. The method for fusion processing of geographical information multi-source data for ecological impact assessment according to claim 1, characterized in that: The step six also includes the following steps: calculating the area of various types of land use, vegetation, ecosystem, and vegetation coverage level in the area of interest, and generating a statistical table.
Citation Information
Patent Citations
Ecological space network optimization method
CN116629436A
Power transmission and transformation project ecological influence space-time boundary determination method and system based on multi-source remote sensing information deep fusion interpretation
CN119514870A