Administrative village residence identification method
By combining remote sensing image data with random forest algorithm and BP neural network model, the location of administrative villages can be identified, which solves the problems of inaccurate information and limited coverage of administrative village locations in existing technologies, and achieves accurate identification over a large area.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies suffer from problems such as large workload, limited coverage, data redundancy, and slow updates when acquiring information on the location of administrative villages, making it difficult to accurately identify the location of administrative villages.
Using remote sensing image data, we identified characteristic indices of administrative village locations by combining Sentinel-1 SAR, Sentinel-2 MSI, NPP VIIRS, CNBH building height, OSM road distribution, and DEM elevation data with random forest algorithm and BP neural network model, and trained land use type and administrative village location identification model.
It has enabled accurate identification of administrative village locations over a wide area, overcoming the problems of large workload in field surveys, limited coverage of authoritative data, and redundancy of map data, and providing relatively accurate data support for administrative village locations.
Smart Images

Figure CN121746912A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of remote sensing information extraction technology and discloses a method for identifying the location of administrative villages. Background Technology
[0002] The administrative village seat is generally located in the center of the administrative village area, characterized by a large population, good infrastructure, and well-developed industries. With the advancement of urbanization, rural populations have continuously migrated to cities, and administrative villages have undergone one or more mergers and reorganizations. Accurately identifying the administrative village seat provides basic data for understanding the rural settlement system, optimizing the village and town layout, formulating regional policies for rural revitalization, and conducting effectiveness assessments of rural revitalization efforts.
[0003] Currently, obtaining information on the locations of administrative villages relies heavily on methods such as on-site surveys, official data collection, and navigation map scraping. On-site surveys typically involve visiting villages in person to obtain relatively accurate locations, but this method is labor-intensive and has limited coverage. Official data collection often involves collaborating with relevant government departments, providing authoritative data, but this is limited by the channels through which these departments work. Navigation map scraping can yield data from a wide range of sources, but it suffers from data redundancy and slow updates. Summary of the Invention
[0004] This invention provides a method for identifying the location of administrative villages based on remote sensing images, which can effectively obtain relatively accurate information on the location of administrative villages over a large area.
[0005] This invention discloses a method for identifying the location of administrative villages based on remote sensing image data. The method includes the following steps:
[0006] Step S1: Obtain remote sensing image data of the year of interest for the area to be identified;
[0007] Step S2: Based on the remote sensing image data of the area to be identified, determine the list of land use types to be classified and mark typical samples of each land use type;
[0008] Step S3: Determine the characteristic index that distinguishes rural construction land from other land use types, and train the land use type classification model based on typical samples of various land use types to identify and select rural construction land patches.
[0009] Step S4: Mark typical administrative village seats in the rural construction land patches;
[0010] Step S5: Determine the characteristic index of the administrative village seat, and train the administrative village seat identification model based on the calibrated typical administrative village seat samples to identify the administrative village seat.
[0011] Furthermore, in step 1, the remote sensing image data includes Sentinel-1 SAR, Sentinel-2 MSI, NPP VIIRS image data, CNBH building height, OSM road distribution, and DEM elevation data;
[0012] Sentinel-1 SAR is VV and VH polarization band data of Sentinel-1 SAR image time series during the year of interest;
[0013] Sentinel-2 MSI consists of spectral band data at 10m spatial resolution and spectral band data at 20m spatial resolution from the Sentinel-2 multispectral instrument sensor during the year of interest, with cloud cover below 20%.
[0014] The NPP VIIRS imagery data consists of monthly visible light datasets (VIIRS-DNB) for the year of interest.
[0015] CNBH building heights are raster data of building heights at a resolution of 10 meters for the years of interest.
[0016] OSM road distribution is Open Street Map road data for the year of interest.
[0017] The DEM elevation data is SRTMGL1_003 data with a resolution of 30 meters for the year of interest.
[0018] Furthermore, in step 2, for Sentinel-2 MSI image data, based on the types of land features in the image, the land use type of the area to be identified is set as rural construction land, urban construction land, forest land, grassland, cultivated land, water body, industrial and mining land, and transportation land.
[0019] Typical samples of urban construction land were determined using a combination of area sampling and point sampling, while typical samples of other land use types were determined using point sampling.
[0020] Furthermore, step 3 also includes the following steps:
[0021] Step S31: Obtain the characteristic indices that distinguish rural construction land from other land use types, including spectral characteristics, temporal characteristics, textural characteristics, and socio-economic characteristics;
[0022] Step S32: Use four types of feature indices as input features for the random forest algorithm, and use the labeled land use types as training labels. The land use types include rural construction land, urban construction land, forest land, grassland, cultivated land, water bodies, industrial and mining land, and transportation land. Train the random forest algorithm model to generate a robust random forest land use classification model, determine the land use type of map patches, and finally select rural construction land patch data for further analysis.
[0023] Furthermore, in step 5, the following are selected as input features: distance from the main road, time cost to the county seat, patch area, building density within the patch, maximum building height within the patch, patch density, and patch perimeter-to-area ratio. The BP neural network model is trained using labeled administrative village seats and non-administrative village seats as output labels to generate an administrative village seat identification model.
[0024] Furthermore, in step 31, the spectral characteristic indices include the normalized vegetation index, the ratio of residential areas index, the vegetation red-edge water body index, and the enhanced impermeable surface index.
[0025] The vegetation index NDVI is:
[0026] ;
[0027] This refers to Sentinel-2. Band data, B8 band wavelength is 785~900 nm; This refers to Sentinel-2. Band data, B4 band wavelength is 650~680 nm;
[0028] The ratio of residential population index (RRI) is:
[0029] ;
[0030] This refers to Sentinel-2. Band data, B2 band wavelength is 458~523 nm. This refers to Sentinel-2. Band data, B8 band wavelength is 785~900 nm;
[0031] The vegetation red-edge water index (RWI) is:
[0032] ;
[0033] This refers to Sentinel-2. Band data, B3 band wavelength is 543~578 nm; This refers to Sentinel-2. Band data, B5 band wavelength is 698~713 nm; This refers to Sentinel-2. Band data, B8A band wavelength is 855~875 nm; This refers to Sentinel-2. Band data, B12 band wavelength is 2100~2280nm;
[0034] The enhanced impermeability index (ENDISI) is:
[0035] ;
[0036] This refers to Sentinel-2. Band data, B11 band wavelength is 1565~1655 nm.
[0037] Furthermore, in step 31, the time characteristic index includes the maximum vegetation index and the vegetation index range;
[0038] Maximum vegetation index for:
[0039] ;
[0040] This represents the NDVI dataset within the study period, where i represents the data from the first period within the study period, and j represents the data from the last period within the study period.
[0041] Extremely poor vegetation index for:
[0042] ;
[0043] This represents the maximum NDVI value during the study period. This represents the minimum NDVI value during the study period.
[0044] Furthermore, in step 31, the texture feature index includes the correlation and contrast of the gray-level co-occurrence matrix calculated using the VV and VH polarization bands of the Sentinel-1 SAR image, respectively.
[0045] The correlation of the VV band is:
[0046] ;
[0047] The correlation of the VH band is:
[0048] ;
[0049] The contrast ratio of the VV band is:
[0050] ;
[0051] The contrast of the VH band is:
[0052] ;
[0053] and This represents any grayscale value of a pixel in VV polarization band data. For an 8-level grayscale image... and The value range is 0~255; This represents the average value of the gray-level co-occurrence matrix in the VV polarization band; This represents the probability of pixel pairs with gray values i and j appearing in the gray-level co-occurrence matrix of the VV polarization band. The variance of the gray-level co-occurrence matrix in the VV polarization band is represented. This represents the maximum value of a k-level grayscale image.
[0054] Furthermore, in step 31, socioeconomic characteristics are reflected using nighttime light data, specifically the composite nighttime light data VIIRS-DNB data.
[0055] The beneficial effects achieved by this invention are:
[0056] This invention provides a method for identifying the location of administrative villages based on remote sensing images. It can effectively obtain relatively accurate information on the location of administrative villages over a large area, and to a certain extent overcomes the problems of large workload in field surveys, limited coverage of authoritative data, redundancy and offset of map crawling data. Attached Figure Description
[0057] Figure 1 This is a flowchart illustrating a method for identifying the location of an administrative village provided by the present invention. Detailed Implementation
[0058] The present invention will be further described below with reference to specific embodiments, and the advantages and features of the present invention will become clearer as a result. However, these embodiments are merely exemplary and do not constitute any limitation on the scope of the present invention. Those skilled in the art should understand that modifications or substitutions can be made to the details and form of the technical solutions of the present invention without departing from the spirit and scope of the present invention, but all such modifications and substitutions fall within the protection scope of the present invention.
[0059] As attached Figure 1As shown, the purpose of this invention is to provide a method for identifying the location of administrative villages based on remote sensing image data. This invention includes the following steps:
[0060] Step S1: Obtain remote sensing image data of the year of interest for the area to be identified;
[0061] The remote sensing image data includes Sentinel-1 SAR, Sentinel-2 MSI, NPP VIIRS image data, as well as CNBH building height, OSM road distribution, DEM elevation, and other data.
[0062] In one embodiment, Sentinel-1 SAR consists of VV and VH polarization band data from Sentinel-1 SAR image time series from January 1, 2022 to December 31, 2022; Sentinel-2 MSI consists of spectral band data (bands 2, 3, 4, and 8) and spectral band data (bands 8A, 11, and 12) at 10m spatial resolution from the Sentinel-2 Multispectral Instrument (MSI) sensor from January 1, 2022 to December 31, 2022, with cloud cover below 20%, and clouds and shadows removed using QA60; NPP VIIRS consists of monthly visible light dataset VIIRS-DNB data from January 1, 2022 to December 31, 2022; CNBH consists of China's first 10-meter resolution building height raster data (CNBH-10m) from 2020; and OSM consists of Open Street... Map data is road data; DEM is a 30-meter resolution digital elevation model SRTMGL1_003.
[0063] Step S2: Based on the remote sensing image data of the area to be identified, determine the list of land use types to be classified and mark typical samples of each land use type;
[0064] For Sentinel-2 MSI imagery data, based on the land cover types in the images, the land use types of the area to be identified were set into eight categories: rural construction land, urban construction land, forest land, grassland, cultivated land, water bodies, industrial and mining land, and transportation land. Urban construction land was sampled using a combination of area and point sampling methods, while other land use types were sampled using point sampling to determine typical samples for each of the eight land use types. The sample selection balanced uniformity and randomness.
[0065] Step S3: Determine the characteristic index that distinguishes rural construction land from seven other land use types, including urban construction land, forest land, grassland, cultivated land, water bodies, industrial and mining land, and transportation land. Train the land use type classification model based on typical samples of various land use types to identify rural construction land patches.
[0066] Specifically, it also includes the following steps:
[0067] Step S31: The purpose of this step is to obtain the characteristic index of land type patches that distinguish between rural construction land and urban construction land, forest land, grassland, cultivated land, water bodies, industrial and mining land, and transportation land, so as to more accurately identify various land type patches and obtain rural construction land from them. The characteristic index includes four categories: spectral characteristics, temporal characteristics, texture characteristics, and socio-economic characteristics.
[0068] Spectral characteristic indices include Normalized Difference Vegetation Index (NDVI), Ratio Residential Index (RRI), Red Edge Vegetation Water Index (RWI), and Enhanced Impermeable Surface Index (ENDISI).
[0069] Among them, the NDVI vegetation index reflects vegetation coverage and growth vitality; a higher value indicates more lush vegetation, while this index is lower in the construction land area. The Ratio Residential Land Index (RRI) enhances remote sensing information on impermeable surfaces such as residential areas and building land, while mitigating interference from other land types such as bare land, vegetation, and water bodies, achieving high accuracy in extracting residential land information. The RWI red-edge water body index amplifies the distinction between water bodies and non-water bodies, enabling water body identification. The ENDISI impermeable surface index efficiently distinguishes impermeable surfaces from bare soil, accurately extracting impermeable surfaces.
[0070] The vegetation index NDVI is:
[0071]
[0072] This refers to Sentinel-2. Band data, B8 band wavelength is 785~900 nm; This refers to Sentinel-2. Band data, B4 band wavelength is 650~680 nm;
[0073] The ratio of residential population index (RRI) is:
[0074]
[0075] This refers to Sentinel-2. Band data, B2 band wavelength is 458~523 nm. This refers to Sentinel-2. Band data, B8 band wavelength is 785~900 nm;
[0076] The Red Edge Water Index (RWI) is:
[0077]
[0078] This refers to Sentinel-2. Band data, B3 band wavelength is 543~578 nm; This refers to Sentinel-2. Band data, B5 band wavelength is 698~713 nm; This refers to Sentinel-2. Band data, B8A band wavelength is 855~875 nm; This refers to Sentinel-2. Band data, B12 band wavelength is 2100~2280nm;
[0079] The impermeability index (ENDISI) is:
[0080]
[0081] This refers to Sentinel-2. Band data, B11 band wavelength is 1565~1655 nm;
[0082] Time characteristic indices include the maximum value of NDVI ( ) and NDVI range ( ).
[0083] in, and It represents the maximum value and range of variation of the vegetation index, used to help distinguish between construction land and vegetation.
[0084] Maximum vegetation index for:
[0085]
[0086] This represents the NDVI dataset within the study period, where i represents the data from the first period within the study period, and j represents the data from the last period within the study period.
[0087] Extremely poor vegetation index for:
[0088]
[0089] This represents the maximum NDVI value during the study period. This represents the minimum NDVI value during the study period;
[0090] Texture feature indices include the correlation and contrast of the gray-level co-occurrence matrix (GLCM) of Sentinel-1 SAR images calculated using VV and VH polarization bands, respectively.
[0091] Among them, VV and VH contrast reflect the "separation degree" of the differences in ground object scattering. The higher the contrast, the more complex the ground object scattering mechanism, such as vegetation and rough ground. The lower the contrast, the closer the polarization scattering signals are and the simpler the ground object scattering mechanism, such as smooth water bodies, flat bare soil, and mean-value building areas.
[0092] The correlation between VV and VH reflects the "consistency" of ground object scattering. The higher the correlation, the stronger the correlation of the scattered signals and the more uniform the ground object surface; the weaker the correlation, the rougher the ground object surface.
[0093] The correlation of the VV band is:
[0094]
[0095] and This represents any grayscale value of a pixel in VV polarization band data. For an 8-level grayscale image... and The value range is 0~255; This represents the average value of the gray-level co-occurrence matrix in the VV polarization band; This represents the probability of pixel pairs with gray values i and j appearing in the gray-level co-occurrence matrix of the VV polarization band. The variance of the gray-level co-occurrence matrix in the VV polarization band is represented. This represents the maximum value of a k-level grayscale image.
[0096] The correlation of the VH band is:
[0097]
[0098] The main parameters in the formula are similar to those calculated for the VV band.
[0099] The contrast ratio of the VV band is:
[0100]
[0101] The main parameters in the formula are calculated similarly to those in the VV band.
[0102] The contrast of the VH band is:
[0103]
[0104] The main parameters in the formula are calculated similarly to those in the VV band.
[0105] Socioeconomic characteristics are reflected using nighttime light data. The nighttime light data is composite nighttime light data (VIIRS-DNB data) obtained from the monthly VIIRS-DNB VI dataset using the median synthesis method.
[0106] Nighttime lighting reflects nighttime brightness. Building areas experience high levels of human and socio-economic activity at night, resulting in higher brightness levels.
[0107] Spectral features reflect the physicochemical properties of land features, temporal features capture the seasonal variation patterns of land features, texture features describe the spatial arrangement patterns of land features and distinguish structural complexity, and socioeconomic features can further distinguish between urban and non-urban areas and improve the identification accuracy of rural construction land categories.
[0108] Step S32: Using 11 indicators from 4 types of feature indices as input features for the random forest algorithm, and using known rural construction land, urban construction land, forest land, grassland, cultivated land, water bodies, industrial and mining land, and transportation land as training labels, the random forest algorithm model is trained and a robust random forest land use classification model is generated to determine the land use type of map patches. Finally, rural construction land map patch data for further analysis is selected.
[0109] The Random Forest algorithm constructs optimal discrimination rules by automatically evaluating the importance of various features, efficiently transforming a complex feature space into classification results. The calculation of various indicators is a necessary means to capture the multidimensional characteristics of ground objects and improve classification accuracy.
[0110] Step S4: Mark typical administrative village seats in the rural construction land patches;
[0111] Rural construction land patches were extracted from the land use type identification results, and administrative village seat patches and non-administrative village seat patches were labeled using high-resolution images of the corresponding years. The sample selection took into account both regional uniformity and randomness.
[0112] Rural construction land patches include patches located at the administrative village seats and patches not located at administrative village seats. Patches located at administrative village seats are generally larger in area, have better construction conditions, and a higher degree of housing concentration than other surrounding patches. Based on the patch characteristics on the map and the marking of village seats, the sample patches are distinguished.
[0113] Step S5: Determine the characteristic index of the administrative village seat, and train the administrative village seat identification model based on the calibrated typical administrative village seat samples to identify the administrative village seat.
[0114] The purpose of this step is to select the land use patches containing administrative village seats from the identified rural construction land use patches. To more accurately identify administrative village seat patches, seven feature indices are selected as input data: distance from main roads, time cost to the county seat, patch area, building density within the patch, maximum building height within the patch, patch density, and patch perimeter-to-area ratio. Using the labeled administrative village seats and non-administrative village seats as output labels, a BP neural network model is trained to generate an administrative village seat identification model.
[0115] A BP neural network model was selected as the administrative village location identification model. The sample data was randomly divided into a training group and a validation group at a ratio of 8:2. The former was used to train the model, and the latter was used to evaluate the classification accuracy.
[0116] The neural network utilizes the differences in seven feature indices of the samples to construct the optimal discrimination rule for distinguishing between administrative village resident patches and non-administrative village resident patches, and then classifies and identifies all rural construction land patches in the study area.
[0117]
[0118] The above are merely preferred embodiments of the present invention and do not constitute any limitation on the scope of protection of the present invention; all technical solutions formed by equivalent transformations or equivalent substitutions fall within the scope of protection of the present invention; the parts of the present invention not described in detail are well-known technologies to those skilled in the art.
Claims
1. A method for identifying the location of administrative villages based on remote sensing image data, characterized in that, The method for identifying the location of administrative villages based on remote sensing image data includes the following steps: Step S1: Obtain remote sensing image data of the year of interest for the area to be identified; Step S2: Based on the remote sensing image data of the area to be identified, determine the list of land use types to be classified and mark typical samples of each land use type; Step S3: Determine the characteristic index that distinguishes rural construction land from other land use types, and train the land use type classification model based on typical samples of various land use types to identify and select rural construction land patches. Step S4: Mark typical administrative village seats in the rural construction land patches; Step S5: Determine the characteristic index of the administrative village seat, and train the administrative village seat identification model based on the calibrated typical administrative village seat samples to identify the administrative village seat.
2. The method for identifying the location of administrative villages based on remote sensing image data according to claim 1, characterized in that, In step 1, the remote sensing image data includes Sentinel-1 SAR, Sentinel-2 MSI, NPP VIIRS image data, CNBH building height, OSM road distribution, and DEM elevation data; Sentinel-1 SAR is VV and VH polarization band data of Sentinel-1 SAR image time series during the year of interest; Sentinel-2 MSI consists of spectral band data at 10m spatial resolution and spectral band data at 20m spatial resolution from the Sentinel-2 multispectral instrument sensor during the year of interest, with cloud cover below 20%. The NPP VIIRS imagery data consists of monthly visible light datasets (VIIRS-DNB) for the year of interest. CNBH building heights are raster data of building heights at a resolution of 10 meters for the years of interest. OSM road distribution is Open Street Map road data for the year of interest. The DEM elevation data is SRTMGL1_003 data with a resolution of 30 meters for the year of interest.
3. The method for identifying the location of administrative villages based on remote sensing image data according to claim 2, characterized in that, In step 2, for Sentinel-2 MSI image data, the land use type of the area to be identified is set as rural construction land, urban construction land, forest land, grassland, cultivated land, water body, industrial and mining land, and transportation land according to the land cover type in the image; Typical samples of urban construction land were determined using a combination of area sampling and point sampling, while typical samples of other land use types were determined using point sampling.
4. The method for identifying the location of administrative villages based on remote sensing image data according to claim 1, characterized in that, Step 3 also includes the following steps: Step S31: Obtain the characteristic indices that distinguish rural construction land from other land use types, including spectral characteristics, temporal characteristics, textural characteristics, and socio-economic characteristics; Step S32: Use four types of feature indices as input features for the random forest algorithm, and use the labeled land use types as training labels. The land use types include rural construction land, urban construction land, forest land, grassland, cultivated land, water bodies, industrial and mining land, and transportation land. Train the random forest algorithm model to generate a robust random forest land use classification model, determine the land use type of map patches, and finally select rural construction land patch data for further analysis.
5. The method for identifying the location of administrative villages based on remote sensing image data according to claim 1, characterized in that, In step 5, the following are selected as input features: distance from the main road, time cost to the county seat, patch area, building density within the patch, maximum building height within the patch, patch density, and patch perimeter-to-area ratio. The BP neural network model is trained using labeled administrative village seats and non-administrative village seats as output labels to generate an administrative village seat identification model.
6. The method for identifying the location of administrative villages based on remote sensing image data according to claim 4, characterized in that, In step 31, the spectral characteristic indices include the normalized vegetation index, the ratio of residential area index, the vegetation red-edge water body index, and the enhanced impermeable surface index. The vegetation index NDVI is: ; This refers to Sentinel-2. Band data, B8 band wavelength is 785~900 nm; This refers to Sentinel-2. Band data, B4 band wavelength is 650~680 nm; The ratio of residential population index (RRI) is: ; This refers to Sentinel-2. Band data, B2 band wavelength is 458~523 nm. This refers to Sentinel-2. Band data, B8 band wavelength is 785~900 nm; The vegetation red-edge water index (RWI) is: ; This refers to Sentinel-2. Band data, B3 band wavelength is 543~578 nm; This refers to Sentinel-2. Band data, B5 band wavelength is 698~713 nm; This refers to Sentinel-2. Band data, B8A band wavelength is 855~875 nm; This refers to Sentinel-2. Band data, B12 band wavelength is 2100~2280 nm; The enhanced impermeability index (ENDISI) is: ; This refers to Sentinel-2. Band data, B11 band wavelength is 1565~1655 nm.
7. The method for identifying the location of administrative villages based on remote sensing image data according to claim 4, characterized in that, In step 31, the time characteristic index includes the maximum vegetation index and the vegetation index range; Maximum vegetation index for: ; This represents the NDVI dataset within the study period, where i represents the data from the first period within the study period, and j represents the data from the last period within the study period. Extremely poor vegetation index for: ; This represents the maximum NDVI value during the study period. This represents the minimum NDVI value during the study period.
8. The method for identifying the location of administrative villages based on remote sensing image data according to claim 4, characterized in that, In step 31, the texture feature index includes the correlation and contrast of the gray-level co-occurrence matrix calculated using the VV and VH polarization bands of Sentinel-1 SAR images, respectively. The correlation of the VV band is: ; The correlation of the VH band is: ; The contrast ratio of the VV band is: ; The contrast of the VH band is: ; and This represents any grayscale value of a pixel in VV polarization band data. For an 8-level grayscale image... and The value range is 0~255; This represents the average value of the gray-level co-occurrence matrix in the VV polarization band; This represents the probability of pixel pairs with gray values i and j appearing in the gray-level co-occurrence matrix of the VV polarization band. The variance of the gray-level co-occurrence matrix in the VV polarization band is represented. This represents the maximum value of a k-level grayscale image.
9. The method for identifying the location of administrative villages based on remote sensing image data according to claim 4, characterized in that, In step 31, socioeconomic characteristics are reflected using nighttime light data, specifically the composite nighttime light data VIIRS-DNB.