A global remote sensing image-text retrieval sample construction method and system
By combining spatial autocorrelation analysis and landscape indices of nighttime light remote sensing data with Open Street Map tags, remote sensing image and text retrieval samples are automatically generated. This solves the problem of uneven sampling of remote sensing datasets worldwide and provides efficient sample acquisition and training datasets for deep learning models.
Patent Information
- Application Number
- CN202310674367.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-07
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2043-06-07
AI Technical Summary
Existing remote sensing dataset sampling methods suffer from uneven spatial distribution of samples globally, complex and diverse terrains, and significant differences in urban construction styles, making sample acquisition difficult. Furthermore, deep learning model training relies heavily on manual annotation and computation, lacking reliable data support.
By analyzing the spatial autocorrelation of nighttime light remote sensing data, significant development areas are delineated using global and local Moran indices. Combined with landscape indices and Open Street Map labels, remote sensing image retrieval samples are automatically generated, reducing manual annotation workload and improving the uniformity and complexity of sample distribution.
It achieves a reliable distribution of sample points globally, reduces the workload of manual annotation, provides a high-quality training dataset for deep learning models, and supports multimodal research on remote sensing images.
Smart Images

Figure CN116861024B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and system for constructing global land use map and text retrieval samples, aiming to provide reliable data sample collection points for tasks such as remote sensing image retrieval and classification. Background Technology
[0002] In recent years, with the continuous development of remote sensing technology, the acquisition and application of remote sensing data have become increasingly widespread. In many remote sensing applications, such as land cover classification, urbanization monitoring, and water resource management, dataset construction is a crucial step. Especially for tasks utilizing deep learning for remote sensing land cover classification, target detection, scene retrieval, and change monitoring, a sufficient quantity and quality of training samples are required to train the model. Sample sampling is a vital step in dataset construction; the quality of sample sampling directly affects the quality of the dataset and its application effectiveness, thus impacting the interpretation of remote sensing images.
[0003] Although remote sensing imagery data is readily available and a wide variety of datasets are available, most existing datasets rely on sampling in localized areas by production units or research projects. This results in uneven spatial distribution of samples and insufficient representation of the geographical characteristics of different regions globally. Furthermore, obtaining reliable samples from diverse terrains and urban development styles across the globe presents significant challenges.
[0004] Currently, there are two main methods for constructing datasets in the field of remote sensing. The first is to select sampling points, acquire remote sensing satellite images of the corresponding areas, and then create a dataset through manual annotation. Traditional sampling methods in this category include random sampling, uniform sampling, and balanced sampling. These methods can quickly and directly generate sampling areas, but the data obtained relies heavily on manual annotation. Additionally, some sampling methods based on prior information, such as land cover classification and landscape index-based sampling, can provide some image category information, reducing the workload of manual annotation. However, globally, the amount of data and computation required to directly determine a large number of sampling points based on this prior information is enormous. The second method is sampling based on deep learning models, acquiring or generating remote sensing images using pre-trained deep learning models. Typical methods in this category include generative methods and transfer learning methods. These methods can automatically annotate remote sensing images or directly generate images based on a pre-trained deep learning model. However, training deep learning models requires a large amount of reliable data. Furthermore, remote sensing images generated through deep learning often lack geographic coordinate information and suffer from poor image quality; annotation of images through transfer learning is limited by the sample categories in the training dataset. Currently, there is no practical deep learning model that can sample images globally. Summary of the Invention
[0005] To overcome the above problems, this invention proposes a method and system for constructing global land use map and text retrieval samples.
[0006] The present invention proposes a method for constructing a global land use map and text retrieval sample, comprising the following steps:
[0007] Step 1: Obtain global nighttime light remote sensing data, perform spatial autocorrelation analysis based on the global administrative region nighttime light remote sensing index, divide administrative regions with significant development levels, and obtain regions with significant clustering or dispersion.
[0008] Step 2: Based on global land cover, first calculate the landscape index LSI, and then calculate the regional landscape index rLSI, the regional land cover category landscape index cLSI, and the landscape index uLSI for each geographic sampling unit in the selected salient areas. This will give us the number of sampling points in each area, the number of samples for each land cover type in each area, and the location of each sampling point, thus obtaining the final sampling distribution of the sample points.
[0009] Step 3: Based on the selected sample points, crop out the image data that matches the geographical location of the sample points to generate sample data, and use the Open Street Map labels of the corresponding areas as text to construct a text-image retrieval sample dataset.
[0010] Furthermore, based on the global Moran index, regions with spatial clustering or dispersion in urban development worldwide are identified. When the global Moran index I > 0 for global regional nighttime light remote sensing data and passes the Z-statistic test, it indicates a significant clustering trend. Then, the local Moran index and Z-statistic test for each region are calculated. Regions that pass the Z-statistic test are considered to have significant clustering or dispersion.
[0011] The formula for the global Moran index is as follows.
[0012]
[0013] Among them, z i The nighttime light remote sensing value x for region i i Its average The deviation, where n is the total number of regions, w i,j This refers to the spatial weight between regions i and j, which is generated based on adjacency relationships. When regions i and j are adjacent, w... i,j =1, otherwise w i,j =0; S0 is the aggregation of all spatial weights, as shown in the following formula.
[0014]
[0015] Furthermore, the local Moran index I i The calculation formula is as follows:
[0016]
[0017] Where, x i It is the nighttime light remote sensing value of region i. w represents the average value of nighttime light remote sensing data for the entire region. i,j Let be the spatial weight between regions i and j, and n be the total number of regions. Furthermore,
[0018]
[0019] The significance of the local Moran's index is tested using the Z-statistic, which is calculated using the following formula.
[0020]
[0021] Among them, E(I i ) represents I i The expectation, VAR(I) i ) represents I i To determine whether the null hypothesis holds, we calculate the variance of Z. i P statistic i value and significance level Compare and take P i The formula for calculating the value is as follows:
[0022]
[0023] in, If the standard normal distribution function is given, then... This indicates that the region exhibits a certain clustering or dispersion trend;
[0024] For region i that passes the significance test, according to Z i and spatial lag value S i Region i can be divided into four patterns: high-high clustering, high-low clustering, low-high clustering, and low-low clustering, wherein the spatial lag value S i The calculation formula is as follows:
[0025]
[0026] Among them, Z j Let w be the Z value of region j. i,j It represents the spatial weight between regions i and j.
[0027] When Z i >0 and S i >0, region i is a high-level cluster; Zi <0 and S i >0, region i is a low-high cluster; Z i >0 and S i <0, region i is a high-low cluster; Z i <0 and S i <0, region i is a low-low cluster.
[0028] Furthermore, the Landscape Index (LSI) is the ratio of landscape perimeter to area within a certain range, quantitatively representing the landscape heterogeneity of that area. The formula is as follows:
[0029]
[0030] Where q is the number of pixels, b j Let be the number of sides of pixel j. The image is raster data, and each pixel is equivalent to a small square. If pixel i and pixel j are adjacent and of different types, there is an edge between the two pixels. If pixel i and pixel j are adjacent but of the same type, there is no edge between the pixels.
[0031] Furthermore, the landscape index of a certain area is denoted as rLSI, and the number of sampling points in the area is N. i Related to the regional landscape index rLSI, the calculation formula is as follows:
[0032]
[0033] Among them, S i and S j Let i and j represent the areas of regions i and j respectively, N be the total number of samples, and n be the total number of regions.
[0034] Furthermore, the landscape index for a specific land cover type in a given region is denoted as cKSI. The sampling quantity for each sample is determined based on the landscape index cLSI for each land cover type. The calculation formula for the sampling quantity for category k in region i is as follows:
[0035]
[0036] Among them, cN i,k N represents the number of samples of category k in region i. i W represents the number of sampling points in the region. i,k Let m be the area percentage of category k in region i, and m be the total number of categories.
[0037] Furthermore, the uLSI geographic sampling unit adaptively selects the location of sample points. Assuming that region i consists of A×B geographic units, landscape indices for various land cover categories can be calculated within each geographic unit. The uLSI formula for the k-th class in row a and column b is as follows.
[0038]
[0039] Calculate the uLSI of each unit, then sort the geographic units according to the uLSI to obtain a distribution curve. The x-axis represents the geographic unit coordinates, and the y-axis represents... For category k in region i, the distribution curve Remove regions with values of 0, then divide cN on the x-axis. i,k Divide the region into equal parts, and randomly select a sampling point in each interval to form the sampling points for category k in region i.
[0040] This invention also provides a global remote sensing image and text retrieval sample construction system, comprising the following modules:
[0041] The data acquisition module is used to acquire global nighttime light remote sensing data, perform spatial autocorrelation analysis based on the global administrative region nighttime light remote sensing index, divide administrative regions with significant development levels, and obtain regions with significant clustering or dispersion.
[0042] The sample point determination module is used to first calculate the landscape index LSI based on global land surface cover, and then calculate the regional landscape index rLSI, the regional land cover category landscape index cLSI, and the landscape index uLSI for each geographic sampling unit in the selected salient areas. This yields the number of sampling points in each area, the number of samples for each land cover type in each area, and the location of each sampling point, resulting in the final sampling distribution of the sample points.
[0043] The dataset construction module is used to generate sample data by cropping image data that matches the geographical location of the selected sample points, and to construct a text-image retrieval sample dataset using the Open Street Map labels of the corresponding areas as text.
[0044] Furthermore, the landscape index of a certain area is denoted as rLSI, and the number of sampling points in the area is N. i Related to the regional landscape index rLSI, the calculation formula is as follows:
[0045]
[0046] Among them, L i and S j Let i and j represent the areas of regions i and j respectively, N be the total number of samples, and n be the total number of regions.
[0047] Furthermore, the landscape index for a specific land cover type in a given region is denoted as cLSI. The sampling quantity for each sample is determined based on the cLSI. The calculation formula for the sampling quantity of category k in region i is as follows:
[0048]
[0049] Among them, cN i,k N represents the number of samples of category k in region i. i W represents the number of sampling points in the region. i,k Let m be the area percentage of category k in region i, and m be the total number of categories.
[0050] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:
[0051] 1. On a global scale, based on prior information such as nighttime light remote sensing data and global land cover, spatial autocorrelation analysis can be used to better delineate regions with different development patterns, providing a reliable basis for the selection of sampling areas worldwide.
[0052] 2. Using three landscape indices to provide a basis for the selection of sample points avoids uneven sample types and can obtain more complex spatial features, which is convenient for training deep learning models.
[0053] 3. Combining Open Street Map (OSM) as image labels reduces the huge amount of manual annotation work and provides a feasible dataset construction model for multimodal research of remote sensing images and text. Attached Figure Description
[0054] Figure 1 A global Moran's index test report generated for an embodiment of the present invention;
[0055] Figure 2 This is a schematic diagram (a) showing that region i is divided into A×B geographic units in an embodiment of the present invention, and the landscape index of each land cover category calculated in each geographic unit.
[0056] Figure 3 This is a schematic diagram of random sampling of sampling points in an embodiment of the present invention;
[0057] Figure 4 This is a schematic diagram of the sampling location (a) and sampling point (b) of category k in region i in an embodiment of the present invention. Detailed Implementation
[0058] To better understand the technical solution of the present invention, the technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0059] This invention provides a method for constructing a global land use map and text retrieval sample, comprising the following steps:
[0060] Step 1 involves dividing the acquired global nighttime light remote sensing data into regions based on global administrative boundaries. Nighttime light remote sensing data is generally proportional to regional development levels; therefore, spatial autocorrelation analysis can be performed on the nighttime light remote sensing indices of global administrative regions to identify administrative regions with significant development levels. This spatial autocorrelation analysis includes the global Moran index and local Moran indices. The global Moran index reveals spatial clustering or dispersion in urban development globally. To identify which specific administrative regions are significantly clustered or dispersed, the local Moran index and Z-statistic test are calculated for each region. Regions that pass the Z-statistic test are considered to have significant clustering or dispersion.
[0061] The formula for the global Moran index is as follows.
[0062]
[0063] Among them, z i The nighttime light remote sensing value x for region i i Its average The deviation, where n is the total number of regions, w i,j This is the spatial weight between regions i and j. The spatial weight is generated based on adjacency relationships; when regions i and j are adjacent, w... i,j =1, otherwise w i,j =0. S0 is the aggregation of all spatial weights, as shown in the following formula.
[0064]
[0065] The significance of the Moran's index is generally tested using the Z-statistic, which is calculated using the following formula.
[0066]
[0067] Where E(I) represents the expected value of I, and VAR(I) represents the variance of I. To determine whether the null hypothesis holds, the p-value of the Z-statistic needs to be calculated and compared with the significance level. Comparison. Generally, take... The formula for calculating the P-value is as follows:
[0068]
[0069] in, This is the standard normal distribution function. If... This indicates the existence of a certain clustering or dispersion trend in the space. According to calculations, the global Moran's index I > 0 for global regional nighttime light remote sensing data and passed the Z-statistic test, indicating a significant clustering trend. Based on this, to further identify regions with significant clustering, local Moran's index calculations are needed.
[0070] The formula for calculating the local Moran index is as follows.
[0071]
[0072] Where, x i It is the nighttime light remote sensing value of region i. w represents the average value of nighttime light remote sensing data for the entire region. i,j Let be the spatial weight between regions i and j, and n be the total number of regions. Furthermore,
[0073]
[0074] The significance of the local Moran's index is tested using the Z-statistic, which is calculated using the following formula.
[0075]
[0076] Among them, E(I i ) represents I i The expectation, VAR(I) i ) represents I i The variance of Z. To determine whether the null hypothesis is true, Z needs to be calculated. i P statistic i value and significance level Comparison. Generally, take... P i The formula for calculating the value is as follows:
[0077]
[0078] in, This is the standard normal distribution function. If... This indicates that the region exhibits a certain clustering or dispersion trend.
[0079] For the local Moran index, if Z i If Z is a positive value, it indicates that the surrounding elements have similar values (high or low). i If it is a negative value, it indicates that there is a statistically significant outlier (high value surrounding low value or low value surrounding high value).
[0080] The local Moran index can more clearly identify the specific clustering type for each region. For region i that passes the significance test, according to Z... i and spatial lag value S i Region i can be divided into four patterns: high-high clustering, high-low clustering, low-high clustering, and low-low clustering. The spatial lag value S... i The calculation formula is as follows:
[0081]
[0082] Among them, Z j Let w be the Z value of region j. i,j It represents the spatial weight between regions i and j.
[0083] When Z i >0 and i >0, region i is a high-level cluster; Z i <0 and i >0, region i is a low-high cluster; Z i >0 and i <0, region i is a high-low cluster; Z i <0 and i <0, region i is a low-low cluster.
[0084] Step 2: Based on global land cover, calculate the regional landscape index rLSI, the regional land cover category landscape index cLSI, and the landscape index uLSI for each geographic sampling unit in the selected significant areas. This will yield the number of sampling points in each area, the number of samples for each land cover type in each area, and the location of each sampling point, thus obtaining the final sampling distribution of the samples.
[0085] The Landscape Index (LSI) is the ratio of landscape perimeter to area within a certain range, quantitatively representing the landscape heterogeneity of that area. The formula is as follows.
[0086]
[0087] Where q is the number of pixels, b j Let be the number of sides of pixel j. The image is raster data, and each pixel can be considered as a small square. If pixels i and j are adjacent and of different types, there is an edge between them; if pixels i and j are adjacent but of the same type, there is no edge between them.
[0088] The landscape index of a certain area is denoted as rLSI, and the number of sampling points in the area is N. i Related to the regional landscape index rLSI, the calculation formula is as follows:
[0089]
[0090] Among them, S i and S j Let i and j represent the areas of regions i and j respectively, and N be the total number of samples.
[0091] The landscape index for a specific land cover type in a region is denoted as cLSI. The sampling quantity for each sample is determined based on the cLSI of the land cover category. The calculation formula for the sampling quantity of category k in region i is as follows.
[0092]
[0093] Among them, cN i,k N represents the number of samples of category k in region i. i W represents the number of sampling points in the region. i,k Let m be the area percentage of category k in region i, and m be the total number of categories.
[0094] The landscape index uLSI of the geographic sampling unit adaptively selects the location of the sample points. Assuming that region i consists of A×B geographic units, the landscape index of each land cover category can be calculated in each geographic unit. The uLSI formula for the k-th class in row a and column b is as follows.
[0095]
[0096] Calculate the uLSI of each unit, then sort the geographic units according to the uLSI to obtain a distribution curve. The x-axis represents the geographic unit coordinates, and the y-axis represents... For category k in region i, the distribution curve Remove regions with values of 0, then divide cN on the x-axis. i,k Divide the region into equal parts, and randomly select a sampling point in each interval to form the sampling points for category k in region i.
[0097] Step 3: Based on the selected sampling points, crop out the image data that matches the geographical location of the sample points to generate sample data, and use the Open Street Map (OSM) labels of the corresponding areas as text to construct a text-image retrieval sample dataset.
[0098] The specific implementation steps of this invention are illustrated below through a concrete example:
[0099] Step 1. Identify regions with significant global development.
[0100] First, prepare global nighttime light remote sensing data, such as VIIRS nighttime light remote sensing data, and global city administrative area boundary vector data, such as Databse of Global Administrative Areas (GADM) data. Import the nighttime light remote sensing raster data V and the global city administrative area boundary vector data B into ArcGIS Pro. After converting the raster data to polygon data and performing raster-to-polygon operations, perform an intersection operation with the administrative boundary data to obtain I, which yields the nighttime light remote sensing values for each administrative area. Select the Spatial Autocorrelation Analysis (Global Moran's I) tool in the ArcGIS Pro toolbox to perform a global Moran's index test on the nighttime light remote sensing value field of the input feature class I, generating a report as attached. Figure 1It can be seen that global nighttime light remote sensing data exhibits a significant clustering pattern, thus allowing for the calculation of the local Moran's index. Using the Clustering and Outlier Analysis (Anselin Local Moran's I) tool in the ArcGIS Pro toolbox, a local Moran's index analysis is performed on the input feature class I, outputting the Z-index for the local region. i P i and I i With Z i x-axis Using the y-axis as the axis, the planar region is divided into four quadrants, corresponding to four spatial structures. When And Z i When >0, the corresponding region i is a high-gathering region; when And Z i When <0, the corresponding region i is a low-high aggregation region; when And Z i When >0, the corresponding region i is a high-low clustering region; when And Z i When <0, the corresponding region i is a low-low clustering region; therefore, each administrative region is determined based on the output Z. i P i and I i This allows us to obtain the spatial structure of development in each region. Regions with significant spatial structures are selected as representative sampling areas.
[0101] Step 2. Determining the sampling points
[0102] Based on the sampling area determined in step 1, prepare the corresponding land use cover data, such as the FROM-GLC10m (2017) global 10m land use cover dataset. Assuming a total of N sampling points are needed, first calculate the regional landscape index rLSI for each region. i Determine the number of samples N for each region. i Within each region, the land cover category landscape index (cLSI) is calculated. i Determine the number of samples of category k in region i, cN i,k Assume region i can be divided into A×B geographical units, such as... Figure 2 As shown in (a), landscape indices for each land cover category can be calculated within each geographic unit. like Figure 2 As shown in (b). The categories k in region i are sorted from largest to smallest. By sorting, a distribution curve can be obtained. The x-axis represents the geographic unit coordinates, and the y-axis represents... For category k in region i, the distribution curve Remove regions with values of 0, then divide cN on the x-axis. i,kDivide the interval equally, and randomly select a sampling point in each interval, such as Figure 3 As shown. This forms the sampling location and sampling point for category k in region i, as follows: Figure 4 As shown in (a)(b).
[0103] Step 3. Construction of Image and Text Retrieval Samples
[0104] Based on the acquired sampling points, remote sensing images of the corresponding regions and Open StreetMap (OSM) labels of the corresponding regions were selected as sample data to form a global image and text retrieval remote sensing dataset.
[0105] In practice, the present invention can be implemented using computer software technology to automate the process, and the device that runs the process of the present invention should also be within the scope of protection.
[0106] This invention also provides a global remote sensing image and text retrieval sample construction system, comprising the following modules:
[0107] The data acquisition module is used to acquire global nighttime light remote sensing data, perform spatial autocorrelation analysis based on the global administrative region nighttime light remote sensing index, divide administrative regions with significant development levels, and obtain regions with significant clustering or dispersion.
[0108] The sample point determination module is used to first calculate the landscape index LSI based on global land surface cover, and then calculate the regional landscape index rLSI, the regional land cover category landscape index cLSI, and the landscape index uLSI for each geographic sampling unit in the selected salient areas. This yields the number of sampling points in each area, the number of samples for each land cover type in each area, and the location of each sampling point, resulting in the final sampling distribution of the sample points.
[0109] The dataset construction module is used to generate sample data by cropping image data that matches the geographical location of the selected sample points, and to construct a text-image retrieval sample dataset using the Open Street Map labels of the corresponding areas as text.
[0110] The specific implementation methods of each module are the same as those of each step, and will not be described in this invention.
[0111] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.
Claims
1. A method for constructing a global remote sensing image and text retrieval sample, characterized in that, Includes the following steps: Step 1: Obtain global nighttime light remote sensing data, perform spatial autocorrelation analysis based on the global administrative region nighttime light remote sensing index, divide administrative regions with significant development levels, and obtain regions with significant clustering or dispersion. The spatial autocorrelation analysis consists of the global Moran index and the local Moran index; Step 2: Based on global land cover, first calculate the landscape index. Then, the regional landscape index is calculated for each of the selected prominent areas. , regional land cover category landscape index Landscape index for each geographic sampling unit The number of sampling points in each region, the number of samples for each type of land cover in each region, and the location of each sampling point are obtained respectively, thus obtaining the final sampling distribution of the sample points; The landscape index of a certain area is denoted as... Number of sampling points in the area With the aforementioned regional landscape index The relevant calculation formula is as follows: in, and Representing regions and area, For the total sample size, Total number of regions; The landscape index of a certain land cover type in a certain region is denoted as... According to the landscape index of land cover category Determine the sampling quantity for each sample, for the region Medium category The number of samples is calculated using the following formula. in, For the region Medium category Number of samples, The number of sampling points in the area. For category In the region area percentage Total number of categories; Step 3: Based on the selected sample points, crop out the image data that matches the geographical location of the sample points to generate sample data, and use the Open Street Map labels of the corresponding areas as text to construct a text-image retrieval sample dataset.
2. The method for constructing a global remote sensing image and text retrieval sample as described in claim 1, characterized in that: Based on the global Moran index, regions worldwide exhibiting spatial clustering or dispersion in urban development are identified. This is determined by the global Moran index of global regional nighttime light remote sensing data. Furthermore, the Z-statistic test indicates a significant clustering trend. Then, the local Moran index and Z-statistic test are calculated for each region. Regions that pass the Z-statistic test are considered to have significant clustering or dispersion. The formula for the global Moran index is as follows. in, It is a region Nighttime light remote sensing value Its average deviation, For the total number of regions, It is a region and Spatial weights, which are generated based on adjacency relationships, are applied between regions. and Adjacent ,on the contrary ; It is the aggregation of all spatial weights, as shown in the following formula. 。 3. The method for constructing a global remote sensing image and text retrieval sample as described in claim 1, characterized in that: Local Moran Index The calculation formula is as follows: in, It is a region The value of nighttime light remote sensing, This represents the average value of nighttime light remote sensing data for the entire region. It is a region and Spatial weights between them Given the total number of regions, and having... pass The significance of the local Moran's index was tested using a statistical test. The formula for calculating the statistic is as follows: in, express Expectations express The variance is calculated to determine whether the null hypothesis is true. Statistic value and significance level Compare and take , The formula for calculating the value is as follows: in, If the standard normal distribution function is given, then... This indicates that the region exhibits a certain clustering or dispersion trend; For regions that pass the significance test ,according to and spatial lag value Region i can be divided into four patterns: high-high clustering, high-low clustering, low-high clustering, and low-low clustering. The spatial lag value... The calculation formula is as follows: in, For the region of value, It is a region and Spatial weights between; when and ,area High-level clustering; and ,area It is a low-high cluster; and ,area High-low clustering; and ,area It belongs to the low-low clustering.
4. The method for constructing a global remote sensing image and text retrieval sample as described in claim 1, characterized in that: The landscape index It is the ratio of landscape perimeter to area within a certain range, quantitatively representing the landscape heterogeneity of that area. The formula is as follows. in, Number of pixels For pixels The number of edges is given. The image is raster data, and each pixel is equivalent to a small square. If pixel i and pixel j are adjacent and of different types, there is an edge between the two pixels. If pixel i and pixel j are adjacent but of the same type, there is no edge between the pixels.
5. The method for constructing a global remote sensing image and text retrieval sample as described in claim 1, characterized in that: Geographic sampling unit Adaptive selection of sample point locations, hypothetical region Depend on It consists of several geographic units, and landscape indices for various land cover categories can be calculated within each geographic unit. , No. Class in OK, Columns The formula is as follows: Calculate each unit And then according to By sorting the geographic units, a distribution curve is obtained. , The axis represents the coordinates of a geographic unit. The axis is For the region Medium category Distribution curve Remove regions with values of 0, then... Axial division Divide the data into equal parts, and randomly select a sampling point in each interval to form a category. In the region The sampling points.
6. A global remote sensing image and text retrieval sample construction system, characterized in that, Includes the following modules: The data acquisition module is used to acquire global nighttime light remote sensing data, perform spatial autocorrelation analysis based on the global administrative region nighttime light remote sensing index, divide administrative regions with significant development levels, and obtain regions with significant clustering or dispersion. The spatial autocorrelation analysis consists of the global Moran index and the local Moran index; The sample point determination module is used to first calculate the landscape index based on global land cover. Then, the regional landscape index is calculated for each of the selected prominent areas. , regional land cover category landscape index Landscape index for each geographic sampling unit The number of sampling points in each region, the number of samples for each type of land cover in each region, and the location of each sampling point are obtained respectively, thus obtaining the final sampling distribution of the sample points; The landscape index of a certain area is denoted as... Number of sampling points in the area With the aforementioned regional landscape index The relevant calculation formula is as follows: in, and Representing regions and area, For the total sample size, Total number of regions; The landscape index of a certain land cover type in a certain region is denoted as... According to the landscape index of land cover category Determine the sampling quantity for each sample, for the region Medium category The number of samples is calculated using the following formula. in, For the region Medium category Number of samples, The number of sampling points in the area. For category In the region area percentage Total number of categories; The dataset construction module is used to generate sample data by cropping image data that matches the geographical location of the selected sample points, and to construct a text-image retrieval sample dataset using the Open Street Map labels of the corresponding areas as text.