A method, system, equipment, and medium for global spatial extrapolation of indicators in data-sparse villages.

By dividing the village area into grid units and utilizing transfer learning and spatial random forest models, the problems of feature extraction accuracy and boundary constraints caused by data sparsity in village-level geospatial information processing are solved, achieving high-precision full-domain spatial extrapolation and providing reliable data support for rural governance.

CN122489672APending Publication Date: 2026-07-31湖南工商大学
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
湖南工商大学
Filing Date
2026-05-09
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies, when processing geospatial information at the village level, rely on large-scale labeled data, which limits the accuracy of feature extraction. They also lack effective geospatial boundary constraints and full-domain analysis capabilities, making it difficult to meet the needs of rural spatial governance for refined and interpretable data analysis.

Method used

The target village area is divided into multiple grid units, multi-source heterogeneous data is acquired, and visual features of remote sensing images are extracted using transfer learning. Topographic data is introduced to construct a spatial prior mask, and interest point and road network data are fused. The model is trained through a spatial random forest model, and a geospatial distribution penalty splitting criterion is adopted to achieve multi-source fusion representation and global extrapolation.

Benefits of technology

It improves the reliability of feature recognition, avoids "enclave" errors, establishes a robust mapping relationship from local samples to the whole domain, generates logically consistent index projection values, and supports rural spatial governance and precise resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489672A_ABST
    Figure CN122489672A_ABST
Patent Text Reader

Abstract

This invention relates to a method, system, device, and medium for full-domain spatial extrapolation of indicators for data-sparse village areas. The method includes: gridding the target village area; acquiring multi-source heterogeneous data and extracting data from each grid cell; extracting the measured values ​​of target indicators for the sampled grid cells; extracting visual features from remote sensing images using a visual representation model; adjusting the visual features based on spatial prior masks; fusing the adjusted visual features, spatial distribution data of points of interest, and road network vector data to obtain a multi-source fusion representation; training a spatial random forest model using the measured values ​​of the target indicators and the multi-source fusion representations corresponding to the sampled grid cells; inputting the multi-source fusion representations of each grid cell into the trained spatial random forest model, performing spatial extrapolation prediction, and outputting the full-domain extrapolated values ​​of the target indicators corresponding to each grid cell. Thus, this invention achieves refined extraction and high-precision inversion of rural spatial information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of urban and rural planning and information processing technology, and in particular to a method, system, device and medium for full-domain spatial extrapolation of indicators for data-sparse village areas. Background Technology

[0002] Geospatial information processing typically relies on the regularized identification of buildings or farmland scenes in specific areas. Such methods often depend heavily on massive amounts of field-labeled data to support feature learning. However, village-level geospatial environments are extremely complex, making it very difficult to obtain high-quality labeled data. This leads to a cold start dilemma caused by data scarcity when dealing with such scenarios. Furthermore, due to the high coupling between natural and human elements in rural environments and the blurred boundaries of features, existing methods struggle to effectively distinguish irregularly shaped rural buildings from surrounding steep slopes, water bodies, and other complex terrain features. This results in generally low accuracy in extracting underlying spatial elements and an inability to adapt to the heterogeneity of rural spaces.

[0003] Problems arising from inaccurate extraction of underlying spatial elements directly impact subsequent numerical inversion modeling processes. Current inversion methods often employ conventional algorithms such as global regression or standard random forests. These algorithms typically rely solely on the numerical similarity of sample node feature values ​​when performing spatial clustering or numerical prediction. However, due to the extremely sparse sample size in rural areas and the widespread presence of natural barriers such as rivers and mountains, conventional algorithms based purely on feature values ​​cannot effectively perceive true geographical barriers. Therefore, it is highly likely that villages separated by natural barriers will be forcibly grouped into one category simply because the feature values ​​of two village nodes are similar. This leads to frequent instances of "enclaves" in the inverted socioeconomic indicators that violate actual geographical boundaries, resulting in significant inversion errors.

[0004] Meanwhile, due to the limitations in low-level feature extraction and sample acquisition, existing technologies face significant bottlenecks in handling the connection between sparse samples and global prediction. Most current methods rely solely on passive inference based on local observation data, making it difficult to effectively establish a robust mapping relationship from a small number of sampled data to the entire spatial domain. Lacking an effective representation of spatial structure, existing models often exhibit severe prediction blind spots in unsampled areas when predicting continuous values ​​covering the entire village, resulting in the inability to generate a complete indicator distribution covering the entire village. This local inference approach makes it difficult to construct global geospatial relationships, making it difficult to guarantee the logical consistency and spatial continuity of the final inference results over a wide area, thus failing to provide a complete evaluation basis for rural spatial governance and scientific decision-making over a wide area. Summary of the Invention

[0005] (a) Technical problems to be solved In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a method, system, device and medium for full-domain spatial extrapolation of indicators for data-sparse village areas. It solves the technical problems of the prior art in dealing with village-level geospatial information processing, which are limited in feature extraction accuracy due to strong reliance on large-scale labeled data, and which are difficult to meet the needs of rural spatial governance for refined and interpretable data analysis due to the lack of effective geospatial boundary constraints and full-domain analysis capabilities.

[0006] (II) Technical Solution To achieve the above objectives, the main technical solutions adopted by the present invention include: In a first aspect, embodiments of the present invention provide a method for global spatial extrapolation of indicators for data-sparse rural areas, including: The target village area is divided into multiple grid units. Multi-source heterogeneous data of the target village area are obtained and the remote sensing images, topographic data, spatial distribution data of points of interest and road network vector data corresponding to each grid unit are extracted. The measured values ​​of target indicators of the grid units used for sampling are also extracted. Visual features of remote sensing images are extracted using a visual representation model obtained through transfer learning. A spatial prior mask constructed from topographic data is introduced to adjust the response intensity of the visual features. The adjusted visual features, spatial distribution data of points of interest, and road network vector data are fused to obtain a multi-source fusion representation of each grid cell. The constructed spatial random forest model is trained using the measured values ​​of the target index and the multi-source fusion representations corresponding to the grid cells used for sampling. Furthermore, a splitting criterion that introduces a geospatial distribution penalty is adopted to optimize between attribute similarity and spatial continuity, resulting in the trained spatial random forest model. The multi-source fusion representation of each grid cell is input into the trained spatial random forest model to perform spatial extrapolation prediction covering the target village area, and outputs the global extrapolation value of the target index corresponding to each grid cell.

[0007] Optionally, the target village area is divided into multiple grid cells, multi-source heterogeneous data of the target village area is acquired, and remote sensing images, topographic data, spatial distribution data of points of interest, and road network vector data corresponding to each grid cell are extracted from them. Measured values ​​of target indicators for the grid cells used for sampling are also extracted, including: Establish a unified spatial coordinate system covering the target village area, and divide the target village area into multiple grid units based on the unified spatial coordinate system; Obtain multi-source heterogeneous data of the target village area from at least one data source, and extract the corresponding remote sensing images, topographic data, spatial distribution data of points of interest and road network vector data within each grid cell from the multi-source heterogeneous data; Extracted measured values ​​of target indicators from partial grid cells used for sampling from multi-source heterogeneous data; Perform preprocessing including normalization on remote sensing imagery, topographic data, spatial distribution data of points of interest, and road network vector data; The target indicators include any one or more of the following: rural resilience index, economic activity level, poverty level, industrial advantage level, infrastructure level, and production / living / ecological development level.

[0008] Optionally, the transfer learning steps of the visual representation model include: Obtain remote sensing images of reference villages with labeled data, as well as measured values ​​of target indicators corresponding to the reference villages. Combine the remote sensing images of the reference villages with the measured values ​​of target indicators corresponding to the reference villages to form reference sample data. The remote sensing images from the reference sample data are input into a visual representation model that includes a feature extraction network and a classifier network, and the measured values ​​of the target indicators from the reference sample data are used as pre-training supervision signals. The visual representation model is driven by pre-trained supervised signals to learn the mapping relationship between remote sensing images and the measured values ​​of target indicators corresponding to reference villages. The weight parameters of the feature extraction network and the classifier network are iteratively updated until the model converges, and the pre-trained visual representation model is output. Keep the weight parameters of the feature extraction network in the pre-trained visual representation model constant; The remote sensing images corresponding to the grid cells used for sampling within the target village area are input into the pre-trained visual representation model. The measured values ​​of the target indicators of the grid cells used for sampling are used as the supervision signal for transfer update. The weight parameters of the classifier network are updated only through error backpropagation, and the visual representation model that has completed transfer learning is output.

[0009] Optionally, visual features of remote sensing images are extracted using a visual representation model obtained through transfer learning. A spatial prior mask constructed from topographic data is introduced to adjust the response intensity of the visual features. The adjusted visual features, spatial distribution data of points of interest, and road network vector data are then fused to obtain a multi-source fused representation for each grid cell, including: The remote sensing image corresponding to each grid unit is divided into multiple image blocks; Based on the topographic data corresponding to each grid cell, identify topographically abnormal areas with slopes greater than the angle threshold or located within water bodies; Construct a spatial prior mask matrix corresponding to the positions of multiple image patches, assign a preset mask value to the corresponding position of the terrain anomaly region in the spatial prior mask matrix, and assign a value of zero to the corresponding position of the region outside the terrain anomaly region in the spatial prior mask matrix. Multiple image patches are input into the visual representation model, and the spatial prior mask matrix is ​​introduced into the self-attention mechanism of the visual representation model. This reduces the attention weight of the image patch corresponding to the terrain anomaly region when the self-attention mechanism participates in the calculation, and outputs the adjusted visual features. Based on the spatial distribution data of interest points and road network vector data corresponding to each grid unit, calculate the socio-economic feature data including the estimated kernel density of interest points and the road network density, and map it into a socio-economic feature vector that matches the adjusted visual feature dimension. Cross-modal attention interaction calculation is performed based on the adjusted visual features and socioeconomic feature vectors to obtain cross-modal contextual features; The adjusted visual features are weighted and fused with cross-modal contextual features to output the comprehensive geospatial feature vector corresponding to each grid unit as a multi-source fusion representation.

[0010] Optionally, the constructed spatial random forest model is trained using the measured values ​​of the target index and the multi-source fusion representations corresponding to the grid cells used for sampling. A splitting criterion incorporating geospatial distribution penalties is employed to optimize between attribute similarity and spatial continuity, resulting in the trained spatial random forest model, including: The multi-source fusion representations corresponding to the grid cells used for sampling are used as training features, and the measured values ​​of the target indicators of the grid cells used for sampling are used as training labels to construct the decision tree in the spatial random forest model. At the current split node of the decision tree, based on the multi-source fusion representation corresponding to the grid cell used for sampling that reaches the current split node, a splitting scheme that generates left and right child nodes is generated in the feature domain. According to the splitting scheme, the grid cells for sampling that arrive at the current splitting node are assigned to the corresponding left and right child nodes. Based on the measured values ​​of the target index of the grid cells assigned to each child node, the change in Gini impurity of the current splitting node before and after splitting is calculated, and the decrease in Gini impurity corresponding to the splitting scheme is obtained. Based on the geographical adjacency status, natural barriers, or administrative boundaries between the grid cells used for sampling that reach the current split node, determine the geospatial distribution penalty term value corresponding to the splitting scheme; The local weighted Gini index is obtained by calculating the difference between the Gini impurity decrease and the geospatial distribution penalty term. The splitting scheme with the largest local weighted Gini index is selected as the final splitting criterion for the current splitting node, and the node is split until the preset stopping condition is met. The trained spatial random forest model is then output.

[0011] Optionally, the multi-source fusion representation of each grid cell is input into the trained spatial random forest model to perform spatial extrapolation prediction covering the target village area, outputting the global extrapolation value of the target index corresponding to each grid cell, including: Extract grid cells within the target village area that do not contain measured values ​​of the target indicator as grid cells to be predicted; The multi-source fusion representation corresponding to the grid cell to be predicted is input into the trained spatial random forest model. The multi-decision trees in the trained spatial random forest model are used to perform parallel inference calculation on the multi-source fusion representation corresponding to the grid cell to be predicted. The output results of the multi-decision trees are aggregated to obtain the inferred value corresponding to the grid cell to be predicted. The inferred values ​​corresponding to the grid cells to be predicted are summarized with the measured values ​​of the target indicators of the grid cells used for sampling to generate an indicator inversion array covering the target village area. Map the index inversion array to the corresponding grid cell position within the target village area, and output the global inversion value of the target index corresponding to each grid cell.

[0012] Optionally, after inputting the multi-source fusion representation of each grid cell into the trained spatial random forest model, performing spatial extrapolation prediction covering the target village region, and outputting the global extrapolation value of the target index corresponding to each grid cell, the method further includes: The global inference value of the target index is decomposed into the basic expected value of the spatial random forest model and the bias corresponding to each input feature in the multi-source fusion representation, and the bias is used as the local marginal contribution value of each input feature in each grid cell. Based on the spatial adjacency relationship between each grid cell, the spatial contribution clustering degree of the local marginal contribution value in the spatial distribution is calculated; Calculate the global average of all local marginal contribution values ​​corresponding to each input feature and take the absolute value to obtain the absolute value of the global average marginal contribution. When the absolute value of the global average marginal contribution of the input feature is lower than the preset contribution threshold, and the corresponding spatial contribution clustering is lower than the preset clustering threshold, the corresponding input feature is removed, and the output is a retained feature matrix composed of the retained input features. Calculate the local spatial variance of the local marginal contribution value of each input feature in the retained feature matrix within the spatial adjacency range. Divide the grid cells with local spatial variance greater than the preset fluctuation threshold into heterogeneous regions and divide the grid cells with local spatial variance not greater than the fluctuation threshold into homogeneous regions. The spatial attenuation coefficient is adaptively increased for heterogeneous regions and adaptively decreased for homogeneous regions to update the spatial weights; whereby the spatial attenuation coefficient is a control parameter that determines the mapping relationship of spatial weight attenuation with increasing distance. The retained feature matrix and spatial weights are then input back into the spatial random forest model for iterative training until the preset convergence condition is met. The optimal adaptive spatial random forest model and the global inference value of the target index after iteration are then output.

[0013] Secondly, embodiments of the present invention provide a system for global spatial extrapolation of indicators for data-sparse rural areas, including: The data acquisition module is used to divide the target village area into multiple grid units, acquire multi-source heterogeneous data of the target village area, and extract remote sensing images, topographic data, spatial distribution data of points of interest and road network vector data corresponding to each grid unit, and extract the measured values ​​of target indicators of the grid units used for sampling. The feature extraction and fusion module is used to extract visual features from remote sensing images using a visual representation model obtained through transfer learning. It also introduces a spatial prior mask constructed from topographic data to adjust the response intensity of the visual features. The adjusted visual features, spatial distribution data of points of interest, and road network vector data are fused to obtain a multi-source fusion representation of each grid cell. The model training module is used to train the constructed spatial random forest model using the measured values ​​of the target index and the multi-source fusion representations corresponding to the grid cells used for sampling. It also adopts a splitting criterion that introduces a geospatial distribution penalty to optimize between attribute similarity and spatial continuity, thus obtaining the trained spatial random forest model. The global prediction module is used to input the multi-source fusion representation of each grid cell into the trained spatial random forest model, perform spatial extrapolation prediction covering the target village area, and output the global extrapolation value of the target index corresponding to each grid cell.

[0014] Thirdly, embodiments of the present invention provide a device for global spatial extrapolation of indicators for sparse data villages, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute the global spatial extrapolation method for indicators based on sparse data villages as described above.

[0015] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method for global spatial extrapolation of indicators for sparse data domains.

[0016] (III) Beneficial Effects The beneficial effects of this invention are: First, by dividing the target village area into multiple grid units and acquiring multi-source heterogeneous data, this invention establishes a standardized spatial basis for subsequent refined analysis. On this basis, the visual representation model obtained by transfer learning is used to extract visual features from remote sensing images, effectively overcoming the dependence of traditional machine learning on massive field observation samples. Even when there is a lack of labeled data in the target village, it still maintains a high level of model generalization ability, solving the bottleneck of model cold start and feature extraction caused by data sparsity from the bottom up.

[0017] Secondly, by introducing a spatial prior mask constructed from topographic data and integrating multi-source spatial elements such as points of interest and road networks, this invention achieves highly robust multi-source fusion representation of grid cells. This scheme can accurately identify and eliminate false signal interference in areas with abnormal terrain, enabling dynamic weighted fusion of visual features and socioeconomic factors. It effectively solves the problem that traditional algorithms easily extract false features in complex rural terrains, thus improving the reliability of feature recognition.

[0018] Next, by introducing a splitting criterion based on geospatial distribution penalty into the spatial random forest model, this invention forces the model to simultaneously weigh attribute similarity and geospatial continuity when splitting nodes. This mechanism endows the model with significant terrain perception and geographic barrier recognition capabilities, avoiding the "enclave" error caused by directly applying urban distance models, ensuring that the index inversion results strictly follow natural geographic boundaries, and fundamentally eliminating the illogical jumps in the spatial distribution of inversion data.

[0019] Finally, by inputting multi-source fusion representations into a trained spatial random forest model to perform spatial extrapolation predictions covering the target village area, this invention effectively establishes a robust mapping relationship from local samples to the entire spatial domain. This method can utilize the constructed global geospatial associations to generate continuous and logically consistent index extrapolation values ​​in unsampled areas, successfully eliminating prediction blind spots in the global extrapolation. It achieves high-precision spatial inversion from local sampling to wide-area coverage, providing reliable data support for rural spatial governance and precise resource allocation. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the overall process of the method provided in the embodiments of the present invention; Figure 2 This is a schematic diagram illustrating the specific process of step S1 of the method provided in this embodiment of the invention; Figure 3 This is a schematic diagram illustrating the specific process prior to step S2 in the method provided in this embodiment of the invention; Figure 4 This is a detailed flowchart illustrating step S2 of the method provided in this embodiment of the invention; Figure 5This is a detailed flowchart illustrating step S3 of the method provided in this embodiment of the invention; Figure 6 This is a detailed flowchart illustrating step S4 of the method provided in this embodiment of the invention; Figure 7 A schematic diagram of the specific process of step S5 of the method provided in the embodiment of the present invention. Detailed Implementation

[0021] To better explain and facilitate understanding of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] like Figure 1 As shown in the embodiment of the present invention, a method for full-domain spatial extrapolation of indicators for sparsely dataed villages is proposed. This method includes: dividing the target village into multiple grid units; acquiring multi-source heterogeneous data of the target village and extracting remote sensing images, topographic data, spatial distribution data of points of interest, and road network vector data corresponding to each grid unit; extracting the measured values ​​of the target indicators for the grid units used for sampling; extracting visual features of the remote sensing images using a visual representation model obtained through transfer learning, and introducing a spatial prior mask constructed from topographic data to adjust the response intensity of the visual features; fusing the adjusted visual features, spatial distribution data of points of interest, and road network vector data to obtain the multi-source fusion representation of each grid unit; training a constructed spatial random forest model using the measured values ​​of the target indicators and the multi-source fusion representations corresponding to the grid units used for sampling, and employing a splitting criterion that introduces a geospatial distribution penalty to optimize between attribute similarity and spatial continuity to obtain the trained spatial random forest model; inputting the multi-source fusion representations of each grid unit into the trained spatial random forest model, performing spatial extrapolation prediction covering the target village, and outputting the full-domain extrapolated values ​​of the target indicators corresponding to each grid unit.

[0023] First, by dividing the target village area into multiple grid units and acquiring multi-source heterogeneous data, this invention establishes a standardized spatial basis for subsequent refined analysis. On this basis, the visual representation model obtained by transfer learning is used to extract visual features from remote sensing images, effectively overcoming the dependence of traditional machine learning on massive field observation samples. Even when there is a lack of labeled data in the target village, it still maintains a high level of model generalization ability, solving the bottleneck of model cold start and feature extraction caused by data sparsity from the bottom up.

[0024] Secondly, by introducing a spatial prior mask constructed from topographic data and integrating multi-source spatial elements such as points of interest and road networks, this invention achieves highly robust multi-source fusion representation of grid cells. This scheme can accurately identify and eliminate false signal interference in areas with abnormal terrain, enabling dynamic weighted fusion of visual features and socioeconomic factors. It effectively solves the problem that traditional algorithms easily extract false features in complex rural terrains, thus improving the reliability of feature recognition.

[0025] Next, by introducing a splitting criterion based on geospatial distribution penalty into the spatial random forest model, this invention forces the model to simultaneously weigh attribute similarity and geospatial continuity when splitting nodes. This mechanism endows the model with significant terrain perception and geographic barrier recognition capabilities, avoiding the "enclave" error caused by directly applying urban distance models, ensuring that the index inversion results strictly follow natural geographic boundaries, and fundamentally eliminating the illogical jumps in the spatial distribution of inversion data.

[0026] Finally, by inputting multi-source fusion representations into a trained spatial random forest model to perform spatial extrapolation predictions covering the target village area, this invention effectively establishes a robust mapping relationship from local samples to the entire spatial domain. This method can utilize the constructed global geospatial associations to generate continuous and logically consistent index extrapolation values ​​in unsampled areas, successfully eliminating prediction blind spots in the global extrapolation. It achieves high-precision spatial inversion from local sampling to wide-area coverage, providing reliable data support for rural spatial governance and precise resource allocation.

[0027] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present invention can be understood more clearly and thoroughly, and that the scope of the present invention can be fully conveyed to those skilled in the art.

[0028] Specifically, embodiments of the present invention provide a method for global spatial extrapolation of indicators for data-sparse rural areas, comprising: S1. Divide the target village area into multiple grid units, acquire multi-source heterogeneous data of the target village area, and extract the remote sensing images, topographic data, spatial distribution data of points of interest and road network vector data corresponding to each grid unit, and extract the measured values ​​of target indicators of the grid units used for sampling.

[0029] Furthermore, such as Figure 2 As shown, step S1 includes: S11. Establish a unified spatial coordinate system covering the target village area, and divide the target village area into multiple grid units based on the unified spatial coordinate system.

[0030] In this step, a unified spatial coordinate system (such as CGCS2000) is first established. Then, the village area is divided into n regular square grid cells with side length L. To accurately capture village-level micro-heterogeneity, the parameter L is set to 30 meters to 100 meters, so the i-th grid cell is U. i It serves as the smallest spatial physical unit for carrying subsequent data.

[0031] S12. Obtain multi-source heterogeneous data of the target village area from at least one data source, and extract the corresponding remote sensing images, topographic data, spatial distribution data of points of interest and road network vector data within each grid cell from the multi-source heterogeneous data.

[0032] Specifically, the required data sources include: high-resolution multi-temporal remote sensing imagery (such as Sentinel-2 or the Gaofen series, requiring a spatial resolution better than 10m), digital elevation models (DEMs), spatial distribution data of points of interest (POIs), and road network vector data (such as road network CAD or GIS Shapefiles). For each grid cell, spatial analysis is used to extract the corresponding remote sensing image slices, topographical DEM data, the number of various POIs falling within that grid cell, and road network density characteristics from the obtained multi-source heterogeneous data.

[0033] S13. Extract the measured values ​​of the target index of some grid cells for sampling from multi-source heterogeneous data.

[0034] Next, stratified random sampling was used to extract the measured values ​​of the target indicators from a subset of grid cells in the multi-source heterogeneous data. Specifically, different spatial levels were defined based on the spatial distance of the grid cells from the main roads of villages and towns or the terrain elevation. Within each spatial level, grid cells were randomly selected according to the proportion of the total number of grid cells to ensure that the sampled data is representative of the entire region under complex terrain. The target indicators cover indicators that characterize the social and economic development of rural areas, such as rural resilience index, economic activity, poverty level, industrial advantage, infrastructure level, and the degree of production / living / ecological development. These measured values ​​of the target indicators will serve as baseline labels for subsequent model validation and training.

[0035] S14. Perform preprocessing including normalization on remote sensing imagery, topographic data, spatial distribution data of points of interest (POIs), and road network vector data. Min-Max normalization is performed on all extracted numerical variables (such as elevation, number of POIs, road network parameters, etc.), mapping all variable values ​​to the [0,1] interval to eliminate the influence of different physical dimensions on model calculations. Finally, the preprocessed multidimensional features are integrated to form standardized spatial base data, providing a unified data input format for subsequent identification and analysis.

[0036] S2. Visual features of remote sensing images are extracted using the visual representation model obtained by transfer learning. A spatial prior mask constructed from topographic data is introduced to adjust the response intensity of the visual features. The adjusted visual features, spatial distribution data of points of interest, and road network vector data are fused to obtain the multi-source fusion representation of each grid unit.

[0037] Furthermore, prior to step S2, to address the cold start problem that may arise in some rural areas due to a lack of labeled data or extremely small sample sizes, such as... Figure 3 As shown, the transfer learning steps of the visual representation model include: P21. Acquire remote sensing images of reference villages with labeled data, along with the measured values ​​of the corresponding target indicators. Combine the remote sensing images and the measured values ​​of the target indicators of the reference villages to form reference sample data. In this step, a rural-urban mapping transfer learning mechanism is established. Specifically, urbanized fringe areas or large central villages with abundant labeled data and multi-source geographic information are identified as source domains, while ordinary villages or remote small villages with sparse samples are identified as target domains. High-quality remote sensing images and supporting indicators from the source domains are used to construct a high-quality reference sample library, providing a sufficient learning foundation for subsequent models.

[0038] P22. The remote sensing images from the reference sample data are input into a visual representation model containing a feature extraction network and a classifier network, and the measured values ​​of the target indicators from the reference sample data are used as pre-training supervision signals. Specifically, the feature extraction network uses a Transformer encoder to perform full-field perceptual feature extraction from the input remote sensing images through a self-attention mechanism, and the measured values ​​of the target indicators from the reference sample data are used as pre-training supervision signals. By introducing supervision signals from the source domain, the visual representation model can initially perceive the feature distribution in complex geographical environments and learn how to extract semantically informative visual features from heterogeneous images, providing initial weights with strong generalization capabilities for subsequent fine-tuning of the model in the target domain.

[0039] P23. Based on pre-trained supervised signals, the visual representation model learns the mapping relationship between remote sensing images and the measured values ​​of target indicators corresponding to the reference village area. The weight parameters of the feature extraction network and the classifier network are iteratively updated until the model converges, outputting the pre-trained visual representation model. Specifically, rich sample data within the source domain is used to drive the feature extraction network in the visual representation model to learn feature space representations. The weight parameters of the feature extraction network and the classifier network are iteratively updated until the model loss function converges, thus outputting the pre-trained visual representation model. During this process, deep iterative optimization enables the model to fully learn the fundamental mapping relationship between macro-built environment and socio-economic indicators. The core objective of this stage is to leverage the Transformer encoder's ability to capture global features, allowing the model to accurately identify and capture the universal spatial logic between urban and rural construction characteristics and indicators such as rural resilience and economic activity.

[0040] P24. Keep the weight parameters of the feature extraction network in the pre-trained visual representation model constant. Before transferring to a specific target region, freeze the weight parameters of the feature extraction network in the pre-trained visual representation model. This preserves the model's ability to recognize general spatial structures learned in the source domain and avoids overfitting when samples are sparse.

[0041] P25. The remote sensing images corresponding to the grid cells used for sampling within the target village area are input into the pre-trained visual representation model. The measured values ​​of the target indicators in the grid cells used for sampling are used as the transfer update supervision signal. Through error backpropagation, only the weight parameters of the classifier network are updated, and the visual representation model that has completed transfer learning is output. In the fine-tuning stage, only a small number of measured indicator values ​​within the target domain are used as supervision signals, and the classifier network is fine-tuned through error backpropagation. This fully utilizes the gradual continuity of spatial morphological evolution in the urban-rural transition zone, and fine-tunes the parameters through a small amount of gradient backpropagation, thereby effectively improving the model's generalization ability in data-sparse villages and solving the modeling problem in small sample environments.

[0042] Furthermore, addressing the issue that the lack of urban-level planning and the highly intertwined nature of houses and farmland in rural landscapes make it easy for conventional networks to extract false features, this step proposes a method that introduces a spatial prior mask, such as... Figure 4 As shown, step S2 includes: S21. Divide the remote sensing image corresponding to each grid cell into multiple image patches. First, divide the i-th grid U... iThe corresponding remote sensing image slices are segmented into a series of fixed-size image blocks and flattened. Then, they are mapped to the feature space through linear projection, embedded with position encoding, and input into the Transformer encoder. The aim is to initially establish the spatial context association between heterogeneous surface coverings and architectural space through an attention mechanism.

[0043] S22. Based on the topographic data corresponding to each grid cell, identify anomalous terrain areas with slopes greater than the angle threshold or located within water bodies. Specifically, identify areas with extremely steep slopes (e.g., greater than 25 degrees) or in the middle of water bodies where it is objectively impossible to deploy dense buildings, and define them as anomalous terrain areas to eliminate interference from unreasonable geographical features on model recognition.

[0044] S23. Construct a spatial prior mask matrix corresponding to the positions of multiple image patches. Assign a preset mask value to the corresponding position in the spatial prior mask matrix for terrain anomaly regions, and assign a value of zero to the corresponding position in the spatial prior mask matrix for regions outside the terrain anomaly regions. The constructed spatial prior mask matrix M prior The matrix is ​​a mapping structure corresponding to the image block sequence index. In the mask matrix, masking values ​​(such as negative infinity or 10) are assigned to the locations corresponding to terrain anomaly regions. -9 The model assigns a value of 0 to reasonable areas outside of abnormal terrain areas (such as extremely negative values), thereby imposing hard constraints on geographical limitations in the feature space through a masking mechanism. The masking value makes the model automatically ignore the visual contribution of abnormal areas during feature extraction.

[0045] S24. Input multiple image patches into the visual representation model, introduce the spatial prior mask matrix into the self-attention mechanism of the visual representation model, so that when the self-attention mechanism participates in the calculation, the attention weight of the image patch corresponding to the terrain anomaly region is reduced, and the adjusted visual features are output.

[0046] In this step, the spatial prior mask matrix M is... prior Introducing multi-head self-attention mechanisms (such as overlay, fusion, or as part of the input) allows for spatial suppression of image patches corresponding to terrain anomalies when calculating attention weights, resulting in an adjusted visual feature vector V. i .

[0047] Specifically, the spatial prior mask matrix M prior Introducing a self-attention calculation process, the calculation formula for single-head attention is reconstructed as follows: ; Where Q, K, and V are the query vector, key vector, and value vector, respectively, and d kThe scaling factor is the square root of the query / key vector dimension, used to prevent the gradient from vanishing due to an excessively large dot product. This is achieved by introducing a spatial prior mask matrix M. prior During Softmax normalization, the model automatically reduces the attention weights of building clusters in unreasonable locations to near zero, thereby eliminating false signal interference from areas with abnormal terrain. The Transformer encoder outputs a visual feature vector V that has been adjusted for geospatial logic. i .

[0048] S25. Based on the spatial distribution data of interest points and road network vector data corresponding to each grid cell, calculate the socioeconomic feature data including the estimated kernel density of interest points and the road network density, and map it into a socioeconomic feature vector that matches the adjusted visual feature dimensions. Extract grid cell U. i The road network density (road network density = total length of road centerlines within an administrative village ÷ area of ​​the administrative village) and the kernel density estimate of the points of interest (kernel density estimate f(x) = ΣK((xx))) i ) / h) / nh, where K represents the kernel function, x is the current point of interest, and x i (where h is the bandwidth parameter and n is the number of interest points) are used as socioeconomic feature data. To avoid the feature space heterogeneity problem caused by direct splicing, these indicators are nonlinearly projected and dimensionally transformed using a multilayer perceptron (MLP) to map them into a socioeconomic feature vector S with the same dimension as the visual features. i .

[0049] S26. Based on the adjusted visual features and socioeconomic feature vectors, perform cross-modal attention interaction calculation to obtain cross-modal context features. Specifically, the visual feature vector V output by the Transformer encoder is used. i As the query vector, the socioeconomic feature vector S i As key and value vectors, the response weights of visual morphology to specific socioeconomic factors are calculated: ; Among them, W q W k W vThese are learnable linear projection weight matrices for the query, key, and value vectors, respectively. The initial values ​​of these learnable linear projection weight matrices are randomly generated based on a regularized distribution. During model training, the matrix parameters are iteratively corrected using the measured index bias of the target region through backpropagation until the model converges, thereby obtaining the optimal weight distribution that maps the semantic association between visual features and socioeconomic features. Through this cross-modal interaction mechanism, the model dynamically adjusts the association strength between heterogeneous features by calculating the response weights of visual spatial morphology to specific socioeconomic factors, thus outputting a semantically consistent cross-modal contextual feature C. i .

[0050] S27. The adjusted visual features are weighted and fused with cross-modal contextual features to output the comprehensive geospatial feature vector corresponding to each grid cell as a multi-source fusion representation. Finally, to address the complexity differences of elements in different village areas, an adaptive gating mechanism is introduced to dynamically adjust the fusion ratio of visual and socioeconomic features. The core of this mechanism lies in dynamically adjusting the fusion ratio of visual representation and socioeconomic representation through a gating network. The specific process is as follows: First, concatenate vectors [V] i C i Input the gated network and calculate the gate weight vector g of the i-th grid cell. i This represents the proportion of visual features retained during the fusion process. ; in, It is the Sigmoid activation function. This is the learnable weight matrix for the gated network. Its initialization and parameter update logic is consistent with that of Wq, Wk, and Wv mentioned above. g This is the bias of the gated network.

[0051] Subsequently, using g i The adjusted visual features and cross-modal context features are summed element-wise with weights to output the comprehensive geospatial feature vector F of the i-th grid cell. i : ; in, It is the Hadamard product of matrices (element-wise multiplication).

[0052] This fusion mechanism enables the model to adaptively assign higher fusion weights to socio-economic feature channels based on the actual data quality of different village-level units (e.g., when the remote sensing image of a village is obscured by clouds, resulting in low confidence of visual features, but its road network and POI data are complete). This improves the robustness of feature extraction in complex rural environments and ultimately outputs a comprehensive geospatial feature vector for each grid unit as a multi-source fusion representation.

[0053] S3. The constructed spatial random forest model is trained using the measured values ​​of the target index and the multi-source fusion representations corresponding to the grid cells used for sampling. The splitting criterion that introduces geospatial distribution penalty is used to optimize between attribute similarity and spatial continuity to obtain the trained spatial random forest model.

[0054] When traditional algorithms directly apply urban-scale distance decay models to rural areas, they only fine-tune the sample weights in terms of quantity (i.e., only giving higher weights to samples that are closer together), without changing the essential logic of random forest partitioning the space. In rural areas with sparse data, this method is prone to forcibly classifying villages separated by rivers or mountains into one category simply because the feature values ​​of two nodes are similar.

[0055] To address the aforementioned shortcomings, this embodiment trains the spatial random forest model using the measured values ​​of the target index and the multi-source fusion representations corresponding to the grid cells used for sampling. To address the "spatial blind spot" problem caused by data sparsity in rural areas, this step does not employ the traditional distance decay model. Instead, it upgrades the Gini impurity during node splitting in a conventional random forest to a locally weighted Gini index with spatial constraints, optimizing between attribute similarity and spatial continuity to obtain the trained spatial random forest model.

[0056] Furthermore, such as Figure 5 As shown, step S3 includes: S31. Using the multi-source fusion representations corresponding to the grid cells used for sampling as training features, and the measured values ​​of the target index of the grid cells used for sampling as training labels, we construct decision trees in the spatial random forest model. Let D be the set of grid cells contained in the current node, and we construct multiple decision trees in the spatial random forest model accordingly.

[0057] S32. At the current split node of the decision tree, based on the multi-source fusion representation corresponding to the grid cells used for sampling that reach the current split node, a splitting scheme for generating left and right child nodes is generated in the feature domain. At the current split node of the decision tree, based on the features of the grid cells reaching that node, all possible splitting dimensions and splitting points are traversed in the feature domain to divide the grid cell set D into two mutually exclusive subsets: left child node D... L Includes grid cells that meet the current candidate segmentation threshold, along with their corresponding multi-source fusion representations and measured values ​​of the target index; right child node D R This includes the remaining grid cell samples. Through this recursive partitioning of spatial samples, hierarchical modeling of complex geographic features in the feature space is achieved.

[0058] S33. According to the partitioning scheme, the grid cells used for sampling that arrive at the current split node are assigned to the corresponding left and right child nodes. Based on the measured values ​​of the target index of the grid cells assigned to each child node, the change in Gini impurity of the current split node before and after splitting is calculated, and the decrease in Gini impurity corresponding to the partitioning scheme is obtained. Specifically, according to the partitioning scheme, the grid cells are assigned, and the change in Gini impurity of the current split node before and after splitting is calculated, and the decrease in Gini impurity ΔGini corresponding to the partitioning scheme is obtained. ; ; Among them, y i Let be the measured value of the target index for the i-th grid cell within the grid cell set D. Let |D| be the average of the measured values ​​of the target index of all grid cells in set D, where |D| is the total number of grid cells in set D.

[0059] S34. Based on the geographical adjacency status, natural barriers, or administrative boundaries between the grid cells used for sampling to reach the current split node, determine the geospatial distribution penalty term value corresponding to the splitting scheme. Based on the geographical adjacency matrix between grid cells, natural barriers (such as rivers, mountains), or administrative boundaries, determine the spatial distribution penalty term P corresponding to the splitting scheme by judging the topological association status of grid cells within the child node. spatial .

[0060] It should be noted that when the partitioning scheme divides two geographically adjacent grid cells used for sampling into D, and there are no natural barriers (such as rivers) or administrative boundaries between them, the partitioning scheme will result in the following: L and D R When the partitioning scheme is implemented in different child nodes, the corresponding geospatial distribution penalty term P is calculated. spatial The penalty value is higher than the preset penalty threshold. When the partitioning boundary corresponding to the partitioning scheme is composed of natural barriers or administrative boundaries, that is, when the partitioning boundary exactly matches the spatial jump caused by the real natural barriers or administrative boundaries, the value of the geospatial distribution penalty item corresponding to the partitioning scheme approaches zero.

[0061] S35. Calculate the difference between the Gini impurity reduction and the geospatial distribution penalty term to obtain the locally weighted Gini index. Select the splitting scheme with the largest locally weighted Gini index as the final splitting criterion for the current splitting node and perform node splitting until a preset stopping condition is met. Output the trained spatial random forest model. Introduce a spatial penalty adjustment coefficient λ, the value of which is set based on expert prior knowledge, and calculate the improved node splitting gain evaluation function Gain. spatial The formula is as follows: ; Select Gainspatial The largest splitting scheme is used as the final splitting criterion. By imposing this spatial constraint, the decision tree is forced to weigh attribute similarity and spatial continuity simultaneously when splitting. This effectively locks in the spatial jumps caused by real geographical boundaries, thereby eliminating the inversion "enclave" phenomenon caused by the sparsity of rural data. The above process is repeated until the preset stopping condition is met, and the trained spatial random forest model is output.

[0062] S4. Input the multi-source fusion representation of each grid cell into the trained spatial random forest model, perform spatial extrapolation prediction covering the target village area, and output the global extrapolation value of the target index corresponding to each grid cell.

[0063] Furthermore, such as Figure 6 As shown, step S4 includes: S41. Extract grid cells within the target village area that do not contain the measured values ​​of the target indicators as grid cells to be predicted.

[0064] S42. Input the multi-source fusion representation corresponding to the grid cell to be predicted into the trained spatial random forest model. Utilize multiple decision trees within the trained spatial random forest model to perform parallel inference computation on the multi-source fusion representation corresponding to the grid cell to be predicted. Aggregate the outputs of multiple decision trees to obtain the inferred value corresponding to the grid cell to be predicted. Alternatively, input the multi-source fusion representation corresponding to the grid cell to be predicted into the trained spatial random forest model, and utilize multiple decision trees within the model for parallel inference. By aggregating (e.g., by averaging or weighted voting) the independent outputs of multiple decision trees, the inferred value corresponding to the grid cell to be predicted is obtained.

[0065] S43. Summarize the inferred values ​​corresponding to the grid cells to be predicted and the measured values ​​of the target indicators of the grid cells used for sampling to generate an indicator inversion array covering the target village area.

[0066] S44. Map the index inversion array to the corresponding grid cell location within the target village area, and output the global inversion value of the target index corresponding to each grid cell. Based on the index identifier of the grid cell, accurately map the values ​​in the index inversion array to the corresponding geographical coordinate location within the target village area, and output the global inversion value of the target index corresponding to each grid cell, thereby realizing the refined, full-space-dimensional inversion of socio-economic indicators in complex rural environments.

[0067] Furthermore, to address the issue of traditional attribution algorithms neglecting the differences in the driving mechanisms of geographical elements across different village areas, this step proposes an attribution and two-layer feedback optimization mechanism combining spatial autocorrelation analysis, aiming to perform closed-loop correction on the inversion model. After step S4, as... Figure 7As shown, it also includes: S5, using an attribution algorithm to calculate the local marginal contribution of each feature in the inverted value, and combining the local spatial autocorrelation index to quantify the spatial clustering effect of the local marginal contribution, thereby performing feature dimensionality reduction and dynamically adjusting the spatial perception bandwidth of the spatial random forest, and outputting the final village-level multidimensional index inversion map after performing modeling training again. Its specific steps include: S51. The global inference value of the target index is decomposed into the basic expected value of the spatial random forest model and the bias corresponding to each input feature in the multi-source fusion representation, and the bias is used as the local marginal contribution value of each input feature in each grid cell. Using the SHAP attribution algorithm, the global inference value of the target index is decomposed into the basic expected value of the spatial random forest model and the bias corresponding to each input feature in the multi-source fusion representation. Its spatial additive interpretation model is defined as: ; in, For grid cell i, the local inversion prediction value is... Let K be the basic expectation of the model, which is the global average prediction level of the target metric on the training set without any feature input, and K be the total dimensionality of the input features. The bias is the local marginal contribution value of the j-th input feature in grid cell i. In this step, the bias is used as the value of each input feature in grid cell i. i The corresponding local marginal contribution value represents the positive or negative driving force of this feature under specific geographic coordinates. A binary mapping variable (which can only be 0 or 1) is used to indicate whether the j-th feature exists in the spatial additive interpretation model of the i-th spatial cell. When its value is 1, it means that the j-th spatial feature exists when calculating the inversion prediction value of the i-th grid cell. When its value is 0, it means that the j-th spatial feature is missing, that is, the value of the feature is replaced by the baseline expected value or the background distribution.

[0068] S52. Based on the spatial adjacency relationships between grid cells, calculate the spatial contribution clustering degree of local marginal contribution values ​​in spatial distribution. Use the Local Spatial Autocorrelation Index (Local Moran's I) to quantify and analyze the spatial clustering effect and heterogeneity of local marginal contribution values. Then, the spatial contribution clustering degree I of the j-th input feature at grid cell i is... i,j for: ; in, Let be the local marginal contribution value of the j-th input feature at grid cell i. Let $j$ be the global average marginal contribution value of feature $j$. The variance of the local marginal contribution value of this input feature. Let be the spatial adjacency weight matrix between grid cell i and grid cell v. Based on the village's walkability, the initial setting is that when the geographical center distance between grid cells i and v is less than or equal to 500 meters, =1. When the geographic center distance between grid cells i and v is greater than 500 meters, =0, where N is the total number of grid cells.

[0069] S53. Calculate the global average of all local marginal contribution values ​​corresponding to each input feature and take its absolute value to obtain the absolute value of the global average marginal contribution. Summarize the local marginal contribution values ​​of each input feature across all grid cells and calculate its global average. Then take the absolute value to obtain the absolute value of the global average marginal contribution corresponding to this feature. .

[0070] S54. When the absolute value of the global average marginal contribution of an input feature is lower than a preset contribution threshold, and the corresponding spatial contribution clustering degree is lower than a preset clustering threshold, the corresponding input feature is removed, and the retained feature matrix composed of the retained input features is output. Global feature filtering and optimization are performed. When the absolute value of the global average marginal contribution of a certain input feature is lower than a preset contribution threshold, the corresponding input feature is removed, and the retained feature matrix is ​​output. If the contribution is below the preset contribution threshold τ, and its spatial contribution clustering degree I i,j When the clustering degree is below the preset clustering threshold, it reflects that when the clustering degree is not significant in the entire domain (i.e., it is manifested as random noise), it is removed and the retained feature matrix is ​​output.

[0071] S55. Calculate the local spatial variance of the local marginal contribution value of each input feature in the retained feature matrix within the spatial adjacency range. Divide the grid cells with local spatial variance greater than the preset fluctuation threshold (i.e., exhibiting strong local heterogeneity and frequent alternation of hot and cold spots) into heterogeneous regions, and divide the grid cells with local spatial variance not greater than the fluctuation threshold into homogeneous regions.

[0072] S56. The spatial decay coefficient is adaptively increased for heterogeneous regions and adaptively decreased for homogeneous regions to update the spatial weights. Dynamic reconstruction of distance decay weights is triggered for regions with drastic changes in feature contribution and differences in consistency with essential attributes. The spatial decay coefficient is a control parameter that determines the mapping relationship between spatial weights and distance. For heterogeneous regions, the spatial decay coefficient is adaptively increased to reduce the spatial perception bandwidth of that region and avoid excessive smoothing of distant, irrelevant samples; for homogeneous regions, the spatial decay coefficient is decreased to incorporate more neighboring samples, thus updating the spatial weights.

[0073] S57. The retained feature matrix and spatial weights are re-inputted into the spatial random forest model for iterative training until the preset convergence condition is met. The optimal adaptive spatial random forest model and the iterated global inference value of the target index are output. The retained feature matrix and dynamically updated spatial weights are re-inputted into the spatial random forest model for iterative training. This closed-loop mechanism breaks the limitations of static parameters in traditional inversion models until the root mean square error of the model on the validation set converges. Finally, the optimal adaptive inversion model and the iterated global inference value of the target index are output, and a geospatial attribution map reflecting the driving strength of the input features is generated based on the iterated local marginal contribution value.

[0074] Finally, the data output from the optimal adaptive inversion model is transformed into a raster layer recognizable by a Geographic Information System (GIS), and village administrative boundaries are overlaid to output a color spatial heat map and a SHAP feature contribution rose diagram. This processing workflow visualizes the calculation results of complex nonlinear models, intuitively presenting the indicator intensity and dominant causes of different areas of a village, providing intelligent decision support with centimeter- to meter-level accuracy for village consolidation, industrial resource allocation, and public facility site selection.

[0075] Compared to traditional methods that only provide result data, this invention, through a two-layer closed loop of SHAP values ​​and spatial autocorrelation analysis, not only eliminates more than 30% of interference noise, but also enables the algorithm to automatically switch spatial perception bandwidth between regions of dramatic feature changes (such as canyons) and regions of homogeneous features (such as plains and farmland). The final output feature contribution rose diagram clearly quantifies the positive and negative driving forces of various spatial factors on rural development, providing direct and reliable data for rural spatial reconstruction and precise resource allocation. Based on the above steps S1~S5, to verify the inversion accuracy and attribution ability of this method in a typical densely populated village environment, the following specific embodiments are provided: (1) The first embodiment includes the following specific steps: In Phase S1, a comprehensive basic geospatial database for Lujia Village, Anji County, Huzhou City, Zhejiang Province was constructed. Data sources included: Sentinel-2 multispectral imagery with a spatial resolution of 10m, DEM data with a resolution of 30m, 1250 POI data points, and road network CAD data. The village area was divided into regular square grid cells with a side length of L=30 meters, generating a total of 450 effective grids. To address the limited number of field survey samples in the village, a rural-urban mapping transfer learning mechanism was initiated. First, a county-level urban periphery area with a large dataset, located 15 kilometers away from the village, was selected as the source domain for pre-training. Then, the underlying feature extractor was frozen, and the model was transferred to Lujia Village, Anji County, Huzhou City, Zhejiang Province (the target domain), where it was fine-tuned using only 45 sampled grid labels.

[0076] In stage S2, spatial feature extraction is performed. The image patch size of the improved visual representation model is set to 16×16, and the number of layers in the multi-layer Transformer encoder is set to 4. To eliminate interference from ponds and waterways around the village, a spatial prior mask matrix is ​​constructed using DEM and topographic data, and the attention mask corresponding to the grid in the center of the water body is assigned a masking value (-∞). Combined with an adaptive gating mechanism, external covariates such as road network density are fused with visual features in a fully connected layer, outputting a 128-dimensional comprehensive geospatial feature vector.

[0077] In phase S3, inversion modeling is performed. The community resilience index of 45 sampled grids is used as the target inversion value. A spatial random forest model is constructed, with the number of decision trees set to M=150 and the initial spatial decay coefficient set to 0.03. When generating decision trees to find the optimal split point, a locally weighted Gini index with spatial constraints is introduced. A spatial penalty adjustment coefficient λ=0.1 is set. When the splitting scheme forcibly divides continuous grids on both sides of the same main road within the village without natural barriers into different categories, a penalty term is applied to ensure a smooth spatial transition of the resilience index within the built-up area.

[0078] In stage S4, SHAP loop closure optimization was performed. The first-layer loop closure removal threshold r=0.08 was set, removing 12 low-contribution features. Local Moran's I analysis revealed strong local heterogeneity in the grid of the northwestern mountainous region. The second-layer loop closure was triggered, adaptively increasing the spatial attenuation coefficient of this region to 0.05 (reducing the sensing bandwidth), and then re-inputting the data into stage S3 for iterative training until convergence.

[0079] In stage S5, the evaluation results are derived. After the above optimization mechanism, the root mean square error of the model on the test set decreased from 0.142 of the baseline model to 0.091, and the coefficient of determination (R²) reached 0.87. Finally, the spatial thermal distribution map and the ranking of feature contributions are output (see Table 1).

[0080] Table 1. SHAP Contribution Values ​​of the Top Three Key Spatial Elements in Resilience Index of Lujia Village, Anji County, Huzhou City, Zhejiang Province

[0081] (2) In the second embodiment, the grid parameters and model structure are adjusted for a town scale with a wider scope that includes multiple natural villages.

[0082] In Phase S1, a geographic information database of all villages in Zhouzhuang Town, Kunshan City, Suzhou, Jiangsu Province was constructed. High-resolution 0.8m Gaofen-2 remote sensing imagery, nighttime light slices, and 3420 Points of Interest (POI) data were acquired. The spatial grid side length L was set to 50 meters, generating 1200 effective grids. In this phase, complete data from neighboring highly commercialized towns were used as the source domain for model pre-training. Subsequently, the model was fine-tuned using estimated mobile payment transaction frequencies from 120 sampled grids in Zhouzhuang Town, Kunshan City, Suzhou, Jiangsu Province, as labels.

[0083] In stage S2, the image patch size of the Transformer model is set to 32×32, and the number of encoder layers is increased to 6. By introducing a spatial prior mask matrix, a mask value is assigned to the continuous large area of ​​basic farmland within the town, thus cutting off the meaningless exploration of the model in areas where high-frequency commercial activities are unlikely to occur in advance, and outputting a 256-dimensional comprehensive feature vector.

[0084] In stage S3, inversion modeling is performed. The number of decision trees is set to M=200, and the initial spatial decay coefficient is 0.015. A spatial distribution penalty term (λ=0.05) is introduced into the Gini index. Since the plain town area lacks natural barriers such as mountains and rivers, the penalty term mainly plays a role based on the administrative village boundaries, avoiding the appearance of illogical "enclaves" of economic inversion indicators at the boundaries of different administrative villages.

[0085] In stage S4, SHAP loop closure optimization was performed. A removal threshold of r=0.10 was set, eliminating 28 invalid features. Moran's I analysis showed that the features of the town center plain area were highly homogeneous; therefore, in the second loop closure, the spatial attenuation coefficient of the central region was adaptively reduced to 0.008, expanding the spatial sensing range.

[0086] In the S5 stage, the evaluation results are derived. The closed-loop mechanism significantly reduces the inversion calculation time from 58 minutes to 31 minutes. The inversion result R² reaches 0.91. The system terminal output shows that "path distance to the town-level logistics distribution center" with a high SHAP value of 0.34 has become the most core spatial factor driving the microeconomic vitality of the town.

[0087] (3) In the third embodiment, the extreme performance of the various improvement mechanisms of the present invention is verified for remote mountainous environments with rugged terrain, extremely high spatial impedance and severe data scarcity.

[0088] In Phase S1, a database of remote villages in Zhaojue County, Liangshan Yi Autonomous Prefecture, Sichuan Province, was constructed. Landsat 8 imagery at 30m resolution, high-precision DEM at 12.5m resolution, and mobile signaling trajectory data were acquired. The grid cell side length L was set to 100 meters, generating 850 effective grids. Field sampling in this village was extremely difficult (only 85 grid labels were obtained). The "Urban-Rural Correspondence Migration Learning" module played a crucial role here: large-scale pre-training was completed using detailed census data on population loss from a major city-level village over five years. Subsequently, only shallow network fine-tuning was performed when migrating to this mountain village, successfully overcoming the cold start obstacle.

[0089] In stage S2, the Transformer model image patch size was set to 8×8, and the number of encoder layers was set to 4. Slope data was extracted using the DEM to construct a strong prior mask matrix, which accurately identified and completely masked the attention features of steep cliffs and uninhabited ridges with slopes greater than 25 degrees, ensuring that the model was fully focused on the habitable area of ​​the valley and outputting a 64-dimensional feature vector.

[0090] In stage S3, the percentage decrease in permanent residents (0-100%) is used as the inversion value. The decision tree M=300, and the spatial decay coefficient is set to 0.05. To address the severe physical barriers in mountainous areas, the penalty coefficient for the spatial constraint Gini index is increased to λ=0.2. When the random forest attempts to forcibly cluster the grids on the north and south banks of the canyon separated by a rushing river, the spatial distribution penalty term P... spatial The highly activated and rejected splitting scheme gives the inversion model a strong ability to perceive terrain and identify fault zones.

[0091] In stage S4, SHAP closed-loop optimization is performed. A removal threshold of r=0.05 is set. Local Moran's I analysis accurately captures the abrupt differences in the river canyon area. In the second closed-loop layer, the spatial attenuation coefficient of this canyon area is locally amplified to 0.08, cutting off the weight interference from the opposite bank.

[0092] In stage S5, the evaluation results are output. With the support of strong spatial constraints and transfer learning, the model's mean absolute error on the validation set is controlled within 4.2%. Attribution analysis clearly indicates that the synergistic effect of the two indicators—"average terrain slope greater than 15 degrees" and "walking impedance to the nearest medical facility exceeding 3 kilometers"—contributes more than 65% of the predicted population loss risk in this village area.

[0093] On the other hand, embodiments of the present invention provide a system for full-domain spatial extrapolation of indicators for data-sparse village areas, comprising: a data acquisition module, used to divide the target village area into multiple grid units, acquire multi-source heterogeneous data of the target village area, and extract remote sensing images, topographic data, spatial distribution data of points of interest, and road network vector data corresponding to each grid unit, and extract the measured values ​​of target indicators for sampling grid units; and a feature extraction and fusion module, used to extract visual features of remote sensing images using a visual representation model obtained by transfer learning, and introduce a spatial prior mask constructed from topographic data to adjust the response intensity of visual features, and fuse the adjusted features. Visual features, spatial distribution data of points of interest, and road network vector data are used to obtain multi-source fusion representations of each grid unit. The model training module is used to train the constructed spatial random forest model using the measured values ​​of the target index and the multi-source fusion representations corresponding to the grid units used for sampling. A splitting criterion with a geospatial distribution penalty is used to optimize between attribute similarity and spatial continuity to obtain the trained spatial random forest model. The global prediction module is used to input the multi-source fusion representations of each grid unit into the trained spatial random forest model, perform spatial extrapolation prediction covering the target village area, and output the global extrapolation values ​​of the target index corresponding to each grid unit.

[0094] Then, this embodiment of the invention provides an index global spatial extrapolation device for sparse data villages, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the index global spatial extrapolation method for sparse data villages as described above.

[0095] Furthermore, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method for global spatial extrapolation of indicators for sparse data regions.

[0096] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0097] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions.

[0098] Furthermore, it should be noted that in the description of this specification, the terms "one embodiment," "some embodiments," "embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0099] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the claims should be interpreted to include both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0100] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, then this invention should also include these modifications and variations.

Claims

1. A method for deriving a global space of indicators oriented to data sparse villages, characterized in that, include: The target village area is divided into multiple grid units. Multi-source heterogeneous data of the target village area are obtained and the remote sensing images, topographic data, spatial distribution data of points of interest and road network vector data corresponding to each grid unit are extracted. The measured values ​​of target indicators of the grid units used for sampling are also extracted. Visual features of remote sensing images are extracted using a visual representation model obtained through transfer learning. A spatial prior mask constructed from topographic data is introduced to adjust the response intensity of the visual features. The adjusted visual features, spatial distribution data of points of interest, and road network vector data are fused to obtain a multi-source fusion representation of each grid cell. The constructed spatial random forest model is trained using the measured values ​​of the target index and the multi-source fusion representations corresponding to the grid cells used for sampling. Furthermore, a splitting criterion that introduces a geospatial distribution penalty is adopted to optimize between attribute similarity and spatial continuity, resulting in the trained spatial random forest model. The multi-source fusion representation of each grid cell is input into the trained spatial random forest model to perform spatial extrapolation prediction covering the target village area, and outputs the global extrapolation value of the target index corresponding to each grid cell.

2. The method of claim 1, wherein the method is a method of deriving a global space of indicators for data sparsity. The target village area was divided into multiple grid cells. Multi-source heterogeneous data of the target village area was acquired, and remote sensing imagery, topographic data, spatial distribution data of points of interest, and road network vector data corresponding to each grid cell were extracted. Measured values ​​of target indicators for the grid cells used for sampling were also extracted, including: Establish a unified spatial coordinate system covering the target village area, and divide the target village area into multiple grid units based on the unified spatial coordinate system; Obtain multi-source heterogeneous data of the target village area from at least one data source, and extract the corresponding remote sensing images, topographic data, spatial distribution data of points of interest and road network vector data within each grid cell from the multi-source heterogeneous data; Extracted measured values ​​of target indicators from partial grid cells used for sampling from multi-source heterogeneous data; Perform preprocessing including normalization on remote sensing imagery, topographic data, spatial distribution data of points of interest, and road network vector data; The target indicators include any one or more of the following: rural resilience index, economic activity level, poverty level, industrial advantage level, infrastructure level, and production / living / ecological development level.

3. The method of claim 1, wherein the data sparse village-oriented indicator global space derivation method is characterized by, The transfer learning steps for visual representation models include: Obtain remote sensing images of reference villages with labeled data, as well as measured values ​​of target indicators corresponding to the reference villages. Combine the remote sensing images of the reference villages with the measured values ​​of target indicators corresponding to the reference villages to form reference sample data. The remote sensing images from the reference sample data are input into a visual representation model that includes a feature extraction network and a classifier network, and the measured values ​​of the target indicators from the reference sample data are used as pre-training supervision signals. The visual representation model is driven by pre-trained supervised signals to learn the mapping relationship between remote sensing images and the measured values ​​of target indicators corresponding to reference villages. The weight parameters of the feature extraction network and the classifier network are iteratively updated until the model converges, and the pre-trained visual representation model is output. Keep the weight parameters of the feature extraction network in the pre-trained visual representation model constant; The remote sensing images corresponding to the grid cells used for sampling within the target village area are input into the pre-trained visual representation model. The measured values ​​of the target indicators of the grid cells used for sampling are used as the supervision signal for transfer update. The weight parameters of the classifier network are updated only through error backpropagation, and the visual representation model that has completed transfer learning is output.

4. The method of claim 2, wherein the data sparse village-oriented indicator global space derivation method is characterized by, Visual features of remote sensing images are extracted using a visual representation model obtained through transfer learning. A spatial prior mask constructed from topographic data is introduced to adjust the response intensity of the visual features. The adjusted visual features, spatial distribution data of points of interest, and road network vector data are then fused to obtain a multi-source fused representation for each grid cell, including: The remote sensing image corresponding to each grid unit is divided into multiple image blocks; Based on the topographic data corresponding to each grid cell, identify topographically abnormal areas with slopes greater than the angle threshold or located within water bodies; Construct a spatial prior mask matrix corresponding to the positions of multiple image patches, assign a preset mask value to the corresponding position of the terrain anomaly region in the spatial prior mask matrix, and assign a value of zero to the corresponding position of the region outside the terrain anomaly region in the spatial prior mask matrix. Multiple image patches are input into the visual representation model, and the spatial prior mask matrix is ​​introduced into the self-attention mechanism of the visual representation model. This reduces the attention weight of the image patch corresponding to the terrain anomaly region when the self-attention mechanism participates in the calculation, and outputs the adjusted visual features. Based on the spatial distribution data of interest points and road network vector data corresponding to each grid unit, calculate the socio-economic feature data including the estimated kernel density of interest points and the road network density, and map it into a socio-economic feature vector that matches the adjusted visual feature dimension. Cross-modal attention interaction calculation is performed based on the adjusted visual features and socioeconomic feature vectors to obtain cross-modal contextual features; The adjusted visual features are weighted and fused with cross-modal contextual features to output the comprehensive geospatial feature vector corresponding to each grid unit as a multi-source fusion representation.

5. The method for global spatial extrapolation of indicators for sparsely dataed village areas as described in claim 1, characterized in that, The constructed spatial random forest model is trained using the measured values ​​of the target index and the multi-source fusion representations corresponding to the grid cells used for sampling. A splitting criterion incorporating geospatial distribution penalties is employed to optimize the relationship between attribute similarity and spatial continuity, resulting in the trained spatial random forest model, which includes: The multi-source fusion representations corresponding to the grid cells used for sampling are used as training features, and the measured values ​​of the target indicators of the grid cells used for sampling are used as training labels to construct the decision tree in the spatial random forest model. At the current split node of the decision tree, based on the multi-source fusion representation corresponding to the grid cell used for sampling that reaches the current split node, a splitting scheme that generates left and right child nodes is generated in the feature domain. According to the splitting scheme, the grid cells for sampling that arrive at the current splitting node are assigned to the corresponding left and right child nodes. Based on the measured values ​​of the target index of the grid cells assigned to each child node, the change in Gini impurity of the current splitting node before and after splitting is calculated, and the decrease in Gini impurity corresponding to the splitting scheme is obtained. Based on the geographical adjacency status, natural barriers, or administrative boundaries between the grid cells used for sampling that reach the current split node, determine the geospatial distribution penalty term value corresponding to the splitting scheme; The local weighted Gini index is obtained by calculating the difference between the Gini impurity decrease and the geospatial distribution penalty term. The splitting scheme with the largest local weighted Gini index is selected as the final splitting criterion for the current splitting node, and the node is split until the preset stopping condition is met. The trained spatial random forest model is then output.

6. The method of claim 1, wherein the data sparse village-oriented indicator global space derivation method is characterized by, The multi-source fusion representations of each grid cell are input into the trained spatial random forest model to perform spatial extrapolation prediction covering the target village area, outputting the global extrapolation values ​​of the target indicators corresponding to each grid cell, including: Extract grid cells within the target village area that do not contain measured values ​​of the target indicator as grid cells to be predicted; The multi-source fusion representation corresponding to the grid cell to be predicted is input into the trained spatial random forest model. The multi-decision trees in the trained spatial random forest model are used to perform parallel inference calculation on the multi-source fusion representation corresponding to the grid cell to be predicted. The output results of the multi-decision trees are aggregated to obtain the inferred value corresponding to the grid cell to be predicted. The inferred values ​​corresponding to the grid cells to be predicted are summarized with the measured values ​​of the target indicators of the grid cells used for sampling to generate an indicator inversion array covering the target village area. Map the index inversion array to the corresponding grid cell position within the target village area, and output the global inversion value of the target index corresponding to each grid cell.

7. The method of claim 1-6, wherein, After inputting the multi-source fusion representations of each grid cell into the trained spatial random forest model, performing spatial extrapolation prediction covering the target village region, and outputting the global extrapolation value of the target index corresponding to each grid cell, the model also includes: The global inference value of the target index is decomposed into the basic expected value of the spatial random forest model and the bias corresponding to each input feature in the multi-source fusion representation, and the bias is used as the local marginal contribution value of each input feature in each grid cell. Based on the spatial adjacency relationship between each grid cell, the spatial contribution clustering degree of the local marginal contribution value in the spatial distribution is calculated; Calculate the global average of all local marginal contribution values ​​corresponding to each input feature and take the absolute value to obtain the absolute value of the global average marginal contribution. When the absolute value of the global average marginal contribution of the input feature is lower than the preset contribution threshold, and the corresponding spatial contribution clustering is lower than the preset clustering threshold, the corresponding input feature is removed, and the output is a retained feature matrix composed of the retained input features. Calculate the local spatial variance of the local marginal contribution value of each input feature in the retained feature matrix within the spatial adjacency range. Divide the grid cells with local spatial variance greater than the preset fluctuation threshold into heterogeneous regions and divide the grid cells with local spatial variance not greater than the fluctuation threshold into homogeneous regions. The spatial attenuation coefficient is adaptively increased for heterogeneous regions and adaptively decreased for homogeneous regions to update the spatial weights; whereby the spatial attenuation coefficient is a control parameter that determines the mapping relationship of spatial weight attenuation with increasing distance. The retained feature matrix and spatial weights are then input back into the spatial random forest model for iterative training until the preset convergence condition is met. The optimal adaptive spatial random forest model and the global inference value of the target index after iteration are then output.

8. A data sparse village-oriented index universal space deduction system, characterized in that, include: The data acquisition module is used to divide the target village area into multiple grid units, acquire multi-source heterogeneous data of the target village area, and extract remote sensing images, topographic data, spatial distribution data of points of interest and road network vector data corresponding to each grid unit, and extract the measured values ​​of target indicators of the grid units used for sampling. The feature extraction and fusion module is used to extract visual features from remote sensing images using a visual representation model obtained through transfer learning. It also introduces a spatial prior mask constructed from topographic data to adjust the response intensity of the visual features. The adjusted visual features, spatial distribution data of points of interest, and road network vector data are fused to obtain a multi-source fusion representation of each grid cell. The model training module is used to train the constructed spatial random forest model using the measured values ​​of the target index and the multi-source fusion representations corresponding to the grid cells used for sampling. It also adopts a splitting criterion that introduces a geospatial distribution penalty to optimize between attribute similarity and spatial continuity, thus obtaining the trained spatial random forest model. The global prediction module is used to input the multi-source fusion representation of each grid cell into the trained spatial random forest model, perform spatial extrapolation prediction covering the target village area, and output the global extrapolation value of the target index corresponding to each grid cell.

9. An index global space deduction device oriented to data sparse village areas, characterized in that, include: At least one processor; and memory that is communicatively connected to at least one processor; The memory stores instructions that can be executed by at least one processor, which are executed by at least one processor to enable the at least one processor to execute the index global spatial extrapolation method based on data sparse village domains as described in any one of claims 1-7.

10. A computer-readable storage medium having stored thereon computer- executable instructions, the computer-executable instructions comprising instructions for: receiving a request for a resource; determining whether the request is for a resource that is subject to a policy; and if the request is for a resource that is subject to a policy, then determining whether the request is from a client that is subject to the policy. When the executable instructions are executed by the processor, they implement the index global spatial extrapolation method for sparse data villages as described in any one of claims 1-7.