Hollow village identification method and system based on multi-source data fusion
Through the hollow village identification method of multi-source data fusion, static and dynamic features are extracted using remote sensing images, village scene pictures and time series night light data, the problems of inaccurate human and material resources consumption and identification in hollow village identification are solved, and efficient and accurate hollow village monitoring is achieved.
Patent Information
- Application Number
- CN202510284834.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-04
AI Technical Summary
The identification of hollow villages in the existing technology relies on field surveys and statistical data, consumes a lot of manpower and material resources, and is difficult to achieve large-scale and real-time monitoring. It is difficult for a single data source to fully capture the spatial characteristics and population dynamics of hollow villages, resulting in inaccurate identification results.
Using the hollow village recognition method of multi-source data fusion, static features are extracted from the remote sensing image through the first feature extraction unit, the second feature extraction unit extracts microscopic features from the village scene pictures, and the third feature extraction unit extracts dynamic features from the time series night light data, and fuses the three into a comprehensive feature vector to output the prediction probability of the hollow village.
Accurate identification of hollow villages is achieved, identification efficiency and accuracy are improved, multi-source data fusion makes up for the bias of a single data source, adapts to the non-standardized characteristics of rural data, and is suitable for large-scale rural monitoring.
Smart Images

Figure CN120259874A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hollow village identification, and more specifically, to a method and system for identifying hollow villages based on multi-source data fusion. Background Art
[0002] The identification of hollow villages is crucial for rural governance and revitalization. Traditional methods mainly rely on on-site investigations and statistical data to infer the degree of hollowness, but these methods require a large amount of manpower and material resources and are difficult to achieve large-scale and real-time monitoring.
[0003] In recent years, emerging data sources such as remote sensing images, street view pictures, and nighttime light data have been used for community environmental assessment. However, existing research mainly focuses on single data sources, making it difficult to comprehensively capture the spatial characteristics and population dynamics of hollow villages, resulting in inaccurate identification results of hollow villages. For example, remote sensing images can reflect land vacancy but cannot reflect population activities; nighttime light data reveals the intensity of human activities but lacks spatial details; village view pictures provide micro-environmental information, but the data is disordered and the coverage is limited. Summary of the Invention
[0004] To overcome the defect that it is difficult to accurately identify hollow villages based on a single data source in the above-mentioned prior art, the present invention provides a method and system for identifying hollow villages based on multi-source data fusion that integrates multi-source data and comprehensively considers static and dynamic characteristics.
[0005] To solve the above technical problems, the technical solution of the present invention is as follows:
[0006] A hollow village identification model, comprising: a first feature extraction unit, a second feature extraction unit, a third feature extraction unit, and a multi-source data fusion and identification unit;
[0007] The first feature extraction unit is configured to convert the received remote sensing image into a high-dimensional feature map, and after performing attention weighting on the high-dimensional feature map in the channel and spatial dimensions, convert the weighted feature map into a static feature vector;
[0008] The second feature extraction unit is configured to flatten the received village view picture into a feature vector of a preset dimension, and aggregate the feature vector into a micro feature vector based on an attention mechanism;
[0009] The third feature extraction unit is configured to capture the long-term dependencies in the time series from the received time series nighttime light data, and capture local features from the time series nighttime light data, and generate a dynamic feature vector based on the long-term dependencies and the local features;
[0010] The multi-source data fusion and recognition unit is used to splice the static feature vector, the microscopic feature vector, and the dynamic feature vector into a comprehensive feature vector, and is used to output the prediction probability that the village corresponding to the remote sensing image, the village scene picture, and the time series night light data is a hollow village based on the comprehensive feature vector.
[0011] The present invention also proposes a method for identifying hollow villages based on multi-source data fusion, including the following steps:
[0012] Obtain multi-source data composed of remote sensing images, village scene pictures, and time series night light data of the village to be identified;
[0013] Input the multi-source data into the hollow village recognition model;
[0014] The hollow village recognition model outputs the prediction probability that the village to be identified is a hollow village.
[0015] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0016] Based on multi-source data composed of remote sensing images, village scene pictures, and time series night light data, the present invention uses the first feature extraction unit, the second feature extraction unit, and the third feature extraction unit to obtain the static feature vector, the microscopic feature vector, and the dynamic feature vector respectively, and comprehensively describes the static environment and dynamic population characteristics of the hollow village by fusing these three feature vectors, so as to obtain an accurate hollow village recognition result. Description of the Drawings
[0017] Figure 1 It is a schematic structural diagram of the hollow village recognition model proposed in Embodiment 1;
[0018] Figure 2 It is a sample example diagram of the remote sensing image proposed in Embodiment 3;
[0019] Figure 3 It is a sample example diagram of the village scene picture proposed in Embodiment 3;
[0020] Figure 4 It is a sample example diagram of the time series night light data proposed in Embodiment 3;
[0021] Figure 5 It is a comparison diagram of the prediction results of different single data sources and the combination of multi-source data proposed in Embodiment 3;
[0022] Figure 6 It is a distribution diagram of the prediction results of hollow villages in four counties proposed in Embodiment 3. Detailed Embodiments
[0023] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0024] To better illustrate this embodiment, some components in the drawings are omitted, enlarged, or reduced, which does not represent the size of the actual product;
[0025] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0026] The technical solution of the present invention will be further described below in conjunction with the drawings and embodiments.
[0027] Embodiment 1
[0028] This embodiment proposes an idle village recognition model. Figure 1 is a schematic structural diagram of the idle village recognition model proposed in this embodiment; among them, Figure 1 The first fully connected layer is not drawn in the third feature extraction unit of.
[0029] As Figure 1 shown, the idle village recognition model of this embodiment includes: a first feature extraction unit, a second feature extraction unit, a third feature extraction unit, and a multi-source data fusion recognition unit;
[0030] The first feature extraction unit is used to convert the received remote sensing image into a high-dimensional feature map, and after performing attention weighting on the high-dimensional feature map in the channel and spatial dimensions, convert the weighted feature map into a static feature vector;
[0031] The second feature extraction unit is used to flatten the received village scene picture into a feature vector of a preset dimension, and aggregate the feature vector into a microscopic feature vector based on the attention mechanism;
[0032] The third feature extraction unit is used to capture the long-term dependencies in the time series from the received time series night light data, and capture local features from the time series night light data, and generate a dynamic feature vector based on the long-term dependencies and the local features;
[0033] The multi-source data fusion recognition unit is used to splice the static feature vector, the microscopic feature vector, and the dynamic feature vector into a comprehensive feature vector, and output the prediction probability that the village corresponding to the remote sensing image, the village scene picture, and the time series night light data is an idle village based on the comprehensive feature vector.
[0034] In the specific implementation process, the present invention is based on multi-source data composed of remote sensing images, village scene pictures, and time series night light data. The first feature extraction unit, the second feature extraction unit, and the third feature extraction unit are used to obtain static feature vectors, microscopic feature vectors, and dynamic feature vectors respectively. By fusing these three feature vectors, the static environment and dynamic population characteristics of hollow villages are comprehensively characterized, so as to obtain accurate hollow village recognition results.
[0035] In an optional embodiment, a first residual network module and a convolutional block attention module are arranged in the first feature extraction unit;
[0036] The first residual network module is used to convert the received remote sensing image into a high-dimensional feature map;
[0037] The convolutional block attention module is used to generate channel weights of the high-dimensional feature map based on channel attention, generate spatial weights of the high-dimensional feature map based on spatial attention, and use the channel weights and spatial weights to weight the high-dimensional feature map to obtain a weighted feature map; the weighted feature map is subjected to global average pooling and fully connected operations to obtain a static feature vector.
[0038] As an exemplary illustration, the first residual network module implements the functions of this module based on a pre-trained ResNet18 model.
[0039] In an optional embodiment, a second residual network module and a SetTransformer module are arranged in the second feature extraction unit;
[0040] The second residual network module is used to flatten the received village scene picture into a feature vector of a preset dimension;
[0041] The Set Transformer module is used to perform two self-attention calculations on the feature vector through a multi-head self-attention mechanism, capture the global relationship between features, and aggregate the feature vector by introducing induced points based on the global relationship to generate a microscopic feature vector of a preset dimension.
[0042] As an exemplary illustration, the second residual network module implements the functions of this module based on a pre-trained ResNet18 model.
[0043] In an optional embodiment, a Block LSTM module, a Block FCN module, and a first fully connected layer are arranged in the third feature extraction unit;
[0044] The Block LSTM module is used to capture long-term dependencies in the received time-series nightlight data; the Block FCN module is used to capture local features from the time-series nightlight data; the output features of the Block LSTM module and the Block FCN module are concatenated and then input into the first fully connected layer, and the first fully connected layer generates a dynamic feature vector based on the concatenated output features;
[0045] The Block LSTM module at least includes a dimension shuffle layer, an LSTM layer, and a Dropout layer connected in sequence, and the output of the Dropout layer is the output of the Block LSTM module;
[0046] The Block FCN module at least includes a plurality of convolutional sub-modules and a global pooling layer connected in sequence. Among them, each convolutional sub-module at least includes a one-dimensional convolutional layer, a batch normalization layer, and an activation function layer. The activation function layer of the last convolutional sub-module is connected to the global pooling layer, and the output of the global pooling layer is the output of the Block FCN module.
[0047] In an optional embodiment, the multi-source data fusion recognition unit at least includes a feature concatenation layer, a second fully connected layer, and a Softmax layer connected in sequence;
[0048] The feature concatenation layer is used to concatenate the static feature vector, the microscopic feature vector, and the dynamic feature vector into a comprehensive feature vector;
[0049] After being processed by the second fully connected layer and the Softmax layer, the comprehensive feature vector is transformed into the predicted probability that the village corresponding to the remote sensing image, the village scene picture, and the time-series nightlight data is an abandoned village.
[0050] As an exemplary illustration, when the feature concatenation layer concatenates the static feature vector, the microscopic feature vector, and the dynamic feature vector into a comprehensive feature vector, it can adopt a direct concatenation method or a method that achieves a concatenation effect by fusing based on the Mixer attention mechanism.
[0051] Embodiment 2:
[0052] This embodiment proposes a method for identifying abandoned villages based on multi-source data fusion as described in Embodiment 1.
[0053] The method for identifying abandoned villages based on multi-source data fusion includes the following steps:
[0054] S1: Obtain multi-source data composed of the remote sensing image, the village scene picture, and the time-series nightlight data of the village to be identified;
[0055] S2: Input the multi-source data into the hollow village recognition model;
[0056] S3: The hollow village recognition model outputs the prediction probability that the village to be recognized is a hollow village.
[0057] In an alternative embodiment, before inputting the multi-source data into the hollow village recognition model, preprocess the multi-source data. The steps for preprocessing the multi-source data include:
[0058] Crop the remote sensing image into grids of a preset size;
[0059] Map the village scene pictures to the grids and perform size adjustment and color correction;
[0060] Combine the rural housing density data to filter the noise of the time-series night light data.
[0061] As an illustrative example, the remote sensing image data is sourced from Google Earth with a spatial resolution of 0.3 meters per pixel.
[0062] The village scene pictures are collected by volunteers through the "Village-by-Village Shooting" platform and contain information such as rural houses, roads, and signs of human activities.
[0063] The time-series night light data is sourced from NOAA, covering from January 2022 to January 2024, with a spatial resolution of approximately 500 meters per pixel.
[0064] As an illustrative example, during preprocessing, crop and resample the remote sensing image and uniformly adjust it to a size of 224×224 pixels; perform quality screening and standardization processing (such as adjusting brightness, contrast, and color) on the village scene pictures, and use Set-Transformer to aggregate features; perform band operations on the night light data to remove the interference of moonlight reflected by vegetation and integrate it into time-series data.
[0065] In an alternative embodiment, before inputting the multi-source data into the hollow village recognition model, train the hollow village recognition model. The training steps include:
[0066] Obtain the multi-source data of several villages with recognition labels as the training set. Among them, the recognition label of a village with a population outflow rate greater than or equal to a preset percentage is a hollow village label, and the recognition label of a village with a population outflow rate less than the preset percentage is a non-hollow village label;
[0067] Input the training set into the hollow village recognition model and train the hollow village recognition model. During the training process, iteratively solve the preset weighted cross-entropy loss function. When the number of iterations reaches the preset value or the weighted cross-entropy loss function reaches the minimum value, end the training to obtain a trained hollow village recognition model;
[0068] When inputting the multi-source data into the hollow village recognition model, input the multi-source data into the trained hollow village recognition model.
[0069] In an optional embodiment, the expression of the population outflow rate includes:
[0070]
[0071] In the formula, POR i represents the population outflow rate of the i-th village in the training set, RP i represents the registered population of the i-th village in the training set, and PP i represents the permanent population of the i-th village in the training set.
[0072] This embodiment also proposes a hollow village recognition system based on multi-source data fusion for implementing the method for recognizing hollow villages based on multi-source data fusion according to this embodiment, including:
[0073] A data acquisition module for acquiring multi-source data composed of remote sensing images, village scene pictures, and time series night light data of the village to be recognized;
[0074] A recognition result output module configured with a hollow village recognition model for inputting the multi-source data into the hollow village recognition model, and the hollow village recognition model outputs the prediction probability that the village to be recognized is a hollow village.
[0075] This embodiment also proposes a computer device, including a memory and a processor. When the computer-readable instructions stored in the memory are executed by the processor, the processor executes the steps of the method for recognizing hollow villages based on multi-source data fusion according to this embodiment.
[0076] Embodiment 3:
[0077] This embodiment is based on the hollow village recognition model proposed in Embodiment 1 and the method for recognizing hollow villages based on multi-source data fusion proposed in Embodiment 2, Figure 2 which is a sample example diagram of remote sensing images proposed for this embodiment; Figure 3 which is a sample example diagram of village scene pictures proposed for this embodiment; Figure 4 which is a sample example diagram of time series night light data proposed for this embodiment; Figure 5Comparison chart of prediction results for different single data sources and combined multi-source data proposed in this embodiment; Figure 6 Distribution map of the prediction results of hollow villages in four counties proposed in this embodiment, as Figures 2 to 6 described, this embodiment proposes a specific implementation example:
[0078] When the method for identifying hollow villages based on multi-source data fusion is specifically applied, samples as Figures 2 to 4 described are used for specific application, and the steps for preprocessing these sample data include:
[0079] Obtain multi-source data: including Google Earth high-resolution remote sensing images in 2020 (resolution 0.3 meters), village scene pictures from the "Village-by-Village Shooting" platform from 2022 to 2024, and NOAA monthly nighttime light data from January 2022 to January 2024 (resolution about 500 meters).
[0080] Data standardization: Crop the remote sensing images into 500×500-meter grids, map the village scene pictures to the grids and perform size adjustment and color correction, and filter noise for the time-series nighttime light data in combination with the rural housing density data.
[0081] Data augmentation: Randomly rotate, flip, and adjust the brightness of the remote sensing images and village scene pictures, and use the dynamic time warping averaging method to augment the sample size for the nighttime light data.
[0082] The content of feature extraction includes:
[0083] The first feature extraction unit: Use the ResNet18 model to extract the high-dimensional feature map of the remote sensing image, and combine the convolutional block attention module (CBAM) to perform attention weighting on the feature map in the channel and spatial dimensions.
[0084] The CBAM (convolutional block attention) module consists of two parts: channel attention (SEBlock) and spatial attention (SpatialAttention):
[0085] Channel attention (SEBlock): Calculate the importance weights of each channel through global average pooling, and use a fully connected layer and a Sigmoid activation function to generate channel weights.
[0086] Spatial attention (SpatialAttention): Generate a spatial weight map by calculating the spatial mean and maximum value of the feature map, and use a convolutional layer and a Sigmoid activation function to generate spatial weights.
[0087] The weighted feature map generates a 256-dimensional static feature vector through global average pooling and fully connected operations.
[0088] Second Feature Extraction Unit: Use the pre-trained ResNet18 model to extract the features of each image from the village scene pictures and flatten them into feature vectors with a fixed dimension.
[0089] Use Set-Transformer to aggregate the feature vectors, which specifically includes the following steps:
[0090] Self-Attention Calculation (SAB): Perform self-attention calculation on the feature vectors twice through the multi-head self-attention mechanism (MultiheadAttention) to capture the global relationships between features.
[0091] Multi-Head Attention Pooling (PMA): Introduce inducing points to aggregate the feature vectors and generate a feature representation with a fixed length.
[0092] Finally, output a 256-dimensional microscopic feature vector.
[0093] Third Feature Extraction Unit: Split the time series of nighttime light data into two subsequences and input them into the Block LSTM module respectively.
[0094] The Block LSTM module consists of multiple layers of LSTM units and is used to capture the long-term dependencies in the time series.
[0095] Meanwhile, use the Block FCN module to perform convolutional feature extraction on the time series data to capture local features.
[0096] The output features of the Block LSTM module and the Block FCN module are concatenated and mapped through a fully connected layer to generate a 256-dimensional dynamic feature vector.
[0097] Multi-Source Data Fusion and Recognition Unit: Directly concatenate the 256-dimensional features of the remote sensing images, village scene pictures, and nighttime light data to form a 768-dimensional comprehensive feature vector.
[0098] Output the classification probabilities of hollow villages (HV) and non-hollow villages (None-HV) through a fully connected layer and the Softmax function.
[0099] Contents of Model Training and Evaluation:
[0100] Based on the 2023 rural construction evaluation questionnaire data, define that the population outflow rate exceeding 50% is a hollow village, and lower than 50% is a non-hollow village.
[0101]
[0102] Among them, PORi is the population outflow rate of village i, RPi is the registered population of village i, and PRi is the resident population of village i.
[0103] Use the weighted cross-entropy loss function to handle class imbalance, train the model in combination with the stochastic gradient descent optimizer, and set the early stopping mechanism and learning rate scheduling.
[0104] The evaluation metrics include overall accuracy (OA), Kappa index, and F1 score.
[0105] Content of technical effects:
[0106] In the experiments in four counties (Huaiji County, Lianping County, Wengyuan County, Yangxi County), the test accuracy of the present invention reaches 0.8451, the Kappa index is 0.6391, and the F1 score is 0.8880, which is significantly better than the single data source models (Remote sensing (RS): 0.7441, Nighttime light (NTL): 0.8425, Village view image (VVI): 0.7270).
[0107] Temporal nighttime light data plays a key role in identifying hollow villages, and remote sensing images and village view pictures, as supplementary data, improve the robustness of the model.
[0108] The effect of aggregating the features of village view pictures using Set-Transformer is better than that of average pooling and Vision-LSTM.
[0109] Direct feature concatenation is superior to attention fusion (Mixer), avoiding noise interference between low-correlation data.
[0110] Table 1 is a comparison table of the results of single-source and multi-source data proposed in this embodiment, and the content of Table 1 is as follows:
[0111] Table 1 Comparison of the results of single-source and multi-source data
[0112]
[0113] Table 2 is a comparison table of the results of different fusion methods proposed in this embodiment, and the content of Table 2 is as follows:
[0114] Table 2 Comparison of the results of different fusion methods
[0115]
[0116] The experiment was conducted on the data of four counties. Combining Figure 5 、 Figure 6 、Table 1 and Table 2, it can be seen that the experimental results show that:
[0117] The overall accuracy of the multi-source data fusion model reaches 0.8451, which is better than the single data source model.
[0118] Nighttime light data plays a key role in identifying hollow villages, while remote sensing images and village scene pictures provide important supplementary information.
[0119] The present invention realizes the spatio-temporal coupling operation of multi-source heterogeneous data by integrating three types of data, namely high-resolution remote sensing images (0.3 m), disordered village scene pictures, and time-series nighttime light data (500 m) simultaneously, and the data time window covers 2020 - 2024. By using a feature extraction architecture composed of three feature extraction units, it solves the problem of single data source bias. The test accuracy is improved by 10.2% (from 0.7441 to 0.8451), and the Kappa index is improved by 0.2363 (from 0.4028 to 0.6391).
[0120] Based on Set-Transformer, it captures the global relationship between village scene pictures through self-attention calculation (SAB), introduces inducing points for multi-head attention pooling (PMA), and generates a fixed-length feature vector. Compared with traditional average pooling (OA = 0.8438, Kappa = 0.6325) and Vision-LSTM (OA = 0.8373, Kappa = 0.6218), Set-Transformer achieves OA = 0.8438, Kappa = 0.6391, and adapts to the unstructured data characteristics of rural scenes (such as disordered shooting angles and unfixed number of images).
[0121] Among them, 256-dimensional static (remote sensing), 256-dimensional microscopic (village scene), and 256-dimensional dynamic (nighttime light) features are directly concatenated into 768-dimensional comprehensive features, which are mapped to the classification space through a fully connected layer, avoiding attention interference between low-correlation features, and is superior to the attention fusion based on Mixer (the OA of direct concatenation is improved by 2.89%), realizing feature complementarity between village scene and nighttime light data.
[0122] The present invention aims to solve the problems of low accuracy in identifying hollow villages and single data source in the prior art, and provides a hollow village identification model and identification method to improve the identification efficiency and accuracy. It solves three core problems existing in the prior art: spatio-temporal coupling modeling of static environmental features and dynamic population activities, semantic feature aggregation of disordered village scene pictures, and feature-level fusion optimization of multi-source heterogeneous data.
[0123] Compared with the prior art, the advantages and beneficial effects of this patent are as follows:
[0124] Multi-source data fusion compensates for the bias of a single data source and comprehensively depicts the static environment and dynamic population characteristics of hollow villages; the Set-Transformer effectively processes unordered village scene pictures and adapts to the non-standard characteristics of rural data; the method is efficient and scalable and is applicable to large-scale rural area monitoring.
[0125] The same or similar reference numerals correspond to the same or similar components;
[0126] The terms describing the positional relationship in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent;
[0127] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not intended to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. An idle village identification model, characterized in that, Including: A first feature extraction unit, a second feature extraction unit, a third feature extraction unit, and a multi-source data fusion and recognition unit; The first feature extraction unit is configured to convert the received remote sensing image into a high-dimensional feature map, and after performing attention weighting on the high-dimensional feature map in terms of channels and spatial dimensions, convert the weighted feature map into a static feature vector; The second feature extraction unit is configured to flatten the received village scene picture into a feature vector of a preset dimension, and aggregate the feature vector into a microscopic feature vector based on an attention mechanism; The third feature extraction unit is configured to capture long-term dependencies in the time series from the received time series of nighttime light data, and capture local features from the time series of nighttime light data, and generate a dynamic feature vector based on the long-term dependencies and the local features; The multi-source data fusion and recognition unit is configured to splice the static feature vector, the microscopic feature vector, and the dynamic feature vector into a comprehensive feature vector, and output a prediction probability that the village corresponding to the remote sensing image, the village scene picture, and the time series of nighttime light data is a hollow village based on the comprehensive feature vector.
2. The hollow village recognition model according to claim 1, wherein A first residual network module and a convolutional block attention module are provided inside the first feature extraction unit; The first residual network module is configured to convert the received remote sensing image into a high-dimensional feature map; The convolutional block attention module is configured to generate channel weights of the high-dimensional feature map based on channel attention, generate spatial weights of the high-dimensional feature map based on spatial attention, and use the channel weights and spatial weights to weight the high-dimensional feature map to obtain a weighted feature map; The weighted feature map undergoes global average pooling and fully connected operations to obtain a static feature vector.
3. The hollow village recognition model according to claim 1, characterized in that, A second residual network module and a Set Transformer module are provided inside the second feature extraction unit; The second residual network module is configured to flatten the received village scene picture into a feature vector of a preset dimension; The Set Transformer module is configured to perform self-attention calculation on the feature vector twice through a multi-head self-attention mechanism, capture the global relationship between features, and introduce induced points based on the global relationship to aggregate the feature vector to generate a microscopic feature vector of a preset dimension.
4. The hollow village recognition model according to claim 1, characterized in that A Block LSTM module, a Block FCN module, and a first fully connected layer are provided inside the third feature extraction unit; The Block LSTM module is configured to capture long-term dependencies in the time series from the received time series of nighttime light data; the Block FCN module is configured to capture local features from the time series of nighttime light data; the output features of the Block LSTM module and the Block FCN module are spliced and then input into the first fully connected layer, and the first fully connected layer generates a dynamic feature vector based on the spliced output features; The Block LSTM module at least includes a dimension shuffle layer, an LSTM layer, and a Dropout layer connected in sequence, and the output of the Dropout layer is the output of the Block LSTM module; The Block FCN module at least includes a plurality of convolutional sub-modules and a global pooling layer connected in sequence, where each convolutional sub-module at least includes a one-dimensional convolutional layer, a batch normalization layer, and an activation function layer, and the activation function layer of the last convolutional sub-module is connected to the global pooling layer, and the output of the global pooling layer is the output of the Block FCN module.
5. The hollow village recognition model according to any one of claims 1 to 4, characterized in that The multi-source data fusion and recognition unit at least includes a feature splicing layer, a second fully connected layer, and a Softmax layer connected in sequence; The feature splicing layer is used to splice the static feature vector, the microscopic feature vector, and the dynamic feature vector into a comprehensive feature vector; After being processed by the second fully connected layer and the Softmax layer, the comprehensive feature vector is converted into the prediction probability that the village corresponding to the remote sensing image, the village scene picture, and the time series night light data is a hollow village.
6. A method for identifying hollow villages based on multi-source data fusion, applying the hollow village identification model according to any one of claims 1 to 5, characterized in that, Including the following steps: Obtain multi-source data composed of the remote sensing image, the village scene picture, and the time series night light data of the village to be recognized; Input the multi-source data into the hollow village recognition model; The hollow village recognition model outputs the prediction probability that the village to be recognized is a hollow village.
7. The method for identifying hollow villages based on multi-source data fusion according to claim 6, characterized in that Before inputting the multi-source data into the hollow village recognition model, preprocess the multi-source data. The steps for preprocessing the multi-source data include: Crop the remote sensing image into grids of a preset size; Map the village scene picture to the grid and perform size adjustment and color correction; Combined with the rural housing density data, filter the noise of the time series night light data.
8. The method for identifying hollow villages based on multi-source data fusion according to claim 6 or 7, characterized in that, Before inputting the multi-source data into the hollow village recognition model, train the hollow village recognition model. The training steps include: Obtain multi-source data of several villages with recognition labels as the training set, where the recognition label of a village with a population outflow rate greater than or equal to a preset percentage is a hollow village label, and the recognition label of a village with a population outflow rate less than the preset percentage is a non-hollow village label; Input the training set into the hollow village recognition model and train the hollow village recognition model. During the training process, iteratively solve the preset weighted cross-entropy loss function. When the number of iterations reaches the preset value or the weighted cross-entropy loss function reaches the minimum value, end the training to obtain a trained hollow village recognition model; When inputting the multi-source data into the hollow village recognition model, input the multi-source data into the trained hollow village recognition model.
9. The method for identifying hollow villages based on multi-source data fusion according to claim 6 or 7, wherein The expression of the population outflow rate includes: where POR i represents the population outflow rate of the \(i\)-th village in the training set, RP i represents the registered population of the \(i\)-th village in the training set, PP i represents the permanent population of the \(i\)-th village in the training set.
10. A hollow village recognition system based on multi-source data fusion, which is used to implement the hollow village recognition method based on multi-source data fusion according to any one of claims 6 to 9, characterized in that, Including: A data acquisition module for obtaining multi-source data composed of the remote sensing image, the village scene picture, and the time series night light data of the village to be recognized; A recognition result output module configured with a hollow village recognition model for inputting the multi-source data into the hollow village recognition model, and the hollow village recognition model outputs the prediction probability that the village to be recognized is a hollow village.