Low-efficiency living space studying and judging method and system based on multi-source spatio-temporal data
By preprocessing multi-source spatiotemporal data and improving deep learning algorithms, combined with spatially constrained machine learning clustering models, the subjective dependence and heterogeneity problems of existing methods for identifying inefficient residential land have been solved, and accurate identification and judgment of inefficient residential spaces have been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN UNIV
- Filing Date
- 2026-03-27
- Publication Date
- 2026-04-28
AI Technical Summary
Existing methods for identifying inefficient residential land use suffer from problems such as strong model subjectivity and insufficient consideration of heterogeneity when integrating multi-source spatiotemporal big data, making it difficult to accurately identify the heterogeneity of different types of residential land use in cities.
By acquiring multi-source spatiotemporal data, performing preprocessing and index screening, and utilizing improved deep learning algorithms and spatially constrained machine learning clustering models, a multi-dimensional evaluation index for residential land efficiency is constructed to identify inefficient residential spaces.
It enables unified analysis of multi-source spatiotemporal data of the target city space, constructs high-quality indicator features, and accurately identifies inefficient living spaces through clustering models, thereby improving the discrimination accuracy and robustness.
Smart Images

Figure CN121935567A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of geographic information technology, and in particular to an inefficient method and system for assessing residential space based on multi-source spatiotemporal data. Background Technology
[0002] The urban development model is shifting from extensive "incremental expansion" to refined "stock optimization," and revitalizing inefficient land use is a prerequisite for achieving refined urban land use.
[0003] The existence of inefficient land use affects the normal development of cities. In recent years, the emergence of multi-source geospatial big data such as points of interest, mobile phone signaling, and street view images has provided multi-dimensional data support for land use efficiency assessment, from building attributes, population vitality, transportation accessibility to environmental quality, and promoted the identification of inefficient land use from subjective qualitative to quantitative analysis.
[0004] However, existing methods for identifying inefficient residential land have two main limitations: First, multi-criteria decision-making analyses rely on expert experience or pre-defined function forms in weight setting, introducing subjective biases, and the model structure is linear, making it difficult to characterize the non-linear relationships between multi-dimensional indicators. Second, cities have two distinct residential types: "residential communities" and "urban villages," which differ significantly in land use patterns, building forms, and socio-economic functions. Most existing studies treat urban residential land as a homogeneous whole, using the same indicators and models for identification, ignoring this inherent heterogeneity. This limits the accuracy, robustness, and practical guiding value of the identification models for specific land types. In summary, existing methods for identifying inefficient residential land suffer from strong model subjectivity and insufficient consideration of heterogeneity when integrating multi-source spatiotemporal big data.
[0005] Existing technologies still suffer from problems such as strong model subjectivity and insufficient consideration of heterogeneity when integrating multi-source spatiotemporal big data, resulting in inefficient residential land identification methods. Therefore, existing technologies need further improvement. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method and system for judging inefficient residential space based on multi-source spatiotemporal data, in order to solve the problems of strong subjective dependence and insufficient consideration of heterogeneity in the existing methods for judging inefficient residential land when integrating multi-source spatiotemporal big data.
[0007] The technical solution adopted by this invention to solve the technical problem is as follows: In a first aspect, the present invention provides a method for inefficient residential space assessment based on multi-source spatiotemporal data, including: Acquire multi-source spatiotemporal data of the target city space, preprocess the multi-source spatiotemporal data, and obtain the indicator feature vector; Based on indicator screening and improved deep learning algorithms, feature extraction is performed on the multi-source spatiotemporal data to obtain indicator features; A spatially constrained machine learning clustering model is constructed, and the index features are input into the machine learning clustering model to obtain the clustering results; The clustering results are used to identify inefficient residential spaces, thereby obtaining inefficient residential spaces in the target city space.
[0008] In one implementation, acquiring multi-source spatiotemporal data of the target city space includes: Select the target urban space; The target urban space is divided into a predetermined number of plots; Acquire multi-source spatiotemporal data for each of the plots; wherein the multi-source spatiotemporal data includes building data, population data, environmental data, economic data, and functional data.
[0009] In one implementation, the preprocessing of the multi-source spatiotemporal data to obtain the indicator feature vector includes: Data cleaning is performed on the multi-source spatiotemporal data; The cleaned multi-source spatiotemporal data were subjected to coordinate system transformation and spatial calibration. Integrate spatially calibrated multi-source spatiotemporal data; The integrated multi-source spatiotemporal data is transformed into indicators, and indicator feature vectors are constructed based on the obtained indicators.
[0010] In one implementation, the feature extraction of the multi-source spatiotemporal data based on the indicator screening and improved deep learning algorithm to obtain indicator features includes: The feature vector of the index is standardized. Based on the correlation coefficient, the standardized indicator feature vector is subjected to correlation analysis and screening to remove redundant indicators from the indicator feature vector; wherein, the redundant indicators are indicators whose correlation coefficient is higher than a preset correlation coefficient threshold. An autoencoder network is constructed based on an improved deep learning algorithm. The autoencoder network is used to extract features from the feature vectors of indicators after correlation analysis and screening, thereby obtaining indicator features.
[0011] In one implementation, the step of constructing an autoencoder network based on an improved deep learning algorithm, and extracting features from the indicator feature vectors filtered by the indicator correlation analysis through the autoencoder network to obtain indicator features, includes: Based on improved deep learning algorithms, network weights, and preset activation functions, an autoencoder network containing an encoder and a decoder is constructed. Calculate the loss function of the autoencoder network, and train and optimize the autoencoder network using the loss function; The autoencoder network is used to extract features from the feature vectors of the indicators after correlation analysis, thereby obtaining the indicator features.
[0012] In one implementation, constructing a spatially constrained machine learning clustering model, inputting the indicator features into the machine learning clustering model, and obtaining clustering results includes: Obtain the spatial constraints of the target city; Construct a Gaussian mixture model, optimize the Gaussian mixture model using the spatial constraint term, and obtain a machine learning clustering model based on spatial constraints; The index features are input into the machine learning clustering model to obtain clustering results; wherein, the clustering results are clusters based on multi-source spatiotemporal data.
[0013] In one implementation, the step of identifying inefficient residential spaces from the clustering results to obtain inefficient residential spaces in the target urban space includes: Obtain the index features corresponding to the clustering results; The characteristics of the indicators are analyzed based on multi-source spatiotemporal data, and the characteristics of the indicators are compared among different clustering results; Based on the analysis and comparison results of the indicator characteristics, the clustering results are used to identify inefficient living spaces. The land parcels in the target urban space that are identified as inefficient residential spaces are obtained.
[0014] Secondly, the present invention provides an inefficient residential space assessment system based on multi-source spatiotemporal data, comprising: The data acquisition module is used to acquire multi-source spatiotemporal data of the target city space, and preprocess the multi-source spatiotemporal data to obtain indicator feature vectors. Based on indicator screening and improved deep learning algorithms, feature extraction is performed on the multi-source spatiotemporal data to obtain indicator features; A spatially constrained machine learning clustering model is constructed, and the index features are input into the machine learning clustering model to obtain the clustering results; The result discrimination module is used to discriminate inefficient living spaces in the clustering results to obtain inefficient living spaces in the target urban space.
[0015] Thirdly, the present invention provides a terminal, comprising: a processor and a memory, wherein the memory stores an inefficient living space assessment program based on multi-source spatiotemporal data, and the inefficient living space assessment program based on multi-source spatiotemporal data is executed by the processor to implement the operation of the inefficient living space assessment method based on multi-source spatiotemporal data as described in the first aspect.
[0016] Fourthly, the present invention also provides a computer-readable storage medium storing an inefficient living space assessment program based on multi-source spatiotemporal data. When executed by a processor, the inefficient living space assessment program based on multi-source spatiotemporal data is used to implement the operation of the inefficient living space assessment method based on multi-source spatiotemporal data as described in the first aspect.
[0017] The present invention, by employing the above technical solution, has the following effects: This invention provides a method and system for identifying inefficient residential spaces based on multi-source spatiotemporal data, comprising: acquiring multi-source spatiotemporal data of a target city space; preprocessing the multi-source spatiotemporal data to obtain indicator feature vectors; extracting features from the multi-source spatiotemporal data based on an indicator screening and improved deep learning algorithm to obtain indicator features; constructing a spatially constrained machine learning clustering model; inputting the indicator features into the machine learning clustering model to obtain clustering results; and identifying inefficient residential spaces in the clustering results to obtain inefficient residential spaces in the target city space. This invention enables unified analysis and indicator transformation of multi-source spatiotemporal data in a target city space, constructing multi-dimensional residential land efficiency evaluation indicators; constructing high-quality indicator features through feature extraction using an indicator screening and improved deep learning method; and accurately identifying inefficient residential spaces in the target city space by performing clustering using the constructed spatially constrained machine learning clustering model and identifying inefficient residential spaces based on the clustering results. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0019] Figure 1 This is a flowchart of the inefficient residential space assessment method based on multi-source spatiotemporal data in this invention.
[0020] Figure 2 This is a schematic diagram of the structure of the autoencoder network in one implementation of the present invention.
[0021] Figure 3 This is a flowchart of a machine learning clustering model clustering method in one implementation of the present invention.
[0022] Figure 4 This is a functional schematic diagram of the terminal in one implementation of the present invention.
[0023] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0025] Exemplary methods The existence of inefficient land use affects the normal development of cities. In recent years, the emergence of multi-source geospatial big data such as points of interest, mobile phone signaling, and street view images has provided multi-dimensional data support for land use efficiency assessment, from building attributes, population vitality, transportation accessibility to environmental quality, and promoted the identification of inefficient land use from subjective qualitative to quantitative analysis.
[0026] However, existing methods for identifying inefficient residential land use have two main limitations: First, multi-criteria decision-making analyses rely on expert experience or pre-defined function forms in weight setting, introducing subjective biases, and the model structure is linear, making it difficult to characterize the non-linear relationships between multi-dimensional indicators. Second, cities contain two distinct residential types: "residential communities" and "urban villages," which differ significantly in land use patterns, building forms, and socio-economic functions. Most existing studies treat urban residential land as a homogeneous whole, using the same indicators and models for identification, ignoring this inherent heterogeneity. This limits the accuracy, robustness, and practical guiding value of the identification models for specific land use types.
[0027] Existing methods for identifying inefficient residential land use suffer from problems such as strong model subjectivity and insufficient consideration of heterogeneity when integrating multi-source spatiotemporal big data. Therefore, existing technologies need further improvement.
[0028] To address the above-mentioned technical problems, this invention provides a method for identifying inefficient residential spaces based on multi-source spatiotemporal data. The method includes: acquiring multi-source spatiotemporal data of a target city space; preprocessing the multi-source spatiotemporal data to obtain indicator feature vectors; extracting features from the multi-source spatiotemporal data using an indicator screening and improved deep learning algorithm to obtain indicator features; constructing a spatially constrained machine learning clustering model; inputting the indicator features into the machine learning clustering model to obtain clustering results; and identifying inefficient residential spaces in the clustering results to obtain inefficient residential spaces within the target city space. This invention enables unified analysis and indicator transformation of multi-source spatiotemporal data in a target city space, constructing multi-dimensional residential land efficiency evaluation indicators; constructing high-quality indicator features through feature extraction using an indicator screening and improved deep learning method; and accurately identifying inefficient residential spaces within the target city space by performing clustering using the constructed spatially constrained machine learning clustering model and identifying inefficient residential spaces based on the clustering results.
[0029] like Figure 1 As shown, this embodiment of the invention provides a method for inefficient residential space assessment based on multi-source spatiotemporal data, comprising the following steps: Step S100: Obtain multi-source spatiotemporal data of the target city space, preprocess the multi-source spatiotemporal data, and obtain the indicator feature vector.
[0030] It should be noted that, in order to objectively characterize the utilization efficiency of residential land and avoid the limitations of traditional methods that only focus on a single building attribute, this embodiment acquires multi-source spatiotemporal data, including building data, population data, environmental data, economic data, and functional data. Specifically, it acquires building data in the building attribute dimension, population data in the human settlement vitality dimension, economic data in the economic efficiency dimension, and environmental data and functional data in the livability attribute dimension. The multi-source spatiotemporal data in these four dimensions constitute the basis for judging inefficient living spaces from the perspectives of "intrinsic attributes", "usage status", "economic benefits" and "external environment".
[0031] Specifically, in one implementation of this embodiment, step S100 includes the following steps: Step S101: Obtain multi-source spatiotemporal data of the target city space.
[0032] In this embodiment, multi-source spatiotemporal data related to the efficiency of urban residential land use are acquired. The main data sources include: building survey data, point of interest (POI) data, area of interest (AOI) data, mobile signaling data, third-party platform data, government open data, road network data, etc. These data have significant differences in format, scale and spatiotemporal resolution, constituting a typical heterogeneous data source.
[0033] Further, step S101 includes the following steps: Step S101a: Select the target city space; Step S101b: Divide the target urban space into a preset number of plots; Step S101c: Obtain multi-source spatiotemporal data for each of the plots; wherein, the multi-source spatiotemporal data includes building data, population data, environmental data, economic data, and functional data.
[0034] Step S102: Perform data cleaning on the multi-source spatiotemporal data.
[0035] In this embodiment, cleaning is performed based on the characteristics of different data. For mobile phone signaling trajectory data, a spatiotemporal constraint-based denoising method is used, employing an adaptive threshold model to filter records with reasonable dwell time and displacement, avoiding over-cleaning that may result from hard threshold processing. For POI / AOI data, a spatiotemporally adaptive deduplication algorithm is applied, using spatiotemporal similarity-based matching rules to identify and correct duplicate and misclassified data. For building census, greening rate, and housing rental reference price data, a multi-scale outlier detection method is used, employing outlier identification based on local spatial information to perform trend prediction and interpolation for housing rental price data missing from a small number of plots. For cases where data attribute information is missing, a spatiotemporal interpolation method based on adjacent similar plots is introduced, combining the spatial structure and similarity of the data to fill in the gaps, avoiding information bias caused by simple spatial interpolation methods.
[0036] Step S103: Perform coordinate system transformation and spatial calibration on the cleaned multi-source spatiotemporal data.
[0037] In this embodiment, the cleaned multi-source spatiotemporal data undergoes coordinate system transformation and spatial calibration. Since geographic data from different sources use inconsistent coordinate systems, this embodiment uses the Geopandas library in Python and ArcGIS Pro software to uniformly transform the cleaned multi-source spatiotemporal data to the same spatial reference system, namely the CGCS2000 projected coordinate system.
[0038] Meanwhile, considering the multi-scale nature of spatial data, this embodiment adopts a hierarchical spatial calibration method to perform calibration at different scales, ensuring that data sources of different scales can be compatible and integrated with each other.
[0039] Step S104: Integrate the spatially calibrated multi-source spatiotemporal data.
[0040] In this embodiment, calibrated multi-source spatiotemporal data are integrated into a unified analysis unit using an adaptive data integration method that dynamically adjusts the data integration strategy based on the data source. During the analysis unit construction phase, considering potential spatial discrepancies between residential land use planning data and actual construction, this embodiment introduces AOI data to assist in correcting the residential plot boundaries. By spatially overlaying AOI data with land use planning data, the differences between actual built-up residential areas and planned residential plots are identified, and the boundaries of residential plots are corrected or refined, thus forming an analysis unit that better reflects the actual built environment. Based on the corrected residential plots, spatial operations such as spatial connectivity, regional statistics, and kernel density analysis are used to assign attribute information from the multi-source data to the corresponding target plots, forming a full attribute list for each plot. The contents of the full attribute list can be referenced from the indicator names in step S105. This process combines multi-scale spatial analysis methods to ensure optimal data integration results at different scales.
[0041] Step S105: The integrated multi-source spatiotemporal data is converted into indicators, and an indicator feature vector is constructed based on the obtained indicators.
[0042] In this embodiment, the integrated multi-source spatiotemporal data is converted into indicators. Based on the acquired multi-source spatiotemporal data of four dimensions—building attributes, human settlement vitality, economic efficiency, and livability attributes—an indicator system for assessing inefficient living spaces is constructed, as shown in Table 1 below. Table 1. Indicator system for assessing inefficient living spaces;
[0043] As shown in the table above, the building attribute dimension includes building age, plot ratio, and building density; the human settlement vitality dimension includes nighttime population density, population marginal-core index, and population profile; the livability attribute dimension mainly includes greening rate; the economic efficiency dimension mainly includes regional economic level; and the livability attribute dimension includes transportation convenience, public service convenience, and functional mixing.
[0044] Specifically, in building data, the building age is obtained by converting the completion year from building census data into the building age; Floor Area Ratio (FAR) is calculated using the following formula: ; in, For the first The base area of the building. Its floor number, The total area of the target plot; Building density (BD) is calculated using the following formula: .
[0045] In population data, nighttime population density (NPD) is calculated using the following formula: ; in, This refers to the number of resident users within the target site at night, which is defined as the period from 1 a.m. to 6 a.m.
[0046] The Edge-Core Index for Night Population Density (ECI-NPD) is calculated using the following formula: ; in, The nighttime population density of the target site. The median nighttime population density of all plots within 500m.
[0047] The population profile is obtained by directly acquiring the average education level index and age group aggregation labels of permanent residents in the target area through mobile signaling data.
[0048] In environmental data, the greening rate of the target plot can be obtained directly from data from a third-party platform.
[0049] In economic data, the regional economic level is measured by directly obtaining reference rental prices for target plots of housing through open government data.
[0050] In the functional data, transportation convenience is based on road network analysis, which is obtained by calculating the number of bus stops and subway stations that can be reached within a 15-minute walk of the target plot; Public service convenience is based on road network analysis. It is obtained by calculating the number of POIs (points of interest) such as schools, hospitals, supermarkets, and parks that can be reached within a 15-minute walk of the target plot. In other words, public service convenience can be further divided into multiple data points. Shannon's Diversity Index (SHDI) data, which is the Shannon diversity index of the target site, is obtained by calculating the number of all POIs within a 15-minute walk of the target site. The calculation formula is as follows: ; in, Number of POI types For the first The proportion of POI-like items.
[0051] In this embodiment, an indicator feature vector is constructed based on the obtained indicators, and each target land parcel is represented as a unified and standardized indicator feature vector. The reference format of this indicator feature vector is as follows: Plot 001 [Building year, plot ratio, building density, greening rate, nighttime population density, population periphery-core index, population profile, rental price, public transportation accessibility, subway accessibility, school accessibility, hospital accessibility, supermarket accessibility, park accessibility, functional mix, geometry]; Plot 002 [Building year, plot ratio, building density, greening rate, nighttime population density, marginal-core population index, population profile, rental price, public transportation accessibility, subway accessibility, school accessibility, hospital accessibility, supermarket accessibility, park accessibility, functional mix, geometry].
[0052] It should be noted that in the above indicator feature vector reference format, transportation convenience is broken down into more specific public transportation convenience and subway convenience; similarly, public service convenience is broken down into school convenience, hospital convenience, supermarket convenience, and park convenience; while geometry is used to characterize the geometric shape and spatial location of the land parcel, and is the basic spatial attribute data of the land parcel.
[0053] like Figure 1 As shown, this embodiment of the invention provides a method for inefficient residential space assessment based on multi-source spatiotemporal data, comprising the following steps: Step S200: Based on index screening and improved deep learning algorithms, feature extraction is performed on the multi-source spatiotemporal data to obtain index features.
[0054] Specifically, in one implementation of this embodiment, step S200 includes the following steps: Step S201: Standardize the indicator feature vector.
[0055] In this embodiment, the indicator feature vectors are standardized to eliminate the influence of dimensions and ensure that the values of each indicator conform to a unified standard. Specifically, Z-score standardization is used to convert the data into a standard normal distribution with a mean of 0 and a standard deviation of 1, ensuring that each feature is compared at the same scale and meeting the data feature requirements of the subsequent method for identifying inefficient residential land.
[0056] In this embodiment, the samples are residential plots within the target urban space, with each residential plot corresponding to a multi-dimensional indicator feature vector. Let the number of samples be... Then all samples constitute a sample matrix: ; in, For the first The indicator feature vector of each residential plot; This indicates the dimensions of the filtered metrics; This represents a sample matrix, where each row corresponds to a land parcel sample and each column corresponds to an indicator feature.
[0057] Step S202: Based on the correlation coefficient, perform correlation analysis and screening on the standardized indicator feature vector to remove redundant indicators from the indicator feature vector; wherein, the redundant indicators are indicators whose correlation coefficient is higher than a preset correlation coefficient threshold.
[0058] In this embodiment, the standardized indicator feature vector is subjected to indicator correlation analysis and screening based on the correlation coefficient. First, the correlation matrix between each indicator is calculated, and a preset correlation coefficient threshold is obtained. Redundant indicators that are highly correlated with other features are eliminated, and the most business-representative indicators in the indicator feature vector are retained to reduce the dimensionality of the data and improve the efficiency and accuracy of subsequent analysis.
[0059] In this embodiment, calculating the correlation matrix between various indicators can be regarded as calculating the correlation coefficient between each indicator in the indicator feature vector. Based on the correlation coefficient, the standardized indicator feature vector is subjected to indicator correlation analysis and screening. The calculation formula is as follows: ; in, The correlation coefficient is... This represents the difference in rank between two variables. This represents the number of samples.
[0060] In this embodiment, the correlation coefficient threshold is set to 0.8. If the absolute value of the correlation coefficient between any two indicators in the indicator feature vector is greater than this threshold, then a more representative indicator is selected and retained according to a preset rule. The preset rule is as follows: Based on their business implications, priority should be given to retaining indicators that have clear physical meaning and are more interpretable in the evaluation of residential land or urban planning, such as plot ratio and building density; while derivative indicators with weak business implications or that are difficult to interpret can be regarded as redundant indicators and eliminated, such as school accessibility and park accessibility. Based on the amount of information in the indicators, if it is impossible to determine a more representative indicator based on the business meaning, then indicators with richer numerical distribution and greater variance should be retained first to ensure that they have stronger distinguishing ability; indicators with smaller variance and less obvious changes can be regarded as redundant indicators.
[0061] Step S203: Construct an autoencoder network based on an improved deep learning algorithm, and extract features from the indicator feature vectors after indicator correlation analysis and screening through the autoencoder network to obtain indicator features.
[0062] Specifically, step S203 includes the following steps: Step S203a: Based on the improved deep learning algorithm, network weights and preset activation functions, construct an autoencoder network containing an encoder and a decoder.
[0063] In this embodiment, to further reduce data dimensionality and extract key features, Principal Component Analysis (PCA) is used to perform linear feature extraction and network initialization on the selected indicators in the indicator feature vector. The bottleneck dimension q is determined based on 80% of the explained variance and is used as the number of neurons in the bottleneck layer of the autoencoder to ensure the optimal linear baseline of the compressed representation.
[0064] Specifically, the core of principal component analysis is achieved through eigenvalue decomposition, with the corresponding formula as follows: ; in, Let covariance matrix be the variance matrix. The indicator feature vector matrix, It is a diagonal matrix of eigenvalues.
[0065] In this embodiment, based on an improved deep learning algorithm, network weights, and a preset activation function, a system is constructed as follows: Figure 2 The diagram shows the structure of the autoencoder network. The improved deep learning algorithm is an optimized design based on the traditional autoencoder structure, specifically including: introducing PCA for linear feature pre-extraction and network parameter initialization, using the cumulative variance contribution rate to determine the bottleneck layer dimension; using the PCA projection matrix to initialize the encoder weights to avoid training instability caused by random initialization; and combining PReLU to construct a progressive training mechanism from linear to nonlinear, enabling the model to gradually enhance its nonlinear expressive ability while maintaining the robustness of linear features.
[0066] The autoencoder network comprises an encoder and a decoder. The encoder compresses the input data into a low-dimensional feature representation, while the decoder restores it from the low-dimensional representation back to the original data. The encoder path is input layer → hidden layer → bottleneck layer, and the decoder path is the reverse. Data compression and reconstruction are achieved through fully connected layers. The network parameters, including weight parameters and PReLU (Parametric Rectifier Linear Unit) activation function parameters, are jointly optimized using the backpropagation algorithm.
[0067] In this embodiment, principal components with a cumulative variance contribution rate ≥ 80% are retained, and their number is the dimension of the autoencoder's hidden layer. The autoencoder weights are initialized using PCA parameters, and the encoder weights are set as the PCA projection matrix, i.e. The encoder bias is set to: ,in This is the sample mean vector.
[0068] The decoder weights and biases are solved by least-squares fitting, and the corresponding formulas are as follows: ; in, For encoder output, This is the original input sample matrix.
[0069] In this embodiment, the preset activation function is the PReLU activation function, that is, the parameterized ReLU activation function. This activation function enables the autoencoder to gradually transition from initial linear feature extraction to nonlinear learning of complex data.
[0070] Specifically, the formula for the PReLU activation function is as follows: ; In the initial stage, the slope of the PReLU negative region is set to... 1. At this point, the activation function is linear, and the autoencoder behaves identically to PCA; during training, As a learnable parameter, it is gradually adjusted with gradient descent, so that the activation function gradually acquires nonlinear characteristics; this mechanism ensures that the model starts from a robust linear starting point and gradually adapts to the complex nonlinear structure in the data, effectively avoiding local optima and training instability.
[0071] Step S203b: Calculate the loss function of the autoencoder network, and train and optimize the autoencoder network using the loss function.
[0072] It should be noted that the training objective of an autoencoder network is to optimize the network weights and activation function parameters by minimizing the reconstruction error, so that it can gradually adapt to the complex nonlinear structure in the data while maintaining the robustness of linear features.
[0073] Specifically, the autoencoder, after initial feature extraction and model initialization, is trained end-to-end on the training set. The optimization objective is to minimize the reconstruction error, which is used as the loss function, and the corresponding formula is as follows: ; Among them, For the first The original input indicator feature vector of each sample, The output of the autoencoder is the reconstructed feature vector of the input index. This is the loss function, used to measure the error between the original input and the reconstructed output.
[0074] During training, the network weights and PReLU parameters are optimized simultaneously, specifically by updating the parameters through the backpropagation algorithm, as shown in the following formula: ; in, For learning rate, For the first Model parameters at the next iteration loss function Regarding parameters The gradient is used to indicate the direction of parameter optimization.
[0075] Step S203c: The feature vectors of the indicators after the correlation analysis of the indicators are filtered by the autoencoder network to extract the indicator features.
[0076] In this embodiment, the network weights and PReLU parameters are jointly optimized to obtain an autoencoder network that combines the robustness of a linear backbone with the ability to fit nonlinear data. This autoencoder network extracts features from the feature vectors of the indicators after correlation analysis and outputs low-dimensional features, i.e., indicator features. These features provide high-quality indicator features as input to subsequent clustering model algorithms and can be used for the accurate identification of inefficient living spaces.
[0077] like Figure 1 As shown, this embodiment of the invention provides a method for inefficient residential space assessment based on multi-source spatiotemporal data, comprising the following steps: Step S300: Construct a machine learning clustering model based on spatial constraints, input the index features into the machine learning clustering model, and obtain the clustering results.
[0078] It should be noted that, considering the first law of geography, the efficiency of residential land use is not only determined by relevant attribute indicators, but adjacent spatial units often have stronger attribute similarities and functional correlations. Therefore, based on the traditional Spatially Constrained Gaussian Mixture Model (SC-GMM), a spatial constraint term is introduced to construct a spatially constrained machine learning clustering model for clustering and discrimination of urban residential land.
[0079] Specifically, in one implementation of this embodiment, step S300 includes the following steps: Step S301: Obtain the spatial constraints of the target city space.
[0080] In this embodiment, the spatial constraints of the target city space include a spatial constraint strength coefficient, which is used to control the consistency of adjacent spatial units in clustering assignment; an adjacency weight between spatial units, which is constructed based on the adjacency relationship or distance decay function of the spatial units; and a spatial unit category consistency indicator function.
[0081] Step S302: Construct a Gaussian mixture model, optimize the Gaussian mixture model using the spatial constraint term, and obtain a machine learning clustering model based on spatial constraints.
[0082] In this embodiment, the low-dimensional nonlinear features, i.e. index features, extracted by the autoencoder are used as model inputs. Independent spatially constrained Gaussian mixture clustering models are established for two types of residential land: residential communities and urban villages. Spatial relationships are modeled using a Graph Convolution Neural Network (GCN) to enhance the spatial dependence of clustering. The optimal number of clusters for each type of land is determined using the Bayesian information criterion, and the parameters of each Gaussian component are estimated iteratively using the expectation-maximization algorithm.
[0083] In this embodiment, the introduced spatial constraint term is used to construct the objective function of the machine learning clustering model. The formula for the objective function is as follows: ; in, , , They represent the first The weights, mean vectors, and covariance matrices of each Gaussian component; This refers to the spatial constraint strength coefficient; spatial unit and The adjacency weights between them; This is a spatial cell category consistency indicator function. It takes a value of 1 when two adjacent cells are assigned to the same cluster, and 0 otherwise.
[0084] Step S303: Input the indicator features into the machine learning clustering model to obtain clustering results; wherein, the clustering results are clusters based on multi-source spatiotemporal data.
[0085] like Figure 3 The diagram shown is a flowchart of the clustering method of the machine learning clustering model in this embodiment. Figure 3 In Chinese, the underscore (_) is used to represent a subscript. The low-dimensional feature shape, i.e., the index feature As input, its shape is First, the optimal number of clusters is determined based on the Bayesian information criterion. Iterative estimation is performed using the EM algorithm (Expectation-Maximization algorithm). , , Introducing spatial constraint terms and spatial coordinate shape Corresponding shape Clustering results are obtained by maximizing the objective function of a machine learning clustering model. The shape of the clustering results is as follows: The range of values is .
[0086] In this embodiment, the clustering result is a clustering group based on multi-source spatiotemporal data, and the maximum number of clusters is [number missing]. .
[0087] like Figure 1 As shown, this embodiment of the invention provides a method for inefficient residential space assessment based on multi-source spatiotemporal data, comprising the following steps: Step S400: The clustering results are used to identify inefficient living spaces to obtain inefficient living spaces in the target city space.
[0088] In this embodiment, the index characteristics corresponding to the clustering results are obtained, and inefficient living spaces are identified based on the clustering results. Variance analysis and descriptive statistics are performed on each output cluster, and the standardized mean of each cluster on all efficiency indicators is calculated. By comparing the performance characteristics of different clusters on various positive and negative indicators, a comprehensive comparison of multiple indicators is conducted to identify cluster categories with typical characteristics of inefficient land use.
[0089] For example, if a cluster exhibits low positive indicators such as green space ratio, transportation convenience, public service convenience, and functional mixing compared to other clusters, while showing high negative indicators such as plot ratio, building density, nighttime population density, and aging population profile, it can be determined that the residential land corresponding to this cluster suffers from inefficiency characteristics such as insufficient spatial resource allocation, low living environment quality, and poor population vitality, classifying it as an inefficient residential community or inefficient urban village. Similarly, if a cluster exhibits low rental prices, insufficient functional mixing, lack of public facilities, and continuous spatial distribution, it can be further identified as an inefficient residential area requiring renovation and redevelopment.
[0090] Specifically, in one implementation of this embodiment, step S400 includes the following steps: Step S401: Obtain the index features corresponding to the clustering results; Step S402: Analyze the indicator characteristics based on multi-source spatiotemporal data and compare the indicator characteristics between different clustering results; Step S403: Based on the analysis and comparison results of the indicator features, inefficient living spaces are identified in the clustering results; Step S404: Obtain the plots in the target urban space that are identified as inefficient residential spaces.
[0091] In one implementation of this embodiment, high-definition remote sensing images, street view panoramic images, and land parcel indicators are comprehensively used for comprehensive interpretation based on expert experience, and verified by limited field investigations. Cross-validation and F1-score are used to evaluate the accuracy of the discrimination results to ensure the accuracy and generalization ability of each model used in the inefficient residential space assessment method based on multi-source spatiotemporal data.
[0092] This embodiment achieves the following technical effects through the above technical solution: 1. This embodiment acquires multi-source spatiotemporal data of urban space, and performs spatial calibration and data integration through data cleaning and multi-scale processing to construct a unified analysis unit; subsequently, index conversion is performed to realize the construction of multi-dimensional residential land use efficiency evaluation indicators.
[0093] 2. High-quality indicator features were constructed by using deep learning methods for indicator selection and improvement.
[0094] 3. By constructing a spatially constrained machine learning clustering model, clustering was performed, and inefficient living spaces were identified based on the clustering results, thus accurately identifying inefficient living spaces in the target city.
[0095] Exemplary device Based on the above embodiments, the present invention also provides an inefficient residential space assessment system based on multi-source spatiotemporal data, comprising: The data acquisition module is used to acquire multi-source spatiotemporal data of the target city space, and preprocess the multi-source spatiotemporal data to obtain indicator feature vectors. Based on indicator screening and improved deep learning algorithms, feature extraction is performed on the multi-source spatiotemporal data to obtain indicator features; A spatially constrained machine learning clustering model is constructed, and the index features are input into the machine learning clustering model to obtain the clustering results; The result discrimination module is used to discriminate inefficient living spaces in the clustering results to obtain inefficient living spaces in the target urban space.
[0096] Based on the above embodiments, the present invention also provides a terminal, the principle block diagram of which can be as follows: Figure 4 As shown.
[0097] The terminal includes: a processor, a memory, an interface, a display screen, and a communication module connected via a system bus; wherein, the processor of the terminal provides computing and control capabilities; the memory of the terminal includes a computer-readable storage medium and internal memory; the computer-readable storage medium stores an operating system and computer programs; the internal memory provides an environment for the operation of the operating system and computer programs in the computer-readable storage medium; the interface is used to connect to external devices; the display screen is used to display relevant information; and the communication module is used to communicate with a cloud server or other devices.
[0098] When executed by a processor, this computer program is used to implement an inefficient method for assessing living space based on multi-source spatiotemporal data.
[0099] It will be understood by those skilled in the art that Figure 4 The schematic diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal to which the present invention is applied. A specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0100] In one embodiment, a terminal is provided, comprising: a processor and a memory, the memory storing an inefficient residential space assessment program based on multi-source spatiotemporal data, the inefficient residential space assessment program based on multi-source spatiotemporal data being executed by the processor to implement the operation of the above-described inefficient residential space assessment method based on multi-source spatiotemporal data.
[0101] In one embodiment, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores an inefficient residential space assessment program based on multi-source spatiotemporal data, which, when executed by a processor, is used to implement the operation of the above-described inefficient residential space assessment method based on multi-source spatiotemporal data.
[0102] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, database, or other media used in the embodiments provided by this invention can include both non-volatile and volatile memory.
[0103] In summary, this invention provides a method and system for identifying inefficient residential spaces based on multi-source spatiotemporal data, comprising: acquiring multi-source spatiotemporal data of a target urban space; preprocessing the multi-source spatiotemporal data to obtain indicator feature vectors; extracting features from the multi-source spatiotemporal data based on indicator screening and improved deep learning algorithms to obtain indicator features; constructing a spatially constrained machine learning clustering model; inputting the indicator features into the machine learning clustering model to obtain clustering results; and identifying inefficient residential spaces in the clustering results to obtain inefficient residential spaces in the target urban space. This invention can perform unified analysis and indicator transformation of multi-source spatiotemporal data in a target urban space, constructing multi-dimensional residential land efficiency evaluation indicators; constructing high-quality indicator features through feature extraction using an indicator screening and improved deep learning method; and accurately identifying inefficient residential spaces in the target urban space by performing clustering using the constructed spatially constrained machine learning clustering model and identifying inefficient residential spaces based on the clustering results.
[0104] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. An inefficient method for assessing residential space based on multi-source spatiotemporal data, characterized in that, include: Acquire multi-source spatiotemporal data of the target city space, preprocess the multi-source spatiotemporal data, and obtain the indicator feature vector; Based on indicator screening and improved deep learning algorithms, feature extraction is performed on the multi-source spatiotemporal data to obtain indicator features; A spatially constrained machine learning clustering model is constructed, and the index features are input into the machine learning clustering model to obtain the clustering results; The clustering results are used to identify inefficient residential spaces, thereby obtaining inefficient residential spaces in the target city space.
2. The inefficient residential space assessment method based on multi-source spatiotemporal data according to claim 1, characterized in that, The acquisition of multi-source spatiotemporal data of the target city space includes: Select the target urban space; The target urban space is divided into a predetermined number of plots; Acquire multi-source spatiotemporal data for each of the plots; wherein the multi-source spatiotemporal data includes building data, population data, environmental data, economic data, and functional data.
3. The inefficient residential space assessment method based on multi-source spatiotemporal data according to claim 1, characterized in that, The preprocessing of the multi-source spatiotemporal data to obtain the indicator feature vector includes: Data cleaning is performed on the multi-source spatiotemporal data; The cleaned multi-source spatiotemporal data were subjected to coordinate system transformation and spatial calibration. Integrate spatially calibrated multi-source spatiotemporal data; The integrated multi-source spatiotemporal data is transformed into indicators, and indicator feature vectors are constructed based on the obtained indicators.
4. The inefficient residential space assessment method based on multi-source spatiotemporal data according to claim 1, characterized in that, The feature extraction of the multi-source spatiotemporal data based on the indicator screening and improved deep learning algorithm yields indicator features, including: The feature vector of the index is standardized. Based on the correlation coefficient, the standardized indicator feature vector is subjected to correlation analysis and screening to remove redundant indicators from the indicator feature vector; wherein, the redundant indicators are indicators whose correlation coefficient is higher than a preset correlation coefficient threshold. An autoencoder network is constructed based on an improved deep learning algorithm. The autoencoder network is used to extract features from the feature vectors of indicators after correlation analysis and screening, thereby obtaining indicator features.
5. The inefficient residential space assessment method based on multi-source spatiotemporal data according to claim 4, characterized in that, The method involves constructing an autoencoder network based on an improved deep learning algorithm. This autoencoder network is used to extract features from the feature vectors of the indicators selected through correlation analysis, yielding indicator features, including: Based on improved deep learning algorithms, network weights, and preset activation functions, an autoencoder network containing an encoder and a decoder is constructed. Calculate the loss function of the autoencoder network, and train and optimize the autoencoder network using the loss function; The autoencoder network is used to extract features from the feature vectors of the indicators after correlation analysis, thereby obtaining the indicator features.
6. The inefficient residential space assessment method based on multi-source spatiotemporal data according to claim 1, characterized in that, The construction of a space-constrained machine learning clustering model, by inputting the indicator features into the machine learning clustering model to obtain clustering results, includes: Obtain the spatial constraints of the target city; Construct a Gaussian mixture model, optimize the Gaussian mixture model using the spatial constraint term, and obtain a machine learning clustering model based on spatial constraints; The index features are input into the machine learning clustering model to obtain clustering results; wherein, the clustering results are clusters based on multi-source spatiotemporal data.
7. The inefficient residential space assessment method based on multi-source spatiotemporal data according to claim 2, characterized in that, The step of identifying inefficient residential spaces in the clustering results to obtain inefficient residential spaces in the target urban space includes: Obtain the index features corresponding to the clustering results; The characteristics of the indicators are analyzed based on multi-source spatiotemporal data, and the characteristics of the indicators are compared among different clustering results; Based on the analysis and comparison results of the indicator characteristics, the clustering results are used to identify inefficient living spaces. The land parcels in the target urban space that are identified as inefficient residential spaces are obtained.
8. An inefficient residential space assessment system based on multi-source spatiotemporal data, characterized in that, include: The data acquisition module is used to acquire multi-source spatiotemporal data of the target city space, and preprocess the multi-source spatiotemporal data to obtain indicator feature vectors. Based on indicator screening and improved deep learning algorithms, feature extraction is performed on the multi-source spatiotemporal data to obtain indicator features; A spatially constrained machine learning clustering model is constructed, and the index features are input into the machine learning clustering model to obtain the clustering results; The result discrimination module is used to discriminate inefficient living spaces in the clustering results to obtain inefficient living spaces in the target urban space.
9. A terminal, characterized in that, include: The processor and memory, wherein the memory stores an inefficient living space assessment program based on multi-source spatiotemporal data, and the inefficient living space assessment program based on multi-source spatiotemporal data, when executed by the processor, is used to implement the operation of the inefficient living space assessment method based on multi-source spatiotemporal data as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an inefficient living space assessment program based on multi-source spatiotemporal data. When executed by a processor, the inefficient living space assessment program based on multi-source spatiotemporal data is used to implement the operation of the inefficient living space assessment method based on multi-source spatiotemporal data as described in any one of claims 1-7.