An intercity similarity analysis method based on perceptual hashing and hilbert curve

CN122734541APending Publication Date: 2026-09-11HUAZHONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610564681.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-27
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

这种差异性使得传统面向单一城市或局部区域的数据分析方法,难以直接适用于跨城市的比较研究,尤其是在多城市情景下,原始高维空间数据体量庞大、结构复杂,直接进行存储、计算或相似性分析,不仅计算成本高,而且难以保证结果的稳定性和可比性

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122734541A_ABST
    Figure CN122734541A_ABST
Patent Text Reader

Abstract

The application discloses an intercity similarity analysis method based on perceptual hashing and Hilbert curve. By introducing the perceptual hashing algorithm to reduce the dimension coding of city light data, the robust compression and structured expression of high-dimensional and large-scale night light images are realized, so that the night light data of different cities has good comparability and calculation efficiency under a unified scale. Further combined with the Hilbert space filling curve, the two-dimensional night light two-dimensional code is mapped to a one-dimensional barcode representation which maintains the spatial proximity relationship, thereby reducing the data complexity while effectively preserving the continuity and local features of the spatial structure of the city night light, so that the intercity similarity analysis and fast retrieval have higher stability and reliability. Through the clustering, retrieval and comparison process based on the two-dimensional code and the barcode, the city groups with similar night light spatial structures can be efficiently identified, and the similar city retrieval and fine structure comparison of the target city can be supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intercity similarity analysis technology, and in particular to an intercity similarity analysis method based on perceptual hashing and Hilbert curves, which can be used for cross-city spatial morphology comparison, multi-city analysis and intercity spatial structure analysis. Background Technology

[0002] Due to significant differences in spatial scale, morphological structure, development stage, and natural constraints among different cities, high-dimensional spatial data often exhibits strong imbalance and heterogeneity at the intercity scale. This variability makes traditional data analysis methods, which are geared towards single cities or local areas, difficult to apply directly to comparative studies across cities. Especially in multi-city scenarios, the original high-dimensional spatial data is massive in volume and complex in structure. Directly storing, computing, or performing similarity analysis is not only computationally costly but also makes it difficult to guarantee the stability and comparability of the results. Summary of the Invention

[0003] This invention solves the aforementioned technical problems in the prior art by providing an intercity similarity analysis method based on perceptual hashing and Hilbert curves.

[0004] This invention provides a method for intercity similarity analysis based on perceptual hashing and Hilbert curves, comprising:

[0005] Obtain lighting data for each city;

[0006] The light data of each city is converted into a two-dimensional light matrix using a perceptual hash algorithm;

[0007] The two-dimensional light matrix is ​​converted into one-dimensional light data using a Hilbert space-filling curve.

[0008] Clustering of cities is performed based on the aforementioned one-dimensional lighting data;

[0009] Based on one-dimensional light data of the city, the similarity between the city to be analyzed and cities of the same category is calculated, and a preset number of candidate cities are selected based on the similarity results;

[0010] The two-dimensional light matrix of the city to be analyzed and the two-dimensional light matrix of each candidate city are subjected to element-level bitwise XOR operation to generate a two-dimensional residual matrix that accurately represents the distribution of spatial morphological differences between cities.

[0011] Spatial sliding window statistical processing is applied to the two-dimensional residual matrix. The global residual matrix is ​​traversed sequentially according to the preset window size and movement step size. The density values ​​of dissimilar sites in each local spatial unit are statistically analyzed block by block. The density values ​​of all local blocks are numerically aggregated to obtain a comprehensive score of the two-dimensional local structural differences between the city to be analyzed and each candidate city.

[0012] The intercity similarity analysis results are obtained based on the comprehensive score of the two-dimensional local structural differences.

[0013] Specifically, obtaining the light data of each city includes:

[0014] Obtain the raw raster data of city lights with city identifiers for each city;

[0015] Using the built-up area boundary as a constraint, the original raster data of city lights are clipped and masked to obtain the effective raster data of city lights under the built-up area constraint.

[0016] The effective raster data of lights in each city are standardized to obtain standardized effective raster data of lights in each city.

[0017] Specifically, the process of converting the light data of each city into a two-dimensional light matrix using a perceptual hash algorithm includes:

[0018] Convert the light data of each city into a light grayscale matrix;

[0019] Perform a two-dimensional discrete cosine transform on the light grayscale matrix to map the light intensity distribution in the spatial domain to the frequency domain, and obtain the corresponding frequency coefficient matrix.

[0020] From the frequency coefficient matrix, the low-frequency sub-block located in the upper left corner is extracted as the feature region;

[0021] Compare all frequency coefficients in the feature region with the comparison value;

[0022] If the frequency coefficient is greater than or equal to the comparison value, the corresponding position is assigned a value of 1; if the frequency coefficient is less than the comparison value, the corresponding position is assigned a value of 0, thus obtaining a two-dimensional binary matrix of lights for each city.

[0023] Specifically, performing a two-dimensional discrete cosine transform on the light grayscale matrix to map the light intensity distribution in the spatial domain to the frequency domain, obtaining the corresponding frequency coefficient matrix, includes:

[0024] Through expression Obtain the frequency coefficient matrix ;in, ; is the light grayscale matrix; N is the side length of the matrix.

[0025] Specifically, the process of converting the two-dimensional light matrix into one-dimensional light data using a Hilbert space-filling curve includes:

[0026] The order of the corresponding Hilbert curve is determined based on the side length of the two-dimensional light matrix.

[0027] Based on the order of the Hilbert curve, the two-dimensional light matrix is ​​recursively partitioned using the Hilbert space filling algorithm to generate a continuously traversed two-dimensional coordinate sequence covering the global space of the two-dimensional light matrix.

[0028] The binary values ​​of the two-dimensional light matrix at the corresponding coordinate positions are obtained sequentially according to the two-dimensional coordinate sequence to construct a one-dimensional ordered sequence.

[0029] Specifically, the step of recursively partitioning the two-dimensional light matrix using the Hilbert space-filling algorithm based on the order of the Hilbert curve to generate a continuously traversed two-dimensional coordinate sequence covering the global space of the two-dimensional light matrix includes:

[0030] Based on the order r of the Hilbert curve, construct a mapping function between the one-dimensional index and the two-dimensional coordinates of the Hilbert curve:

[0031]

[0032] Where k represents the k-th one-dimensional position index in the Hilbert curve traversal order. This indicates the corresponding coordinates in the two-dimensional light matrix at that location;

[0033] This results in coverage of the entire Two-dimensional coordinate sequence of a two-dimensional light matrix .

[0034] Specifically, the clustering of cities based on the one-dimensional light data includes:

[0035] The one-dimensional light data is input into a K-means or similar spatial clustering algorithm for computation. The algorithm calculates the structural difference between the one-dimensional light data of each city and each cluster center based on the Hamming distance metric, and divides each city into the category set with the smallest distance. Then, the one-dimensional light data of each category is recalculated and updated based on the current division result. The above division and update process is continuously repeated until the cluster center no longer shifts significantly or the maximum number of iterations is reached.

[0036] Specifically, before clustering cities based on the one-dimensional light data, the process also includes:

[0037] Through formula The average profile coefficient was calculated. ;in, For the sample The average Hamming distance between the sample and other samples in the same cluster. For the sample The average Hamming distance between samples in the nearest neighbor cluster; take the mean Hamming distance between the samples in the nearest neighbor cluster; To obtain the maximum value As the optimal number of clusters under the silhouette coefficient criterion;

[0038] or,

[0039] Through formula Calculate the sum of squared errors within the group. ;in, Indicates sample With the Cluster centers The square of the Hamming distance between them Indicates the first The set of samples contained in each cluster; take The point where the rate of descent changes significantly corresponds to As the optimal number of clusters under the elbow criterion.

[0040] Specifically, the method involves calculating the similarity between the city to be analyzed and cities of the same category based on one-dimensional light data, and selecting a preset number of candidate cities based on the similarity results, including:

[0041] The cities are grouped according to clustering categories, and a corresponding index structure is created for each category of cities;

[0042] Based on the cluster category labels of the cities to be analyzed, one-dimensional light data of candidate cities with the same or adjacent cluster labels are extracted from the established index structure through Boolean index filtering operation, and the one-dimensional light data is assembled into a two-dimensional candidate feature matrix.

[0043] Based on the one-dimensional light data of the city to be analyzed, the dimension of the city is aligned with the two-dimensional candidate feature matrix of the candidate cities by broadcasting and copying. Batch bitwise XOR operation is performed on the two matrices, and the XOR results are summed along the row direction to calculate the Hamming distance between the city to be analyzed and each candidate city.

[0044] Obtain a preset number of candidate cities with the minimum Hamming distance.

[0045] Specifically, it also includes:

[0046] Based on the one-dimensional light data and two-dimensional light matrix corresponding to the city to be analyzed and the candidate cities, the data is reorganized into structured call data units according to the set field order;

[0047] Structured data units of the same type of city are sorted and retrieved according to city identifier and time identifier. Hamming distance is calculated for one-dimensional light data and two-dimensional light matrix between adjacent periods or between specified periods to obtain the corresponding time series difference result table, which quantitatively tracks the evolution trajectory and expansion direction of urban spatial structure.

[0048] One or more technical solutions provided in this invention have at least the following technical effects or advantages:

[0049] 1. By introducing a perceptual hashing algorithm to encode urban nighttime light remote sensing data for dimensionality reduction, robust compression and structured representation of high-dimensional, large-scale nighttime light imagery are achieved, enabling good comparability and computational efficiency of nighttime light data from different cities at a unified scale. Furthermore, by combining Hilbert space-filling curves, two-dimensional nighttime light QR codes are mapped to one-dimensional barcodes that maintain spatial proximity. This effectively preserves the continuity and local features of the spatial structure of urban nighttime lights while reducing data complexity, resulting in higher stability and reliability for intercity similarity analysis and rapid retrieval. Through clustering, retrieval, and comparison processes based on QR codes and barcodes, urban groups with similar nighttime light spatial structures can be efficiently identified. This supports similar city retrieval and fine-grained structural comparison for target cities, providing an intuitive and repeatable method for city type identification and spatial feature comparison at the intercity scale. Overall, this improves the efficiency and applicability of nighttime light data analysis in urban research and related applications, and provides a unified coding foundation for subsequent visualization and interactive analysis. This invention functionally divides and coordinates the improved two-dimensional structure hash and the adjacency-preserving one-dimensional sequence. The one-dimensional sequence is used as a fast index feature to achieve efficient initial screening of large-scale samples. Then, based on the two-dimensional structure matrix, fine comparison and sorting of local spatial structures are performed, thus forming a hierarchical analysis mode of "fast retrieval—fine matching". This framework significantly improves the accuracy of spatial structure recognition while maintaining high computational efficiency, transforming the similarity analysis of high-dimensional urban spatial data from a single static comparison into a searchable, scalable, and reusable engineering analysis system.

[0050] 2. Before hashing, the original spatial data undergoes effective regional constraints (with the built-up area as the spatial core) and standard scale normalization to ensure that different spatial samples have a consistent spatial alignment basis at the structural input level. Based on this, the hash output format is structurally improved, reconstructing the original one-dimensional hash fingerprint into a two-dimensional binary structure matrix. This allows the hash result to directly carry the two-dimensional structural information of the urban spatial pattern, not only enhancing the cross-sample comparability of high-dimensional urban spatial data but also giving the hash result spatial structural interpretability and local comparison capabilities.

[0051] 3. By extracting low-frequency structural features and expressing them through structural hashing, this invention enables the urban spatial morphology evolution process to be stably captured in the form of structural unit changes. This is beneficial for long-term monitoring of urban expansion patterns, functional reorganization paths, and spatial reconstruction trends, significantly improving the monitoring and pattern recognition capabilities of urban spatial pattern evolution and providing a stable structural benchmark for long-term time series analysis.

[0052] 4. The side length parameter of the low-frequency sub-block is used as a controllable parameter in the two-dimensional encoding stage to control the number of QR code grids and the granularity of spatial representation. Simultaneously, the order of the Hilbert curve is set as an adjustable parameter constrained by the resolution of the smallest coding unit of the QR code, used to control the unfolding scale of the two-dimensional to one-dimensional mapping and the barcode length. This invention can be flexibly configured according to the structural preservation requirements, encoding length requirements, and computational resource conditions under different application scenarios, thereby achieving a better balance between spatial structural representation capability, storage compression efficiency, and subsequent analysis performance. Attached Figure Description

[0053] Figure 1 This is a schematic diagram of the intercity similarity analysis method based on perceptual hashing and Hilbert curve provided in the embodiments of the present invention;

[0054] Figure 2 It is a "QR code" of urban nighttime light data generated by two-dimensional dimensionality reduction based on perceptual hashing in this embodiment of the invention;

[0055] Figure 3 This is a schematic diagram illustrating the principle of converting two-dimensional structure encoding to one-dimensional ordered encoding in an embodiment of the present invention;

[0056] Figure 4 This is a "barcode" of urban nighttime light data generated based on one-dimensional dimensionality reduction using Hilbert space-filling curves in this embodiment of the invention.

[0057] Figure 5 This is a flowchart of the intercity similarity analysis method based on perceptual hashing and Hilbert curves provided in the embodiments of the present invention. Detailed Implementation

[0058] like Figure 1 As shown, the technical solution in this embodiment of the invention addresses the technical problems existing in the prior art, and the overall approach is as follows:

[0059] A. Basic Information Input

[0060] This step is used to acquire urban nighttime light data at the intercity scale, and through built-up area clipping and raster standardization, form a unified input that is comparable across cities, providing a data foundation for subsequent generation of two-dimensional dimensionality reduction codes based on hash algorithms.

[0061] The specific steps are as follows:

[0062] A1. Collect urban nighttime light raster data and establish a sample index: Obtain nighttime light raster data for the cities to be analyzed, and establish a unique city identifier and time identifier for each city to construct an intercity nighttime light sample set. Nighttime light data is stored in raster form, with cell values ​​representing nighttime light intensity.

[0063] A2. Obtain and spatially align the built-up area boundaries: Obtain the built-up area boundary data for each city and perform spatial alignment processing on the built-up area boundaries and nighttime light grid data, including coordinate reference system consistency, spatial range matching, and projection transformation, to ensure that the built-up area boundaries can be accurately used for spatial clipping of nighttime light data.

[0064] A3. Based on the built-up area boundary obtained in A2, the nighttime light data collected in A1 is cropped and masked: Based on the built-up area boundary, the nighttime light raster data is cropped and masked, retaining only the pixel values ​​within the built-up area and eliminating non-built-up areas, thereby avoiding the dilution effect of non-built-up areas on the spatial structure characteristics of urban nighttime light.

[0065] A4. Raster Standardization and Output: The cropped urban nighttime light data is standardized to ensure that the nighttime light data of different cities are consistent in spatial resolution, pixel size and matrix dimension, forming standardized urban nighttime light data, which provides a unified input for subsequent dimensionality reduction coding.

[0066] B. Two-dimensional reduction coding and QR code generation based on perceptual hashing algorithm

[0067] This step simplifies and reduces the dimensionality of the standardized urban nighttime light data NTL output from step A4. A perceptual hashing algorithm is used to convert the high-dimensional nighttime light image into a fixed-size two-dimensional binary representation, thereby achieving robust compression, size standardization, and structural representation of the nighttime light data. This provides a unified input for subsequent space filling and one-dimensional mapping based on Hilbert curves. The specific steps are as follows:

[0068] B1. Unified scaling and grayscale matrix construction: The standardized urban nighttime light data is converted into a single-channel grayscale matrix, and the grayscale matrix is ​​scaled according to a preset size so that all urban data have a unified two-dimensional matrix scale.

[0069] B2. Two-dimensional Discrete Cosine Transform: Perform a two-dimensional discrete cosine transform on the grayscale matrix to map the nighttime light data from the spatial domain to the frequency domain, and obtain the corresponding frequency coefficient matrix.

[0070] B3. Low-frequency sub-block extraction: A low-frequency sub-block located in the upper left corner is extracted from the frequency coefficient matrix to characterize the overall features of the urban nighttime light spatial structure. The side length of the low-frequency sub-block is a settable parameter; by adjusting the side length value, the side length of the subsequent two-dimensional binary matrix and its grid resolution can be controlled.

[0071] B4. Low-frequency feature statistics and threshold determination;

[0072] B5. Binarization Generation of Urban Nighttime Light QR Codes: The frequency coefficients in the low-frequency sub-blocks are compared with the comparison values, and a fixed-size two-dimensional binary matrix is ​​generated through binarization. The side length of the two-dimensional binary matrix is ​​determined by the side length of the low-frequency sub-blocks, and the total number of grids can be controlled according to parameter settings, allowing for adjustment between the accuracy of urban spatial structure representation and data compression efficiency.

[0073] B6. Output the "QR code" of the two-dimensional reduction code for city nighttime lights: Output the two-dimensional binary matrix as the two-dimensional reduction result of the city nighttime light data, and associate it with the corresponding city identifier to form a QR code representation of the city nighttime lights, such as... Figure 2 As shown.

[0074] C. Two-dimensional to one-dimensional space filling mapping and "barcode" generation based on Hilbert curves

[0075] like Figure 3 As shown, this step is used to uniformly scale, grayscale, and perform two-dimensional discrete cosine transform on the urban nighttime light "QR code" output in step B6 to extract low-frequency structural features. Then, threshold binarization is used to generate a two-dimensional binary structure matrix with a fixed grid size, forming the urban nighttime light "QR code". Based on this, a Hilbert space-filling curve of the corresponding order is determined according to the grid size of the two-dimensional binary structure matrix. Following the continuous traversal path of this curve from the starting point to the ending point, the binary values ​​of each grid cell in the two-dimensional matrix are read sequentially to generate a fixed-length one-dimensional ordered binary sequence, i.e., the urban nighttime light "barcode". Through the above processing, this embodiment can achieve progressive dimensionality reduction and encoding compression of urban spatial data while preserving the adjacency relationships and overall structural features between two-dimensional spatial cells as much as possible, providing a unified, compact, and stable structured coding foundation for subsequent intercity similarity calculations, cluster analysis, rapid retrieval, and visualization comparison. The specific steps are as follows:

[0076] C1. Determine the relationship between the order of the Hilbert curve and the matrix matching: Based on the matrix size of the city night light QR code, determine the order of the Hilbert curve that matches it, so that the curve traversal path can cover all the units of the QR code matrix.

[0077] C2. Generate Hilbert curve traversal path: Generate a Hilbert curve traversal path based on a determined order, the path consisting of a set of sequentially arranged two-dimensional coordinates.

[0078] C3. Read the two-dimensional binary matrix according to Hilbert traversal order: Based on the fixed-size two-dimensional binary matrix output in step B6 (as the data source input to be processed) and the two-dimensional coordinate sequence generated in step C2 (as the location index input for spatial access), perform step-by-step dimensionality reduction extraction of spatial data. The specific processing logic is as follows: using the two-dimensional coordinate sequence as a navigation pointer, strictly following its arrangement order, access the spatial grid cells corresponding to the coordinates in the two-dimensional binary matrix one by one, and extract the binary bit value (0 or 1) at that position. The extracted binary bit values ​​are linearly concatenated according to the order of reading, and finally output a fixed-length one-dimensional ordered binary sequence. Through this data recombination process, the original two-dimensional spatial shape matrix is ​​mapped into a dimensionality-reduced one-dimensional data carrier. At the same time, relying on the topological properties of the Hilbert mapping, the adjacent local patterns in the two-dimensional urban nighttime light space are transformed into continuous codes with similar positions in the one-dimensional sequence.

[0079] The specific formula is as follows:

[0080] Following the Hilbert coordinate sequence generated in step C2, read the adjusted QR code matrix sequentially. Construct a one-dimensional ordered sequence from the binary values ​​at the corresponding coordinate positions. Its definition is:

[0081]

[0082] The resulting one-dimensional ordered binary sequence can be represented as:

[0083]

[0084] Wherein, the k-th element in the sequence This is the encoded value of the QR code matrix at the k-th position of the Hilbert traversal path.

[0085] C4. One-dimensional sequence expansion and construction of urban nighttime light "barcodes": Following the traversal order of the Hilbert curve, the binary bit values ​​of the corresponding positions in the urban nighttime light QR code matrix are read sequentially to generate a one-dimensional ordered sequence that maintains spatial proximity.

[0086] C5. Output the city's nighttime light "barcode" representation: Unfold the one-dimensional ordered sequence and output it as a fixed-length one-dimensional structure to form the city's nighttime light barcode representation, such as... Figure 4 As shown.

[0087] D. Intercity similarity measurement and cluster analysis based on QR codes and barcodes

[0088] This step performs intercity similarity analysis and clustering on the city nighttime light QR codes output in step B6 and the city nighttime light barcodes output in step C5. It constructs a unified feature representation and uses a clustering algorithm to identify city groups with similar nighttime light spatial structures, thereby outputting the city clustering results. The specific steps are as follows:

[0089] D1. Constructing Urban Feature Representations for Clustering: Using the urban nighttime light QR codes output in step B6 and the urban nighttime light barcodes output in step C5 as inputs, construct urban feature representations for intercity similarity analysis. The feature representation can use a one-dimensional binary sequence of the urban nighttime light barcode as the main feature vector, or it can unfold the urban nighttime light QR code into a one-dimensional vector according to a predetermined order, or combine the QR code features with the barcode features to form a uniform-length urban feature vector, used to characterize the spatial structure features of each city's nighttime lights.

[0090] D2. Feature Vector Standardization and Distance Metric Preparation: Using the city nighttime light QR code feature representation constructed in step D1 as the main input data, format standardization and dimension alignment are performed. Specifically, the standardization operation involves flattening the two-dimensional matrix-shaped nighttime light QR codes of each city according to fixed row and column traversal rules, converting them into one-dimensional binary feature sequences, and strictly verifying the total length of all sequences to ensure absolute consistency in the data bit width of the feature vectors from different cities. Based on the attribute that the feature vectors are all pure binary discrete data, Hamming distance is specified as a metric to characterize the spatial morphological differences between cities. Through the specified distance metric, the structural differences between the standardized binary feature vectors of different cities are transformed into specific numerical discrete distance relationships, thus providing a standardized metric foundation that supports fast bit operations for subsequent clustering analysis.

[0091] D3. Determine Clustering Parameters and Perform Cluster Analysis: Using the standardized city binary feature vector matrix output from step D2 and the specified Hamming distance metric model as input, perform unsupervised cluster analysis to group cities with similar nighttime light spatial structures into the same category, forming city clustering results. First, by calculating the average silhouette coefficient or within-group sum of squared errors under different numbers of clusters, find the values ​​where the evaluation index shows a clear optimal solution, thereby objectively determining the optimal number of clusters. Specifically, calculate the average silhouette coefficient and within-group sum of squared errors under different numbers of clusters K to evaluate the structural compactness and inter-class separation under different numbers of clusters.

[0092] D4. Output the city clustering results;

[0093] D5. Structured storage and interface output of clustering results.

[0094] E. Intercity retrieval and comparison based on clustering results

[0095] This step is used to build an intercity retrieval and comparison mechanism based on nighttime light dimensionality reduction coding, based on the city clustering results obtained in step D. By using structured calls to the city nighttime light QR codes and barcodes, it enables rapid retrieval of similar cities and quantitative comparison of spatial structural differences between cities.

[0096] E1. Construct an index structure for urban nighttime light types: Based on the results of urban clustering, establish an index relationship between urban identifiers and corresponding urban nighttime light QR codes and barcodes.

[0097] E2. Automatic one-dimensional similarity retrieval based on barcodes: Taking the city nighttime light type index structure established in step E1 as input, the cluster category labels of the city to be analyzed and the corresponding nighttime light barcode sequences are read from it. One-dimensional similarity retrieval based on Hamming distance is performed to obtain a set of candidate cities that are similar to the nighttime light spatial structure of the city to be analyzed.

[0098] E3. Automatic fine comparison based on QR codes: Taking the preliminary ranking set of candidate cities output in step E2 as input, retrieve the nighttime light QR code index of the city to be analyzed and each candidate city from the index structure established in step E1, and perform fine comparison of local space based on two-dimensional structure matrix.

[0099] E4. Output Intercity Search and Comparison Results: Extract the two-dimensional local structural fine-grained difference score calculated in step E3, and calculate the final spatial structural similarity comprehensive index between the city to be analyzed and each candidate city. Then, perform numerical comparison based on this comprehensive index to generate the final intercity similarity ranking list in descending order and simultaneously construct a standardized structured search result data table. This data table includes the identifier of the city to be analyzed, the identifier of the candidate city, the comprehensive similarity value, the similarity ranking position, the cluster category label, and the corresponding QR code and barcode data index information.

[0100] E5. Structured storage of retrieval and comparison results: The intercity retrieval and comparison results obtained in step E4 are stored in a structured form so that they can be repeatedly called in subsequent city analysis processes or jointly analyzed with multi-source data from other cities, thereby realizing the continuous application of the intercity retrieval and comparison function based on nighttime light dimensionality reduction coding.

[0101] E6. Retrieval and Comparison Result Retrieval: Retrieve the two-dimensional binary QR code matrix and one-dimensional binary barcode sequence corresponding to the city to be analyzed and candidate cities from step E5, and reorganize them into structured data units according to a fixed field order. For cross-city development planning comparison, structured data units of multiple cities under the same cluster category are retrieved according to cluster category labels, and the QR code matrix of each city is expanded into a one-dimensional binary vector according to a fixed row and column order. By calculating the Hamming distance between vectors of any two cities, the corresponding structural similarity result table is output for the common spatial pattern characteristics between cities. For long-term spatial morphological evolution monitoring applications at the intercity scale, structured data units of the same city can be sorted and retrieved according to city and time identifiers, and the Hamming distance between QR code matrices and barcode sequences between adjacent periods or specified periods is calculated, outputting the corresponding time-series difference result table to quantitatively track the evolution trajectory and expansion direction of urban spatial structure.

[0102] The structured coding system developed in this invention directly characterizes the spatial organization of cities, shifting urban classification from attribute-driven to spatial structure-driven, and providing a more accurate structural basis for the formulation of differentiated urban development strategies, spatial planning zoning, and the identification of renewal patterns.

[0103] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0104] like Figure 5 As shown in the embodiment of the present invention, the intercity similarity analysis method based on perceptual hashing and Hilbert curves includes:

[0105] Step S110: Obtain light data for each city;

[0106] This step will be explained in detail, involving obtaining light data for each city, including:

[0107] Obtain raw raster data of city lights with city identifiers, and create a unique identifier and time identifier for each city to form a city sample index table. Nighttime light data is stored in raster form, with cell values ​​representing light intensity;

[0108] Acquire the boundary data of the built-up areas of each city, and perform spatial alignment processing between the built-up area boundaries and the nighttime light grid data, including coordinate reference system consistency, spatial range matching and projection transformation, so as to ensure that the built-up area boundaries can be used to accurately crop the nighttime light grid, so as to avoid cropping errors caused by inconsistent spatial references.

[0109] Using the built-up area boundary as a constraint, the original raster data of urban lights are cropped and masked, retaining only the pixel values ​​within the built-up area and setting or removing pixels outside the built-up area as invalid values, thus obtaining the effective raster data of urban lights under the built-up area constraint. This process avoids the dilution effect of large areas of low-value pixels in non-built-up areas on the overall brightness structure of the city, thereby improving the ability to express the spatial morphology information of the city.

[0110] The effective raster data of lights in each city are standardized to obtain standardized effective raster data for each city, ensuring a unified input specification for data from different cities in subsequent calculations. Standardization includes unifying spatial resolution, pixel size, and raster matrix size. When the spatial extent of built-up areas differs between cities, it is converted into a two-dimensional matrix of a preset dimension through resampling, scaling, or padding, while preserving the relative spatial distribution structure of pixel values.

[0111] Step S120: Convert the light data of each city into a two-dimensional light matrix using a perceptual hash algorithm;

[0112] This step is explained in detail. It involves converting the light data of each city into a two-dimensional light matrix using a perceptual hashing algorithm, including:

[0113] The two-dimensional matrix of nighttime lights for each city is scaled and adjusted to a square matrix of a preset size, so that the nighttime light data of all cities are consistent in terms of the number of pixels and the matrix dimension.

[0114] The light data of each city is converted into a light grayscale matrix; the pixel value represents the nighttime light intensity, thus forming a unified input matrix G(x,y) for frequency domain transformation, the calculation formula of which is:

[0115]

[0116] Where G(x,y) represents the gray value of the pixel in the x-th row and y-th column, and N is the side length of the pre-scaled square matrix.

[0117] A two-dimensional discrete cosine transform is performed on the light grayscale matrix to map the light intensity distribution in the spatial domain to the frequency domain, resulting in the corresponding frequency coefficient matrix. In the frequency coefficient matrix, the low-frequency components are used to characterize the overall spatial structure of urban nighttime lights, while the high-frequency components are used to characterize local detail changes and noise information.

[0118] Specifically, a two-dimensional discrete cosine transform is performed on the light grayscale matrix to map the light intensity distribution in the spatial domain to the frequency domain, resulting in the corresponding frequency coefficient matrix, including:

[0119] Through expression obtaining a frequency coefficient matrix F(u,v); wherein, ; G(x,y) is a light grayscale matrix; N is the side length of the matrix. F(u,v) represents a discrete cosine transform coefficient at the position of the u-th row and the v-th column of the grayscale matrix in the frequency domain; when u and v take small values, they correspond to low-frequency components, which are used to characterize the overall spatial structure of urban nighttime lights; when u and v take large values, they correspond to high-frequency components, which are used to characterize local detail changes and noise information.

[0120] intercepting a low-frequency sub-block located at the upper left corner from the frequency coefficient matrix as a feature region; the size of the low-frequency sub-block is a preset fixed size. Through this processing, only main frequency information reflecting the overall spatial structure in the nighttime light image is retained, and the influence of high-frequency noise on subsequent representation is suppressed.

[0121] setting the side length of the preset low-frequency sub-block as m, which satisfies m<N, then intercepting a low-frequency sub-block L(i,j) from the upper left corner of the frequency coefficient matrix F(u,v), denoted as:

[0122]

[0123] the low-frequency sub-block L(i,j) is the frequency coefficient matrix that retains main low-frequency structural information characterizing the overall spatial pattern of urban nighttime lights.

[0124] in this embodiment, the side length of the low-frequency sub-block is a controllable parameter, which is used to control the grid resolution of two-dimensional dimension reduction coding. By changing the value of , the side length and the total number of grids of the generated two-dimensional binary matrix Q(i,j) can be directly changed, thereby realizing active control of the expression granularity of urban spatial structure. When the value is large, more overall spatial structure details can be retained, and the resolution ability of differences in spatial forms of different cities can be improved; when the value is small, the total number of two-dimensional grid units can be reduced, and the storage scale and subsequent similarity calculation overhead can be reduced. Correspondingly, the total number of grids of the generated two-dimensional binary matrix is , wherein, represents the total number of grid units in the two-dimensional code. By adjusting the value of the parameter m, control can be performed between the retention accuracy of urban spatial structure and data compression efficiency. Therefore, the present invention provides an adjustable two-dimensional structure coding mechanism for different analysis accuracy requirements and different computing power conditions.

[0125] comparing all frequency coefficients in the feature region L(i,j) with a comparison value ;

[0126] If the frequency coefficient is greater than or equal to the comparison value, the corresponding position is assigned a value of 1; if the frequency coefficient is less than the comparison value, the corresponding position is assigned a value of 0, thus obtaining a two-dimensional binary matrix of lights for each city. Through this binarization process, a fixed-size two-dimensional binary matrix Q(i, j) is generated, whose expression is:

[0127]

[0128] Where Q(i,j) represents the binary code value of the city nighttime light QR code at the i-th row and j-th column position.

[0129] The resulting two-dimensional binary matrix is: .

[0130] The two-dimensional binary matrix Q(i,j) is organized in a regular grid form, and its spatial position corresponds one-to-one with the position of each frequency coefficient in the low-frequency sub-block. It is used to characterize the overall spatial structure characteristics of urban nighttime lights and serves as the two-dimensional dimensionality reduction result of urban nighttime light data.

[0131] Specifically, statistical calculations are performed on all frequency coefficients in the extracted low-frequency sub-block L(i,j) to obtain the feature comparison value T of the low-frequency sub-block, which is used to characterize the overall level of the low-frequency structure. This comparison value serves as the threshold benchmark for binarization processing, used to distinguish the relative high and low frequency features in the nighttime light structure. Its specific representation is as follows:

[0132]

[0133] The generated fixed-size two-dimensional binary matrix is ​​output as a two-dimensional dimensionality-reduced encoding of urban nighttime light data, and associated with the corresponding city identifier to form a QR code representation of urban nighttime lights. The urban nighttime light QR code is stored in the form of a two-dimensional matrix Q(i,j), with its size and encoding rules remaining consistent across different cities, ensuring a unified structural expression for the nighttime light data after dimensionality reduction. This urban nighttime light QR code significantly reduces the dimensionality and storage size of the original nighttime light image data while retaining the main low-frequency features of the urban nighttime light spatial structure. It also serves as a standardized two-dimensional structured input, providing a unified and standardized input basis for subsequent steps such as Hilbert curve-based spatial filling traversal and one-dimensional barcode generation.

[0134] Step S130: Convert the two-dimensional light matrix into one-dimensional light data using the Hilbert space fill curve;

[0135] This step is explained in detail, converting a two-dimensional light matrix into one-dimensional light data using a Hilbert space-filling curve, including:

[0136] The order of the Hilbert curve is determined based on the side length of the two-dimensional lighting matrix, ensuring that the traversal grid generated by the Hilbert curve matches the two-dimensional binary matrix in spatial scale. Specifically, in the two-dimensional binary matrix... middle, Let be the side length of the QR code matrix. For a Hilbert curve of order r, its traversal grid side length is... The total number of grid cells it covers is The order r can be adaptively determined based on the side length of the QR code matrix, satisfying... It can also be preset according to the target encoding length, spatial resolution requirements, or computational complexity constraints. If the side length of the QR code matrix is ​​inconsistent with the set side length of the Hilbert traversal grid, the QR code matrix is ​​padded, cropped, or resampled to achieve scale matching with the Hilbert traversal grid.

[0137] Based on the order of Hilbert curves, the Hilbert space-filling algorithm is used to recursively partition the two-dimensional light matrix, generating a continuously traversed two-dimensional coordinate sequence covering the global space of the two-dimensional light matrix. Specifically, this process establishes a deterministic mathematical mapping from one-dimensional linear indices to two-dimensional grid coordinates, ultimately outputting a set of two-dimensional coordinate sequences strictly ordered according to spatial coherence rules. This output sequence explicitly specifies the unique access order for each spatial unit in the two-dimensional matrix, serving as the physical location index for the one-dimensional unfolding, thereby ensuring that the mapped one-dimensional data can retain the topological adjacency relationships of the original urban spatial morphology to the greatest extent.

[0138] Specifically, based on the order of the Hilbert curve, the Hilbert space-filling algorithm is used to recursively partition the two-dimensional light matrix, generating a continuously traversed two-dimensional coordinate sequence covering the global space of the two-dimensional light matrix, including:

[0139] Based on the order r of the Hilbert curve, construct a mapping function between the one-dimensional index and the two-dimensional coordinates of the Hilbert curve:

[0140]

[0141] Where k represents the k-th one-dimensional position index in the Hilbert curve traversal order. This indicates the corresponding coordinates in the two-dimensional light matrix at that location;

[0142] This results in coverage of the entire Two-dimensional coordinate sequence of a two-dimensional light matrix The two-dimensional coordinate sequence P is arranged in an ordered manner according to the Hilbert space filling rule, and is used to define the unique access order of each spatial unit of the QR code matrix.

[0143] Obtain the binary values ​​of the two-dimensional light matrix at the corresponding coordinate positions in sequence according to the two-dimensional coordinate sequence, and construct a one-dimensional ordered sequence.

[0144] The obtained ordered binary sequences are arranged sequentially along the horizontal dimension according to their generation order to construct a corresponding one-dimensional structured representation. This one-dimensional structured representation is presented as continuous stripes, where each position corresponds to a spatial unit in a two-dimensional binary matrix, and its value reflects the relative state of that spatial unit within the urban nighttime light structure. Through this arrangement, the original two-dimensional dimensionality reduction encoding is further mapped into a one-dimensional structural representation with a clear spatial order, and expressed in the form of a "barcode" to represent the urban nighttime light spatial structure. This "barcode" maintains the sequential characteristics of two-dimensional spatial proximity relationships while achieving a compact representation of nighttime light data in one-dimensional space.

[0145] The specific formula is as follows:

[0146] The obtained one-dimensional ordered binary sequence S is expanded sequentially according to its index order to construct the city's nighttime light barcode B, i.e.

[0147]

[0148] in, If the QR code matrix satisfies The barcode length can then be further expressed as... .

[0149] The constructed city nighttime light barcodes are output as one-dimensional reduced-dimensional codes of nighttime light data and are associated with corresponding city identifiers to form a city nighttime light "barcode" library. The barcodes are fixed-length one-dimensional structures, with their length determined by the matrix size of the two-dimensional reduced-dimensional encoding and the Hilbert curve traversal rule. This city nighttime light barcode, while preserving spatial proximity information in the two-dimensional reduced-dimensional encoding, achieves further compression and expression of the spatial structure of nighttime lights, providing an efficient and unified input for subsequent intercity similarity calculations and cluster analysis based on reduced-dimensional encoding.

[0150] Step S140: Cluster the cities based on one-dimensional light data;

[0151] This step is explained in detail, involving clustering of cities based on one-dimensional light data, including:

[0152] The initial cluster center selection strategy and maximum number of iterations are set. One-dimensional light data is input into a K-means or similar spatial clustering algorithm. The algorithm calculates the structural difference between the one-dimensional light data of each city and its cluster centers using the Hamming distance metric, and assigns each city to the group with the smallest distance. Then, based on the current division results, the one-dimensional light data for each group is recalculated and updated. This division and update process is repeated until the cluster centers no longer shift significantly or the maximum number of iterations is reached. Finally, a stable and convergent city clustering result is output, accurately grouping cities with highly similar nighttime light spatial structures into the same category.

[0153] Specifically, before clustering cities based on one-dimensional light data, the following steps are also included:

[0154] Through formula The average profile coefficient was calculated. This is used to measure the intra-class consistency and inter-class discriminability of the overall clustering results. For the sample The average Hamming distance between a sample and other samples in the same cluster reflects the degree of dispersion of the sample within its cluster. For the sample The average Hamming distance between a sample and samples in its nearest neighbor cluster reflects the degree of separation between that sample and its neighboring clusters. To obtain the maximum value As the optimal number of clusters under the silhouette coefficient criterion;

[0155] or,

[0156] Through formula Calculate the sum of squared errors within the group. This is used to measure the dispersion of samples within each cluster around the cluster center; where, Indicates sample With the Cluster centers The square of the Hamming distance between them Indicates the first The set of samples contained in each cluster; take The point where the rate of descent changes significantly corresponds to As the optimal number of clusters under the elbow criterion.

[0157] The clustering results are organized and output, assigning a corresponding cluster category identifier to each city and forming a city-category correspondence table. Simultaneously, representative features of each category can be output to characterize the overall differences in the nighttime light spatial structure among different city categories.

[0158] The city cluster category identifiers, corresponding QR codes, and barcode index information are linked and stored to form a structured clustering result dataset. This clustering result dataset can be used as input for subsequent city type analysis, similar city retrieval, or further spatial analysis.

[0159] Step S150: Based on the city's one-dimensional light data, calculate the similarity between the city to be analyzed and cities of the same category, and select a preset number of candidate cities based on the similarity results;

[0160] This step is explained in detail as follows: Based on the city's one-dimensional light data, the similarity between the city to be analyzed and cities of the same category is calculated. Based on the similarity results, a predetermined number of candidate cities are selected, including:

[0161] Cities are grouped according to clustering categories, and a corresponding index structure is established for each category. This index structure includes at least a city identifier, a corresponding city night light QR code index, and a city night light barcode index, which are used to support subsequent fast retrieval and comparison operations based on category constraints.

[0162] Based on the cluster category labels of the cities to be analyzed, one-dimensional light data of candidate cities with the same or adjacent cluster labels are extracted from the established index structure through Boolean index filtering operations. The one-dimensional light data is then assembled into a two-dimensional candidate feature matrix, thereby reducing the retrieval space and ensuring the consistency of feature dimensions.

[0163] Using the one-dimensional light data of the city to be analyzed as a benchmark, the dimension of the city is aligned with the two-dimensional candidate feature matrix of the candidate city by broadcasting and copying. Batch bitwise XOR operation is performed on the two matrices, and the XOR results are summed along the row direction to calculate the Hamming distance between the city to be analyzed and each candidate city, thus obtaining a distance vector containing the one-dimensional structural similarity measure of all candidate cities.

[0164] Candidate cities are sorted in ascending order by distance vector, and a predetermined number of candidate cities with the smallest Hamming distance are obtained to obtain a preliminary sorted set of candidate cities that are most similar to the nighttime light spatial structure of the city to be analyzed.

[0165] Step S160: Perform element-level bitwise XOR operations on the two-dimensional light matrix of the city to be analyzed and the two-dimensional light matrices of each candidate city to generate a two-dimensional residual matrix that accurately represents the distribution of spatial morphological differences between cities.

[0166] Step S170: Apply spatial sliding window statistical processing to the two-dimensional residual matrix, traverse the entire residual matrix sequentially according to the preset window size and moving step size, and statistically analyze the density values ​​of dissimilar sites in each local spatial unit block by block; numerically aggregate the density values ​​of all local blocks to quantify and obtain the comprehensive score of the two-dimensional local structural differences between the city to be analyzed and each candidate city.

[0167] Step S180: Obtain the intercity similarity analysis results based on the comprehensive score of two-dimensional local structural differences. Specifically, for all candidate cities in the preliminary candidate city ranking set, calculate the comprehensive score of two-dimensional local structural differences between each city and the city to be analyzed, and perform a secondary fine ranking based on this score; the lower the score, the higher the spatial local morphological matching degree between the two cities, and output a fine-grained intercity comparison ranking list arranged from low to high based on the comprehensive score of two-dimensional local structural differences.

[0168] To facilitate cross-city development plan comparisons, the following are also included:

[0169] Based on the one-dimensional light data and two-dimensional light matrix corresponding to the city to be analyzed and the candidate cities, the data is reorganized into structured call data units according to the set field order;

[0170] Structured data units of the same type of city are sorted and retrieved according to city identifier and time identifier. Hamming distance is calculated for one-dimensional light data and two-dimensional light matrix between adjacent periods or between specified periods to obtain the corresponding time series difference result table, which quantitatively tracks the evolution trajectory and expansion direction of urban spatial structure.

[0171] To achieve automatic generation, rendering, and interactive visual analysis of urban spatial feature codes, this invention also includes the following steps:

[0172] Step 1: Importing Intercity Spatial Data Parameters and Initializing the Visualization Interactive Console: This step aims to build a graphical front-end entry point for processing massive amounts of urban spatial data. The front-end receives the list of cities to be analyzed and basic geographic constraints imported by the user through the interactive interface as initial input parameters, and then calls the underlying spatial rendering engine to initialize the visualization interactive console. This console automatically loads the standardized city-level vector base map module and the core configuration parameter panel, and establishes an interface communication link with the underlying automated algorithm module. Finally, this step outputs an intelligent operation workbench with high-dimensional spatial data access capabilities and panoramic map rendering foundation, along with a corresponding front-end rendering canvas container, providing a human-computer interaction and visual display base for subsequent automated dimensionality reduction encoding generation and intercity comparison.

[0173] Step 2: Automated Visual Mapping and Synchronous Dynamic Rendering of Urban Spatial Feature Codes: Based on the intelligent operation platform and front-end rendering canvas container output from Step 1, this step performs automated visual mapping transformation from underlying numerical values ​​to high-fidelity graphics. The obtained two-dimensional binary matrix and one-dimensional ordered binary sequence of urban nighttime lights serve as underlying data support. The graphics rendering function library is called to assign discrete visual mapping attributes to the aforementioned binary values. Specifically, the numerical values ​​1 and 0 are strictly mapped to spatial pixel patches with contrasting colors or grayscale levels. A clear image of urban nighttime light QR codes and barcode maps are synchronously rendered in the specified canvas container generated in Step 1. This step ultimately dynamically generates and outputs a high-resolution dimensionality reduction feature map and a structural feature status monitoring panel in the front-end display area, allowing users to view the visual evolution results of the transformation from high-dimensional spatial raster to low-dimensional structural encoding in real time.

[0174] Step 3: Intelligent Intercity Similarity Retrieval and Multi-View Linked Comparison Based on Human-Computer Interaction: Building upon the high-resolution dimensionality-reduced feature map and structural feature status monitoring panel dynamically generated on the front end in Step 2, this step transforms the tedious matrix operation results into an intuitive, interactive analysis view. When a user selects a specific target city in the status monitoring panel and issues an intelligent retrieval command, it triggers automated calculations based on two levels of fine-grained comparison using barcodes and QR codes. After the backend completes the calculations, the front-end console receives the returned structured retrieval result data table and automatically generates a multi-view linked comparison dashboard by calling the data visualization component package in conjunction with the visual feature image rendered in Step 2. This dashboard integrates a similarity descending order chart and a two-dimensional residual space heatmap, ultimately outputting a highly integrated visualization analysis interface for intercity spatial structure differences. This allows users to perform immersive quantitative comparisons of spatial detail features between different candidate cities through interactive actions.

[0175] Step 4: Automated Chart Layout and Multimodal Analysis Report Export of Intelligent Comparison Results: To solidify the visualization analysis results and support subsequent spatial decision-making, this step uses the intercity spatial structure difference visualization analysis interface and multi-view linked comparison dashboard output in Step 3 as input for automated archiving. It obtains the status of the currently active linked comparison dashboard in the front-end console, accurately extracts high-resolution QR code and barcode images of the target city and highly matching city sets, and performs similarity calculation and quantitative scoring. Then, it automatically formats the extracted text and image content according to standardized visual templates for spatial data science analysis, and integrates and encapsulates the above multi-source data according to the user's export instructions, directly outputting a multimodal analysis report file containing vector graphic interface screenshots and detailed optimized calculation records, thus forming a technical closed loop from data input to graphic comparison to decision-making output.

[0176] In summary, this invention provides a method for dimensionality reduction, clustering, retrieval, and similarity analysis of intercity spatial structures based on nighttime light data. This method enables unified encoding, similarity analysis, and comparison of urban nighttime light spatial structures at the intercity scale. Based on urban nighttime light remote sensing data, this method combines hash dimensionality reduction with spatial filling mapping to achieve efficient encoding and comparable analysis of nighttime light images. Because this invention focuses on spatial structure rather than the numerical meaning of specific physical quantities, high-dimensional urban spatial data from different sources and with different dimensions can be mapped to a unified structural encoding system. This allows for collaborative analysis and similarity comparison of multi-source spatial data, including nighttime light, land use, building density, and functional intensity. Through two-dimensional structural hashing and adjacency-preserving sequence construction, this invention maintains key structural features of complex spatial patterns in a low-dimensional representation, thereby achieving rapid retrieval and high-precision comparison of spatial structures at the intercity scale, balancing efficiency and structural discrimination capabilities. Furthermore, through low-dimensional structure encoding and hierarchical retrieval mechanisms, this invention can complete the rapid indexing and analysis of large-scale spatial data in ordinary computing environments, avoiding strong dependence on high-performance computing resources. It is suitable for the engineering deployment needs of urban big data platforms, spatial decision support systems, and online spatial analysis services. The structured encoding and mapping method proposed in this invention is also applicable to various types of high-dimensional urban spatial data such as land use patterns, building spatial distribution, and functional intensity fields.

[0177] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0178] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0179] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0180] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0181] Any aspects of this invention not described in detail in the embodiments are well-known techniques to those skilled in the art. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this invention and not to limit it. Although this invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of this invention without departing from the spirit and scope of this invention, and all such modifications and substitutions should be covered within the scope of the claims of this invention.

Claims

1. An intercity similarity analysis method based on perceptual hashing and Hilbert curve, characterized in that, include: Obtain lighting data for each city; The light data of each city is converted into a two-dimensional light matrix using a perceptual hash algorithm; The two-dimensional light matrix is ​​converted into one-dimensional light data using a Hilbert space-filling curve. Clustering of cities is performed based on the aforementioned one-dimensional lighting data; Based on one-dimensional light data of the city, the similarity between the city to be analyzed and cities of the same category is calculated, and a preset number of candidate cities are selected based on the similarity results; The two-dimensional light matrix of the city to be analyzed and the two-dimensional light matrix of each candidate city are subjected to element-level bitwise XOR operation to generate a two-dimensional residual matrix that accurately represents the distribution of spatial morphological differences between cities. Spatial sliding window statistical processing is applied to the two-dimensional residual matrix. The global residual matrix is ​​traversed sequentially according to the preset window size and movement step size. The density values ​​of dissimilar sites in each local spatial unit are statistically analyzed block by block. The density values ​​of all local blocks are numerically aggregated to obtain a comprehensive score of the two-dimensional local structural differences between the city to be analyzed and each candidate city. The intercity similarity analysis results are obtained based on the comprehensive score of the two-dimensional local structural differences. 2.The intercity similarity analysis method based on perceptual hashing and Hilbert curve according to claim 1, wherein, The acquisition of light data for each city includes: Obtain the original raster data of city lights with city identifiers for each city; Using the built-up area boundary as a constraint, the original raster data of city lights are clipped and masked to obtain the effective raster data of city lights under the built-up area constraint. The effective raster data of lights in each city are standardized to obtain standardized effective raster data of lights in each city. 3.The intercity similarity analysis method based on perceptual hashing and Hilbert curve according to claim 1, wherein, The process of converting the light data of each city into a two-dimensional light matrix using a perceptual hash algorithm includes: Convert the light data of each city into a light grayscale matrix; Perform a two-dimensional discrete cosine transform on the light grayscale matrix to map the light intensity distribution in the spatial domain to the frequency domain, and obtain the corresponding frequency coefficient matrix. From the frequency coefficient matrix, the low-frequency sub-block located in the upper left corner is extracted as the feature region; Compare all frequency coefficients in the feature region with the comparison value; If the frequency coefficient is greater than or equal to the comparison value, the corresponding position is assigned a value of 1; if the frequency coefficient is less than the comparison value, the corresponding position is assigned a value of 0, thus obtaining a two-dimensional binary matrix of lights for each city. 4.The intercity similarity analysis method based on perceptual hashing and Hilbert curve according to claim 3, wherein, The step of performing a two-dimensional discrete cosine transform on the light grayscale matrix to map the light intensity distribution in the spatial domain to the frequency domain, obtaining the corresponding frequency coefficient matrix, includes: The frequency coefficient matrix is obtained by the expression ; wherein, ; G(x, y) is the light gray matrix; N is the side length of the matrix.​ 5.The intercity similarity analysis method based on perceptual hashing and Hilbert curve according to claim 1, wherein, The process of converting the two-dimensional light matrix into one-dimensional light data using a Hilbert space-filling curve includes: The order of the corresponding Hilbert curve is determined based on the side length of the two-dimensional light matrix. Based on the order of the Hilbert curve, the two-dimensional light matrix is ​​recursively partitioned using the Hilbert space filling algorithm to generate a continuously traversed two-dimensional coordinate sequence covering the global space of the two-dimensional light matrix. The binary values ​​of the two-dimensional light matrix at the corresponding coordinate positions are obtained sequentially according to the two-dimensional coordinate sequence to construct a one-dimensional ordered sequence. 6.The intercity similarity analysis method based on perceptual hashing and Hilbert curve according to claim 5, wherein, The step of recursively partitioning the two-dimensional light matrix using the Hilbert space-filling algorithm based on the Hilbert curve order to generate a continuously traversed two-dimensional coordinate sequence covering the global space of the two-dimensional light matrix includes: Based on the order r of the Hilbert curve, construct a mapping function between the one-dimensional index and the two-dimensional coordinates of the Hilbert curve: ; wherein k represents the kth one-dimensional position index in the Hilbert curve traversal order, represents the corresponding coordinate in the two-dimensional light matrix of the position index. Thus, a two-dimensional coordinate sequence of a two-dimensional light matrix is obtained which covers the entire two-dimensional coordinate sequence of a two-dimensional light matrix . 7.The intercity similarity analysis method based on perceptual hashing and Hilbert curve according to claim 1, wherein, The clustering of cities based on the one-dimensional light data includes: The one-dimensional light data is input into a K-means or similar spatial clustering algorithm for computation. The algorithm calculates the structural difference between the one-dimensional light data of each city and each cluster center based on the Hamming distance metric, and divides each city into the category set with the smallest distance. Then, the one-dimensional light data of each category is recalculated and updated based on the current division result. The above division and update process is continuously repeated until the cluster center no longer shifts significantly or the maximum number of iterations is reached. 8.The intercity similarity analysis method based on perceptual hashing and Hilbert curve of claim 7, wherein, Before clustering cities based on the one-dimensional light data, the method further includes: The average silhouette coefficient is calculated by the formula The average silhouette coefficient is calculated by the formula wherein is the average Hamming distance between the sample and other samples in the same cluster, is the average Hamming distance between the sample and the samples in the nearest neighbor cluster; and the value of k is taken as the optimal number of clusters under the silhouette coefficient criterion, which makes the maximum value the maximum value or, Through formula Calculate the sum of squared errors within the group. ;in, Indicates sample With the Cluster centers The square of the Hamming distance between them Indicates the first The set of samples contained in each cluster; take The point where the rate of descent changes significantly corresponds to As the optimal number of clusters under the elbow criterion.

9. The intercity similarity analysis method based on perceptual hashing and Hilbert curves as described in claim 1, characterized in that, The method uses one-dimensional light data of cities to calculate the similarity between the city to be analyzed and cities of the same category, and selects a preset number of candidate cities based on the similarity results, including: The cities are grouped according to clustering categories, and a corresponding index structure is created for each category of cities; Based on the cluster category labels of the cities to be analyzed, one-dimensional light data of candidate cities with the same or adjacent cluster labels are extracted from the established index structure through Boolean index filtering operation, and the one-dimensional light data is assembled into a two-dimensional candidate feature matrix. Based on the one-dimensional light data of the city to be analyzed, the dimension of the city is aligned with the two-dimensional candidate feature matrix of the candidate cities by broadcasting and copying. Batch bitwise XOR operation is performed on the two matrices, and the XOR results are summed along the row direction to calculate the Hamming distance between the city to be analyzed and each candidate city. Obtain a preset number of candidate cities with the minimum Hamming distance.

10. The intercity similarity analysis method based on perceptual hashing and Hilbert curves as described in claim 1, characterized in that, Also includes: Based on the one-dimensional light data and two-dimensional light matrix corresponding to the city to be analyzed and the candidate cities, the data is reorganized into structured call data units according to the set field order; Structured data units of the same type of city are sorted and retrieved according to city identifier and time identifier. Hamming distance is calculated for one-dimensional light data and two-dimensional light matrix between adjacent periods or between specified periods to obtain the corresponding time series difference result table, which quantitatively tracks the evolution trajectory and expansion direction of urban spatial structure.