Industrial aggregation analysis method, system and equipment based on real estate registration data spatial clustering algorithm, and medium
Through the spatial clustering algorithm based on real estate registration data, the industrial clustering area is identified, and the problem that traditional methods cannot reveal the spatial distribution characteristics of the industry is solved, and scientific support for industrial layout optimization and resource allocation is achieved.
Patent Information
- Application Number
- CN202510113166.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional industrial analysis methods cannot reveal the spatial distribution characteristics and laws of industrial agglomeration as a whole, and it is difficult to discover industrial agglomeration areas, evaluate industrial agglomeration effects, and optimize industrial layout and resource allocation.
The spatial clustering algorithm based on real estate registration data is used to identify the clustering areas of different industrial types through data collection, preprocessing, spatial feature extraction, industrial attribute correlation and spatial clustering analysis, and obtain clustering results and interpret and apply them.
It has realized the disclosure of the spatial distribution characteristics of the industry, discovered industrial agglomeration areas, evaluated industrial agglomeration effects, optimized industrial layout and resource allocation, and provided scientific data support and decision-making basis.
Smart Images

Figure CN120045965A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cluster analysis and big data processing, and in particular to an industrial cluster analysis method, system, equipment and medium based on a spatial clustering algorithm for real estate registration data. Background Art
[0002] As an important source of geographic information data, real estate registration data contains detailed registration information of real estate such as land, houses, sea area use rights, and mineral resources. These data not only have spatial location attributes, but also contain a wealth of key attributes such as property type, area, and use, providing valuable data support for industrial agglomeration analysis. However, traditional industrial analysis methods are often limited to the statistics and analysis of individual enterprise information, and cannot reveal the spatial distribution characteristics and laws of industrial agglomeration as a whole.
[0003] Therefore, how to use real estate registration data to conduct spatial cluster analysis, reveal the spatial distribution characteristics of the industry, discover industrial agglomeration areas, evaluate the industrial agglomeration effect, and optimize industrial layout and resource allocation is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The technical task of the present invention is to provide an industrial agglomeration analysis method, system, equipment and medium based on the spatial clustering algorithm of real estate registration data, so as to solve the problem of how to use real estate registration data for spatial clustering analysis, reveal the characteristics of industrial spatial distribution, discover industrial agglomeration areas, evaluate industrial agglomeration effects, and optimize industrial layout and resource allocation.
[0005] The technical task of the present invention is achieved in the following way: an industrial clustering analysis method based on a spatial clustering algorithm of real estate registration data, the method is specifically as follows:
[0006] Data collection: Obtaining detailed registration data of real estate from real estate registration agencies; the detailed registration data of real estate includes geographical location information, property type, area and use of real estate;
[0007] Data preprocessing: Clean, remove duplicates, unify the format and standardize the collected detailed registration data of real estate, eliminate outliers and redundant information in the data, and form a data set that can be used for analysis;
[0008] Spatial feature extraction: Based on the preprocessed data set, the spatial features of real estate are extracted to construct a spatial attribute data set; the spatial features of real estate include spatial location, distribution density, adjacent relationship, and spatial relationship with other real estate or geographical elements;
[0009] Industry attribute association: associate the use attribute of real estate with the industry type to obtain the industry category to which each real estate belongs;
[0010] Spatial clustering analysis: Based on the DBSCAN clustering algorithm, adjust the clustering parameters ε (the minimum distance between two samples) and MinPts (the minimum number of samples required to form a cluster), and perform clustering analysis according to the spatial attribute dataset and industrial categories to identify the agglomeration areas of different industrial types and obtain the clustering results;
[0011] Result interpretation and application: Interpret the clustering results, analyze the spatial distribution characteristics, laws and influencing factors of industrial agglomeration; and apply the analysis results to regional economic planning, industrial planning formulation, land resource management and investment guidance.
[0012] Preferably, the data preprocessing further includes verifying the integrity and accuracy of the data, specifically: checking the missing values, outliers and duplicate values of the detailed registration data of the real estate collected, and screening and classifying the detailed registration data of the real estate collected according to the analysis requirements to form an accurate dataset;
[0013] The spatial feature extraction further includes calculating the spatial distance or similarity measure between real estates, constructing a spatial relationship matrix and extracting the spatial distribution pattern of real estates; among them, the real estate distribution pattern is a hot spot area, a business circle area or an agglomeration area.
[0014] Preferably, the industrial attribute association further includes mapping the real estate to the corresponding industrial classification system according to the use attribute of the real estate, and considering the multiple use attributes of the real estate to conduct multi-level industrial classification.
[0015] Preferably, the clustering results are represented by different colors or shapes to reveal the spatial distribution characteristics and laws of industrial agglomeration, and then check the rationality and accuracy of the clustering of the clustering results. According to the verification results, adjust the clustering parameters ε and MinPts values for optimization, and use clustering evaluation indicators to quantitatively evaluate the clustering results and timely adjust the parameters of the clustering algorithm for iterative optimization; among them, the clustering evaluation indicators include the silhouette coefficient and the Calinski-Harabasz index.
[0016] Preferably, the spatial clustering analysis is as follows:
[0017] ① Identification of core points and ε-neighborhood query: For point p, find all points N∈(p) within the ε-neighborhood of point p:
[0018] If |N∈(p)|≥MinPts, then mark p as a core point and create a new cluster, and add p and the points in N∈(p) to the new cluster;
[0019] ② Expansion of clusters: For each core point q in the cluster, repeat the ε-neighborhood query and core point determination:
[0020] If q is a core point, all points within the ε-neighborhood of point q are added to the current cluster;
[0021] If any point has not been visited and the number of points within the corresponding ε-neighborhood satisfies MinPts, the corresponding point is marked as a core point;
[0022] ③ Border point: During the process of expanding the cluster, if it is found that any point does not meet the condition to become a core point but belongs to the ε-neighborhood of any core point, the corresponding point is marked as a border point and added to the current cluster;
[0023] ④ Repeat steps ① to ③, select the next unvisited data point until all points have been visited;
[0024] ⑤ Set parameters: The selection of ε and MinPts is as follows:
[0025] The smaller ε and the larger MinPts are, the more and smaller clusters are generated;
[0026] The larger ε and the smaller MinPts are, the fewer and larger clusters are generated.
[0027] Preferably, the spatial clustering analysis further includes selecting the number of clusters, the distance metric method or the noise processing strategy according to the characteristics of the data and the analysis requirements, and visually displaying the clustering results to intuitively understand the spatial distribution characteristics of industrial clustering.
[0028] An industrial agglomeration analysis system based on a spatial clustering algorithm for real estate registration data, the system includes:
[0029] A data collection module for obtaining detailed registration data of real estate from a real estate registration agency; wherein, the detailed registration data of real estate includes the geographical location information, property type, area and use of the real estate;
[0030] A data preprocessing module for cleaning, de-duplicating, unifying the format and standardizing the collected detailed registration data of real estate, eliminating outliers and redundant information in the data, and forming a dataset available for analysis;
[0031] A spatial feature extraction module for extracting the spatial features of real estate based on the preprocessed dataset and constructing a spatial attribute dataset; wherein, the spatial features of real estate include spatial location, distribution density, adjacency relationship, and spatial relationship with other real estate or geographical elements;
[0032] An industrial attribute association module for associating the use attribute of real estate with the industrial type to obtain the industrial category to which each real estate belongs;
[0033] Spatial clustering analysis module, which is used to adjust the clustering parameters ε (the minimum distance between two samples) and MinPts (the minimum number of samples required to form a cluster) based on the DBSCAN clustering algorithm, perform clustering analysis according to the spatial attribute dataset and industrial categories, identify the aggregation areas of different industrial types, and obtain the clustering results; the clustering results are represented by different colors or shapes, revealing the spatial distribution characteristics and laws of industrial aggregation;
[0034] Result verification and optimization module, which is used to verify the rationality and accuracy of the clustering results, adjust the values of the clustering parameters ε and MinPts for optimization according to the verification results, and quantitatively evaluate the clustering results using clustering evaluation indicators and timely adjust the parameters of the clustering algorithm for iterative optimization; among them, the clustering evaluation indicators include the silhouette coefficient and the Calinski-Harabasz index;
[0035] Result interpretation and application module, which is used to interpret the clustering results, analyze the spatial distribution characteristics, laws and influencing factors of industrial aggregation; and apply the analysis results to regional economic planning, industrial planning formulation, land resource management and investment guidance.
[0036] Preferably, the working process of the spatial clustering analysis module is specifically as follows:
[0037] ① Identification of core points and ε-neighborhood query: For point p, find all points N∈(p) within the ε-neighborhood of point p:
[0038] If |N∈(p)|≥MinPts, then mark p as a core point and create a new cluster, and add p and the points in N∈(p) to the new cluster;
[0039] ② Expand the cluster: For each core point q in the cluster, repeat the ε-neighborhood query and core point determination:
[0040] If q is a core point, then add all points within the ε-neighborhood of point q to the current cluster;
[0041] If any point has not been visited and the number of points within the corresponding ε-neighborhood satisfies MinPts, then mark the corresponding point as a core point;
[0042] ③ Border points: During the process of expanding the cluster, if it is found that any point does not meet the condition of becoming a core point but belongs to the ε-neighborhood of any core point, then mark the corresponding point as a border point and add it to the current cluster;
[0043] ④ Repeat steps ① to ③, select the next unvisited data point until all points have been visited;
[0044] ⑤ Set parameters: The selection of ε and MinPts is specifically as follows:
[0045] The smaller ε is and the larger MinPts is, the more and smaller the generated clusters are;
[0046] The larger ε is and the smaller MinPts is, the fewer and larger the generated clusters are.
[0047] An electronic device, comprising: a memory and at least one processor;
[0048] Wherein, a computer program is stored on the memory;
[0049] The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the industrial agglomeration analysis method based on the spatial clustering algorithm of real estate registration data as described above.
[0050] A computer-readable storage medium, in which a computer program is stored, and the computer program can be executed by a processor to implement the industrial agglomeration analysis method based on the spatial clustering algorithm of real estate registration data as described above.
[0051] The industrial agglomeration analysis method, system, device and medium based on the spatial clustering algorithm of real estate registration data of the present invention have the following advantages:
[0052] (1) By integrating and processing key information such as geographical location, property type, area, and use in real estate registration data, the present invention constructs a basic data set for spatial analysis; subsequently, using advanced spatial feature extraction technology, core spatial features such as the spatial location, distribution density, and adjacent relationship of real estate are extracted from the basic data set to form a spatial attribute data set; finally, through visual display and interaction design, the analysis results are presented to users in an intuitive and easy-to-use manner;
[0053] (2) The present invention can use real estate registration data for spatial clustering algorithm analysis, reveal the spatial distribution characteristics of industries, discover industrial agglomeration areas, evaluate industrial agglomeration effects, optimize industrial layout and resource allocation, and support the formulation and implementation of plans, providing strong data support and scientific basis for industrial analysis and plan formulation;
[0054] (3) The present invention has good visualization effects: by utilizing the detail and spatiality of real estate registration data, the precision and visualization of industrial agglomeration analysis are realized;
[0055] (4) The present invention has accurate analysis: through the application of the spatial clustering algorithm, it can accurately identify the agglomeration areas of different industrial types, reveal the spatial distribution characteristics and laws of industrial agglomeration, and provide a scientific basis for regional economic planning;
[0056] (5) The present invention has strong decision-making support: combined with the regional economic background and policy environment, it provides a scientific basis and decision-making support for decision-making fields such as regional economic planning, industrial planning formulation, and land resource management, and helps to promote the sustainable development of the regional economy;
[0057] (6) The core steps, features, and their combination methods of the industrial agglomeration analysis method based on the spatial clustering algorithm of real estate registration data, as well as the structural composition of the corresponding analysis system, provide accurate and real-time decision-making support for industrial agglomeration analysis, and improve the scientificity and efficiency of planning decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The present invention will be further described below with reference to the accompanying drawings.
[0059] Att Figure 1 is a flow chart of an industrial agglomeration analysis method based on a spatial clustering algorithm of real estate registration data;
[0060] Att Figure 2 is an industrial agglomeration analysis diagram;
[0061] Att Figure 3 is an effect display diagram of industrial clustering analysis. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] The industrial agglomeration analysis method, system, device, and medium of the present invention based on the spatial clustering algorithm of real estate registration data will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.
[0063] Embodiment 1:
[0064] As shown in Att Figure 1 This embodiment provides an industrial agglomeration analysis method based on a spatial clustering algorithm of real estate registration data, and the method is as follows:
[0065] S1. Data collection: Obtain detailed registration data of real estate from the real estate registration agency; among them, the detailed registration data of real estate includes the geographical location information, property type, area, and use of the real estate;
[0066] S2. Data preprocessing: Clean, de-duplicate, format-unify, and standardize the collected detailed registration data of real estate, eliminate outliers and redundant information in the data, and form a data set available for analysis;
[0067] S3. Spatial feature extraction: Based on the preprocessed data set, extract the spatial features of real estate and construct a spatial attribute data set; among them, the spatial features of real estate include spatial position, distribution density, adjacent relationship, and spatial relationship with other real estate or geographical elements;
[0068] S4. Industrial Attribute Association: Associate the usage attributes of real estate with industrial types to obtain the industrial categories to which each piece of real estate belongs;
[0069] S5. Spatial Clustering Analysis: Based on the DBSCAN clustering algorithm, adjust the clustering parameters ε (the minimum distance between two samples) and MinPts (the minimum number of samples required to form a cluster), and perform clustering analysis according to the spatial attribute dataset and industrial categories to identify the aggregation areas of different industrial types and obtain the clustering results;
[0070] S6. Result Interpretation and Application: Interpret the clustering results, analyze the spatial distribution characteristics, rules, and influencing factors of industrial agglomeration; and apply the analysis results to regional economic planning, industrial planning formulation, land resource management, and investment guidance.
[0071] The data preprocessing in step S2 of this embodiment further includes verifying the integrity and accuracy of the data. Specifically, check the missing values, outliers, and duplicate values of the detailed registration data of the collected real estate, and screen and classify the detailed registration data of the collected real estate according to the analysis requirements to form an accurate dataset.
[0072] The spatial feature extraction in step S3 of this embodiment further includes calculating the spatial distance or similarity measure between real estates, constructing a spatial relationship matrix, and extracting the spatial distribution pattern of real estates; among them, the real estate distribution pattern is a hot spot area, a business circle area, or an aggregation area.
[0073] The industrial attribute association in step S4 of this embodiment further includes mapping the real estate to the corresponding industrial classification system according to the usage attributes of the real estate, and considering the multiple usage attributes of the real estate to perform multi-level industrial classification.
[0074] The clustering results in step S5 of this embodiment are represented by different colors or shapes to reveal the spatial distribution characteristics and rules of industrial agglomeration. Then, check the rationality and accuracy of the clustering for the clustering results, adjust the values of the clustering parameters ε and MinPts for optimization according to the verification results, and use clustering evaluation indicators to quantitatively evaluate the clustering results and timely adjust the parameters of the clustering algorithm for iterative optimization; among them, the clustering evaluation indicators include the silhouette coefficient and the Calinski-Harabasz index.
[0075] The spatial clustering analysis in step S5 of this embodiment is specifically as follows:
[0076] ① Identification of core points and ε-neighborhood query: For point p, find all points N ∈ (p) within the ε-neighborhood of point p:
[0077] If |N∈(p)| ≥ MinPts, then mark p as a core point and create a new cluster, adding p and the points in N∈(p) to the new cluster;
[0078] ② Expand the cluster: For each core point q in the cluster, repeat the ε-neighborhood query and core point determination:
[0079] If q is a core point, then add all the points within the ε-neighborhood of point q to the current cluster;
[0080] If any point has not been visited and the number of points within the corresponding ε-neighborhood satisfies MinPts, then mark the corresponding point as a core point;
[0081] ③ Border point: During the process of expanding the cluster, if it is found that any point does not meet the condition of becoming a core point but belongs to the ε-neighborhood of any core point, then mark the corresponding point as a border point and add it to the current cluster;
[0082] ④ Repeat steps ① to ③, select the next unvisited data point until all points have been visited;
[0083] ⑤ Set parameters: The selection of ε and MinPts is as follows:
[0084] When ε is smaller and MinPts is larger, more and smaller clusters are generated;
[0085] When ε is larger and MinPts is smaller, fewer and larger clusters are generated.
[0086] The spatial clustering analysis in step S5 of this embodiment further includes selecting the number of clusters, distance measurement method, or noise processing strategy according to the characteristics of the data and analysis requirements, and visually displaying the clustering results to intuitively understand the spatial distribution characteristics of industrial clustering.
[0087] Embodiment 2:
[0088] This embodiment provides an industrial aggregation analysis system based on a spatial clustering algorithm for real estate registration data. The system includes:
[0089] A data collection module for obtaining detailed registration data of real estate from real estate registration agencies; among them, the detailed registration data of real estate includes the geographical location information, property right type, area, and use of the real estate;
[0090] A data preprocessing module for cleaning, de-duplicating, unifying the format, and standardizing the collected detailed registration data of real estate, eliminating outliers and redundant information in the data, and forming a dataset available for analysis;
[0091] A spatial feature extraction module, which is used to extract the spatial features of real estate based on the preprocessed dataset and construct a spatial attribute dataset; wherein, the spatial features of real estate include spatial location, distribution density, adjacent relationship, and spatial relationship with other real estate or geographical elements;
[0092] An industrial attribute association module, which is used to associate the usage attributes of real estate with industrial types to obtain the industrial category to which each real estate belongs;
[0093] A spatial clustering analysis module, which is used to adjust the clustering parameters ε (the minimum distance between two samples) and MinPts (the minimum number of samples required to form a cluster) based on the DBSCAN clustering algorithm, and perform clustering analysis according to the spatial attribute dataset and industrial categories to identify the aggregation areas of different industrial types and obtain the clustering results; the clustering results are represented by different colors or shapes to reveal the spatial distribution characteristics and laws of industrial aggregation;
[0094] A result verification and optimization module, which is used to verify the rationality and accuracy of the clustering results, adjust the values of the clustering parameters ε and MinPts for optimization according to the verification results, and quantitatively evaluate the clustering results using clustering evaluation indicators and timely adjust the parameters of the clustering algorithm for iterative optimization; wherein, the clustering evaluation indicators include the silhouette coefficient and the Calinski-Harabasz index;
[0095] A result interpretation and application module, which is used to interpret the clustering results, analyze the spatial distribution characteristics, laws and influencing factors of industrial aggregation; and apply the analysis results to regional economic planning, industrial planning formulation, land resource management and investment guidance.
[0096] The working process of the spatial clustering analysis module in this embodiment is specifically as follows:
[0097] ① Identification of core points and ε-neighborhood query: For point p, find all points N∈(p) within the ε-neighborhood of point p:
[0098] If |N∈(p)|≥MinPts, then mark p as a core point and create a new cluster, and add p and the points in N∈(p) to the new cluster;
[0099] ② Expansion of clusters: For each core point q in the cluster, repeat the ε-neighborhood query and core point determination:
[0100] If q is a core point, then add all points within the ε-neighborhood of point q to the current cluster;
[0101] If any point has not been visited and the number of points within the corresponding ε-neighborhood satisfies MinPts, then mark the corresponding point as a core point;
[0102] ③ Border point: During the process of expanding the cluster, if any point is found not to meet the condition of becoming a core point but belongs to the ε-neighborhood of any core point, then the corresponding point is marked as a border point and added to the current cluster;
[0103] ④ Repeat steps ① to ③, select the next unvisited data point until all points have been visited;
[0104] ⑤ Set parameters: The selection of ε and MinPts is as follows:
[0105] When ε is smaller and MinPts is larger, more and smaller clusters are generated;
[0106] When ε is larger and MinPts is smaller, fewer and larger clusters are generated.
[0107] Embodiment 3:
[0108] This embodiment also provides an electronic device, including: a memory and a processor;
[0109] Wherein, the memory stores computer execution instructions;
[0110] The processor executes the computer execution instructions stored in the memory, so that the processor executes the industrial cluster analysis method based on the real estate registration data space clustering algorithm in any embodiment of the present invention.
[0111] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0112] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, memory, plug-in hard disks, smart media cards (SMCs), secure digital (SD) cards, flash memory cards, at least one magnetic disk storage period, flash memory devices, or other volatile solid-state storage devices.
[0113] Embodiment 4:
[0114] This embodiment also provides a computer-readable storage medium storing multiple instructions that are loaded by a processor to cause the processor to execute the industrial agglomeration analysis method based on the real estate registration data spatial clustering algorithm in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.
[0115] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments, so the program code and the storage medium storing the program code constitute a part of the present invention.
[0116] Examples of the storage medium for providing the program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.
[0117] In addition, it should be clear that not only can the functions of any one of the above embodiments be implemented by executing the program code read by the computer, but also by the operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.
[0118] Furthermore, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit executes part and all of the actual operations, thereby implementing the functions of any one of the above embodiments.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An industrial agglomeration analysis method based on real estate registration data spatial clustering algorithm, characterized in that: The method is as follows: Data collection: Obtaining detailed registration data of real estate from real estate registration agencies; the detailed registration data of real estate includes geographical location information, property type, area and use of real estate; Data preprocessing: Clean, remove duplicates, unify the format and standardize the collected detailed registration data of real estate, eliminate outliers and redundant information in the data, and form a data set that can be used for analysis; Spatial feature extraction: Based on the preprocessed data set, the spatial features of real estate are extracted to construct a spatial attribute data set; the spatial features of real estate include spatial location, distribution density, adjacent relationship, and spatial relationship with other real estate or geographical elements; Industry attribute association: associate the use attribute of real estate with the industry type to obtain the industry category to which each real estate belongs; Spatial clustering analysis: Based on the DBSCAN clustering algorithm, the clustering parameters ε and MinPts are adjusted, and clustering analysis is performed according to the spatial attribute data set and industry categories to identify the clustering areas of different industry types and obtain clustering results; Interpretation and application of results: Interpret the clustering results, analyze the spatial distribution characteristics, laws and influencing factors of industrial agglomeration; and apply the analysis results to regional economic planning, industrial planning, land resource management and investment guidance.
2. The industrial clustering analysis method based on the real estate registration data spatial clustering algorithm according to claim 1 is characterized in that: Data preprocessing also includes verifying the completeness and accuracy of the data, specifically: checking the missing values, outliers and duplicate values of the collected detailed registration data of real estate, and screening and classifying the collected detailed registration data of real estate according to analysis requirements to form an accurate data set; Spatial feature extraction also includes calculating the spatial distance or similarity measure between real estates, constructing a spatial relationship matrix and extracting the spatial distribution pattern of real estates; wherein the real estate distribution pattern is a hot spot area, a commercial circle area or a clustered area.
3. The industrial clustering analysis method based on the real estate registration data spatial clustering algorithm according to claim 1 is characterized in that: Industrial attribute association also includes mapping real estate to the corresponding industrial classification system based on its use attributes, and conducting multi-level industrial classification by considering the various use attributes of real estate.
4. The industrial clustering analysis method based on the spatial clustering algorithm of real estate registration data according to claim 1 is characterized in that: The clustering results are represented by different colors or shapes to reveal the spatial distribution characteristics and laws of industrial agglomeration. The clustering results are then checked to verify the rationality and accuracy of the clustering. According to the verification results, the clustering parameters ε and MinPts value are adjusted for optimization. The clustering evaluation indicators are used to quantitatively evaluate the clustering results and timely adjust the parameters of the clustering algorithm for iterative optimization. Among them, the clustering evaluation indicators include silhouette coefficient and Calinski-Harabasz index.
5. The industrial clustering analysis method based on the spatial clustering algorithm of real estate registration data according to claim 1 is characterized in that: The spatial cluster analysis is as follows: ① Identification of core points and ε-neighborhood query: For point p, find all points N∈(p) in the ε-neighborhood of point p: If |N∈(p)|≥MinPts, then mark p as a core point, create a new cluster, and add the points in p and N∈(p) to the new cluster; ② Expand the cluster: For each core point q in the cluster, repeat the ε-neighborhood query and core point determination: If q is a core point, all points in the ε-neighborhood of point q are added to the current cluster; If any point has not been visited and the number of points in the corresponding ε-neighborhood satisfies MinPts, the corresponding point is marked as a core point; ③ Boundary point: In the process of expanding the cluster, if any point is found not to meet the conditions of becoming a core point, but belongs to the ε-neighborhood of any core point, the corresponding point will be marked as a boundary point and added to the current cluster; ④ Repeat steps ① to ③, select the next unvisited data point, until all points have been visited; ⑤Set parameters: ε and MinPts selection, as follows: When ε is smaller and MinPts is larger, more and smaller clusters are generated; The larger ε is and the smaller MinPts is, the fewer and larger clusters are generated.
6. The industrial clustering analysis method based on the spatial clustering algorithm of real estate registration data according to any one of claims 1 to 5, characterized in that: Spatial cluster analysis also includes selecting the number of clusters, distance measurement methods or noise processing strategies according to the characteristics of the data and analysis requirements, as well as visualizing the clustering results to intuitively understand the spatial distribution characteristics of industrial clusters.
7. An industrial clustering analysis system based on real estate registration data spatial clustering algorithm, characterized in that: The system includes: The data collection module is used to obtain detailed registration data of real estate from the real estate registration agency; wherein the detailed registration data of real estate includes geographical location information, property type, area and use of the real estate; The data preprocessing module is used to clean, remove duplicates, unify the format and standardize the collected detailed registration data of real estate, eliminate outliers and redundant information in the data, and form a data set that can be used for analysis; A spatial feature extraction module is used to extract the spatial features of real estate based on the preprocessed data set and construct a spatial attribute data set; wherein the spatial features of real estate include spatial location, distribution density, adjacent relationship, and spatial relationship with other real estate or geographical elements; The industrial attribute association module is used to associate the use attribute of real estate with the industrial type and obtain the industrial category to which each real estate belongs; The spatial clustering analysis module is used to adjust the clustering parameters ε and MinPts based on the DBSCAN clustering algorithm, perform clustering analysis based on the spatial attribute data set and industry categories, identify clustering areas of different industry types, and obtain clustering results; the clustering results are represented in different colors or shapes to reveal the spatial distribution characteristics and laws of industrial clusters; The result verification and optimization module is used to verify the rationality and accuracy of clustering results, adjust the clustering parameters ε and MinPts value for optimization according to the verification results, and use clustering evaluation indicators to quantitatively evaluate the clustering results and timely adjust the parameters of the clustering algorithm for iterative optimization; among them, the clustering evaluation indicators include the silhouette coefficient and the Calinski-Harabasz index; The result interpretation and application module is used to interpret the clustering results, analyze the spatial distribution characteristics, laws and influencing factors of industrial agglomeration; and apply the analysis results to regional economic planning, industrial planning, land resource management and investment guidance.
8. The industrial clustering analysis system based on the real estate registration data spatial clustering algorithm according to claim 7 is characterized in that: The working process of the spatial clustering analysis module is as follows: ① Identification of core points and ε-neighborhood query: For point p, find all points N∈(p) in the ε-neighborhood of point p: If |N∈(p)|≥MinPts, then mark p as a core point, create a new cluster, and add the points in p and N∈(p) to the new cluster; ② Expand the cluster: For each core point q in the cluster, repeat the ε-neighborhood query and core point determination: If q is a core point, all points in the ε-neighborhood of point q are added to the current cluster; If any point has not been visited and the number of points in the corresponding ε-neighborhood satisfies MinPts, the corresponding point is marked as a core point; ③ Boundary point: In the process of expanding the cluster, if any point is found not to meet the conditions of becoming a core point, but belongs to the ε-neighborhood of any core point, the corresponding point will be marked as a boundary point and added to the current cluster; ④ Repeat steps ① to ③, select the next unvisited data point, until all points have been visited; ⑤Set parameters: ε and MinPts selection, as follows: When ε is smaller and MinPts is larger, more and smaller clusters are generated; The larger ε is and the smaller MinPts is, the fewer and larger clusters are generated.
9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores a computer program; The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the industrial clustering analysis method based on the spatial clustering algorithm of real estate registration data as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which can be executed by a processor to implement the industrial clustering analysis method based on the spatial clustering algorithm of real estate registration data as described in any one of claims 1 to 6.