Vector data clustering method and system integrating spatial constraints
By constructing and pruning the similarity matrix in the spectral clustering algorithm and considering the adjacency relationship of vector data, the problems of large computational complexity and insufficient utilization of spatial features in the spectral clustering algorithm when processing vector data are solved, and more accurate clustering results are achieved.
Patent Information
- Application Number
- CN202510659719.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-10-03
AI Technical Summary
Existing spectral clustering algorithms fail to effectively utilize spatial features when processing vector data, resulting in the clustering results being difficult to reflect spatial relationship characteristics. In addition, the amount of computation is large when processing large amounts of data, making it impossible to continue clustering.
By determining the number of clusters and the similarity threshold, the initial similarity matrix is constructed, abnormal nodes are removed, and a pruned similarity matrix is generated. The pruned similarity matrix is then input into the spectral clustering algorithm for clustering. The adjacency and attributes of vector elements are considered to reduce the amount of computation and take into account both spatial continuity and attribute similarity.
The spatial characteristics of vector data are effectively utilized to reduce the amount of spectral clustering operations, improve the spatial continuity and attribute similarity of clustering results, and avoid the defects of traditional spectral clustering methods.
Smart Images

Figure CN120744557A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of spatial data processing, and in particular to a vector data clustering method and system integrating spatial constraints. Background Art
[0002] Spatial clustering, also known as spatial partitioning, is a traditional research topic in geography. It not only reveals common processes and characteristics within clusters but also highlights differences between clusters, giving new clusters richer geographical meaning and providing a foundation for further integrative research. In spatial clustering methods, spatial units with similar attributes are grouped into the same cluster, and spatial continuity constraints are imposed to ensure the accuracy of the clustering results. This constraint can be strict, requiring that the spatial units in a cluster be spatially continuous, or soft, requiring only that the majority of spatial units be continuous.
[0003] Spectral clustering is an unsupervised clustering algorithm based on graph theory. It clusters data by representing data samples as nodes on a graph and then clustering them within that graph. Traditional spectral clustering algorithms fail to consider the spatial characteristics of the vector data and the spatial relationships between vector elements when clustering vector data. Consequently, their clustering results fail to reflect the spatial relationships between the objects under study. Furthermore, when processing a large number of vector elements simultaneously, similarity graphs can become so large that clustering becomes impossible.
[0004] However, weighing the geographic coordinates or location differences is crucial for the clustering results. Too large a location weight may lead to a large difference in attributes in the clustering results, while too small a location weight may lead to spatial dispersion and poor continuity in the clustering results. Summary of the Invention
[0005] The present invention provides a vector data clustering method and system integrating spatial constraints, which are used to solve the problems of large computational complexity of existing clustering operations and inability to effectively utilize spatial features of vector data.
[0006] The present invention provides a vector data clustering method integrating spatial constraints, comprising: Get the vector data to be classified; Determining the number of clusters based on the vector data and calculating a similarity threshold; Based on the similarity threshold, a clustering network is constructed through vector attributes and spatial adjacency relationships and an initial similarity matrix is generated; Reducing the similarity graph constructed by the initial similarity matrix and removing abnormal nodes to obtain a pruned similarity matrix; The pruned similarity matrix is input as a parameter into the preset spectral clustering algorithm for clustering to obtain the clustering results.
[0007] According to a vector data clustering method with integrated spatial constraints provided by the present invention, determining the number of clusters based on the vector data and calculating the similarity threshold specifically includes: Determining the number of clusters based on the vector data by a specified or set algorithm; Based on the number of clusters The algorithm clusters the vector data to obtain different clusters, calculates the similarity between all elements in each cluster, and determines the similarity threshold by sorting.
[0008] According to a vector data clustering method with integrated spatial constraints provided by the present invention, the method of constructing a clustering network based on a similarity threshold through vector attributes and spatial adjacency relationships and generating an initial similarity matrix specifically includes: Map at least one vector feature in the vector data into a node set, define the neighborhood range based on the central vector feature, create a vector feature buffer, find all other vector features that intersect with the current vector feature buffer, and mark all vector features that intersect with the buffer as adjacent to the current vector feature; All nodes are traversed to calculate the similarity of nodes in the neighborhood, and an edge set is constructed based on the similarity threshold to form a clustering network, and an initial similarity matrix is constructed simultaneously.
[0009] According to a vector data clustering method with integrated spatial constraints provided by the present invention, the similarity graphs in the initial similarity matrix are reduced and abnormal nodes are removed to obtain a pruned similarity matrix, specifically comprising: Traversing the initial similarity matrix to detect isolated nodes; When the isolated node meets the preset conditions, it is added to the abnormal node set; Node pruning is performed based on the abnormal node set to remove abnormal nodes and generate a pruned similarity matrix.
[0010] According to a vector data clustering method with integrated spatial constraints provided by the present invention, the pruned similarity matrix is input as a parameter into a preset spectral clustering algorithm to perform clustering to obtain a clustering result, specifically comprising: The pruned similarity matrix is input into the spectral clustering algorithm to obtain the clustering results of all non-abnormal nodes; According to the correspondence between vector elements and data points, all cluster labels are assigned to the classification result field of vector data, and abnormal nodes are marked at the same time, and the clustering results of the vector data spectral clustering method with integrated spatial constraints are obtained.
[0011] According to a vector data clustering method with integrated spatial constraints provided by the present invention, obtaining vector data to be classified specifically includes: The attribute data for clustering corresponding to each vector element is extracted from the vector data to be classified.
[0012] The present invention also provides a vector data clustering system integrating spatial constraints, the system comprising: A vector data acquisition module, used to acquire vector data to be classified; A similarity calculation module, configured to determine the number of clusters based on the vector data and calculate a similarity threshold; An initial similarity matrix generation module is used to construct a clustering network and generate an initial similarity matrix based on a similarity threshold through vector attributes and spatial adjacency relationships; a matrix pruning module, which reduces the similarity graph constructed by the initial similarity matrix and removes abnormal nodes to obtain a pruned similarity matrix; The clustering module inputs the pruned similarity matrix as a parameter into the preset spectral clustering algorithm to perform clustering and obtain the clustering results.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the vector data clustering method with integrated spatial constraints as described above is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described vector data clustering methods with integrated spatial constraints.
[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any one of the above-mentioned vector data clustering methods with integrated spatial constraints.
[0016] The present invention provides a vector data clustering method and system integrating spatial constraints. The method determines the number of clusters and calculates the similarity threshold through vector data. The spatial constraints are integrated by considering the adjacency relationship between vector elements. When creating a similarity graph, screening is performed to reduce and eliminate abnormal nodes to reduce the amount of spectral clustering operations. The pruned similarity matrix is input as a parameter into a preset spectral clustering algorithm for clustering to obtain clustering results. The method avoids the defect of the spectral clustering method that cannot take into account spatial continuity, and takes into account both spatial continuity and attribute similarity of the vector data clustering results. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 It is a flow chart of the vector data clustering method with integrated spatial constraints provided by the present invention.
[0019] Figure 2 This is a schematic diagram of vector data attribute data provided by the present invention.
[0020] Figure 3 This is a flowchart of the vector element neighborhood search provided by the present invention.
[0021] Figure 4 This is a schematic diagram of the vector element neighborhood search provided by the present invention.
[0022] Figure 5 It is a schematic diagram of the vector data clustering result provided by the present invention.
[0023] Figure 6 This is a schematic diagram of module connections of the vector data clustering system with integrated spatial constraints provided by the present invention.
[0024] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention.
[0025] Reference numerals: 110: vector data acquisition module; 120: similarity calculation module; 130: initial similarity matrix generation module; 140: matrix pruning module; 150: clustering module; 710: processor; 720: communication interface; 730: memory; 740: communication bus. DETAILED DESCRIPTION
[0026] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0027] The following combination Figure 1 The present invention describes a vector data clustering method integrating spatial constraints, which includes: step 100, obtaining vector data to be classified.
[0028] In the present invention, vector data to be classified and attribute data corresponding to each vector element for clustering are obtained.
[0029] In a specific embodiment, the Guerry dataset created by Andre-Michel Guerry in 1833 is used as an example. This dataset contains socioeconomic data of various administrative regions in France at that time, including demographics, crime rates, education levels and other indicators. This paper selects Crime_prop, Donations and Infants data as clustering fields to illustrate the implementation plan of the vector data spectral clustering method with integrated spatial constraints. In this embodiment, the attribute data of different fields used for classification are as follows: Figure 2 shown.
[0030] Step 200: Determine the number of clusters based on the vector data and calculate a similarity threshold.
[0031] Specifically, determining the number of clusters based on the vector data by specifying or setting an algorithm; Based on the number of clusters The algorithm clusters the vector data to obtain different clusters, calculates the similarity between all elements in each cluster, and determines the similarity threshold by sorting.
[0032] In the present invention, the number of clusters is directly specified Or use the elbow rule based on the k-means algorithm to determine the number of clusters for attribute data . use The algorithm clusters the vector data and obtains clusters, calculate the similarity between all elements in each cluster, sort the similarities between elements in each cluster from small to large, and select the first quartile of the similarity; The first quartile of the clusters is averaged to obtain the similarity threshold .
[0033] In a specific embodiment, the number of clusters is directly specified. Or determine the number of clusters for vector data based on the elbow rule In this embodiment, the elbow rule is used to determine the number of clusters in this embodiment. First, the range of the number of clusters is set to ; Then for each given number of clusters , using the K-means algorithm to get Clusters, calculate the distance from all samples to the center of the cluster to which they belong; square sum (SSE); draw a curve with the number of clusters as the horizontal axis and SSE as the vertical axis. Observe the curve and find an obvious inflection point. The inflection point is the determined number of clusters. In this embodiment, the number of clusters is 6. Use the K-means algorithm to cluster the example vector data to obtain 6 clusters; calculate the similarity between all elements in each cluster separately, and the similarity can be obtained by formula 1; sort the similarity between the elements in each cluster, select the first quartile of the similarity, and obtain the first quartile of the 6 clusters. Take the average of these six numbers and regard it as the similarity threshold , and in this embodiment, the similarity threshold is .
[0034] (1) in, is the node attribute vector, is the Euclidean distance between the attribute vectors of two nodes. Note that the similarity calculation method is not limited to this, and other methods can also be used to calculate the similarity.
[0035] Step 300: construct a clustering network based on a similarity threshold through vector attributes and spatial adjacency relationships and generate an initial similarity matrix.
[0036] Specifically, at least one vector element in the vector data is mapped into a node set, a neighborhood range is defined based on the central vector element, a vector element buffer is created, all other vector elements intersecting with the buffer of the current vector element are found, and all vector elements intersecting with the buffer are marked as adjacent to the current vector element; All nodes are traversed to calculate the similarity of nodes in the neighborhood, and an edge set is constructed based on the similarity threshold to form a clustering network, and an initial similarity matrix is constructed simultaneously.
[0037] In the present invention, all vector elements are mapped into node sets , with the center vector The feature is the base, and the creation range is Buffer, find all other vector features that intersect with the buffer of the current vector feature, mark all vector features that intersect with the buffer as adjacent to the current vector feature; recommended distance Take one fifth of the average width and height of the cluster area; traverse all nodes , calculate the nodes in its neighborhood Similarity: (2) in is the similarity calculation function, is the node attribute vector.
[0038] Retention satisfaction edge Construct edge sets , forming a graph structure Synchronous construction Order similarity matrix , where row and column indices correspond to vector numbers, and non-null elements store valid similarity values. Specifically, when searching for adjacent features within a buffer zone, if no feature intersects the buffer zone, there are two solutions. First, mark the current feature as isolated and proceed with the next step. Second, expand the buffer zone and search again. If a feature intersects the current buffer zone or the buffer zone exceeds a specified value, stop the search and proceed with the next step.
[0039] In one embodiment, all vector elements are mapped into node sets. ,initialization Order sparse matrix . Iterate through all vectors and create a range of Buffer, find all other vector features that intersect with the current vector feature buffer, mark all vector features that intersect with the buffer as adjacent to the current vector feature, and , calculate its adjacent elements according to formula 1 Similarity .
[0040] Retention satisfaction edge Construct edge sets , forming a graph structure Synchronous construction Order similarity matrix , whose row and column indices correspond to vector numbers, and non-empty elements store valid similarity values.
[0041] Specifically, when finding adjacent features using a buffer zone, there are two solutions when no features intersect the buffer zone: First, mark the current feature as an isolated feature and proceed to the subsequent steps; Second, expand the buffer range, and the buffer expansion step is , in this embodiment , re-search for features that intersect with the current buffer. When features that intersect with the current buffer appear, or the buffer range is greater than the maximum buffer range In this embodiment, , stop searching and proceed to the next steps. The process is as follows Figure 3 shown.
[0042] Step 400: Reduce the similarity graph constructed by the initial similarity matrix and remove abnormal nodes to obtain a pruned similarity matrix.
[0043] Specifically, the initial similarity matrix is traversed to detect isolated nodes; When the isolated node meets the preset conditions, it is added to the abnormal node set; Node pruning is performed based on the abnormal node set to remove abnormal nodes and generate a pruned similarity matrix.
[0044] In the present invention, detection Isolated node set ,in Is an indicator function, which takes the value 1 when the condition is true and 0 otherwise, through matrix pruning operation Generate refined matrix , is a set difference operation, which represents the node set after removing abnormalities, and the clustering network is updated synchronously: .
[0045] In a specific embodiment, by traversing Detect isolated nodes. When a node satisfy When it is added to the exception collection . Perform node pruning operations , obtain the refined similarity matrix , the process is as follows Figure 4 shown.
[0046] Step 500: Input the pruned similarity matrix as a parameter into a preset spectral clustering algorithm to perform clustering and obtain a clustering result.
[0047] Specifically, the pruned similarity matrix is input into the spectral clustering algorithm to obtain the clustering results of all non-abnormal nodes; According to the correspondence between vector elements and data points, all cluster labels are assigned to the classification result field of vector data, and abnormal nodes are marked at the same time, and the clustering results of the vector data spectral clustering method with integrated spatial constraints are obtained.
[0048] In this specific embodiment, the similarity matrix Construct the non-regularized Laplacian matrix ( ), Laplace matrix ( ) is calculated as in formula (3).
[0049] (3) in, is the Laplace matrix, is the degree matrix, is a diagonal matrix with diagonal elements , is a similarity matrix.
[0050] After constructing the Laplace matrix, solve the generalized eigenvalue problem ,get The eigenvector corresponding to the smallest eigenvalue .
[0051] According to the feature vector Constructing a Matrix . The matrix Each row is considered as a data point and the k-means algorithm is used to Perform clustering to obtain a cluster label for each data point.
[0052] Obtain the clustering results of all non-abnormal nodes, assign all cluster labels to the classification result field of the vector data according to the correspondence between vector elements and data points, and mark the abnormal nodes at the same time to obtain the clustering results of the vector data spectral clustering method with integrated spatial constraints. In this embodiment Figure 5 shown.
[0053] A vector data clustering method with integrated spatial constraints provided by the present invention determines the number of clusters and calculates the similarity threshold through vector data. The spatial constraints are integrated by considering the adjacency relationship between vector elements. When creating a similarity graph, screening is performed to reduce and eliminate abnormal nodes to reduce the amount of spectral clustering calculations. The pruned similarity matrix is input as a parameter into a preset spectral clustering algorithm for clustering to obtain clustering results. The method avoids the defect of the spectral clustering method that cannot take into account spatial continuity and takes into account both spatial continuity and attribute similarity of the vector data clustering results.
[0054] refer to Figure 6 The present invention also discloses a vector data clustering system integrating spatial constraints, the system comprising: A vector data acquisition module 110 is used to acquire vector data to be classified; A similarity calculation module 120 is configured to determine the number of clusters based on the vector data and calculate a similarity threshold; An initial similarity matrix generation module 130 is used to construct a clustering network based on a similarity threshold through vector attributes and spatial adjacency relationships and generate an initial similarity matrix; A matrix pruning module 140 is configured to reduce the similarity graph constructed by the initial similarity matrix and remove abnormal nodes to obtain a pruned similarity matrix; The clustering module 150 inputs the pruned similarity matrix as a parameter into a preset spectral clustering algorithm to perform clustering and obtain a clustering result.
[0055] The process of obtaining the vector data to be classified specifically includes: The attribute data for clustering corresponding to each vector element is extracted from the vector data to be classified.
[0056] The determining the number of clusters and calculating the similarity threshold based on the vector data specifically includes: Determining the number of clusters based on the vector data by a specified or set algorithm; Based on the number of clusters The algorithm clusters the vector data to obtain different clusters, calculates the similarity between all elements in each cluster, and determines the similarity threshold by sorting.
[0057] The method of constructing a clustering network based on a similarity threshold through vector attributes and spatial adjacency relationships and generating an initial similarity matrix specifically includes: Map at least one vector feature in the vector data into a node set, define the neighborhood range based on the central vector feature, create a vector feature buffer, find all other vector features that intersect with the current vector feature buffer, and mark all vector features that intersect with the buffer as adjacent to the current vector feature; All nodes are traversed to calculate the similarity of nodes in the neighborhood, and an edge set is constructed based on the similarity threshold to form a clustering network, and an initial similarity matrix is constructed simultaneously.
[0058] The similarity graphs in the initial similarity matrix are reduced and abnormal nodes are removed to obtain a pruned similarity matrix, specifically including: Traversing the initial similarity matrix to detect isolated nodes; When the isolated node meets the preset conditions, it is added to the abnormal node set; Node pruning is performed based on the abnormal node set to remove abnormal nodes and generate a pruned similarity matrix.
[0059] The pruned similarity matrix is input as a parameter into the preset spectral clustering algorithm to obtain clustering results, including: The pruned similarity matrix is input into the spectral clustering algorithm to obtain the clustering results of all non-abnormal nodes; According to the correspondence between vector elements and data points, all cluster labels are assigned to the classification result field of vector data, and abnormal nodes are marked at the same time, and the clustering results of the vector data spectral clustering method with integrated spatial constraints are obtained.
[0060] The present invention provides a vector data clustering system integrating spatial constraints. The system determines the number of clusters and calculates the similarity threshold based on vector data. The spatial constraints are integrated by considering the adjacency relationship between vector elements. When creating a similarity graph, the system performs screening, reduces and removes abnormal nodes to reduce the amount of spectral clustering operations. The pruned similarity matrix is input as a parameter into a preset spectral clustering algorithm to obtain clustering results. The system avoids the defect of the spectral clustering method that cannot take into account spatial continuity and takes into account both spatial continuity and attribute similarity of the vector data clustering results.
[0061] Figure 7 An example of a physical structure diagram of an electronic device is shown below. Figure 7 As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 may call logic instructions in the memory 730 to execute a vector data clustering method with integrated spatial constraints, the method comprising: obtaining vector data to be classified; determining the number of clusters and calculating a similarity threshold based on the vector data; constructing a clustering network based on the vector attributes and spatial adjacency relationships based on the similarity threshold and generating an initial similarity matrix; reducing the similarity graph constructed from the initial similarity matrix and removing abnormal nodes to obtain a pruned similarity matrix; and inputting the pruned similarity matrix as a parameter into a preset spectral clustering algorithm for clustering to obtain a clustering result.
[0062] Furthermore, the logic instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0063] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a vector data clustering method with integrated spatial constraints provided by the above methods, the method including: obtaining vector data to be classified; determining the number of clusters based on the vector data and calculating a similarity threshold; constructing a clustering network based on the similarity threshold through vector attributes and spatial adjacency relationships and generating an initial similarity matrix; reducing the similarity graph constructed by the initial similarity matrix and removing abnormal nodes to obtain a pruned similarity matrix; inputting the pruned similarity matrix as a parameter into a preset spectral clustering algorithm for clustering to obtain a clustering result.
[0064] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute a vector data clustering method with integrated spatial constraints provided by the above methods, the method comprising: obtaining vector data to be classified; determining the number of clusters based on the vector data and calculating a similarity threshold; constructing a clustering network based on the similarity threshold through vector attributes and spatial adjacency relationships and generating an initial similarity matrix; reducing the similarity graph constructed by the initial similarity matrix and removing abnormal nodes to obtain a pruned similarity matrix; and inputting the pruned similarity matrix as a parameter into a preset spectral clustering algorithm for clustering to obtain a clustering result.
[0065] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0066] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A vector data clustering method integrating spatial constraints, characterized in that: include: Get the vector data to be classified; Determining the number of clusters based on the vector data and calculating a similarity threshold; Based on the similarity threshold, a clustering network is constructed through vector attributes and spatial adjacency relationships and an initial similarity matrix is generated; Reducing the similarity graph constructed by the initial similarity matrix and removing abnormal nodes to obtain a pruned similarity matrix; The pruned similarity matrix is input as a parameter into the preset spectral clustering algorithm for clustering to obtain the clustering results.
2. The vector data clustering method with integrated spatial constraints according to claim 1, characterized in that: The determining the number of clusters and calculating the similarity threshold based on the vector data specifically includes: Determining the number of clusters based on the vector data by a specified or set algorithm; Based on the number of clusters The algorithm clusters the vector data to obtain different clusters, calculates the similarity between all elements in each cluster, and determines the similarity threshold by sorting.
3. The vector data clustering method with integrated spatial constraints according to claim 1, characterized in that: The method of constructing a clustering network based on a similarity threshold through vector attributes and spatial adjacency relationships and generating an initial similarity matrix specifically includes: Map at least one vector feature in the vector data into a node set, define the neighborhood range based on the central vector feature, create a vector feature buffer, find all other vector features that intersect with the current vector feature buffer, and mark all vector features that intersect with the buffer as adjacent to the current vector feature; All nodes are traversed to calculate the similarity of nodes in the neighborhood, and an edge set is constructed based on the similarity threshold to form a clustering network, and an initial similarity matrix is constructed simultaneously.
4. The vector data clustering method with integrated spatial constraints according to claim 1, characterized in that: The reducing the similarity graphs in the initial similarity matrix and removing abnormal nodes to obtain a pruned similarity matrix specifically includes: Traversing the initial similarity matrix to detect isolated nodes; When the isolated node meets the preset conditions, it is added to the abnormal node set; Node pruning is performed based on the abnormal node set to remove abnormal nodes and generate a pruned similarity matrix.
5. The vector data clustering method with integrated spatial constraints according to claim 1, characterized in that: The pruned similarity matrix is input as a parameter into a preset spectral clustering algorithm to perform clustering to obtain a clustering result, specifically including: The pruned similarity matrix is input into the spectral clustering algorithm to obtain the clustering results of all non-abnormal nodes; According to the correspondence between vector elements and data points, all cluster labels are assigned to the classification result field of vector data, and abnormal nodes are marked at the same time, and the clustering results of the vector data spectral clustering method with integrated spatial constraints are obtained.
6. The vector data clustering method with integrated spatial constraints according to claim 1, characterized in that: The obtaining of the vector data to be classified specifically includes: The attribute data for clustering corresponding to each vector element is extracted from the vector data to be classified.
7. A vector data clustering system integrating spatial constraints, characterized in that: The system comprises: A vector data acquisition module, used to acquire vector data to be classified; A similarity calculation module, configured to determine the number of clusters based on the vector data and calculate a similarity threshold; An initial similarity matrix generation module is used to construct a clustering network and generate an initial similarity matrix based on a similarity threshold through vector attributes and spatial adjacency relationships; a matrix pruning module, which reduces the similarity graph constructed by the initial similarity matrix and removes abnormal nodes to obtain a pruned similarity matrix; The clustering module inputs the pruned similarity matrix as a parameter into the preset spectral clustering algorithm to perform clustering and obtain the clustering results.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the vector data clustering method with integrated spatial constraints as claimed in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the vector data clustering method with integrated spatial constraints as claimed in any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the vector data clustering method with integrated spatial constraints as claimed in any one of claims 1 to 6 is implemented.