Space-time density peak value clustering method based on local covariance and space-time shared neighbor
By defining spatiotemporal nearest neighbors and local covariance Mahalanobis distance in spatiotemporal data, and optimizing density calculation and two-step allocation strategies, the problems of cluster center identification and outlier identification in spatiotemporal clustering are solved, achieving more accurate clustering and outlier removal.
Patent Information
- Application Number
- CN202610004344.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2046-01-05
AI Technical Summary
Existing spatiotemporal density peak clustering methods struggle to accurately identify cluster centers when processing spatiotemporal data, easily overlooking key features in the time dimension, leading to unreasonable sample allocation and an inability to effectively identify anomalous samples.
By defining the spatiotemporal nearest neighbors of samples, calculating the local covariance matrix and Mahalanobis distance, constructing a spatiotemporal hybrid similarity matrix, and employing a two-step allocation strategy and local outlier factor to identify anomalous samples, the cluster center identification and sample attribution are optimized.
It achieves accurate clustering in spatiotemporal data, reduces error propagation, and improves the reliability of cluster center identification and the ability to identify anomalous samples.
Smart Images

Figure CN121456519A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a spatiotemporal density peak clustering method based on local covariance and spatiotemporal shared neighbors. BACKGROUND
[0002] At present, the information infrastructure such as perception, transmission and storage is continuously improved, and under the drive of technologies such as big data, cloud computing and Internet of Things, various perception terminals and business systems gradually converge into a large data set with diverse sources and huge scale. Under this data environment, the existing data processing and analysis methods face higher requirements to further realize the effective extraction of key information. Cluster analysis relies on the basic assumption of "similarity within clusters and difference between clusters", and can sort out potential patterns in unlabeled data. This method does not require prior labels, can extract internal rules from massive data, and is widely used in fields such as earthquake monitoring, public health, traffic management and environmental monitoring, helping researchers and decision makers to provide data support and practical basis.
[0003] In the process of continuous development of data form, its dimension is continuously expanded, and attribute association is also closer, and the traditional clustering method begins to face new adaptation problems. Taking spatiotemporal data as an example, it integrates spatial position and time series information, and can present the dynamic distribution and evolution law of geographical entities. Such data needs to consider spatial correlation, time continuity and attribute consistency at the same time, and its complexity is far beyond that of traditional data form, which puts higher requirements on the comprehensive processing ability of clustering methods.
[0004] In the existing spatiotemporal clustering method, the density peak clustering (DPC) algorithm has fewer parameters and is simple and efficient, and shows good application prospect, but when facing spatiotemporal data, DPC still has obvious limitations: the application of DPC algorithm for classification is insufficient to represent the real spatiotemporal relationship between samples, and the key features of the time dimension are easily ignored, it is difficult to accurately identify the cluster center, the division of the cluster is easy to expand the error range and cannot correct the previous deviation, resulting in unreasonable sample allocation.
[0005] Therefore, it is necessary to provide a new spatiotemporal data clustering method to obtain balanced expression between time synchronization and space connectivity of spatiotemporal data, and accurately cluster spatiotemporal data. SUMMARY
[0006] The purpose of the present application is to provide a spatiotemporal density peak clustering method based on local covariance and spatiotemporal shared neighbors, to obtain balanced expression between time synchronization and space connectivity of spatiotemporal data, and accurately cluster spatiotemporal data.
[0007] In a first aspect, the present application provides a spatiotemporal density peak clustering method based on local covariance and spatiotemporal shared neighbors, comprising: defining spatiotemporal Near neighbors, in the spatio-temporal The covariance matrix is calculated within the near neighbor range, the local covariance Mahalanobis distance between samples is calculated, the spatio-temporal shared near neighbors between samples are measured, a spatio-temporal hybrid similarity matrix between samples is constructed, the spatio-temporal shared near neighbor local density of the sample is calculated based on the spatio-temporal hybrid similarity between samples, the decision graph is constructed based on the spatio-temporal shared near neighbor local density of the sample and the local covariance Mahalanobis distance between samples, and the class cluster center is selected, the maximum similarity assignment is performed based on the spatio-temporal hybrid similarity matrix to obtain the preliminary classified class cluster and the edge sample, the edge sample is assigned to the nearest high-density cluster based on the local covariance Mahalanobis distance between samples, and the sample is assigned to the nearest high-density cluster based on the spatio-temporal The local outlier factor of the sample is calculated based on the local covariance Mahalanobis distance of the near neighbor sample, the abnormal sample is identified based on the local outlier factor, the abnormal sample is removed from the class cluster to obtain the final clustering result.
[0008] The spatio-temporal density peak clustering method based on local covariance and spatio-temporal shared near neighbors provided by the application has the beneficial effects that the covariance matrix is calculated within the local spatio-temporal neighborhood, the local covariance Mahalanobis distance between samples is defined based on this, the scale difference between time and space is automatically unified to improve the accuracy of similarity measurement, the local density calculation mode is optimized by using the spatio-temporal shared near neighbor number and the local covariance Mahalanobis distance to enhance the reliability of class cluster center identification, the double-step assignment and abnormal identification strategy is adopted, the sample attribution is gradually determined and improved through the assignment in the two stages of maximum similarity and nearest high density, the abnormal sample is effectively identified by using the local outlier factor, the error propagation of the sample is reduced, and accurate clustering is realized.
[0009] In a possible embodiment, the spatio-temporal Near neighbors, in the spatio-temporal The covariance matrix is calculated within the near neighbor range, including: defining the relationship between the sample The covariance matrix is calculated within the near neighbor range, including: defining the relationship between the sample The sample is defined as the sample The spatio-temporal Near neighbor set: , The time Euclidean distance between the sample And is represented by t, The time truncation distance is represented by T, The number of near neighbor samples is represented by K; the calculation of the covariance matrix satisfies the following formula: The local covariance matrix of the sample is represented by C, The total number of samples in the spatio-temporal Near neighbor set of the sample is represented by N, The sample The three-dimensional spacetime vector, Indicates sample The local mean vector, This represents the ridge regression regularization coefficient. It is a 3x3 identity matrix. This represents the transpose symbol.
[0010] In another possible embodiment, the Mahalanobis distance of the local covariance between samples is calculated according to the following formula:
[0011] ,in, The Mahalanobis distance represents the local covariance between samples. Represents a symmetric mean metric matrix. Indicates sample The inverse of the local covariance matrix.
[0012] In other possible embodiments, the spatiotemporal shared nearest neighbor local density of samples is calculated according to the following formula: ,in, Indicates sample The spatiotemporal shared nearest neighbor local density, where N represents the sample and Spatiotemporal shared nearest neighbor count, Indicates sample and Spatiotemporal mixing similarity, Indicates sample spacetime The total number of samples in the nearest neighbor set.
[0013] The process involves constructing a decision graph based on the spatiotemporal shared nearest neighbor local density and the Mahalanobis distance of the local covariance between samples, and selecting cluster centers. This includes: calculating the relative distance between samples based on the Mahalanobis distance of the local covariance between samples; constructing a decision graph with the spatiotemporal shared nearest neighbor local density and the relative distance as the horizontal and vertical axes; calculating decision values based on the spatiotemporal shared nearest neighbor local density and the relative distance; and selecting cluster centers based on the decision values of the samples.
[0014] The process of obtaining preliminary cluster and edge samples by maximizing similarity allocation based on the spatiotemporal hybrid similarity matrix includes: using the cluster center sample as the initial labeled node, traversing the spatiotemporal hybrid similarity matrix to find the maximum similarity path between the initial labeled node and the unlabeled node; assigning the unlabeled node to a cluster based on the similarity value between nodes in the spatiotemporal hybrid similarity matrix; clearing the assigned nodes from the spatiotemporal hybrid similarity matrix, and expanding the allocation along the maximum similarity path until all cross-cluster similarity values are zero to complete the maximum similarity allocation; and recording the remaining unclassified samples after the maximum similarity allocation as edge samples.
[0015] The edge sample is assigned to the nearest high-density cluster based on the local covariance Mahalanobis distance between samples, including: selecting a neighbor sample with the minimum distance according to the local covariance Mahalanobis distance between the neighbor sample and the corresponding edge sample; and assigning the class cluster label of the neighbor sample to the edge sample to assign the edge sample to the nearest high-density cluster.
[0016] The local outlier factor of the sample is calculated according to the local covariance Mahalanobis distance of the sample and the spatio-temporal neighbor sample, including: searching for the spatio-temporal neighbor sample of the sample to be identified in the spatio-temporal The local outlier factor of the sample is calculated according to the local covariance Mahalanobis distance of the sample and the spatio-temporal neighbor sample, including: searching for the spatio-temporal neighbor sample of the sample to be identified in the spatio-temporal The local outlier factor of the sample is calculated according to the local covariance Mahalanobis distance of the sample and the spatio-temporal neighbor sample, including: searching for the spatio-temporal neighbor sample of the sample to be identified in the spatio-temporal The local outlier factor of the sample is calculated according to the local covariance Mahalanobis distance of the sample and the spatio-temporal neighbor sample, including: searching for the spatio-temporal neighbor sample of the sample to be identified in the spatio-temporal
[0017] The abnormal sample is identified according to the local outlier factor, including: comparing the local outlier factor with a preset abnormal threshold, and determining the corresponding sample as an abnormal sample when the local outlier factor is greater than the abnormal threshold.
[0018] In a second aspect, the present application further provides a spatio-temporal density peak clustering device based on local covariance and spatio-temporal shared neighbor, including: a covariance matrix calculation unit, configured to calculate a covariance matrix of a sample in a spatio-temporal neighbor range of the sample; a Mahalanobis distance calculation unit, configured to calculate a local covariance Mahalanobis distance between samples; a local density calculation unit, configured to measure a spatio-temporal shared neighbor between samples, construct a spatio-temporal hybrid similarity matrix between samples, and calculate a spatio-temporal shared neighbor local density of the sample based on a spatio-temporal hybrid similarity between samples; a class cluster center selection unit, configured to construct a decision graph and select a class cluster center based on the spatio-temporal shared neighbor local density of the sample and the local covariance Mahalanobis distance between samples; and a class cluster division unit, configured to perform maximum similarity assignment based on the spatio-temporal hybrid similarity matrix to obtain a preliminary division class cluster and an edge sample, assign the edge sample to the nearest high-density cluster based on a local covariance Mahalanobis distance between samples, and identify an abnormal sample according to a local outlier factor of the sample and the spatio-temporal neighbor range of the sample; a Mahalanobis distance calculation unit, configured to calculate a local covariance Mahalanobis distance between samples; a local density calculation unit, configured to measure a spatio-temporal shared neighbor between samples, construct a spatio-temporal hybrid similarity matrix between samples, and calculate a spatio-temporal shared neighbor local density of the sample based on a spatio-temporal hybrid similarity between samples; a class cluster center selection unit, configured to construct a decision graph and select a class cluster center based on the spatio-temporal shared neighbor local density of the sample and the local covariance Mahalanobis distance between samples; and a class cluster division unit, configured to perform maximum similarity assignment based on the spatio-temporal hybrid similarity matrix to obtain a preliminary division class cluster and an edge sample, assign the edge sample to the nearest high-density cluster based on a local covariance Mahalanobis distance between samples, and identify an abnormal sample according to a local outlier factor of the sample and the spatio-temporal neighbor range of the sample; a Mahalanobis distance calculation unit, configured to calculate a local covariance Mahalanobis distance between samples; a local density calculation unit, configured to measure a spatio-temporal shared neighbor between samples, construct a spatio-temporal hybrid similarity matrix between samples, and calculate a spatio-temporal shared neighbor local density of the sample based on a spatio-temporal hybrid similarity between samples; a class cluster center selection unit, configured to construct a decision graph and select a class cluster center based on the spatio-temporal shared neighbor local density of the sample and the local covariance Mahalanobis distance between samples; and a class cluster division unit, configured to perform maximum similarity assignment based on the spatio-temporal hybrid similarity matrix to obtain a preliminary division class cluster and an edge sample, assign the edge sample to the nearest high-density cluster based on a local covariance Mahalanobis distance between samples, and identify an abnormal sample according to a local outlier factor of the sample and the spatio-temporal
[0019] The beneficial effects of the above-mentioned second aspect can be referred to the description of the first aspect. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1A flowchart of a spatiotemporal density peak clustering method based on local covariance and spatiotemporal shared neighbors is provided for the embodiment of the present application.
[0021] Figure 2a A data set composed of four spatiotemporal clusters with different densities is provided for the embodiment of the present application for identification and comparison A distribution state diagram in three-dimensional space is shown.
[0022] Figure 2b A cluster center identification result diagram obtained by applying the DPC algorithm is provided for the embodiment of the present application.
[0023] Figure 2c A cluster center identification result diagram obtained by applying the spatiotemporal density peak clustering method based on local covariance and spatiotemporal shared neighbors is provided for the embodiment of the present application.
[0024] Figure 3 A distribution diagram of experimental seismic spatiotemporal data is provided for the embodiment of the present application.
[0025] Figure 4 A seismic data clustering result diagram obtained by applying the spatiotemporal density peak clustering method based on local covariance and spatiotemporal shared neighbors is provided for the embodiment of the present application.
[0026] Figure 5 A schematic diagram of a spatiotemporal density peak clustering device based on local covariance and spatiotemporal shared neighbors is provided for the embodiment of the present application.
[0027] Figure 6 An electronic device structure diagram is provided for the embodiment of the present application. DETAILED DESCRIPTION
[0028] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application. Unless otherwise defined, the technical terms or scientific terms used herein should be understood as the general meanings understood by those skilled in the art in the field of the present application. The words such as “comprise” and similar words used herein mean that the elements or objects before the words cover the elements or objects listed after the words and their equivalents, without excluding other elements or objects.
[0029] The present embodiment provides a spatiotemporal density peak clustering method based on local covariance and spatiotemporal shared neighbors. Referring to the drawings in the descriptionFigure 1 The method comprises:
[0030] S101: defining the spatiotemporal of the sample nearest neighbors, in the spatiotemporal of the sample calculating the covariance matrix within the nearest neighbor range.
[0031] In one possible embodiment, the spatiotemporal of the sample is defined nearest neighbors, in the spatiotemporal of the sample calculating the covariance matrix within the nearest neighbor range, comprising:
[0032] defining the relationship between the sample satisfies the following formula samples are the samples in the spatiotemporal of the sample nearest neighbor set: , denotes the time Euclidean distance between the sample and denotes the time truncation distance, denotes the number of nearest neighbor samples.
[0033] The calculation of the covariance matrix satisfies the following formula: denotes the local covariance matrix of the sample denotes the total number of samples in the spatiotemporal nearest neighbor set of the sample denotes the three-dimensional spatiotemporal vector of the sample denotes the local mean vector of the sample denotes the ridge regression regularization coefficient, is a 3-order unit matrix, denotes the transposition symbol.
[0034] In one specific embodiment, according to the definition samples satisfying the time constraint are screened , and then the sample set obtained by screening is sorted in ascending order of the spatial distance from the sample , and the nearest samples are sequentially selected as the spatiotemporal nearest neighbor set of the sample , if there are less than samples satisfying the screening condition, all the samples obtained by screening are regarded as the spatiotemporal nearest neighbor set of the sample .
[0035] The local covariance matrix is constructed based on the spatio-temporal neighbors of the sample, which only utilizes the covariance information within the region to capture the local correlation in both temporal and spatial dimensions. The calculation of the local covariance matrix specifically satisfies the following formula: wherein, represents the local covariance matrix of the sample ; represents the spatio-temporal neighbors of the sample ; represents the total number of samples in the spatio-temporal neighbors of the sample ; represents the three-dimensional spatio-temporal vector of the sample , the three-dimensional spatio-temporal vector contains the temporal, longitude and latitude attributes of the sample ; represents the local mean vector of the sample ; represents the ridge regression regularization coefficient, the value of which is ; is a 3-order unit matrix; represents the transpose symbol, which is a basic operation of a vector or a matrix, used to interchange the row and column positions of the vector.
[0036] The local covariance matrix is constructed based on the spatio-temporal neighbors of the sample , the local vector mean of the samples in the region is first calculated, then the deviation outer product of each sample and the mean is calculated and summed, and finally the normalization is realized by using as the denominator. The introduction of the ridge regression regularization coefficient ensures the positive definite and invertible of the local covariance matrix, and the combination of the unit matrix performs a small amplitude correction. The local covariance matrix calculation method designed by the present application can not only avoid the problem of singular matrix to ensure the stability of subsequent calculation, but also will not damage the key spatio-temporal information originally recorded by the matrix.
[0037] S102: Calculate the local covariance Mahalanobis distance between samples.
[0038] In a possible embodiment, the calculation of the local covariance Mahalanobis distance between samples satisfies the following formula: wherein, represents the local covariance Mahalanobis distance between samples, represents the symmetric average metric matrix, represents the inverse matrix of the local covariance matrix of the sample .
[0039] The calculation of local covariance Mahalanobis distance introduces the local covariance matrix of the samples to measure the distance of similarity between two samples in their local spatiotemporal distribution. Its calculation... In the sample The inverse of the local covariance matrix and samples The inverse of the local covariance matrix Provides a local scaling factor in the metric space, and a symmetric average metric matrix. make sure It satisfies positive definiteness and symmetry. The metric weights of the aforementioned local covariance Mahalanobis distance can reflect the local fluctuation characteristics of time and space coordinates. When the local variance of a certain dimension is large, its proportion in the metric weight will decrease accordingly, avoiding misjudging the natural extended distribution characteristics of the data as a separation state between samples. When the local variance of a certain dimension is small and there is a clear main distribution direction, its proportion in the metric weight will increase significantly, so as to both amplify the differences of samples in that dimension and highlight the anisotropy of the data distribution.
[0040] The metric weights of the aforementioned local covariance Mahalanobis distance are entirely derived from the statistical information of the sample’s spatiotemporal nearest neighbors, without the need for manual pre-setting of spatiotemporal dimension weights. It can achieve dynamic adaptive matching of local features in different data regions. Compared with the fixed-weight distance calculation method in traditional density peak clustering, the local covariance Mahalanobis distance designed in this invention can dynamically adjust the metric scale according to the local density and distribution direction of the data, and better characterize the spatiotemporal relative positional relationship between samples.
[0041] S103: Measure the spatiotemporal shared nearest neighbors between samples, construct the spatiotemporal hybrid similarity matrix between samples, and calculate the local density of spatiotemporal shared nearest neighbors of samples based on the spatiotemporal hybrid similarity between samples.
[0042] In one possible implementation, the spatiotemporal shared nearest neighbor local density of samples is calculated according to the following formula: ,in, Indicates sample The spatiotemporal shared nearest neighbor local density, where N represents the sample and Spatiotemporal shared nearest neighbor count, Indicates sample and Spatiotemporal mixing similarity, Indicates sample spacetime The total number of samples in the nearest neighbor set.
[0043] The application introduces time dimension information in local density estimation, and defines a spatiotemporal shared neighbor local density. The number of spatiotemporal shared neighbors of a sample pair is counted, and the similarity of samples in the spatiotemporal dimension is calculated based on the local covariance Mahalanobis distance. A new local density calculation method is constructed by combining the similarity sum, the effective neighbor size, and the average similarity of the neighborhood, and a spatiotemporal shared neighbor local density is obtained, which takes into account the time synchronization and the space connectivity.
[0044] In a specific embodiment, the spatiotemporal shared neighbor local density is calculated by the following steps:
[0045] Measuring the spatiotemporal shared neighbors between samples: , denotes the shared neighbor set of samples and , denotes the spatiotemporal neighbor set of sample , denotes the spatiotemporal neighbor set of sample , denotes the spatiotemporal neighbor set of sample .
[0046] According to the number of spatiotemporal shared neighbors between samples and the distance measure between samples, the spatiotemporal hybrid similarity between samples is calculated: denotes the spatiotemporal hybrid similarity between samples and , denotes the number of shared neighbors between samples and , denotes the median of the local covariance Mahalanobis distance between all samples. A spatiotemporal hybrid similarity matrix is constructed according to the calculation results of the spatiotemporal hybrid similarity.
[0047] Calculating the spatiotemporal shared neighbor local density: .
[0048] The time-space shared near-neighbor set of the sample fuses a large amount of local structure information, which is helpful to reveal the structure overlapping characteristics between local regions, so as to more effectively capture the distribution rule of the sample in the time-space dimension. The time-space hybrid similarity is calculated from the number of time-space shared near neighbors between samples and the distance metric between samples: the distance metric between samples is derived from the exponential decay of the local covariance Mahalanobis distance; the local covariance Mahalanobis distance is constructed using the local covariance estimation combined with the time and space coordinates, and the covariance matrix adds a small positive regularity to ensure reversibility, and the covariance inverse matrix between samples is averaged to obtain the metric tensor; the decay bandwidth takes the median of the global distance distribution to adapt to the differences between dense and sparse regions; the time-space hybrid similarity is used to ensure that the samples with time-space proximity and close distance obtain higher weight, and enhance the distinguishability between class clusters. The time-space shared near-neighbor local density is composed of the similarity sum, the number of time-space shared near neighbors and the neighborhood average similarity: the similarity sum reflects the global similarity; the effective neighbor size reflects the local sample density; the neighborhood average similarity describes the local similarity, the time-space shared near-neighbor local density combines the time constraint and the spatial connectivity, describes the anisotropic structure through the local covariance Mahalanobis distance, and can stably highlight the high-density area, and improve the accuracy of class cluster center identification.
[0049] In a specific embodiment, the relative distance between samples can be obtained according to the local covariance Mahalanobis distance as follows: wherein, denotes the relative distance between samples, denotes the local covariance Mahalanobis distance of the sample and denotes the time-space shared near-neighbor local density of the sample, denotes the maximum value in the time-space shared near-neighbor local density corresponding to all samples.
[0050] S104: constructing a decision graph based on the time-space shared near-neighbor local density of the sample and the local covariance Mahalanobis distance between samples and selecting a class cluster center.
[0051] In a possible embodiment, constructing a decision graph based on the time-space shared near-neighbor local density of the sample and the local covariance Mahalanobis distance between samples and selecting a class cluster center comprises: calculating the relative distance of the sample based on the local covariance Mahalanobis distance between samples, and constructing a decision graph taking the time-space shared near-neighbor local density of the sample and the relative distance as the horizontal and vertical coordinates; calculating a decision value based on the time-space shared near-neighbor local density and the relative distance, and selecting a class cluster center according to the decision value of the sample.
[0052] Exemplarily, the two indexes of the spatiotemporal shared near-neighbor local density and the relative distance of the sample are taken as the horizontal and vertical coordinates to construct a decision graph. Since the local density and the relative distance of the density peak are more prominent, the sample located in the upper right area of the decision graph is more likely to be the cluster center. To reduce subjective interference, a decision value variable is further calculated to quantitatively determine the selection of the cluster center. The decision value calculation satisfies the following formula: , represents the decision value, represents the spatiotemporal shared near-neighbor local density, represents the relative distance, and the samples with the maximum decision value are selected as the cluster centers.
[0053] In a specific embodiment, the results of identifying the cluster centers by the method of the present application and the DPC algorithm commonly used in the prior art are compared. Referring to Figure 2a , a data set composed of four spatiotemporal clusters with different densities is used for identification comparison The distribution state in a three-dimensional space. In the cluster identification process, the DPC algorithm calculates the local density of the sample by truncating the kernel and Gaussian kernel function, but this algorithm has obvious limitations, that is, it is not sensitive enough to the sparse cluster with low density, and this defect directly leads to the problem that the sparse cluster is often not fully identified. From Figure 2b the identification result of the DPC algorithm can be seen, it successfully identifies two centers in the lower cluster with higher density, but fails to detect the sparse cluster on the upper side. The calculation method of the spatiotemporal shared near-neighbor local density proposed in the present application performs better in the identification task of the four clusters. As shown in Figure 2c , not only can all clusters be accurately identified, but also the appropriate cluster center can be selected in each cluster.
[0054] S105: Based on the spatiotemporal hybrid similarity matrix, the maximum similarity is allocated to obtain the preliminary classified clusters and edge samples. Based on the local covariance Mahalanobis distance between samples, the edge samples are allocated to the nearest high-density cluster. According to the local covariance Mahalanobis distance between the sample and the spatiotemporal neighbor sample, the local outlier factor of the sample is calculated; according to the local outlier factor, the abnormal sample is identified and removed from the cluster to obtain the final clustering result.
[0055] In a possible embodiment, the maximum similarity assignment based on the spatio-temporal hybrid similarity matrix obtains preliminary clustering and edge samples, including: taking the cluster center sample as an initial labeled node, traversing the spatio-temporal hybrid similarity matrix to find a maximum similarity path between the initial labeled node and an unlabeled node; assigning the unlabeled node to a cluster based on the similarity value between the nodes in the spatio-temporal hybrid similarity matrix; removing the assigned nodes in the spatio-temporal hybrid similarity matrix, expanding the assignment along the maximum similarity path until all the cross-cluster similarity values are zero to complete the maximum similarity assignment; and recording the remaining unclassified samples after the maximum similarity assignment as edge samples.
[0056] In a possible embodiment, the edge samples are assigned to the nearest high-density cluster based on the local covariance Mahalanobis distance between the samples, including: selecting a neighbor sample with the minimum distance according to the local covariance Mahalanobis distance between the neighbor sample and the corresponding edge sample; and assigning the cluster label of the neighbor sample to the edge sample to assign the edge sample to the nearest high-density cluster.
[0057] In a possible embodiment, the local outlier factor of a sample is calculated according to the local covariance Mahalanobis distance between the sample and the spatio-temporal neighbor sample of the sample, including: searching for the farthest distance between the to-be-identified sample and the spatio-temporal neighbor sample of the sample in the spatio-temporal neighbor sample of the sample; comparing the farthest distance with the local covariance Mahalanobis distance from each neighbor sample to the to-be-identified sample, and taking the larger value as the reachable distance from the neighbor sample to the to-be-identified sample; calculating the reachable density of the to-be-identified sample according to the reachable distance, and calculating the local outlier factor according to the reachable density.
[0058] In a specific embodiment, a two-step assignment strategy including a maximum similarity assignment stage and a nearest high-density assignment stage is designed to assign the samples to different clusters.
[0059] Specifically, in the maximum similarity assignment stage, the selected cluster center sample is taken as an initial labeled node, and the maximum similarity path between the initial labeled node and unlabeled nodes is found by traversing the spatio-temporal hybrid similarity matrix. At this time, the similarity values between each node and other nodes in the spatio-temporal hybrid similarity matrix reflect their mutual relationship, and the maximum similarity value represents the closest relationship. Based on these relationships, the unlabeled nodes are assigned to the corresponding clusters. After the assignment is completed, the corresponding column of the assigned nodes in the spatio-temporal hybrid similarity matrix is cleared to avoid repeated processing. With the advancement of the iteration process, the maximum similarity path calculated by the spatio-temporal shared neighbor local density and the local covariance Mahalanobis distance is expanded until all cross-cluster similarity values are zero. The remaining unclassified samples after the maximum similarity assignment are recorded as edge samples. Through the expansion of the maximum similarity path, the area with strong similarity can be quickly covered, thereby efficiently completing the preliminary cluster division.
[0060] The density around the edge samples is usually low and difficult to be directly classified. In the recent high-density assignment stage, the local covariance Mahalanobis distance between all neighbors of the edge sample and the corresponding edge sample is calculated. The neighbor sample with the smallest distance is selected as a reference, and the cluster label of the neighbor sample is assigned to the current edge sample. The local covariance Mahalanobis distance considers the local structure in the spatio-temporal dimension of the sample and can more accurately reflect the distance relationship between the samples. In this way, the edge sample is classified into the nearest high-density area, improving the assignment accuracy and ensuring more accurate cluster boundaries.
[0061] In a possible embodiment, all neighbors of the edge sample are the spatio-temporal neighbors of the edge sample in the neighbor set of the edge sample. All samples in the neighbor set; step S102 calculates the local covariance Mahalanobis distance between each sample. The local covariance Mahalanobis distance between the edge sample and its neighbor sample can be obtained according to the calculation result of step S102.
[0062] In a specific embodiment, after the assignment of all samples is completed, an outlier identification is performed to identify abnormal samples that are incorrectly assigned, thereby improving the reliability of the clustering result. The outlier identification strategy designed by the present application includes: calculating the reachable distance of the sample through the local covariance Mahalanobis distance, and calculating the reachable density of the sample based on the reachable distance. According to the density difference between the sample and its neighbor samples, a local outlier factor is obtained. Based on whether the local outlier factor exceeds a set abnormal threshold, the outlier is determined.
[0063] The calculation of the reachable distance satisfies the following formula: wherein, the reachable distance of the sample and ; and the reachable distance of the sample Its farthest The Mahalanobis distance of the local covariance of the nearest neighbors can quantify the samples at a local scale. The density relative to the surrounding environment; Indicates sample and The local covariance Mahalanobis distance.
[0064] The reachability density can be calculated using the following formula: ,in, This represents the reachability density, which is calculated by averaging the reachable distances and then taking the reciprocal. The larger the reachable distance, the smaller the reachability density. Indicates sample spacetime The total number of samples in the nearest neighbor set.
[0065] The technique for local outlier factors satisfies the following formula: ,in, Indicates sample Neighbor density, a measure of neighbor samples The degree of congestion in the local area; Indicates sample Its own density; Indicates sample The local outlier factor is essentially the expected density ratio. If the average density of the nearest neighbor samples is significantly higher than the density of the sample itself, the sample is considered an outlier.
[0066] In one possible embodiment, identifying anomalous samples based on local outliers includes: comparing the local outlier with a preset outlier threshold; when the local outlier is greater than the outlier threshold, the corresponding sample is determined to be an anomalous sample.
[0067] In a specific embodiment, outlier samples are identified based on local outlier factors, satisfying the following formula: ,in, This indicates the set abnormal threshold; samples exceeding this threshold are considered abnormal.
[0068] In the anomaly identification process, the distance farthest from the sample is found in its neighborhood, which is taken as the neighborhood scale, ensuring that all relatively distant samples are covered and capturing the global relationship between samples. Then the scale is compared with the local covariance Mahalanobis distance of each neighbor to the sample, and the larger value is taken as the reachable distance. The average of all reachable distances is then calculated to avoid misjudging edge samples as anomalies. The reciprocal of the reachable distance between samples is used to calculate the reachable density, which measures the tightness of the range. A lower reachable density indicates that the sample has a greater difference with other samples in the neighborhood and can be considered a potential anomaly. Finally, the sum of the reachable density ratios of the sample and its neighbors is calculated, and the average is the local outlier factor. The larger the local outlier factor, the sparser the neighborhood environment, and the more prominent the distribution difference with other samples in the neighborhood. When the local outlier factor exceeds the set threshold, the corresponding sample is determined as an anomaly point and removed from the cluster, thereby enhancing the cluster boundary and reducing the impact of false assignments.
[0069] To evaluate the clustering effect of the spatiotemporal density peak clustering method based on local covariance and spatiotemporal shared neighbors (LCSN-STDPC) of the present application, a classic spatiotemporal clustering algorithm and three latest spatiotemporal clustering algorithms based on DPC improvement (ST-ADPTC, STSNN-DPC, ST-CFSFDP, ST-DBSCAN) were selected for comparative experiments. The comparative experiments were carried out on four groups of artificially generated data sets. To further test the applicability and actual value of the spatiotemporal density peak clustering method based on local covariance and spatiotemporal shared neighbors on real spatiotemporal data, it was applied to the identification of foreshocks and aftershocks in earthquake event sequences.
[0070] Table 1 Key information of synthetic data sets
[0071]
[0072] Table 1 contains four groups of spatiotemporal data sets with different sizes and cluster structures - Key parameters. Contains four clusters of different sizes and linear abnormal samples connected to each other, and each cluster has significant differences in the time attribute. Composed of two internally sparse and morphologically different clusters. Composed of two density approximate "L" shaped clusters, which can be regarded as a combination of spatially extended clusters and temporally extended clusters; the left cluster structure is complete, and the right cluster is extremely sparse at the joint. Simultaneously presents four spatiotemporal clusters with different densities and morphologies. In practical applications, the domestic earthquake sequence data set covering the region of 73°-125° east longitude and 18°-48° north latitude was selected as the case study object, and its basic statistical characteristics are shown in Table 2.
[0073] Table 2 Basic characteristics of seismic data sets
[0074]
[0075] To evaluate the consistency between clustering results and true labels, three common external evaluation indicators are adopted: adjusted mutual information (AMI), adjusted Rand index (ARI), and FM index (FMI). The values of these three indicators are usually between 0 and 1, and the closer the value is to 1, the higher the matching degree between the clustering labels and the actual categories. When clustering seismic sequences, DB index (DBI) and CH index (CHI) are introduced as quantitative evaluation means of internal structure. DBI is used to measure the compactness within the class cluster and the separability between the class clusters. The smaller the value, the more concentrated the class cluster division and the clearer the boundary. CHI reflects the recognition degree of clustering structure by calculating the ratio of the variance between class clusters and the variance within class clusters. The larger the CHI value, the more explicit the clustering structure and the better the aggregation effect. To comprehensively analyze the clustering quality, abnormal samples are regarded as independent class clusters and included in the index calculation.
[0076] Table 3 Indicators of 5 algorithms on synthetic data sets
[0077]
[0078] According to the results in Table 3, the optimal values of AMI, ARI, and FMI indicators of 5 spatiotemporal clustering algorithms on 4 synthetic data sets are marked in bold. It can be seen that the LCSN-STDPC algorithm achieves the highest score on all data sets and all evaluation indicators, significantly outperforming other comparison algorithms in clustering accuracy and result consistency. This indicates that the algorithm can stably recover the true cluster structure under different synthetic scenarios and has robust and superior clustering performance. In contrast, the overall performance of the ST-DBSCAN algorithm ranks second, achieving suboptimal results on most data sets, while the overall performance of the ST-ADPTC, STSNN-DPC, and ST-CFSFDP algorithms is relatively weak, further highlighting the advantages of the LCSN-STDPC algorithm.
[0079] Based on the seismic dataset, the LCSN-STDPC algorithm is used to carry out the spatio-temporal clustering experiment, aiming to mine its potential spatio-temporal structure features, identify the main shock location and assist in inferring the foreshock and aftershock activity pattern related to it. Among them, the foreshock usually refers to the earthquake event that has occurred before the main shock, while the aftershock is a series of earthquake activities after the main shock. The spatio-temporal clustering analysis of such earthquake sequence helps to deeply understand the triggering mechanism of strong earthquake, so as to improve the practicability of earthquake prediction. Although the existing spatial clustering technology has been widely used in the identification of foreshock and aftershock of earthquake, these methods fail to fully integrate the time information of earthquake events. Therefore, the LCSN-STDPC algorithm is used to reveal the clustering pattern of earthquake in the comprehensive spatio-temporal dimension.
[0080] The experimental seismic data is from the National Earthquake Scientific Data Center, and the recording time range is from January 2008 to August 2009, covering the earthquake activities in China. The analysis focuses on the spatio-temporal clustering feature mining in the short period, in order to explore the structure and sequence relationship of the earthquake swarm. Figure 3 The spatio-temporal distribution of 1329 earthquakes with a magnitude of 4.0 or above in this period is shown.
[0081] The clustering effect of five algorithms on the earthquake dataset is evaluated by DBI and CHI indexes. Table 4 shows the evaluation results of each algorithm, and the optimal value is marked in bold. The results show that the LCSN-STDPC algorithm performs best in both indexes, which can not only maintain the high consistency within the cluster, but also has good separability between clusters, which helps to clearly reveal the sequence characteristics of foreshock, main shock and aftershock. In contrast, although the ST-DBSCAN algorithm performs better in DBI, indicating that the clustering structure is more compact, it is not as good as ST-ADPTC and STSNN-DPC algorithms in CHI value, reflecting that there is still room for improvement in its cluster separation ability. STSNN-DPC algorithm ranks second in CHI index, showing that it has certain cluster separation ability, but its DBI value is the highest among all methods, indicating that the internal structure of the cluster is relatively loose, and the cohesion between samples is insufficient, which affects the clustering quality.
[0082] Table 4 Indexes of five algorithms on the earthquake dataset
[0083]
[0084] When the LCSN-STDPC algorithm is used to analyze the spatio-temporal clustering of earthquake data, five typical earthquake clusters are successfully identified Figure 4 The statistical information of each cluster is shown in Table 5. Specifically, the cluster The main shock event of the cluster occurred on June 10, 2008, with the epicenter located in Golmud, Qinghai. There were only two foreshocks, and most of the events were aftershocks, showing a clear main-aftershock structure. The seismic event corresponding to the density peak of the cluster occurred on October 5, 2008, with the epicenter located in Akto, Xinjiang. The spatial and temporal distribution characteristics showed a high concentration of aftershocks. Cluster The main shock of the cluster occurred on August 25, 2008, in Zhongba, Tibet, and was followed by 61 aftershocks. The seismic events were highly superimposed in space on the main shock location, showing strong local concentration. Cluster The main shock of the cluster occurred on March 20, 2008, with the epicenter located in Cele, Xinjiang. In addition to the 16 aftershocks recorded locally, 28 and 36 aftershocks were recorded in Yutian County and Ritusi County, respectively, showing a clear cross-regional distribution characteristic. Cluster - Although there is a certain continuity in the time dimension, the spatial distribution is relatively dispersed, and there is a lack of significant spatial clustering characteristics, making it difficult to effectively divide through pure spatial clustering methods. The largest cluster The cluster covers the aftershock sequence after the 8.0 magnitude Wenchuan earthquake on May 12, 2008, with a total of 790 events, and the aftershocks continued until August 2009, showing a long-term decay trend. The cluster is large in size and long in duration, fully reflecting the profound impact of strong earthquake events and providing important data support for in-depth study of the mechanism of strong earthquakes.
[0085] Table 5 Statistical information of clustering results of seismic data sets
[0086]
[0087] The spatiotemporal density peak clustering method based on local covariance and spatiotemporal shared neighbors provided by the application calculates the covariance matrix in the local spatiotemporal neighborhood, defines the local covariance Mahalanobis distance between samples based on this, automatically unifies the scale difference between time and space to improve the accuracy of similarity measurement, optimizes the calculation method of local density using the number of spatiotemporal shared neighbors and the local covariance Mahalanobis distance to enhance the reliability of cluster center identification, adopts a two-step assignment and anomaly identification strategy to gradually determine and improve the attribution of samples through the assignment of the two stages of maximum similarity and nearest high density, and then uses the local outlier factor to effectively identify abnormal samples to reduce the error propagation of samples. Experimental results show that the spatiotemporal density peak clustering method based on local covariance and spatiotemporal shared neighbors performs well in clustering on multiple data sets and has good prospects in the application of seismic data sets.
[0088] The spatiotemporal density peak clustering method based on local covariance and spatiotemporal shared neighbor of the application solves a series of technical problems. The introduction of local covariance Mahalanobis distance involves a series of technical obstacles that traditional methods cannot overcome in the context of spatiotemporal data: Mahalanobis distance usually relies on stable overall covariance structure to ensure the reliability of scale adjustment and distance measurement, but spatiotemporal data presents significant local heterogeneity at different locations and different time periods, and its local distribution often does not have uniform covariance characteristics, making it difficult for Mahalanobis distance directly applied to global or fixed covariance structure to accurately reflect the true spatiotemporal relationship. In addition, the estimation of local covariance is limited by the limited amount of neighborhood data, inconsistent time span, and large spatial density difference, which can easily form a degenerate matrix or lead to an extremely uneven distribution of eigenvalues, making the distance calculation numerically unstable. More importantly, when a local covariance matrix is independently constructed for each sample, the metric space changes dynamically with the sample location, and the distance between different samples no longer has a uniform measurement basis that can be directly compared, which destroys the consistency requirement of global ordering of density peak clustering, leading to conflicts in the value of δ and the density ordering. The method of the application systematically designs the stability, symmetrization strategy, scale balance and metric consistency of local covariance to eliminate the structural contradictions brought by spatiotemporal heterogeneity, otherwise the core discrimination mechanism of the algorithm will not hold.
[0089] The local density calculation of the sample has the fundamental problem that traditional density estimation methods cannot simultaneously depict the spatiotemporal structure: time and space are inconsistent in terms of change speed, scale difference and distribution structure; if time is simply added to the calculation as an additional dimension, the neighborhood structure will often be distorted: the spatial relationship in areas with a large time span is weakened, and the local spatial noise in time-intensive areas may be overemphasized, making it difficult to correctly identify the cluster center through density characteristics. And the shared neighbor relationship does not have stable symmetry - two samples may share neighbors in time, but not necessarily have equal connectivity in space, leading to irregular jumps in the shared neighbor matrix, making it difficult for density estimation to maintain continuity and orderability. The shared neighbor number, local covariance Mahalanobis distance and neighborhood average similarity have different dimensions and large variation amplitudes, which cannot be simply added for density calculation, and simple addition calculation can easily cause density ordering distortion, making the key step of center point identification chaotic. The method of the application constructs a similarity matrix based on local covariance Mahalanobis distance, a neighborhood size correction mechanism and an average similarity adjustment strategy, so that the local density can be balanced between time synchronization and spatial connectivity, thereby solving the fundamental problem that traditional density estimation methods cannot simultaneously depict the spatiotemporal structure.
[0090] The similarity matrix based on the local covariance Mahalanobis distance is constructed, the neighborhood size correction mechanism and the average similarity adjustment strategy are used, so that the local density can be balanced between time synchronization and space connectivity,
[0091] Referring to the drawings Figure 5 The embodiment also provides a spatiotemporal density peak clustering device based on local covariance and spatiotemporal shared neighbors, which is used to implement the method embodiment. The device comprises:
[0092] A covariance matrix calculation unit 201 is configured to define the spatiotemporal neighbors of samples, and calculate the covariance matrix in the spatiotemporal neighbor range of the samples.
[0093] A Mahalanobis distance calculation unit 202 is configured to calculate the local covariance Mahalanobis distance between the samples.
[0094] A local density calculation unit 203 is configured to measure the spatiotemporal shared neighbors between the samples, construct the spatiotemporal hybrid similarity matrix between the samples, and calculate the spatiotemporal shared neighbor local density of the samples based on the spatiotemporal hybrid similarity between the samples.
[0095] A cluster center selection unit 204 is configured to construct a decision graph and select a cluster center based on the spatiotemporal shared neighbor local density of the samples and the local covariance Mahalanobis distance between the samples.
[0096] A cluster division unit 205 is configured to perform maximum similarity allocation based on the spatiotemporal hybrid similarity matrix to obtain preliminary divided clusters and edge samples, allocate the edge samples to the nearest high-density cluster based on the local covariance Mahalanobis distance between the samples, calculate the local outlier factor of the sample based on the local covariance Mahalanobis distance between the sample and the spatiotemporal neighbor sample of the sample, identify abnormal samples based on the local outlier factor, and remove the abnormal samples from the clusters to obtain a final clustering result.
[0097] All related contents of the steps involved in the above method embodiment can be cited to the function description of the corresponding function module, which will not be repeated here.
[0098] In some other embodiments of the present application, the embodiments of the present application disclose an electronic device, as shown in the figure, which can include: one or more processors 301; a memory 302; a display 303; one or more application programs (not shown); and one or more computer programs 304, wherein the above devices can be connected through one or more communication buses 305. The one or more computer programs 304 are stored in the above memory and are configured to be executed by the one or more processors 301, the one or more computer programs 304 include instructions which can be used to perform the steps of the above method embodiments. Figure 6 Figure 1 and respective embodiments.
[0099] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0100] The functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0101] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application essentially or in other words the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: a flash memory, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, and various media that can store program codes.
[0102] The above is only a specific implementation of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any change or replacement within the technical scope disclosed in the embodiments of the present application should be covered in the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be subject to the protection scope of the claims.
Claims
1. A spatio-temporal density peak clustering method based on local covariance and spatio-temporal shared neighbors, characterized in that, The method comprises the following steps: Defining a spatiotemporal sample Nearest neighbors, in the spatiotemporal sample Computing the covariance matrix over the neighborhood range; calculating the local covariance Mahalanobis distance between samples; measuring the spatiotemporal shared neighbors between samples, constructing a spatiotemporal hybrid similarity matrix between samples, and calculating the spatiotemporal shared neighbor local density of samples based on the spatiotemporal hybrid similarity between samples; constructing a decision graph based on the spatiotemporal shared neighbor local density of samples and the local covariance Mahalanobis distance between samples, and selecting class cluster centers; The maximum similarity is allocated based on the spatio-temporal hybrid similarity matrix to obtain preliminary clustering and edge samples, the edge samples are allocated to the nearest high-density cluster based on the local covariance Mahalanobis distance between the samples, the local outlier factor of the sample is calculated according to the local covariance Mahalanobis distance of the nearest neighbor sample, the abnormal sample is identified according to the local outlier factor, and the abnormal sample is removed from the cluster to obtain a final clustering result. The local outlier factor of the sample is calculated according to the local covariance Mahalanobis distance of the nearest neighbor sample, the abnormal sample is identified according to the local outlier factor, and the abnormal sample is removed from the cluster to obtain a final clustering result.
2. The method of claim 1, wherein, Defining a spatiotemporal sample Nearest neighbors, in the spatiotemporal sample Computing the covariance matrix over the neighborhood, including: Definition and sample The relationship between the definition and the sample satisfies the following formula The number of samples is the sample The space-time of The neighborhood set: , Indicates the time Euclidean distance between the sample And Indicates the time cut-off distance, Indicates the number of neighbor samples; The covariance matrix is calculated according to the following formula: Indicates sample The local covariance matrix, Indicates sample spacetime The total number of samples in the nearest neighbor set, Indicates sample The three-dimensional spacetime vector, Indicates sample The local mean vector, This represents the ridge regression regularization coefficient. It is a 3x3 identity matrix. This represents the transpose symbol.
3. The method of claim 2, wherein, the local covariance Mahalanobis distance between samples satisfies the following formula: wherein, denotes the local covariance Mahalanobis distance between samples, denotes the symmetric average metric matrix, denotes the inverse of the local covariance matrix of the samples .
4. The method of claim 1, wherein, the spatiotemporal shared neighbor local density of samples satisfies the following formula: wherein, denotes the spatio-temporal shared nearest neighbors local density of a sample denotes the total number of samples in the spatio-temporal shared nearest neighbors set of a sample and denotes the spatio-temporal shared nearest neighbors number of a sample denotes the spatio-temporal hybrid similarity of a sample and denotes the spatio-temporal hybrid similarity of a sample denotes the spatio-temporal nearest neighbors set of a sample denotes the total number of samples in the spatio-temporal nearest neighbors set of a sample 5. The method of claim 1, wherein, constructing a decision graph based on the spatiotemporal shared neighbor local density of samples and the local covariance Mahalanobis distance between samples, and selecting class cluster centers, comprising: calculating the relative distance of samples based on the local covariance Mahalanobis distance between samples, and constructing a decision graph with the spatiotemporal shared neighbor local density of samples and the relative distance as the horizontal and vertical coordinates; calculating the decision value based on the spatiotemporal shared neighbor local density and the relative distance, and selecting the class cluster center according to the decision value of the sample.
6. The method of claim 1, wherein, The maximum similarity assignment based on the spatiotemporal hybrid similarity matrix obtains the preliminary classification clusters and edge samples, comprising: taking the class cluster center sample as an initial labeled node, traversing the spatiotemporal hybrid similarity matrix to find the maximum similarity path between the initial labeled node and the unlabeled node; assigning the unlabeled node to the class cluster based on the similarity value between the nodes in the spatiotemporal hybrid similarity matrix; clearing the assigned nodes in the spatiotemporal hybrid similarity matrix, expanding the assignment along the maximum similarity path until all the cross-class cluster similarity values are zero to complete the maximum similarity assignment, and recording the remaining unclassified samples after the maximum similarity assignment as edge samples.
7. The method of claim 6, wherein, Assigning the edge samples to the nearest high-density cluster based on the local covariance Mahalanobis distance between samples, comprising: selecting the neighbor sample with the smallest distance according to the local covariance Mahalanobis distance between the neighbor sample of the edge sample and the corresponding edge sample; assigning the class label of the neighbor sample to the edge sample to assign the edge sample to the nearest high-density cluster.
8. The method of claim 1, wherein, According to the sample and the spatio-temporal The local covariance Mahalanobis distance of the neighboring samples is calculated to obtain a local outlier factor of the sample, comprising: In the spatio-temporal In the spatio-temporal The farthest distance between the nearest neighbor samples Comparing the farthest distance with the local covariance Mahalanobis distance from each neighbor sample to the to-be-identified sample, and taking the larger value as the reachable distance of the neighbor sample to the to-be-identified sample; calculating the reachable density of the to-be-identified sample according to the reachable distance, and calculating the local outlier factor according to the reachable density.
9. The method of claim 1, wherein, Identifying abnormal samples according to the local outlier factor, comprising: comparing the local outlier factor with a preset abnormal threshold, and determining the corresponding sample as an abnormal sample when the local outlier factor is greater than the abnormal threshold.
10. A spatio-temporal density peak clustering apparatus based on local covariance and spatio-temporal shared nearest neighbors, characterized in that, The device comprises: a covariance matrix computation unit for defining a spatio-temporal neighborhood of the sample a covariance matrix is computed within the neighborhood a covariance matrix is computed within the neighborhood a Mahalanobis distance calculation unit for calculating the local covariance Mahalanobis distance between samples; a local density calculation unit for measuring the spatiotemporal shared neighbors between samples, constructing a spatiotemporal hybrid similarity matrix between samples, and calculating the spatiotemporal shared neighbor local density of samples based on the spatiotemporal hybrid similarity between samples; a class cluster center selection unit for constructing a decision graph based on the spatiotemporal shared neighbor local density of samples and the local covariance Mahalanobis distance between samples, and selecting class cluster centers; The class cluster division unit is configured to perform maximum similarity allocation based on the spatio-temporal hybrid similarity matrix to obtain preliminary division class clusters and edge samples, allocate the edge samples to the nearest high-density clusters based on local covariance Mahalanobis distances between the samples, and identify abnormal samples according to local outlier factors of the samples, and remove the abnormal samples from the class clusters to obtain a final clustering result. The local covariance Mahalanobis distance of the nearest neighbor samples is used to calculate a local outlier factor of the sample, and the abnormal samples are identified according to the local outlier factor, and the abnormal samples are removed from the class clusters to obtain a final clustering result.
Citation Information
Patent Citations
Seismic event time and space gathering mode extraction method based on shared density
CN103869367A
Multi-stage spatial-temporal clustering method and system based on fused mahalanobis distance
CN121051489A
Space-time density peak value clustering method and system based on shared neighbor weighting
CN121117654A
Method and arrangements for determining information regarding an intensity peak position in a space-time volume of image frames
US20230186510A1