Spatiotemporal Density Peak Clustering Method, Device, Medium and Electronic Device for Earthquake Events

By constructing space-time and space-time neighbor sets and multi-step allocation strategies, the problem of taking into account both the temporal and spatial attributes in earthquake event analysis is solved, and high accuracy and consistency clustering is achieved, which can effectively identify the space-time mode of earthquake events and support strong earthquake prediction.

CN120123802BActive Publication Date: 2025-07-08NANCHANG INST OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510610307.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-08
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The existing technology fails to effectively take into account both temporal and spatial attributes in seismic event analysis, resulting in insufficient cluster accuracy and consistency, especially when dealing with the merging of clusters of class and cluster centers with large density differences, there is a problem of mislabeling anomalies.

Method used

By constructing a space-time neighbor set, the spatial weight coefficient and time weight coefficient of the seismic event sample are calculated, the potential cluster center is determined based on local density and relative distance, and the comprehensive similarity merging rules and multi-step allocation strategy are adopted to identify abnormal points and improve clustering accuracy and consistency.

Benefits of technology

The clustering accuracy and consistency of earthquake event samples are improved, and the foreshocks, main shocks and aftershocks can be effectively identified, and the space-time evolution patterns of earthquake sequences can be quantified, providing clues for strong earthquake prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123802B_ABST
    Figure CN120123802B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, medium and electronic device for clustering the peak of spatio-temporal density of seismic events, including: determining the spatio-temporal neighbor set of samples in the spatio-temporal dataset; calculating the spatial weight coefficient and the time weight coefficient of the samples, and calculating the local density of each sample under spatio-temporal constraints; calculating the relative distance of the samples; calculating the decision value of each sample to select the potential cluster centers of the spatio-temporal dataset; eliminating the potential cluster centers with the potential of merging according to the comprehensive similarity of the cluster centers to obtain the cluster centers of various clusters; allocating the remaining samples after determining the cluster centers; identifying the outliers in various clusters to obtain the clustering result of the final sample dataset. By applying this method, time and space constraints are simultaneously introduced into the cluster center merging rule to obtain a cluster center with higher accuracy and consistency; a more flexible multi-step allocation strategy is adopted to replace the single allocation mode, and combined with the outlier identification mechanism based on dual criteria, the clustering effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of machine learning and data mining, and particularly to a spatio-temporal density peak clustering method, device, medium and electronic device for seismic events. Background Art

[0002] Clustering algorithms are an important unsupervised learning technique in machine learning and data mining, suitable for grouping unlabeled samples. Its main purpose is to classify samples into several clusters with specific distribution patterns by analyzing the similarity between samples. This method ensures that samples in the same cluster have a high degree of correlation, while reducing the similarity between samples in different clusters. Clustering algorithms can automatically discover the potential structures and patterns in the dataset without the need to provide labels in advance, and are widely used and effective in research fields such as information retrieval and pattern recognition.

[0003] With the continuous increase of spatio-temporal sequence data, spatio-temporal data has gradually become a research hotspot in many fields and plays an important role in fields such as environmental monitoring, geographic information systems, and public health. Spatio-temporal data not only contains spatial coordinate information but also involves time tags, making the data exhibit complex spatio-temporal dynamic characteristics. To mine potential patterns and rules from spatio-temporal data, clustering analysis stands out as an effective method and can provide strong support for decision-making in related fields. However, the clustering task of spatio-temporal data is more challenging than that of traditional static data clustering because it not only needs to consider the spatial distribution but also must analyze the evolution and trends in the time series. The spatio-temporal dependence, heterogeneity, and high-dimensionality of spatio-temporal data require clustering algorithms to have high flexibility and adaptability to cope with the complex interaction between space and time.

[0004] In the prior art, spatial clustering algorithms have been applied to the analysis of seismic events and the identification of foreshocks and aftershocks, but the spatial clustering algorithms ignore the time attribute of earthquakes. For the analysis of seismic events, the spatial attribute and time attribute factors are equally important, and a method that can take into account both the time attribute and the spatial attribute is needed for the analysis of seismic events.

[0005] Based on the need for clustering analysis of spatio-temporal data, a variety of methods have been proposed in the prior art for clustering analysis of spatio-temporal data, but there are disadvantages in that the influence of time factors on the clustering process cannot be fully considered, and there are the following defects in the face of spatio-temporal datasets: 1) The definition of local density is not applicable to clusters with large density differences; 2) The merging process of cluster centers does not consider the influence of time factors; 3) The single-step allocation strategy based on the time axis is prone to mislabeling normal samples as outliers when dealing with clusters with a large spatial span between clusters.

[0006] Therefore, it is necessary to provide a clustering method for spatio-temporal data of seismic events with high accuracy and consistency. Summary of the Invention

[0007] The object of the present invention is to provide a method, device, medium and electronic device for clustering the peak values of the spatio-temporal density of seismic events, so as to improve the accuracy and consistency of the clustering method when processing spatio-temporal data.

[0008] In the first aspect, the method for clustering the peak values of the spatio-temporal density of seismic events provided by the present invention includes: obtaining sample data, where the sample data includes a spatio-temporal data set of seismic events and a preset distance truncation threshold, normalizing the spatio-temporal data set of seismic events, constructing a spatial Euclidean distance matrix and a time distance matrix between seismic event samples, and determining a spatio-temporal neighbor set of seismic event samples in the spatio-temporal data set of seismic events; calculating a spatial weight coefficient and a time weight coefficient of each seismic event sample according to the spatio-temporal neighbor set of each seismic event sample in the spatio-temporal data set of seismic events, and calculating the local density of each seismic event sample under spatio-temporal constraints according to the spatial weight coefficient and the time weight coefficient; calculating the relative distance of each seismic event sample according to the time distance between seismic event samples and the local density of each seismic event sample under spatio-temporal constraints; calculating a decision value of each seismic event sample according to the local density and relative distance of the seismic event sample to select potential cluster centers of the spatio-temporal data set of seismic events; calculating the spatial similarity and time similarity of each cluster center in the potential cluster centers, determining the comprehensive similarity of the cluster centers according to the spatial similarity and the time similarity, and removing potential cluster centers with the potential for merging according to the comprehensive similarity of the cluster centers to obtain cluster centers of various clusters; allocating seismic event samples whose spatial distance from the cluster centers of various clusters within a time window is less than the preset distance truncation threshold to the corresponding clusters to achieve window allocation, and performing secondary allocation on the remaining seismic event samples after window allocation according to the spatial distance between the seismic event samples and the adjacent cluster centers to achieve search allocation, and dividing the unallocated seismic event samples after search allocation into the clusters to which the high-density nearest neighbor seismic event samples belong; calculating the core density, core distance and boundary density of each cluster for identifying outliers in each cluster, and obtaining the clustering result of the final sample data.

[0009] The beneficial effects of the method for clustering the peak values of the spatio-temporal density of seismic events provided by the present invention are as follows: when calculating the local density of seismic event samples, reliable potential cluster centers are selected by expanding the calculation range and enhancing the influence degree of the spatio-temporal neighborhood; both time and space constraints are introduced into the cluster center merging rule to obtain cluster centers with higher accuracy and consistency; a more flexible multi-step allocation strategy is adopted to replace the single allocation mode, and an outlier recognition mechanism based on double criteria is combined to improve the clustering effect of the algorithm.

[0010] In a possible embodiment, it is defined that the spatial weight coefficient and the time weight coefficient respectively satisfy the following formulas: , , where, represents the spatial weight coefficient, represents the earthquake event sample of the spatio-temporal nearest neighbor set, represents the total spatial distance from all earthquake event samples in the spatio-temporal nearest neighbor set to ; represents the earthquake event samples in the spatio-temporal nearest neighbor set except ; represents the remaining earthquake event samples in the spatio-temporal nearest neighbor set except to the total spatial distance of ; represents the time weight coefficient, represents the total time distance from all earthquake event samples in the spatio-temporal nearest neighbor set to ; represents the remaining earthquake event samples in the spatio-temporal nearest neighbor set except to the total time distance of ;

[0011] The local density of earthquake event samples under spatio-temporal constraints satisfies the following formula: , where represents the local density of earthquake event sample , represents the preset distance truncation threshold, represents the spatial distance from earthquake event sample to earthquake event sample , represents the preset time truncation threshold, represents the time distance from earthquake event sample to earthquake event sample .

[0012] In another possible embodiment, the calculation of the relative distance satisfies the following formula: , where represents the relative distance of earthquake event sample , represents the relative distance of earthquake event sample , represents the local density of earthquake event sample , represents the time distance from earthquake event sample to earthquake event sample , represents the scale of earthquake event samples in the data set.

[0013] In other possible embodiments, the decision value of each seismic event sample is calculated according to the local density and relative distance of the seismic event samples for selecting potential cluster centers of the spatio-temporal dataset of seismic events, including: sorting the decision values of the seismic event samples in the spatio-temporal dataset of seismic events in descending order, and selecting several positive integer seismic event samples with the largest decision values as potential cluster centers, and setting this positive integer as z, thereby obtaining a set of potential cluster centers 。

[0014] The calculation of the spatial similarity satisfies the following formula: , where represents the spatial similarity between the cluster center and the cluster center ; represents the spatial distance between the cluster center in the potential cluster centers and the cluster center ; represents the cluster center with a decision value higher than the cluster center ; represents a preset distance truncation threshold; represents the decision value of the cluster center ; represents the decision value of the cluster center ; The calculation of the temporal similarity satisfies the following formula: , where represents the temporal similarity between the cluster center and the cluster center ; represents the temporal distance between the cluster center in the potential cluster centers and the cluster center ; represents a preset temporal truncation threshold; The comprehensive similarity satisfies the following formula: , where represents the comprehensive similarity of the cluster center; Eliminating potential cluster centers with the potential to merge according to the comprehensive similarity of the cluster centers to obtain the cluster centers of various clusters, including: when the comprehensive similarity of the cluster center is 1, removing the cluster center from the potential cluster centers.

[0015] The remaining earthquake event samples after window allocation are secondary allocated according to the spatial distance between the earthquake event samples and the adjacent cluster centers to realize search allocation; the earthquake event samples that are not allocated after search allocation are divided into the cluster to which the high-density nearest neighbor earthquake event samples belong, including: comparing the spatial distances between the remaining earthquake event samples after window allocation and the two adjacent cluster centers, and allocating the earthquake event samples to the cluster to which the cluster center with a small spatial distance belongs to realize search allocation; the earthquake event samples that are not allocated after search allocation are divided into the cluster to which the earthquake event samples that have a local density higher than that of the earthquake event samples and are closest in time to the earthquake event samples belong.

[0016] The core density, core distance and boundary density of each cluster are calculated for the identification of outliers in each cluster, including: earthquake event samples whose local density is less than the lower value of the boundary density and core density of the cluster to which they belong, and whose relative distance is less than the core distance of the cluster to which they belong, are determined as outliers.

[0017] In a second aspect, the present invention further provides a device for clustering the spatiotemporal density peak values ​​of earthquake events, comprising:

[0018] A data processing unit is used to obtain earthquake event sample data, wherein the earthquake event sample data includes an earthquake event spatiotemporal data set and a preset distance cutoff threshold, normalize the earthquake event spatiotemporal data set, construct a spatial Euclidean distance matrix and a temporal distance matrix between earthquake event samples, and determine a spatiotemporal neighbor set of earthquake event samples in the earthquake event spatiotemporal data set;

[0019] A local density calculation unit is used to calculate the spatial weight coefficient and the time weight coefficient of the earthquake event sample according to the spatial and temporal neighbor set of each earthquake event sample in the earthquake event spatial and temporal data set, and calculate the local density of each earthquake event sample under the spatial and temporal constraints according to the spatial weight coefficient and the time weight coefficient;

[0020] A relative distance calculation unit, used for calculating the relative distance of the seismic event samples according to the time distance between the seismic event samples and the local density of each seismic event sample under the time and space constraints;

[0021] A potential cluster center selection unit is used to calculate the decision value of each earthquake event sample according to the local density and relative distance of the earthquake event sample to select the potential cluster center of the earthquake event spatiotemporal data set;

[0022] A cluster center determination unit is used to calculate the spatial similarity and temporal similarity of the cluster centers among the potential cluster centers, determine the comprehensive similarity of the cluster centers according to the spatial similarity and the temporal similarity, and eliminate the potential cluster centers with merging potential according to the comprehensive similarity of the cluster centers to obtain the cluster centers of various clusters;

[0023] An earthquake event sample allocation unit, which is used to allocate earthquake event samples within a time window whose spatial distance from the cluster centers of various clusters is less than a preset distance truncation threshold to the corresponding clusters to achieve window allocation, and perform secondary allocation on the remaining earthquake event samples after window allocation according to the spatial distance between the earthquake event samples and the adjacent cluster centers to achieve search allocation, and divide the unallocated earthquake event samples after search allocation into the cluster to which the high-density nearest neighbor earthquake event samples belong;

[0024] An outlier recognition unit, which is used to calculate the core density, core distance and boundary density of various clusters for outlier recognition in various clusters, and obtain the clustering result of the final earthquake event sample data set.

[0025] In a third aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned earthquake event spatio-temporal density peak clustering method is implemented.

[0026] In a fourth aspect, the present invention also provides an electronic device, including: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory, so that the electronic device executes the above-mentioned earthquake event spatio-temporal density peak clustering method.

[0027] For the beneficial effects of the above second to fourth aspects, reference can be made to the description of the above first aspect. Description of the Drawings

[0028] Figure 1 It is a schematic flowchart of a method for clustering earthquake event spatio-temporal density peaks provided by an embodiment of the present invention;

[0029] Figure 2 It is a schematic diagram of the window allocation process of earthquake event samples provided by an embodiment of the present invention;

[0030] Figure 3 It is a basic information table of simulated data provided by an embodiment of the present invention;

[0031] Figure 4 It is a basic information table of real data provided by an embodiment of the present invention;

[0032] Figure 5 It is a clustering result table of five algorithms on a simulated data set provided by an embodiment of the present invention;

[0033] Figure 6 It is a Friedman value table of evaluation indexes of five algorithms on a simulated data set provided by an embodiment of the present invention;

[0034] Figure 7 It is a schematic diagram of a spatio-temporal density peak clustering device provided by an embodiment of the present invention;

[0035] Figure 8 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0036] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings as understood by those of ordinary skill in the art in the field to which the present invention belongs. The words such as "including" used herein mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items.

[0037] This embodiment provides a method for clustering the peak of spatio-temporal density of earthquake events. Refer to the appended Figure 1 description. The method includes:

[0038] S101: Obtain sample data, where the sample data includes a spatio-temporal data set of earthquake events and a preset distance truncation threshold, normalize the spatio-temporal data set of earthquake events, construct a spatial Euclidean distance matrix and a time distance matrix between earthquake event samples, and determine the spatio-temporal near-neighbor set of earthquake event samples in the spatio-temporal data set of earthquake events.

[0039] In a specific embodiment, in the spatio-temporal data set of earthquake events , the spatio-temporal near-neighbor set of the earthquake event sample can be expressed as: , where represents the earthquake event sample in the spatio-temporal data set of earthquake events except , represents the spatial distance from the earthquake event sample to the earthquake event sample , represents the time distance from the earthquake event sample to the earthquake event sample .

[0040] represents the earthquake event sample, and its spatio-temporal near-neighbor set is composed of those satisfying that its spatial distance is less than or equal to the spatial truncation distance , and the time distance is less than or equal to the time truncation distance Composed of earthquake event samples.

[0041] In a specific embodiment, the spatio-temporal dataset of earthquake events is a set of earthquake event samples including time attributes and spatial attributes, and the earthquake event sample information specifically includes information such as the occurrence time, occurrence location, and magnitude of the earthquake event.

[0042] S102: Calculate the spatial weight coefficient and time weight coefficient of each earthquake event sample in the spatio-temporal dataset of earthquake events, and calculate the local density of each earthquake event sample under spatio-temporal constraints according to the spatial weight coefficient and time weight coefficient.

[0043] In a possible embodiment, it is defined that the spatial weight coefficient and time weight coefficient respectively satisfy the following formulas: , , where represents the spatial weight coefficient, represents the earthquake event sample of the spatio-temporal neighbor set, represents the total spatial distance of all earthquake event samples in the spatio-temporal neighbor set to , represents the earthquake event sample except in the spatio-temporal neighbor set, represents the total spatial distance of the remaining earthquake event samples except in the spatio-temporal neighbor set to , represents the time weight coefficient, represents the total time distance of all earthquake event samples in the spatio-temporal neighbor set to , represents the remaining earthquake event samples except in the spatio-temporal neighbor set to the time distance.

[0044] The local density of the earthquake event sample under spatio-temporal constraints satisfies the following formula: , where represents the local density of the earthquake event sample , represents the preset distance truncation threshold, represents the earthquake event sample to the earthquake event sample the spatial distance, represents the preset time truncation threshold, represents the earthquake event sample to the earthquake event sample the time distance.

[0045] The local density weighted by spatio-temporal neighbors aims to quantify the contribution degree of spatio-temporal neighboring earthquake event samples to and evaluate the contribution of earthquake event samples outside the domain to . It enhances the influence of earthquake event samples closely related to and through the weight coefficients , expands the local density gap between different earthquake event samples; takes the sum of the reciprocals of the spatial distance and the time distance as a measurement index, emphasizing the interaction between

[0046] S103: Calculate the relative distance of earthquake event samples according to the time distance between earthquake event samples and the local density of each earthquake event sample under spatio-temporal constraints.

[0047] In a possible embodiment, the calculation of the relative distance satisfies the following formula: , where represents the relative distance of earthquake event sample , represents the relative distance of earthquake event sample , represents the local density of earthquake event sample , represents the time distance from earthquake event sample to earthquake event sample , represents the scale of earthquake event samples in the data set. That is, the relative distance of the non-maximum local density earthquake event sample is defined as the spatial distance between earthquake event sample and the nearest earthquake event sample with a higher local density; the earthquake event sample with the largest local density value has the global maximum relative distance.

[0048] S104: Calculate the decision value of each earthquake event sample according to the local density and relative distance of the earthquake event sample to select potential cluster centers of the earthquake event spatio-temporal data set.

[0049] In a possible embodiment, the calculation of the decision value of the earthquake event sample satisfies the following formula: .

[0050] In a specific embodiment, the decision value of each seismic event sample is calculated based on the local density and relative distance of the seismic event samples for selecting potential cluster centers of the seismic event spatio-temporal dataset, including: arranging the decision values of the seismic event samples in the seismic event spatio-temporal dataset in descending order, and selecting the z seismic event samples with the largest numerical decision values as potential cluster centers (potential class cluster centers of each type of cluster), where z is a positive integer, thereby obtaining a set of potential cluster centers .

[0051] S105: Calculate the spatial similarity and temporal similarity of each cluster center in the potential cluster centers, determine the comprehensive similarity of the cluster centers according to the spatial similarity and temporal similarity, and eliminate the potential cluster centers with the potential for merging according to the comprehensive similarity of the cluster centers to obtain the cluster centers of each type of cluster.

[0052] In a possible embodiment, the calculation of the spatial similarity satisfies the following formula: , where represents the spatial similarity between cluster center and cluster center , represents the spatial distance between cluster center and cluster center in the potential cluster centers, represents the cluster center with a decision value higher than cluster center , represents a preset distance truncation threshold, represents the decision value of cluster center , represents the decision value of cluster center .

[0053] The calculation of the temporal similarity satisfies the following formula: , where represents the temporal similarity between cluster center and cluster center , represents the temporal distance between cluster center and cluster center in the potential cluster centers, represents a preset temporal truncation threshold.

[0054] The comprehensive similarity satisfies the following formula: , where is the comprehensive similarity of the cluster center.

[0055] Eliminating the potential cluster centers with the potential for merging according to the comprehensive similarity of the cluster centers to obtain the cluster centers of each type of cluster, including: when the comprehensive similarity of cluster center is 1, cluster center Removed from potential cluster centers.

[0056] Used to evaluate cluster centers For their proximity to the nearest neighboring cluster centers in the spatial dimension. If a cluster center and a cluster center with a decision value higher than its own have a minimum spatial distance less than a preset threshold , then set the spatial similarity to 1, indicating that the cluster center can be regarded as the nearest neighbor of a cluster center with a higher decision value in space; otherwise, set it to 0, indicating that the cluster center is spatially far from other cluster centers.

[0057] Used to evaluate cluster centers For their proximity to the nearest neighboring cluster centers in the temporal dimension. If a cluster center and a cluster center with a decision value higher than its own have a minimum temporal distance less than a preset threshold , then set the temporal similarity to 1, indicating that the cluster center can be regarded as the nearest neighbor of a cluster center with a higher decision value in time; otherwise, set it to 0, indicating that the cluster center is temporally far from other cluster centers.

[0058] Only when both the spatial similarity and the temporal similarity are 1, the comprehensive similarity of the cluster center is 1. That is, when a cluster center is regarded as a neighbor of a cluster center with a higher decision value in both the spatial and temporal dimensions, the cluster center has the ability to merge with other cluster centers, and then it can be removed from the set of potential cluster centers. This method avoids the possibility of having multiple cluster centers in a single cluster, ensuring that the merged cluster centers can efficiently guide the allocation of the remaining cluster centers.

[0059] S106: Allocate the earthquake event samples within the time window whose spatial distance from the cluster centers of various clusters is less than the preset distance truncation threshold to the corresponding clusters to achieve window allocation. For the remaining earthquake event samples after window allocation, perform secondary allocation based on the spatial distance between the earthquake event samples and the adjacent cluster centers to achieve search allocation. Divide the unallocated earthquake event samples after search allocation into the clusters to which the high-density nearest neighbor earthquake event samples belong.

[0060] In a possible embodiment, the remaining seismic event samples after window allocation are secondarily allocated according to the spatial distance between the seismic event samples and the adjacent cluster centers to achieve search allocation, and the unallocated seismic event samples after search allocation are divided into the cluster to which the high-density nearest neighbor seismic event samples belong, including: comparing the spatial distance between the remaining seismic event samples after window allocation and the two adjacent cluster centers, and allocating the seismic event samples to the cluster to which the cluster center with a smaller spatial distance belongs to achieve search allocation; dividing the unallocated seismic event samples after search allocation into the cluster to which the seismic event samples with a local density higher than that of the seismic event samples and the closest time distance to the seismic event samples belong.

[0061] In a specific embodiment, the three steps of allocating the remaining seismic event samples after determining the cluster centers are specifically executed as follows:

[0062] The first allocation strategy: window allocation. For each cluster center on the time axis , calculate the seismic event samples within the window interval to the cluster center . If to 's spatial distance is less than the preset spatial truncation distance , then classify it into the cluster represented by .

[0063] See the attached drawings of the specification Figure 2 , the correctly allocated seismic event samples in front of the first cluster center and behind the last cluster center , whose spatial distance from or is less than , and the time distance is controlled within . That is, the seismic event samples within this range are in the near-neighbor area of the cluster center and have a high correlation with it, so they should be classified into the cluster to which the corresponding cluster center belongs. For the seismic event samples outside the window range and those with a relatively large spatial distance from the cluster center within the window, they are temporarily regarded as outliers in this step of allocation. As the starting stage of the entire allocation strategy, the range of window allocation is set small, aiming to give priority to processing the seismic event samples with the strongest correlation with each cluster center and providing several stable base clusters composed of a certain number of reliable seismic event samples for subsequent allocation.

[0064] The second allocation strategy: search allocation. Since the time truncation distance The setting value of is usually smaller than the time distance between adjacent cluster centers, so the search range of earthquake event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, the closer cluster is selected for allocation by comparing the spatial distance between the earthquake event sample and the two adjacent cluster centers, without involving the spatial cutoff distance. The comparison can avoid the influence of abnormal parameters on the allocation results and improve the stability of allocation.

[0065] The third allocation strategy: nearby allocation. For earthquake event samples that have not been successfully allocated in the first two allocation processes, each earthquake event sample is traversed in order from high to low local density. , find the earthquake event sample with a local density higher than the current earthquake event sample and the closest time distance to it , classify the current earthquake event sample into the earthquake event sample Belongs to the cluster.

[0066] The core basis of this processing method is that high-density earthquake event samples are usually cluster centers or their neighboring points, so they are more likely to be regarded as representative earthquake event samples of a certain cluster. The nearest allocation strategy searches for global reliable information to ensure that all remaining earthquake event samples are correctly allocated to their respective clusters to the greatest extent possible.

[0067] S107: Calculate the core density, core distance and boundary density of each cluster to identify outliers in each cluster, and obtain the final clustering result of the sample data.

[0068] In a possible embodiment, the core density, core distance and boundary density of each cluster are calculated for identifying outliers in each cluster, including: earthquake event samples whose local density is less than the lower value of the boundary density and core density of the cluster to which they belong, and whose relative distance is less than the core distance of the cluster to which they belong, are determined as outliers.

[0069] In a specific embodiment, the present invention defines a criterion for outlier identification: if the local density of a seismic event sample is Lower than the core density of the cluster to which it belongs , and the relative distance Exceeds the core distance of the cluster to which it belongs , it is considered as an outlier.

[0070] The calculation of core density satisfies the following formula: , the calculation of the core distance satisfies the following formula: .in, Represents the clustering result The core density of each cluster, Represents the clustering result The core distance of each cluster, represents the median of the local densities of earthquake event samples in the th cluster, represents the median of the relative distances of earthquake event samples in the th cluster, , represents the scale factor, represents the upper quartile of the local densities of all earthquake event samples in the th cluster (the value at the position after ascending order), represents the upper quartile of the relative distances of all earthquake event samples in the th cluster (the value at the position after ascending order), represents the lower quartile of the local densities of all earthquake event samples in the th cluster (the value at the position after ascending order), represents the lower quartile of the relative distances of all earthquake event samples in the th cluster (the value at the position after ascending order). The first half of the calculation of the core density and the core distance uses the median to reflect the overall level of the local densities and relative distances of earthquake event samples in the th cluster. The second half calculates the difference between the data at the th cluster of earthquake event samples in the local density and relative distance and the data at the position and the position after ascending order, which is used to measure the dispersion degree of the data distribution. By adjusting the scale factors , , the strictness of the screening can be flexibly controlled. and The larger they are, the fewer outliers detected by the algorithm.

[0071] The scale factor has a great influence on the expected detection effect. If only relying on a single recognition criterion, the fluctuation of the outlier threshold may lead to a significant decline in the detection effect of the algorithm. Therefore, the present invention further defines the second criterion for outlier recognition: If the local density of a certain earthquake event sample is less than the boundary density of its affiliated cluster, then it is determined as an outlier.

[0072] Among them, the upper bound of the average density is divided into the left upper bound of the average density and the right upper bound of the average density . The left upper bound of the average density is the lowest density threshold between the current cluster and the earlier cluster in its time series, and the right upper bound of the average density is the minimum density threshold between the current cluster and the later cluster in its time series. The calculation of the two density upper bounds satisfies the following formula: , .in, After clustering, The upper bound of the left average density of clusters, After clustering, The upper bound of the right average density of clusters is Indicates Clusters of classes, Indicates the time axis and The adjacent previous cluster, Indicates the time axis and The next adjacent cluster, Indicates that it belongs to a cluster A sample of earthquake events, Indicates that it belongs to a cluster or The upper bound of the left average density is , right average density upper bound Clusters and or All satisfying spatial distances Less than The maximum value of the average density of earthquake event sample pairs.

[0073] Class Cluster The boundary density By its left average density upper bound and the upper bound of the right mean density The larger value of determines: .

[0074] In particular, for the first cluster in the time series, its boundary density is determined by the upper bound of its right average density, and the boundary density of the last cluster is equal to its left average density.

[0075] The two outlier identification criteria defined in the present invention can be expressed by the following formula: That is, if the earthquake event sample The local density Failed to reach its cluster The boundary density Core density The lower value in the Does not exceed the cluster Core distance , then the earthquake event sample can be determined as an abnormal point.

[0076] To evaluate the clustering performance of the spatio-temporal density peak clustering method (WNMA-STDPC) proposed in the present invention, simulation datasets and real datasets were used for experiments, and compared with 4 existing clustering methods: ST-DBSCAN (Spatial-Temporal Density Based Spatial Clustering of Applications with Noise), ST-OPTICS (Spatio-Temporal Ordering Points to Identify Clustering Structure), ST-AGNES (Spatial-Temporal Agglomerative Nesting), and ST-CFSFDP (Spatial-Temporal Clustering by Fast Search and Find of Density Peaks).

[0077] Specifically, 5 time-series datasets with different data volumes and numbers of clusters were constructed. - . The generation steps of the simulation dataset are as follows: first, determine the number, shape, and spatial position of the core clusters, and then add noise as needed to make the data closer to the real situation. As Figure 3 shown in Table 1 below, the content covers the basic information of the simulation dataset. In and datasets, there are significant differences in the time dimension among various clusters, so as to test the performance of the WNMA-STDPC algorithm on datasets with significant time attributes. and datasets have the characteristic of containing multiple clusters with different spatial positions within a specific period. Among them, contains clusters with relatively fuzzy geographical boundaries, which is used to evaluate the clustering performance of the WNMA-STDPC algorithm when dealing with datasets with dense spatial distributions. The clusters located below dataset have small differences in the time dimension, while there are significant differences in time characteristics from the clusters above. At the same time, the spatial positions of various clusters are close to each other, making the spatio-temporal clustering task on this dataset take into account the challenges in both time and space dimensions. The real datasets used are domestic earthquake datasets and . The research area covers 73° to 125° east longitude and 18° to 48° north latitude. It has rich information on the distribution of earthquake event samples, which can reflect the diversity of real-life scenarios, so as to further verify the effectiveness of the algorithm in this paper. The selected data information is as Figure 4As shown in Table 2 below. All clustering algorithm experiments were conducted in the following environment: a 64-bit operating system of Windows 11, an 11th Gen Intel(R) Core(TM) i5-1155G7 @ 2.50GHz CPU, 16GB of memory, and MATLAB R2023b.

[0078] The clustering results on the simulated dataset were evaluated using three common external metrics: Adjusted Mutual Information (AMI), Adjusted Rand Index (ARI), and Fowlkes-Mallows Index (FMI). The upper limit of each evaluation metric is 1, and the larger the metric value, the more accurate the clustering result. The real dataset was quantitatively analyzed based on the Davies-Bouldin Index (DBI). DBI comprehensively measures the clustering result by considering the within-cluster similarity and between-cluster dissimilarity. The smaller its value, the better the clustering effect. To comprehensively analyze the clustering quality, outliers were also regarded as independent clusters to participate in the index calculation. Due to the different characteristics of the datasets, it is impossible to determine a common parameter value applicable to all datasets. To achieve the best analysis effect, the range of algorithm parameter values was set in the experiment, and the parameters were adjusted at a predetermined step size. After multiple rounds of testing, the best results were selected. By plotting a two-dimensional scatter plot, the length and width of each cluster were estimated, and the in-circle with a radius of was used to cover all earthquake event samples within the cluster. Define the maximum value in as , and record the maximum value of the vertical height of each cluster as . The spatial neighborhood parameter of the ST-DBSCAN and ST-OPTICS algorithms has a value range of , the temporal neighborhood parameter has a value range of , the step size is 0.1, and the minimum number of points within the neighborhood is the logarithm of the dataset size . The spatial truncation distance of the ST-CFSFDP and WNMA-STDPC algorithms has a value range the same as , and the temporal truncation distance has a value range the same as . The scale factors , of the WNMA-STDPC algorithm have a value range of , and the step size is 0.1. The distance parameter of the ST-AGNES algorithm has a value range of , with a step size of 1.

[0079] As Figure 5 shown in Table 3 in , and the clustering effects of 5 algorithms on the simulated dataset are presented. The bold data represent the optimal clustering effects on each dataset. The experimental results show that the WNMA-STDPC algorithm performs optimally on all 5 datasets. Followed by the ST-OPTICS and ST-DBSCAN algorithms. Among them, the ST-OPTICS algorithm achieves sub-optimal results on the datasets , and

[0080] According to the attached instructions Figure 6 shown, Table 4 presents the statistical significance analysis of the clustering performance of 5 algorithms on the simulated dataset through the Friedman test. This method integrates the evaluation results of multiple datasets to obtain a global ranking of algorithm performance. The higher the average rank of the algorithm, the more significant the clustering effect of the algorithm. Analyzing Table 4, it can be seen that the WNMA-STDPC algorithm has the best average ranking in all evaluation indicators. It can be inferred from this that the WNMA-STDPC algorithm can effectively process datasets with complex spatio-temporal characteristics.

[0081] The earthquake event spatio-temporal density peak clustering method, device, medium and electronic device provided by the present invention utilize a Gaussian kernel function and introduce a spatio-temporal nearest neighbor weighting strategy when selecting cluster centers, so as to reduce the uncertainty of parameter selection and broaden the parameter selection interval, define a locally weighted spatio-temporal density, and improve the accuracy and reliability of density peak selection; during the process of merging cluster centers, time constraints are considered and the spatial and temporal characteristics between different clusters are strengthened, and a merging rule based on comprehensive similarity is designed to avoid mis-merging cluster centers; when processing the remaining earthquake event samples, a multi-step allocation strategy of gradually expanding the search range and improving the allocation accuracy is adopted to avoid the chain effect from damaging the integrity of the cluster. After all earthquake event samples are allocated, a dual-criterion is used to accurately identify outliers to ensure the comprehensive and effective allocation of samples.

[0082] Applying the method of the present invention to analyze earthquake data takes into account both the time attribute and the spatial attribute of earthquakes, can generate a clustering structure with high cohesion and low coupling, and effectively identifies foreshocks, main shocks and aftershocks by quantifying the spatio-temporal evolution pattern of earthquake sequences. This clustering analysis of earthquake events helps to understand the causes of strong earthquakes and even provides clues for predicting strong earthquakes.

[0083] See the attached Figure 7 to the specification. This embodiment also provides an earthquake event spatio-temporal density peak clustering device, which is used to implement the above method embodiment. The device includes:

[0084] A data processing unit 201, configured to obtain earthquake event sample data, where the earthquake event sample data includes an earthquake event spatio-temporal data set and a preset distance truncation threshold, normalize the earthquake event spatio-temporal data set, construct a spatial Euclidean distance matrix and a time distance matrix between earthquake event samples, and determine the spatio-temporal nearest neighbor set of earthquake event samples in the earthquake event spatio-temporal data set.

[0085] A local density calculation unit 202, configured to calculate the spatial weight coefficient and the time weight coefficient of each earthquake event sample according to the spatio-temporal nearest neighbor set of each earthquake event sample in the earthquake event spatio-temporal data set, and calculate the local density of each earthquake event sample under spatio-temporal constraints according to the spatial weight coefficient and the time weight coefficient.

[0086] A relative distance calculation unit 203, configured to calculate the relative distance of each earthquake event sample according to the time distance between earthquake event samples and the local density of each earthquake event sample under spatio-temporal constraints;

[0087] A potential cluster center selection unit 204, configured to calculate the decision value of each earthquake event sample according to the local density and relative distance of the earthquake event sample for selecting the potential cluster center of the earthquake event spatio-temporal data set.

[0088] A cluster center determination unit 205, configured to calculate the spatial similarity and temporal similarity of the cluster centers in the potential clusters, determine the comprehensive similarity of the cluster centers according to the spatial similarity and the temporal similarity, and eliminate the potential clusters with the potential for merging according to the comprehensive similarity of the cluster centers to obtain the cluster centers of various types of clusters.

[0089] An earthquake event sample allocation unit 206, configured to allocate earthquake event samples within a time window whose spatial distance from the cluster centers of various types of clusters is less than a preset distance truncation threshold to the corresponding clusters to achieve window allocation, and perform secondary allocation on the remaining earthquake event samples after window allocation according to the spatial distance between the earthquake event samples and the adjacent cluster centers to achieve search allocation, and divide the unallocated earthquake event samples after search allocation into the clusters to which the high-density nearest neighbor earthquake event samples belong.

[0090] An outlier identification unit 207, configured to calculate the core density, core distance, and boundary density of various types of clusters for outlier identification in various types of clusters, and obtain the clustering result of the final earthquake event sample data set.

[0091] All relevant contents of each step involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be elaborated here.

[0092] In some other embodiments of the present application, embodiments of the present application disclose an electronic device, as Figure 8 shown, the electronic device 300 may include: one or more processors 301; a memory 302; a display 303; one or more applications (not shown); and one or more computer programs 304. The above devices may be connected through one or more communication buses 305. The one or more computer programs 304 are stored in the above memory and are configured to be executed by the one or more processors 301. The one or more computer programs 304 include instructions, and the above instructions may be used to execute as Figure 1 、 Figure 7 and the respective steps in the corresponding embodiments.

[0093] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions may be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described system, device, and unit may refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.

[0094] In each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0095] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk, or optical disc.

[0096] As described above, the above are only the specific implementation manners of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of the present application should be covered by the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be subject to the protection scope of the claims.

Claims

1. A method for clustering the peak spatio-temporal density of seismic events, characterized in that, Including: Obtain sample data, where the sample data includes a spatio-temporal dataset of seismic events and a preset distance truncation threshold, normalize the spatio-temporal dataset of seismic events, construct a spatial Euclidean distance matrix and a time distance matrix between seismic event samples, and determine the spatio-temporal nearest neighbor set of seismic event samples in the spatio-temporal dataset of seismic events; Calculate the spatial weight coefficient and the time weight coefficient of each seismic event sample according to the spatio-temporal nearest neighbor set of each seismic event sample in the spatio-temporal dataset of seismic events, and calculate the local density of each seismic event sample under spatio-temporal constraints according to the spatial weight coefficient and the time weight coefficient; Calculate the relative distance of each seismic event sample according to the time distance between seismic event samples and the local density of each seismic event sample under spatio-temporal constraints; Calculate the decision value of each seismic event sample according to the local density and relative distance of the seismic event sample for selecting potential cluster centers of the spatio-temporal dataset of seismic events; Calculate the spatial similarity and the time similarity of each cluster center in the potential cluster centers, determine the comprehensive similarity of the cluster centers according to the spatial similarity and the time similarity, and eliminate potential cluster centers with the potential for merging according to the comprehensive similarity of the cluster centers to obtain the cluster centers of various clusters; Assign seismic event samples whose spatial distance from the cluster centers of various clusters within the time window is less than the preset distance truncation threshold to the corresponding clusters to achieve window allocation, and perform secondary allocation on the remaining seismic event samples after window allocation according to the spatial distance between the seismic event samples and the adjacent cluster centers to achieve search allocation, and divide the unallocated seismic event samples after search allocation into the clusters to which the high-density nearest neighbor seismic event samples belong; Calculate the core density, core distance, and boundary density of various clusters for identifying outliers in various clusters to obtain the clustering result of the final sample data.

2. The method according to claim 1, wherein Define that the spatial weight coefficient and the temporal weight coefficient respectively satisfy the following formulas: , , where represents the spatial weight coefficient, represents the spatio-temporal nearest neighbor set of earthquake event samples , represents the total spatial distance from all earthquake event samples in the spatio-temporal nearest neighbor set to , represents the earthquake event samples in the spatio-temporal nearest neighbor set except , represents the total spatial distance from the remaining earthquake event samples in the spatio-temporal nearest neighbor set except to ; represents the temporal weight coefficient, represents the total temporal distance from all earthquake event samples in the spatio-temporal nearest neighbor set to , represents the remaining earthquake event samples in the spatio-temporal nearest neighbor set except to ; The local density of the seismic event sample under spatio-temporal constraints satisfies the following formula: , where represents the local density of earthquake event samples , represents a preset distance truncation threshold, represents the earthquake event sample to the earthquake event sample spatial distance, represents a preset time truncation threshold, represents the earthquake event sample to the earthquake event sample time distance.

3. The method according to claim 1, wherein The calculation of the relative distance satisfies the following formula: , where represents the relative distance of the earthquake event sample , represents the relative distance of the earthquake event sample , represents the local density of the earthquake event sample , represents the time distance from the earthquake event sample to the earthquake event sample , represents the scale of the earthquake event sample in the dataset.

4. The method according to claim 1, characterized in that, Calculate the decision value of each seismic event sample according to the local density and relative distance of the seismic event sample for selecting potential cluster centers of the spatio-temporal dataset of seismic events, including: Sort the decision values of the earthquake event samples in the spatio-temporal dataset of earthquake events in descending order, and select several positive integer earthquake event samples with the largest decision values as potential cluster centers. Let this positive integer be z, and then obtain the set of potential cluster centers .

5. The method according to claim 1, characterized in that, The calculation of the spatial similarity satisfies the following formula: , where represents the cluster center and the spatial similarity between the cluster center ; represents the cluster center in the potential cluster center and the spatial distance between the cluster center ; represents the cluster center with a decision value higher than the cluster center ; represents a preset distance truncation threshold represents the decision value of the cluster center ; represents the decision value of the cluster center ; The calculation of the time similarity satisfies the following formula: , where represents the cluster center and the time similarity between the cluster center ; represents the time distance between the cluster center in the potential cluster and the cluster center ; represents a preset time truncation threshold; The comprehensive similarity satisfies the following formula: , where is the comprehensive similarity of the cluster center; Eliminate potential cluster centers with the potential for merging according to the comprehensive similarity of the cluster centers to obtain the cluster centers of various clusters, including: When the comprehensive similarity of the cluster center is 1, remove the cluster center from the potential cluster centers.

6. The method according to claim 1, characterized in that Perform secondary allocation on the remaining seismic event samples after window allocation according to the spatial distance between the seismic event samples and the adjacent cluster centers to achieve search allocation; divide the unallocated seismic event samples after search allocation into the clusters to which the high-density nearest neighbor seismic event samples belong, including: Compare the spatial distances between the remaining seismic event samples after window allocation and the adjacent two cluster centers, and assign the seismic event samples to the clusters to which the cluster centers with smaller spatial distances belong to achieve search allocation; Divide the unallocated seismic event samples after search allocation into the clusters to which the seismic event samples with higher local density than this seismic event sample and the closest time distance to this seismic event sample belong.

7. The method according to claim 1, characterized in that Calculate the core density, core distance, and boundary density of various clusters for identifying outliers in various clusters, including: Determine the seismic event samples with local density less than the lower value of the boundary density and the core density of the cluster to which they belong and relative distance less than the core distance of the cluster to which they belong as outliers.

8. An earthquake event spatio-temporal density peak clustering device, characterized in that, The device includes: A data processing unit, configured to obtain sample data, where the sample data includes a spatio-temporal data set of seismic events and a preset distance truncation threshold, normalize the spatio-temporal data set of seismic events, construct a spatial Euclidean distance matrix and a time distance matrix between seismic event samples, and determine a spatio-temporal nearest neighbor set of seismic event samples in the spatio-temporal data set of seismic events; A local density calculation unit, configured to calculate a spatial weight coefficient and a time weight coefficient of each seismic event sample according to the spatio-temporal nearest neighbor set of the seismic event samples in the spatio-temporal data set of seismic events, and calculate the local density of each seismic event sample under spatio-temporal constraints according to the spatial weight coefficient and the time weight coefficient; A relative distance calculation unit, configured to calculate the relative distance of each seismic event sample according to the time distance between seismic event samples and the local density of each seismic event sample under spatio-temporal constraints; A potential cluster center selection unit, configured to calculate a decision value of each seismic event sample according to the local density and relative distance of the seismic event sample for selecting potential cluster centers of the spatio-temporal data set of seismic events; A cluster center determination unit, configured to calculate the spatial similarity and time similarity of the cluster centers in the potential cluster centers, determine the comprehensive similarity of the cluster centers according to the spatial similarity and the time similarity, and eliminate potential cluster centers with the potential for merging according to the comprehensive similarity of the cluster centers to obtain the cluster centers of various clusters; A seismic event sample allocation unit, configured to allocate seismic event samples whose spatial distance from the cluster centers of various clusters within a time window is less than the preset distance truncation threshold to the corresponding clusters to achieve window allocation, perform secondary allocation on the remaining seismic event samples after window allocation according to the spatial distance between the seismic event samples and adjacent cluster centers to achieve search allocation, and divide the unallocated seismic event samples after search allocation into the clusters to which the high-density nearest neighbor seismic event samples belong; An outlier identification unit, configured to calculate the core density, core distance, and boundary density of each cluster for outlier identification in each cluster, and obtain the clustering result of the final sample data.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the spatio-temporal density peak clustering method for seismic events described in any one of claims 1 to 7.

10. An electronic device, characterized in that, Comprising: A processor and a memory; The memory is used to store a computer program; The processor is used to execute the computer program stored in the memory, so that the electronic device executes the spatio-temporal density peak clustering method for seismic events described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data labeling method based on artificial intelligence, apparatus and storage medium

    US20230316709A1

  • Large-scale data clustering method and apparatus, computer device and computer-readable storage medium

    WO2021042844A1