Earthquake event space-time density peak value clustering method and device, medium and electronic equipment
Through a spatiotemporal density peak clustering method for earthquake events, the problem of not fully considering time factors in the existing technology is solved. Through multi-step allocation strategy and double-criteria anomaly point recognition, the accuracy and consistency of clustering are improved, and more efficient spatiotemporal data clustering of earthquake events is achieved.
Patent Information
- Application Number
- CN202510610307.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The prior art cannot fully consider time factors when processing seismic events spatiotemporal data, resulting in local density definitions not applicable to clusters with large density differences. The cluster center merging process does not consider time factors, and a single-step allocation strategy based on the timeline is prone to mislabel normal samples as abnormal points.
A method of clustering of spatiotemporal density peaks in earthquake events is proposed. By acquiring and normalizing spatiotemporal data, a spatial and temporal distance matrix is constructed, the spatial and temporal distance matrix is calculated, the spatial and temporal neighbor set, local density, relative distance and decision value of the sample is calculated, the potential cluster center is selected, and the cluster center with merge potential is eliminated, and a multi-step allocation strategy and a double-criteria abnormal point recognition mechanism is adopted.
The accuracy and consistency of the clustering method are improved. By expanding the computing range and enhancing the influence of space-time and near neighbors, reliable cluster centers are selected, and flexible multi-step allocation strategy and dual-criteria exception point recognition are adopted, which improves the clustering effect of the algorithm.
Smart Images

Figure CN120123802A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of machine learning and data mining, and particularly to a spatio-temporal density peak clustering method, device, medium and electronic device for seismic events. Background Art
[0002] Clustering algorithms are an important unsupervised learning technique in machine learning and data mining, and are applicable to grouping unlabeled samples. Its main purpose is to classify samples into several clusters with specific distribution patterns by analyzing the similarity between samples. This method ensures that samples in the same cluster have a high degree of correlation, while reducing the similarity between samples in different clusters. Clustering algorithms can automatically discover the potential structures and patterns in the dataset without the need to provide labels in advance, and are widely used and effective in research fields such as information retrieval and pattern recognition.
[0003] With the continuous increase of spatio-temporal sequence data, spatio-temporal data has gradually become a research hotspot in many fields and plays an important role in fields such as environmental monitoring, geographic information systems, and public health. Spatio-temporal data not only contains spatial coordinate information but also involves time tags, making the data exhibit complex spatio-temporal dynamic characteristics. To mine potential patterns and rules from spatio-temporal data, clustering analysis stands out as an effective method and can provide strong support for decision-making in related fields. However, the clustering task of spatio-temporal data is more challenging than that of traditional static data clustering because it not only needs to consider the spatial distribution but also must analyze the evolution and trends in the time series. The spatio-temporal dependence, heterogeneity, and high-dimensionality of spatio-temporal data require clustering algorithms to have high flexibility and adaptability to cope with the complex interaction between space and time.
[0004] In the prior art, spatial clustering algorithms have been applied to the analysis of seismic events and the identification of foreshocks and aftershocks, but the spatial clustering algorithms ignore the time attribute of earthquakes. For the analysis of seismic events, the spatial attribute and time attribute factors are equally important, and a method that can take into account both the time attribute and the spatial attribute is needed for the analysis of seismic events.
[0005] Based on the need for clustering analysis of spatio-temporal data, various methods have been proposed in the prior art for clustering analysis of spatio-temporal data, but there are disadvantages in that the influence of time factors on the clustering process cannot be fully considered, and there are the following defects in the face of spatio-temporal datasets: 1) The definition of local density is not applicable to clusters with large density differences; 2) The merging process of cluster centers does not consider the influence of time factors; 3) The single-step assignment strategy based on the time axis is prone to mislabeling normal samples as outliers when dealing with clusters with a large spatial span between clusters.
[0006] Therefore, it is necessary to provide a spatio-temporal data clustering method for seismic events with high accuracy and consistency. Summary of the Invention
[0007] The object of the present invention is to provide a method, device, medium and electronic device for clustering the peak of spatio-temporal density of seismic events, so as to improve the accuracy and consistency of the clustering method when processing spatio-temporal data.
[0008] In a first aspect, the method for clustering the peak of spatio-temporal density of seismic events provided by the present invention includes: obtaining sample data, where the sample data includes a spatio-temporal data set of seismic events and a preset distance truncation threshold, normalizing the spatio-temporal data set of seismic events, constructing a spatial Euclidean distance matrix and a time distance matrix between seismic event samples, and determining a spatio-temporal neighbor set of seismic event samples in the spatio-temporal data set of seismic events; calculating a spatial weight coefficient and a time weight coefficient of each seismic event sample according to the spatio-temporal neighbor set of the seismic event sample in the spatio-temporal data set of seismic events, and calculating the local density of each seismic event sample under spatio-temporal constraints according to the spatial weight coefficient and the time weight coefficient; calculating the relative distance of each seismic event sample according to the time distance between seismic event samples and the local density of each seismic event sample under spatio-temporal constraints; calculating a decision value of each seismic event sample according to the local density and the relative distance of the seismic event sample to select potential cluster centers of the spatio-temporal data set of seismic events; calculating the spatial similarity and the time similarity of each cluster center in the potential cluster centers, determining the comprehensive similarity of the cluster centers according to the spatial similarity and the time similarity, and removing potential cluster centers with the potential for merging according to the comprehensive similarity of the cluster centers to obtain cluster centers of various clusters; allocating seismic event samples with a spatial distance less than the preset distance truncation threshold from the cluster centers of various clusters within a time window to the corresponding clusters to achieve window allocation, and performing secondary allocation on the remaining seismic event samples after window allocation according to the spatial distance between the seismic event samples and the adjacent cluster centers to achieve search allocation, and dividing the unallocated seismic event samples after search allocation into the clusters to which the high-density nearest neighbor seismic event samples belong; calculating the core density, core distance and boundary density of each cluster for identifying outliers in each cluster, and obtaining the clustering result of the final sample data.
[0009] The beneficial effects of the method for clustering the peak of spatio-temporal density of seismic events provided by the present invention are as follows: when calculating the local density of seismic event samples, reliable potential cluster centers are selected by expanding the calculation range and enhancing the influence degree of the spatio-temporal neighborhood; time and space constraints are simultaneously introduced into the cluster center merging rule to obtain cluster centers with higher accuracy and consistency; a more flexible multi-step allocation strategy is adopted to replace the single allocation mode, and an outlier identification mechanism based on a dual criterion is combined to improve the clustering effect of the algorithm.
[0010] In a possible embodiment, it is defined that the spatial weight coefficient and the time weight coefficient respectively satisfy the following formulas: , , where represents the spatial weight coefficient, represents the spatio-temporal neighbor set of earthquake event samples of represents the sum of the spatial distances from all earthquake event samples in the spatio-temporal neighbor set to ; represents the earthquake event samples in the spatio-temporal neighbor set except ; represents the sum of the spatial distances from the remaining earthquake event samples in the spatio-temporal neighbor set except to ; represents the temporal weight coefficient, represents the sum of the temporal distances from all earthquake event samples in the spatio-temporal neighbor set to ; represents the sum of the temporal distances from the remaining earthquake event samples in the spatio-temporal neighbor set except to ; The local density of earthquake event samples under spatio-temporal constraints satisfies the following formula: , where represents the local density of earthquake event sample ; represents the preset distance truncation threshold, represents the spatial distance from earthquake event sample to earthquake event sample ; represents the preset time truncation threshold, represents the temporal distance from earthquake event sample to earthquake event sample ;
[0011] In another possible embodiment, the calculation of the relative distance satisfies the following formula: , where represents the relative distance of earthquake event sample ; represents the relative distance of earthquake event sample ; represents the local density of earthquake event sample ; represents the temporal distance from earthquake event sample to earthquake event sample ; represents the scale of earthquake event samples in the data set.
[0012] In other possible embodiments, the decision value of each seismic event sample is calculated according to the local density and relative distance of the seismic event samples, and is used to select potential cluster centers of the seismic event spatio-temporal data set, including: arranging the decision values of the seismic event samples in the seismic event spatio-temporal data set in descending order, and selecting several positive integer seismic event samples with the largest decision values as potential cluster centers, and setting this positive integer as z, so as to obtain a set of potential cluster centers .
[0013] The calculation of the spatial similarity satisfies the following formula: , where represents the cluster center and the cluster center of the spatial similarity, represents the cluster center in the potential cluster centers and the cluster center the spatial distance between, represents the decision value higher than the cluster center of the cluster center, represents a preset distance truncation threshold, represents the cluster center of the decision value, represents the cluster center of the decision value; the calculation of the time similarity satisfies the following formula: , where represents the cluster center and the cluster center of the time similarity, represents the cluster center in the potential cluster centers and the cluster center the time distance between, represents a preset time truncation threshold; the comprehensive similarity satisfies the following formula: , where represents the comprehensive similarity of the cluster center; according to the comprehensive similarity of the cluster center, the potential cluster centers with the potential to merge are eliminated to obtain the cluster centers of various clusters, including: when the comprehensive similarity of the cluster center is 1, the cluster center is eliminated from the potential cluster centers.
[0014] The remaining seismic event samples after window allocation are secondarily allocated according to the spatial distance between the seismic event samples and the adjacent cluster centers to achieve search allocation; the unallocated seismic event samples after search allocation are divided into the cluster to which the high-density nearest neighbor seismic event sample belongs, including: comparing the spatial distances between the remaining seismic event samples after window allocation and the adjacent two cluster centers, and allocating the seismic event samples to the cluster to which the cluster center with the smaller spatial distance belongs to achieve search allocation; the unallocated seismic event samples after search allocation are divided into the cluster to which the seismic event sample with a local density higher than this seismic event sample and the closest time distance to this seismic event sample belongs.
[0015] Calculate the core density, core distance, and boundary density of each type of cluster for outlier identification in each type of cluster, including: determining an earthquake event sample with a local density less than the lower value of the boundary density and core density of its affiliated cluster and a relative distance less than the core distance of its affiliated cluster as an outlier.
[0016] In a second aspect, the present invention also provides an earthquake event spatio-temporal density peak clustering device, including: A data processing unit, configured to obtain earthquake event sample data, where the earthquake event sample data includes an earthquake event spatio-temporal data set and a preset distance truncation threshold, normalize the earthquake event spatio-temporal data set, construct a spatial Euclidean distance matrix and a time distance matrix between earthquake event samples, and determine a spatio-temporal neighbor set of earthquake event samples in the earthquake event spatio-temporal data set; A local density calculation unit, configured to calculate a spatial weight coefficient and a time weight coefficient of each earthquake event sample according to the spatio-temporal neighbor set of each earthquake event sample in the earthquake event spatio-temporal data set, and calculate the local density of each earthquake event sample under spatio-temporal constraints according to the spatial weight coefficient and the time weight coefficient; A relative distance calculation unit, configured to calculate the relative distance of each earthquake event sample according to the time distance between earthquake event samples and the local density of each earthquake event sample under spatio-temporal constraints; A potential cluster center selection unit, configured to calculate a decision value of each earthquake event sample according to the local density and relative distance of the earthquake event sample for selecting potential cluster centers of the earthquake event spatio-temporal data set; A cluster center determination unit, configured to calculate the spatial similarity and time similarity of the cluster centers in the potential cluster centers, determine the comprehensive similarity of the cluster centers according to the spatial similarity and the time similarity, and eliminate potential cluster centers with the potential for merging according to the comprehensive similarity of the cluster centers to obtain the cluster centers of each type of cluster; An earthquake event sample allocation unit, configured to allocate earthquake event samples with a spatial distance less than the preset distance truncation threshold from the cluster centers of each type of cluster within a time window to the corresponding clusters to achieve window allocation, perform secondary allocation on the remaining earthquake event samples after window allocation according to the spatial distance between the earthquake event samples and adjacent cluster centers to achieve search allocation, and divide the unallocated earthquake event samples after search allocation into the clusters to which the high-density nearest neighbor earthquake event samples belong; An outlier identification unit, configured to calculate the core density, core distance, and boundary density of each type of cluster for outlier identification in each type of cluster, and obtain the clustering result of the final earthquake event sample data set.
[0017] In a third aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method for clustering the peak of the spatio-temporal density of seismic events is implemented.
[0018] In a fourth aspect, the present invention also provides an electronic device, including: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory, so that the electronic device executes the above-mentioned method for clustering the peak of the spatio-temporal density of seismic events.
[0019] For the beneficial effects of the above second to fourth aspects, reference may be made to the description of the first aspect above. Description of the Drawings
[0020] Figure 1 It is a schematic flowchart of a method for clustering the peak of the spatio-temporal density of seismic events provided by an embodiment of the present invention; Figure 2 It is a schematic diagram of the window allocation process of seismic event samples provided by an embodiment of the present invention; Figure 3 It is a basic information table of simulation data provided by an embodiment of the present invention; Figure 4 It is a basic information table of real data provided by an embodiment of the present invention; Figure 5 It is a clustering result table of five algorithms on a simulation data set provided by an embodiment of the present invention; Figure 6 It is a Friedman value table of evaluation indexes of five algorithms on a simulation data set provided by an embodiment of the present invention; Figure 7 It is a schematic diagram of a spatio-temporal density peak clustering device provided by an embodiment of the present invention; Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0021] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art in the field to which the present invention pertains. The terms such as "including" used herein mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items.
[0022] This embodiment provides a method for clustering the spatio-temporal density peaks of seismic events. Referring to the accompanying drawings of the specification Figure 1 , the method includes: S101: Obtain sample data, where the sample data includes a spatio-temporal dataset of seismic events and a preset distance truncation threshold, normalize the spatio-temporal dataset of seismic events, construct a spatial Euclidean distance matrix and a time distance matrix between seismic event samples, and determine the spatio-temporal nearest neighbor set of the seismic event samples in the spatio-temporal dataset of seismic events.
[0023] In a specific embodiment, in the spatio-temporal dataset of seismic events , the spatio-temporal nearest neighbor set of the seismic event sample can be expressed as: , where represents the seismic event samples in the spatio-temporal dataset of seismic events except , represents the spatial distance from the seismic event sample to the seismic event sample , represents the time distance from the seismic event sample to the seismic event sample .
[0024] represents a seismic event sample, and its spatio-temporal nearest neighbor set consists of seismic event samples that satisfy the condition that its spatial distance is less than or equal to the spatial truncation distance , and the time distance is less than or equal to the time truncation distance .
[0025] In a specific embodiment, the spatio-temporal dataset of seismic events is a set of seismic event samples including time attributes and spatial attributes, and the seismic event sample information specifically includes information such as the occurrence time, occurrence location, and magnitude of the seismic event.
[0026] S102: Calculate the spatial weight coefficient and the temporal weight coefficient of each earthquake event sample according to the spatio-temporal nearest neighbor set of earthquake event samples in the spatio-temporal dataset of earthquake events, and calculate the local density of each earthquake event sample under spatio-temporal constraints according to the spatial weight coefficient and the temporal weight coefficient.
[0027] In a possible embodiment, it is defined that the spatial weight coefficient and the temporal weight coefficient respectively satisfy the following formulas: , , where represents the spatial weight coefficient, represents the earthquake event sample 's spatio-temporal nearest neighbor set, represents the total spatial distance from all earthquake event samples in the spatio-temporal nearest neighbor set to , represents the earthquake event sample except in the spatio-temporal nearest neighbor set, represents the remaining earthquake event samples except in the spatio-temporal nearest neighbor set to 's total spatial distance, represents the temporal weight coefficient, represents the total temporal distance from all earthquake event samples in the spatio-temporal nearest neighbor set to , represents the remaining earthquake event samples except in the spatio-temporal nearest neighbor set to 's total temporal distance.
[0028] The local density of the earthquake event sample under spatio-temporal constraints satisfies the following formula: , where represents the local density of the earthquake event sample , represents a preset distance truncation threshold, represents the earthquake event sample to the earthquake event sample 's spatial distance, represents a preset temporal truncation threshold, represents the earthquake event sample to the earthquake event sample 's temporal distance.
[0029] The locally density weighted by spatio-temporal nearest neighbors aims to quantify the contribution degree of spatio-temporal nearest neighbor earthquake event samples to and evaluate the contribution of earthquake event samples outside the domain to . It enhances the relationship with and , Closely associate with the influence of earthquake event samples, and expand the local density gap between different earthquake event samples; use the sum of the reciprocals of the spatial distance and the time distance as a measurement index, and emphasize the interaction with earthquake event samples from other regions outside the domain. Use the square term as the denominator to effectively reduce the proportion of the influence of earthquake event samples from outside the domain on the overall local density calculation, and provide a more comprehensive reference basis for the evaluation of local density. Equation (11) comprehensively considers the spatio-temporal neighborhood information and global information of earthquake event samples, and improves the discrimination between density peak and non-density peak earthquake event samples.
[0030] S103: Calculate the relative distance of earthquake event samples according to the time distance between earthquake event samples and the local density of each earthquake event sample under spatio-temporal constraints.
[0031] In a possible embodiment, the calculation of the relative distance satisfies the following formula: , where represents the relative distance of earthquake event sample , represents the relative distance of earthquake event sample , represents the local density of earthquake event sample , represents the time distance from earthquake event sample to earthquake event sample , represents the scale of earthquake event samples in the dataset. That is, the relative distance of the non-maximum local density earthquake event sample is defined as the spatial distance between earthquake event sample and the nearest earthquake event sample with a higher local density; the earthquake event sample with the largest local density value has the global maximum relative distance.
[0032] S104: Calculate the decision value of each earthquake event sample according to the local density and relative distance of the earthquake event sample to select potential cluster centers of the earthquake event spatio-temporal dataset.
[0033] In a possible embodiment, the decision value of the earthquake event sample satisfies the following formula: .
[0034] In a specific embodiment, the decision value of each seismic event sample is calculated based on the local density and relative distance of the seismic event samples to select potential cluster centers of the seismic event spatio-temporal data set, including: arranging the decision values of the seismic event samples in the seismic event spatio-temporal data set in descending order, and selecting the z seismic event samples with the largest numerical decision values as potential cluster centers (potential class cluster centers of each class), where z is a positive integer, thereby obtaining a set of potential cluster centers .
[0035] S105: Calculate the spatial similarity and temporal similarity of each cluster center in the potential cluster centers, determine the comprehensive similarity of the cluster centers according to the spatial similarity and temporal similarity, and remove the potential cluster centers with the potential to merge according to the comprehensive similarity of the cluster centers to obtain the cluster centers of each class.
[0036] In a possible embodiment, the calculation of the spatial similarity satisfies the following formula: , where represents the spatial similarity between the cluster center and the cluster center , represents the spatial distance between the cluster center in the potential cluster centers and the cluster center , represents the cluster center whose decision value is higher than the cluster center , represents the preset distance truncation threshold, represents the decision value of the cluster center , represents the decision value of the cluster center .
[0037] The calculation of the temporal similarity satisfies the following formula: , where represents the temporal similarity between the cluster center and the cluster center , represents the temporal distance between the cluster center in the potential cluster centers and the cluster center , represents the preset temporal truncation threshold.
[0038] The comprehensive similarity satisfies the following formula: , where is the comprehensive similarity of the cluster center.
[0039] Removing the potential cluster centers with the potential to merge according to the comprehensive similarity of the cluster centers to obtain the cluster centers of each class includes: when the comprehensive similarity of the cluster center is 1, removing the cluster center from the potential cluster centers.
[0040] For evaluating the cluster center and its proximity to the nearest neighboring cluster center in the spatial dimension. If the cluster center and the cluster center with a decision value higher than its own the nearest spatial distance between them is less than the preset threshold , then the spatial similarity is set to 1, indicating that the cluster center can be regarded as the nearest neighbor of a cluster center with a higher decision value in space; otherwise, it is set to 0, indicating that the cluster center is far away from other cluster centers in space.
[0041] For evaluating the cluster center and its proximity to the nearest neighboring cluster center in the time dimension. If the cluster center and the cluster center with a decision value higher than its own the shortest time distance between them is less than the preset threshold , then the time similarity is set to 1, indicating that the cluster center can be regarded as the nearest neighbor of a cluster center with a higher decision value in time; otherwise, it is set to 0, indicating that the cluster center is far away from other cluster centers in time.
[0042] Only when both the spatial similarity and the time similarity are 1, the comprehensive similarity of the cluster center is 1. That is, when a cluster center is regarded as the neighbor of a cluster center with a higher decision value in both the spatial and time dimensions, the cluster center has the ability to merge with other cluster centers, and then it can be removed from the set of potential cluster centers. This method avoids the possibility of multiple cluster centers existing in a single class cluster, ensuring that the merged cluster centers can efficiently guide the allocation of the remaining cluster centers.
[0043] S106: Allocate the earthquake event samples within the time window whose spatial distance from the cluster centers of various clusters is less than the preset distance truncation threshold to the corresponding clusters to achieve window allocation, and perform secondary allocation on the remaining earthquake event samples after window allocation according to the spatial distance between the earthquake event samples and the adjacent cluster centers to achieve search allocation. Divide the unallocated earthquake event samples after search allocation into the clusters to which the high-density nearest neighbor earthquake event samples belong.
[0044] In a possible embodiment, the remaining seismic event samples after window allocation are secondarily allocated according to the spatial distance between the seismic event samples and the adjacent cluster centers to achieve search allocation, and the unallocated seismic event samples after search allocation are divided into the cluster to which the high-density nearest neighbor seismic event samples belong, including: comparing the spatial distances between the remaining seismic event samples after window allocation and the adjacent two cluster centers, and allocating the seismic event samples to the cluster to which the cluster center with a smaller spatial distance belongs to achieve search allocation; dividing the unallocated seismic event samples after search allocation into the cluster to which the seismic event samples with a local density higher than that of the seismic event samples and the closest time distance to the seismic event samples belong.
[0045] In a specific embodiment, the three steps of allocating the remaining seismic event samples after determining the cluster centers are specifically executed as follows: The first allocation strategy: window allocation. For each cluster center on the time axis , calculate the seismic event samples within the window interval to the cluster center . If to the spatial distance is less than the preset spatial cut-off distance , then it is classified into the cluster represented by .
[0046] See the attached drawings of the specification Figure 2 , the correctly allocated seismic event samples in front of the first cluster center and behind the last cluster center , and their spatial distance from or is less than , and the time distance is controlled within . That is, the seismic event samples within this range are in the near-neighbor region of the cluster center and have a high correlation with it. Therefore, they should be classified into the cluster to which the corresponding cluster center belongs. For the seismic event samples outside the window range and those with a relatively large spatial distance from the cluster center within the window, they are temporarily regarded as outliers in this step of allocation. As the starting stage of the entire allocation strategy, the window allocation range is set relatively small, aiming to preferentially process the seismic event samples with the strongest correlation with each cluster center and provide several stable base clusters composed of a certain number of reliable seismic event samples for subsequent allocation.
[0047] The second allocation strategy: search allocation. Since the set value of the time cut-off distance is usually less than the time distance between adjacent cluster centers, the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the first step. In the search allocation stage, by comparing the spatial distances between the seismic event samples and the adjacent two cluster centers, the cluster of the closer side is selected for allocation, without involving the spatial cut-off distance The set value is usually less than the time distance between adjacent cluster centers, so the search range of seismic event samples in the second allocation process is expanded compared with the The comparison can avoid the influence of abnormal parameters on the allocation results and improve the stability of allocation.
[0048] The third allocation strategy: nearby allocation. For earthquake event samples that have not been successfully allocated in the first two allocation processes, each earthquake event sample is traversed in order from high to low local density. , find the earthquake event sample with a local density higher than the current earthquake event sample and the closest time distance to it , classify the current earthquake event sample into the earthquake event sample Belongs to the cluster.
[0049] The core basis of this processing method is that high-density earthquake event samples are usually cluster centers or their neighboring points, so they are more likely to be regarded as representative earthquake event samples of a certain cluster. The nearest allocation strategy searches for global reliable information to ensure that all remaining earthquake event samples are correctly allocated to their respective clusters to the greatest extent possible.
[0050] S107: Calculate the core density, core distance and boundary density of each cluster to identify outliers in each cluster, and obtain the final clustering result of the sample data.
[0051] In a possible embodiment, the core density, core distance and boundary density of each cluster are calculated for identifying outliers in each cluster, including: earthquake event samples whose local density is less than the lower value of the boundary density and core density of the cluster to which they belong, and whose relative distance is less than the core distance of the cluster to which they belong, are determined as outliers.
[0052] In a specific embodiment, the present invention defines a criterion for outlier identification: if the local density of a seismic event sample is Lower than the core density of the cluster to which it belongs , and the relative distance Exceeds the core distance of the cluster to which it belongs , it is considered as an outlier.
[0053] The calculation of core density satisfies the following formula: , the calculation of the core distance satisfies the following formula: .in, Represents the clustering result The core density of each cluster, Represents the clustering result The core distance of each cluster, Indicates The median of the local density of each earthquake event sample in a cluster, Indicates The median of the relative distances of the earthquake event samples in each cluster, , represents the scale factor, Indicates The upper quartile of the local density of all earthquake event samples in a cluster (arranged in ascending order) The value of the position), Indicates The upper quartile of the relative distances of all earthquake event samples in a cluster (arranged in ascending order) The value of the position), Indicates The lower quartile of the local density of all earthquake event samples in a cluster (after ascending order) The value of the position), Indicates The lower quartile of the relative distance of all earthquake event samples in a cluster (after sorting in ascending order) The first half of the calculation of core density and core distance uses the median to reflect the first The overall level of local density and relative distance of each earthquake event sample in the cluster is calculated in the second half and arranged in ascending order. The local density and relative distance of earthquake event samples in clusters Location and The difference in position data is used to measure the degree of dispersion of data distribution. By adjusting the scale factor , The strictness of the screening can be flexibly controlled. and The larger it is, the fewer outliers the algorithm detects.
[0054] The proportional factor has a great influence on the expected detection effect. If only a single recognition criterion is relied upon, the fluctuation of the anomaly threshold may cause the detection effect of the algorithm to be significantly reduced. Therefore, the present invention further defines a second criterion for outlier recognition: if the local density of a seismic event sample is Less than the boundary density of the cluster to which it belongs , it is judged as an abnormal point.
[0055] Among them, the upper bound of the average density Divided into left mean density upper bound and the upper bound of the right mean density . The upper bound of the left mean density is is the minimum density threshold between the current cluster and the earlier cluster in its time series, and the upper bound of the right average density is the minimum density threshold between the current cluster and the later cluster in its time series. The calculation of the two density upper bounds satisfies the following formula: , .in, After clustering, The upper bound of the left average density of clusters, After clustering, The upper bound of the right average density of clusters is Indicates Clusters of classes, Indicates the time axis and The adjacent previous cluster, Indicates the time axis and The next adjacent cluster, Indicates that it belongs to a cluster A sample of earthquake events, Indicates that it belongs to a cluster or The upper bound of the left average density is , right average density upper bound Clusters and or All satisfying spatial distances Less than The maximum value of the average density of earthquake event sample pairs.
[0056] Class Cluster The boundary density By its left average density upper bound and the upper bound of the right mean density The larger value of determines: .
[0057] In particular, for the first cluster in the time series, its boundary density is determined by the upper bound of its right average density, and the boundary density of the last cluster is equal to its left average density.
[0058] The two outlier identification criteria defined in the present invention can be expressed by the following formula: That is, if the earthquake event sample The local density Failed to reach its cluster The boundary density Core density The lower value in the Does not exceed the cluster Core distance , then the earthquake event sample can be determined as an abnormal point.
[0059] To evaluate the clustering performance of the proposed spatio-temporal density peak clustering method (WNMA-STDPC) of the present invention, simulation datasets and real datasets were used for experiments, and comparisons were made with 4 existing clustering methods, namely ST-DBSCAN (Spatial-Temporal Density Based Spatial Clustering of Applications with Noise), ST-OPTICS (Spatio-Temporal Ordering Points to Identify Clustering Structure), ST-AGNES (Spatial-Temporal Agglomerative Nesting), and ST-CFSFDP (Spatial-Temporal Clustering by Fast Search and Find of Density Peaks).
[0060] Specifically, 5 time series datasets with different data volumes and numbers of clusters were constructed - . The generation steps of the simulation dataset were as follows: first, determine the number, shape, and spatial location of the core clusters, and then add noise as needed to make the data closer to the real situation. As Figure 3 shown in Table 1 below, the content covers the basic information of the simulation dataset. In and the datasets, there are significant differences in the time dimension among various clusters, so as to test the performance of the WNMA-STDPC algorithm on datasets with significant time attributes. and The datasets have the characteristic of containing multiple clusters with different spatial positions within a specific time period. Among them, the geographical boundaries of the contained clusters are relatively fuzzy, which is used to evaluate the clustering performance of the WNMA-STDPC algorithm when dealing with datasets with dense spatial distributions. The clusters located below the dataset have small differences in the time dimension, while there are significant differences in time characteristics from the clusters above. At the same time, the spatial positions of various clusters are close to each other, making the spatio-temporal clustering task on this dataset take into account the challenges of both time and space dimensions. The real dataset used is the domestic earthquake dataset and , and the research area covers 73° to 125° east longitude and 18° to 48° north latitude. It has rich distribution information of earthquake event samples and can reflect the diversity of real-life scenarios to further verify the effectiveness of the algorithm in this paper. The selected data information is as Figure 4As shown in Table 2. The experiments of all clustering algorithms were conducted in the following environment: a 64-bit operating system of Windows 11, an 11th Gen Intel(R) Core(TM) i5-1155G7 @ 2.50GHz CPU, 16GB of memory, and MATLAB R2023b.
[0061] The clustering results on the simulated datasets were evaluated using three common external metrics: Adjusted Mutual Information (AMI), Adjusted Rand Index (ARI), and Fowlkes-Mallows Index (FMI). The upper limit of each evaluation metric is 1, and the larger the metric value, the more accurate the clustering result. The real datasets were quantitatively analyzed based on the Davies-Bouldin Index (DBI). DBI comprehensively measures the clustering result by considering the within-cluster similarity and between-cluster dissimilarity. The smaller its value, the better the clustering effect. To comprehensively analyze the clustering quality, outliers were also regarded as independent clusters to participate in the index calculation. Due to the different characteristics of the datasets, it is impossible to determine a universal parameter value applicable to all datasets. To achieve the best analysis effect, the value range of the algorithm parameters was set in the experiment, and the parameters were adjusted at a predetermined step size. After multiple rounds of testing, the best results were selected. By plotting the two-dimensional spatial scatter plot, the length and width of each cluster were estimated, and the in-circle with a radius of was used to cover all earthquake event samples within the cluster. Define the maximum value in as , and record the maximum value of the vertical height of each cluster as . The spatial neighborhood parameter of the ST-DBSCAN and ST-OPTICS algorithms has a value range of , the temporal neighborhood parameter has a value range of , with a step size of 0.1. The minimum number of points within the neighborhood is the logarithm of the dataset size . The spatial truncation distance of the ST-CFSFDP and WNMA-STDPC algorithms has the same value range as , and the temporal truncation distance has the same value range as . The scaling factors , of the WNMA-STDPC algorithm have a value range of , with a step size of 0.1. The distance parameter of the ST-AGNES algorithm has a value range of , with a step size of 1.
[0062] As Figure 5 shown in Table 3 in , and , the content shows the clustering effects of 5 algorithms on the simulated dataset, where the bold data represents the optimal clustering effect on each dataset. The experimental results show that the WNMA-STDPC algorithm performs optimally on all 5 datasets, followed by the ST-OPTICS and ST-DBSCAN algorithms. Among them, the ST-OPTICS algorithm achieves sub-optimal results on the datasets , and
[0063] According to the attached Figure 6 to the specification, the content shown in Table 4 conducts a statistical significance analysis of the clustering performance of 5 algorithms on the simulated dataset through the Friedman test. This method integrates the evaluation results of multiple datasets to obtain a global ranking of algorithm performance. The higher the average rank of the algorithm, the more significant the clustering effect of the algorithm. Analyzing Table 4, it can be seen that the WNMA-STDPC algorithm has the best average ranking in all evaluation indicators. It can be inferred from this that the WNMA-STDPC algorithm can effectively process datasets with complex spatio-temporal characteristics.
[0064] The earthquake event spatio-temporal density peak clustering method, device, medium and electronic device provided by the present invention utilize a Gaussian kernel function and introduce a spatio-temporal nearest neighbor weighting strategy when selecting cluster centers, so as to reduce the uncertainty of parameter selection and broaden the parameter selection range, define a spatio-temporal nearest neighbor weighted local density, and improve the accuracy and reliability of density peak selection; during the cluster center merging process, time constraints are considered and the spatial and temporal characteristics between different clusters are strengthened, and a merging rule based on comprehensive similarity is designed to avoid mis-merging cluster centers; when processing the remaining earthquake event samples, a multi-step allocation strategy of gradually expanding the search range and improving the allocation accuracy is adopted to avoid the chain effect from destroying the cluster integrity. After all earthquake event samples are allocated, a double criterion is used to accurately identify outliers to ensure the comprehensive and effective allocation of samples.
[0065] Applying the method of the present invention to analyze earthquake data takes into account both the time attribute and the spatial attribute of earthquakes, can generate a clustering structure with high cohesion and low coupling, and effectively identifies foreshocks, main shocks and aftershocks by quantifying the spatio-temporal evolution pattern of earthquake sequences. This clustering analysis of earthquake events helps to understand the causes of strong earthquakes and even provides clues for predicting strong earthquakes.
[0066] See the attached Figure 7 description. This embodiment also provides an earthquake event spatio-temporal density peak clustering device, which is used to implement the above method embodiment. The device includes: A data processing unit 201, configured to obtain earthquake event sample data, where the earthquake event sample data includes an earthquake event spatio-temporal data set and a preset distance truncation threshold, normalize the earthquake event spatio-temporal data set, construct a spatial Euclidean distance matrix and a time distance matrix between earthquake event samples, and determine the spatio-temporal nearest neighbor set of earthquake event samples in the earthquake event spatio-temporal data set.
[0067] A local density calculation unit 202, configured to calculate the spatial weight coefficient and the time weight coefficient of each earthquake event sample according to the spatio-temporal nearest neighbor set of each earthquake event sample in the earthquake event spatio-temporal data set, and calculate the local density of each earthquake event sample under spatio-temporal constraints according to the spatial weight coefficient and the time weight coefficient.
[0068] A relative distance calculation unit 203, configured to calculate the relative distance of each earthquake event sample according to the time distance between earthquake event samples and the local density of each earthquake event sample under spatio-temporal constraints; A potential cluster center selection unit 204, configured to calculate the decision value of each earthquake event sample according to the local density and relative distance of the earthquake event sample to select the potential cluster center of the earthquake event spatio-temporal data set.
[0069] A cluster center determination unit 205, configured to calculate the spatial similarity and temporal similarity of the cluster centers in the potential clusters, determine the comprehensive similarity of the cluster centers according to the spatial similarity and the temporal similarity, and eliminate the potential clusters with the potential to merge according to the comprehensive similarity of the cluster centers to obtain the cluster centers of various types of clusters.
[0070] An earthquake event sample allocation unit 206, configured to allocate earthquake event samples within a time window whose spatial distance from the cluster centers of various types of clusters is less than a preset distance truncation threshold to the corresponding clusters to achieve window allocation, and perform secondary allocation on the remaining earthquake event samples after window allocation according to the spatial distance between the earthquake event samples and the adjacent cluster centers to achieve search allocation, and divide the unallocated earthquake event samples after search allocation into the clusters to which the high-density nearest neighbor earthquake event samples belong.
[0071] An outlier identification unit 207, configured to calculate the core density, core distance, and boundary density of various types of clusters for outlier identification in various types of clusters, and obtain the clustering result of the final earthquake event sample dataset.
[0072] All relevant contents of each step involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be elaborated here.
[0073] In some other embodiments of the present application, embodiments of the present application disclose an electronic device, as Figure 8 shown. The electronic device 300 may include: one or more processors 301; a memory 302; a display 303; one or more applications (not shown); and one or more computer programs 304. The above components may be connected through one or more communication buses 305. Wherein the one or more computer programs 304 are stored in the above memory and are configured to be executed by the one or more processors 301. The one or more computer programs 304 include instructions, and the above instructions may be used to execute as Figure 1 、 Figure 7 and the respective steps in the corresponding embodiments.
[0074] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.
[0075] In each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0076] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk, or optical disc.
[0077] As mentioned above, the above are only the specific implementation manners of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of the present application should be covered by the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be subject to the protection scope of the claims.
Claims
1. A method for clustering the peak values of the spatiotemporal density of earthquake events, characterized in that: include: Acquire sample data, the sample data including a spatiotemporal data set of earthquake events and a preset distance cutoff threshold, normalize the spatiotemporal data set of earthquake events, construct a spatial Euclidean distance matrix and a temporal distance matrix between earthquake event samples, and determine a spatiotemporal neighbor set of earthquake event samples in the spatiotemporal data set of earthquake events; Calculate the spatial weight coefficient and time weight coefficient of the earthquake event sample according to the spatial and temporal neighbor set of each earthquake event sample in the earthquake event spatial and temporal data set, and calculate the local density of each earthquake event sample under spatial and temporal constraints according to the spatial weight coefficient and the time weight coefficient; The relative distance of earthquake event samples is calculated according to the time distance between earthquake event samples and the local density of each earthquake event sample under the space-time constraint; The decision value of each earthquake event sample is calculated according to the local density and relative distance of the earthquake event sample to select the potential cluster center of the earthquake event spatiotemporal data set; Calculating the spatial similarity and the temporal similarity of each cluster center among the potential cluster centers, determining the comprehensive similarity of the cluster centers according to the spatial similarity and the temporal similarity, and eliminating the potential cluster centers with merging potential according to the comprehensive similarity of the cluster centers to obtain the cluster centers of various clusters; The earthquake event samples whose spatial distances to the cluster centers of various clusters within the time window are less than the preset distance cutoff threshold are allocated to the corresponding clusters to realize window allocation. The remaining earthquake event samples after window allocation are secondary allocated according to the spatial distances between the earthquake event samples and the adjacent cluster centers to realize search allocation. The earthquake event samples that are not allocated after search allocation are divided into the clusters to which the high-density nearest neighbor earthquake event samples belong. The core density, core distance and boundary density of each cluster are calculated to identify outliers in each cluster and obtain the final clustering result of the sample data.
2. The method according to claim 1, characterized in that The spatial weight coefficient and the temporal weight coefficient are defined to satisfy the following formulas: , ,in, represents the spatial weight coefficient, Represents a sample of earthquake events The space-time neighbor set of Represents all earthquake event samples in the space-time neighbor set The sum of the spatial distances, Indicates that the space-time neighbor set is The earthquake event samples outside Represents the space-time neighbor set except The remaining earthquake event samples are The sum of the spatial distances, represents the time weight coefficient, Represents all earthquake event samples in the space-time neighbor set The sum of the time distance, Represents the space-time neighbor set except The remaining earthquake event samples are The sum of the time distance; The local density of earthquake event samples under time and space constraints satisfies the following formula: ,in, Represents a sample of earthquake events The local density of Represents the preset distance cutoff threshold, Represents a sample of earthquake events To the earthquake event sample The spatial distance Indicates the preset time cutoff threshold. Represents a sample of earthquake events To the earthquake event sample time distance.
3. The method according to claim 1, characterized in that The calculation of relative distance satisfies the following formula: ,in, Represents a sample of earthquake events The relative distance Represents a sample of earthquake events The relative distance Represents a sample of earthquake events The local density of Represents a sample of earthquake events To the earthquake event sample The time distance, Represents the sample size of earthquake events in the dataset.
4. The method according to claim 1, characterized in that: The decision value of each earthquake event sample is calculated according to the local density and relative distance of the earthquake event sample to select the potential cluster center of the earthquake event spatiotemporal data set, including: The decision values of earthquake event samples in the earthquake event spatiotemporal data set are arranged in descending order, and several positive integer earthquake event samples with the largest decision values are selected as potential cluster centers, where the positive integer is set to z, thereby obtaining a potential cluster center set: .
5. The method according to claim 1, characterized in that The calculation of the spatial similarity satisfies the following formula: ,in, Cluster Heart With cluster heart The spatial similarity of represents the cluster center among potential cluster centers With cluster heart The spatial distance between Indicates that the decision value is higher than the cluster center The heart of cluster, Represents the preset distance cutoff threshold, Cluster Heart The decision value of Cluster Heart The decision value of The calculation of the time similarity satisfies the following formula: ,in, Cluster Heart With cluster heart The temporal similarity of represents the cluster center among potential cluster centers With cluster heart The time distance between Indicates the preset time cutoff threshold; The comprehensive similarity satisfies the following formula: ,in, is the comprehensive similarity of cluster centers; According to the comprehensive similarity of cluster centers, potential cluster centers with merging potential are eliminated to obtain cluster centers of various clusters, including: When the heart When the comprehensive similarity is 1, the cluster center Remove from potential clusters.
6. The method according to claim 1, characterized in that The remaining earthquake event samples after window allocation are secondary allocated according to the spatial distance between the earthquake event samples and the adjacent cluster centers to achieve search allocation; the earthquake event samples that are not allocated after search allocation are divided into the clusters to which the high-density nearest neighbor earthquake event samples belong, including: Compare the spatial distances between the remaining earthquake event samples and the two adjacent cluster centers after window allocation, and allocate the earthquake event samples to the cluster to which the cluster center with the smaller spatial distance belongs to realize search allocation; The earthquake event samples that are not assigned after the search assignment are divided into the cluster to which the earthquake event samples with a local density higher than that of the earthquake event samples and the closest time distance to the earthquake event samples belong.
7. The method according to claim 1, characterized in that Calculate the core density, core distance, and boundary density of each cluster for outlier identification in each cluster, including: The earthquake event samples whose local density is less than the lower value of the boundary density and core density of the cluster to which they belong, and whose relative distance is less than the core distance of the cluster to which they belong, are judged as abnormal points.
8. A device for clustering the spatiotemporal density peaks of earthquake events, characterized in that: The device comprises: A data processing unit is used to obtain sample data, wherein the sample data includes a spatiotemporal data set of earthquake events and a preset distance cutoff threshold, normalize the spatiotemporal data set of earthquake events, construct a spatial Euclidean distance matrix and a temporal distance matrix between earthquake event samples, and determine a spatiotemporal neighbor set of earthquake event samples in the spatiotemporal data set of earthquake events; A local density calculation unit is used to calculate the spatial weight coefficient and the time weight coefficient of the earthquake event sample according to the spatial and temporal neighbor set of each earthquake event sample in the earthquake event spatial and temporal data set, and calculate the local density of each earthquake event sample under the spatial and temporal constraints according to the spatial weight coefficient and the time weight coefficient; A relative distance calculation unit, used for calculating the relative distance of the seismic event samples according to the time distance between the seismic event samples and the local density of each seismic event sample under the time and space constraints; A potential cluster center selection unit is used to calculate the decision value of each earthquake event sample according to the local density and relative distance of the earthquake event sample to select the potential cluster center of the earthquake event spatiotemporal data set; A cluster center determination unit is used to calculate the spatial similarity and temporal similarity of the cluster centers among the potential cluster centers, determine the comprehensive similarity of the cluster centers according to the spatial similarity and the temporal similarity, and eliminate the potential cluster centers with merging potential according to the comprehensive similarity of the cluster centers to obtain the cluster centers of various clusters; An earthquake event sample allocation unit is used to allocate earthquake event samples whose spatial distances from cluster centers of various clusters within a time window are less than a preset distance cutoff threshold to corresponding clusters to realize window allocation, and to perform secondary allocation on the remaining earthquake event samples after window allocation according to the spatial distances between the earthquake event samples and the adjacent cluster centers to realize search allocation, and to divide the earthquake event samples that are not allocated after search allocation into the cluster to which the high-density nearest neighbor earthquake event samples belong; The outlier identification unit is used to calculate the core density, core distance and boundary density of each cluster for outlier identification in each cluster to obtain the final clustering result of the sample data.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for clustering the spatiotemporal density peak values of earthquake events according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: include: Processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory so that the electronic device performs the seismic event spatiotemporal density peak clustering method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Improved density peak clustering method based on extension correlation function
CN110414583A
Space-time data clustering method
CN113886667A
Density peak clustering method based on natural nearest neighbor and multi-cluster merging
CN115510959A
Basic farmland planning method, product, medium and equipment based on improved density peak clustering algorithm
CN118657403A
Dynamic clustering of sparse data utilizing hash partitions
US20210326361A1
Cited By
Multi-stage spatial-temporal clustering method and system based on fused mahalanobis distance
CN121051489A
Operation and maintenance performance assessment data management system based on dynamic weight optimization
CN121387860A
Social event early warning method and device based on complex network and medium
CN121746146A
A spatio-temporal clustering method and system based on shared neighbors and representative point correction
CN122548350A