Method for clustering time series data and apparatus therefor
Patent Information
- Application Number
- CN202280101605.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2025-07-01
AI Technical Summary
The existing technology cannot accurately cluster the actual generated time series signals. It has poor noise robustness and low efficiency, making it difficult to extract information from high-speed and large-scale real-time data streams.
A method based on pre-clustering is proposed, which traverses sample points in time series data, merges clusters based on distance and mean difference, optimizes the nearest neighbor search process of the DBSCAN algorithm, uses KD trees or ball trees to limit the scale, and improves clustering efficiency.
Effectively classify noise points accurately, significantly reduce the time complexity of clustering, improve the accuracy and efficiency of time series data classification, and adapt to scenarios with changing features.
Smart Images

Figure CN120239855A_ABST
Abstract
Description
Method and device for clustering time series data Technical Field
[0001] The present disclosure relates to clustering technology, and more particularly, to a method and device for clustering time series data. Background Art
[0002] With the development of internet computing technology, real-time data streams have become a crucial form of information and are widely used in fields such as network traffic control, data monitoring systems, and internet finance. Extracting information quickly and efficiently from high-speed, high-volume, real-time data streams has become a major challenge in data stream mining. Cluster analysis is a key technique in data mining. Its core goal is to classify data based on similarity. It is an unsupervised machine learning task, and density-based clustering algorithms are the most effective. Machine learning is a data-driven approach that studies and constructs specialized algorithms that enable computers to learn from data and make predictions, thus quickly and efficiently solving this problem.
[0003] However, the existing technology cannot accurately cluster the actual time series signals, has poor robustness to noise, and is inefficient.
[0004] Therefore, a technique that can accurately cluster time series signals is needed.
[0005] Summary of the Invention
[0006] The purpose of the embodiments of the present disclosure is to provide an effective solution for accurately clustering time series signals. Specifically, the embodiments of the present disclosure provide a method and device for clustering time series data.
[0007] According to a first aspect of the present disclosure, a method for clustering time series data is proposed, comprising: pre-clustering sample points in the time series data, wherein the pre-clustering comprises traversing all sample points in the sample points, clustering the sample points into corresponding clusters based on the distance between the current sample point and the previous sample point of the current sample point, and calculating the mean of all sample points in each of all the generated clusters, determining whether to merge the current cluster and its adjacent clusters based on the mean difference between the current cluster and its adjacent clusters; and clustering all the clusters after the pre-clustering based on the mean.
[0008] In some embodiments, the time series data includes nanopore sequence data and is stored in a memory in floating point format, wherein the nanopore sequence data includes a sequencing signal and a pore current signal.
[0009] In some embodiments, the time series data stored in the memory is converted into an array.
[0010] In some embodiments, the distance is obtained by calculating the L1 norm of the current sample point and the sample point before the current sample point.
[0011] In some embodiments, clustering the sample points into corresponding clusters based on the distance between the current sample point and the previous sample point of the current sample point includes: when the L1 norm of the current sample point and the previous sample point of the current sample point is less than the initial minimum distance ε, marking the current sample point and the previous sample point as the same cluster; when the L1 norm of the current sample point and the previous sample point of the current sample point is greater than or equal to the ε, if the total number of sample points in the current cluster is less than the minimum number of sample points MinPts, marking the current sample point and the previous sample point as the same cluster, otherwise setting the current sample point as a new cluster.
[0012] In some embodiments, based on the mean difference between the current cluster and its adjacent clusters, determining whether to merge the current cluster and its adjacent clusters includes: if the mean difference between the current cluster and its adjacent clusters is less than a threshold eps_c, merging the current cluster and its adjacent clusters; otherwise, keeping the current cluster unchanged.
[0013] In some embodiments, the initial minimum distance ε, the minimum number of sample points MinPts and the threshold eps_c are determined using nearest neighbor search.
[0014] In some embodiments, the initial minimum distance ε, the minimum number of sample points MinPts, and the threshold eps_c are adjustable parameters.
[0015] In some embodiments, the nearest neighbor search is implemented by building a KD tree or a ball tree.
[0016] In some embodiments, clustering all the pre-clustered clusters using the DBSCAN algorithm is performed based on the mean and also based on the standard deviation.
[0017] According to the second aspect of the present disclosure, a device for clustering time series data is also provided, comprising: one or more processors; and one or more memories, wherein the one or more memories store computer-executable instructions, and when the computer-executable instructions are executed by the one or more processors, the one or more processors execute any of the above methods.
[0018] According to a third aspect of the present disclosure, a computer storage medium is further provided, on which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor is caused to perform any of the above methods.
[0019] This disclosure creatively pre-clusters one-dimensional time series signals, optimizes the search process, and can be fully applied to time series data clustering. Specifically:
[0020] 1) Addressing the robustness issues of existing technologies
[0021] The method based on pre-clustering design can effectively and accurately classify noise points. In theory, it can solve the clustering of noise points of various forms and can be effectively applied to applications that do not need to cluster noise points into other categories.
[0022] 2) Addressing the low efficiency of existing technologies
[0023] This invention is based on pre-clustering DBSCAN. The nearest neighbor search process can be optimized for scale constraints by establishing a KD tree or a ball tree. Compared with the DBSCAN algorithm, the improved algorithm has a significantly reduced time complexity and significantly shortens the clustering convergence time for large sample sets, effectively improving the accuracy and efficiency of time series data classification and providing better adaptability to scenarios with constantly changing features. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] For a more complete understanding of the present disclosure and its advantages, reference will now be made to the following description taken in conjunction with the accompanying drawings, in which:
[0025] FIG1 is a schematic diagram showing nanopore sequence data as an example of time series data;
[0026] FIG2 is a schematic diagram showing the DBSCAN algorithm;
[0027] FIG3 is a schematic flow chart illustrating a method for clustering time series data according to an embodiment of the present disclosure;
[0028] FIG4 is a flowchart illustrating a specific implementation of a method for clustering time series data according to an embodiment of the present disclosure;
[0029] 5 to 7 are diagrams showing pre-clustering results of a method for clustering time series data according to an embodiment of the present disclosure;
[0030] FIG8 is a diagram showing a clustering result of a method for clustering time series data according to an embodiment of the present disclosure;
[0031] FIG9 is a diagram showing clustering results of simulated data using the DBSCAN algorithm;
[0032] FIG10 is a diagram showing clustering results of simulation data according to a method for clustering time series data according to an embodiment of the present disclosure;
[0033] FIG11 is a diagram showing the clustering results of the experimental data using the DBSCAN algorithm;
[0034] FIG12 is a diagram showing clustering results of experimental data according to a method for clustering time series data according to an embodiment of the present disclosure;
[0035] FIG13 is a table showing a comparison of processing time using the DBSCAN algorithm and a method for clustering time series data according to an embodiment of the present disclosure;
[0036] FIG14 schematically shows a block diagram of a device 1400 for clustering time series data according to an embodiment of the present disclosure.
[0037] In the drawings, the same or similar structures are marked with the same or similar reference numerals. DETAILED DESCRIPTION
[0038] Other aspects, advantages, and salient features of the disclosure will become apparent to those skilled in the art from the following detailed description of exemplary embodiments of the disclosure, taken in conjunction with the accompanying drawings.
[0039] In this disclosure, the terms "include" and "including" and their derivatives mean inclusion without limitation; the term "or" is inclusive, meaning and / or.
[0040] In this specification, the various embodiments described below for describing the principles of the present disclosure are illustrative only and should not be interpreted in any way as limiting the scope of the disclosure. The following description with reference to the accompanying drawings is used to help fully understand the exemplary embodiments of the present disclosure as defined by the claims and their equivalents. The following description includes a variety of specific details to aid understanding, but these details should be considered merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, for the sake of clarity and brevity, descriptions of well-known functions and structures have been omitted. In addition, throughout the drawings, the same figure numbers are used for similar functions and operations.
[0041] FIG. 1 is a schematic diagram showing nanopore sequence data as an example of time-series data.
[0042] The present disclosure relates to a clustering method for nanopore sequencing signals (i.e., nanopore sequence data). Nanopore sequencing is a process in which an artificially synthesized multipolymer membrane is immersed in an ionic solution. The multipolymer membrane is covered with modified transmembrane channel proteins (nanopores). Different voltages are applied on both sides of the membrane to generate a voltage difference. The DNA chain unwinds through the nanopore protein under the traction of the motor protein, and different bases form characteristic ion current change signals. Nanopore electrical signals are one-dimensional time-series signals. Nanopore sequencing signals refer to changes in the perforation current recorded when a DNA molecule passes through a nanopore. This change is mainly caused by different currents of different bases in the nanopore. The clustering method provided by the present disclosure is mainly to solve the clustering of time-series data such as nanopore sequencing signals. As shown in Figure 1, nanopore sequence data may include sequencing signals and pore current signals. Nanopore sequence data of any form can be collected and stored in a memory in floating-point type as sample points for clustering methods.
[0043] FIG2 is a schematic diagram illustrating the DBSCAN algorithm.
[0044] In the existing technology for clustering time series data, in order to detect real-time time series data (for example, nanopore electrical signals) and cluster them, the density-based clustering analysis algorithm examines the continuity between samples from the perspective of sample density, and continuously expands cluster clusters based on continuous samples to obtain the final clustering results. The DBSCAN algorithm is a clustering algorithm based on density space, which is widely used in the fields of machine learning and data mining. Its clustering principle is generally that the density of each cluster is higher than the density around the cluster, and the density of noise is less than the density of any cluster. As shown in Figure 2, the basic idea of the DBSCAN algorithm is: set a minimum distance ε, and for each point in the data set, draw a circle with the minimum distance ε as the radius. The points in the circle are called neighbors. If the number of neighboring points is less than the threshold MinPts, the current point is set as a boundary molecule. If the number of neighboring points is greater than or equal to MinPts, the point is called a core molecule and marked as a cluster with its neighboring points. Points that are neither core points nor boundary points (i.e., points with zero neighboring points) are called outliers. All neighboring points are looped through simultaneously, and the neighboring points of each neighbor are set to the same cluster. When all points that meet the conditions are traversed, the cluster is incremented by 1 and the next loop begins.
[0045] Although the existing DBSCAN algorithm can cluster time series data, it cannot accurately cluster the time series signals generated by actual nanopores, has poor robustness to noise, and has low clustering efficiency when the sample set is large.
[0046] Therefore, in order to solve the above problems, the present disclosure proposes an improved method for clustering time series data. The method is a pre-clustering-based clustering method improved from the above-mentioned existing DBSCAN algorithm.
[0047] FIG3 is a schematic flowchart illustrating a method 300 for clustering time series data according to an embodiment of the present disclosure.
[0048] As shown in FIG3 , a method 300 for clustering time series data according to an embodiment of the present disclosure may include:
[0049] Step S310: pre-clustering the sample points in the time series data; and
[0050] Step S320: clustering all pre-clustered clusters using the DBSCAN algorithm based on the mean of all sample points in the cluster.
[0051] According to an embodiment, the pre-clustering in step S310 may include:
[0052] Step S311: traverse all sample points and cluster the sample points into corresponding clusters based on the distance between the current sample point and the previous sample point of the current sample point; and
[0053] Step S312: Calculate the mean of all sample points in each of all generated clusters, and determine whether to merge the current cluster with its adjacent clusters based on the mean difference between the current cluster and its adjacent clusters.
[0054] 4 , a specific example of performing pre-clustering is shown.
[0055] An initial minimum distance eps_p can be set (S410). For the input one-dimensional time series data, each sample point pt is traversed from the beginning to the end (S420). If the L1 norm of the current sample point pt and the previous sample point is less than eps_p (S430: Yes), the current sample point pt can be marked as the same cluster as the previous sample point (S431). Otherwise (S430: No), it is set as a new cluster and the next sample point is continued. For numerical values, the L1 norm is the absolute value of the difference between the current sample point and the previous sample point. Here, the L1 norm is used as the distance, but those skilled in the art can also envision other forms of distance. The results of the above distance-based pre-clustering operation are shown in Figure 5.
[0056] During the traversal process, even if the L1 norm of the current sample point pt and the previous sample point is greater than eps_p (S430: No), if the total number of sample points in the current cluster is less than the minimum number of sample points MinPts (S432: Yes), the current sample point pt can be marked as the same cluster as the previous sample point (S431). Otherwise (S432: No), it is set as a new cluster (S433) and the next sample point is continued. The result of the above pre-clustering operation based on the total number of sample points is shown in Figure 6.
[0057] The clustering result of the above operation may be stored in a database (S440), and the clustering result may include a cluster start index, a cluster end index, a cluster, the mean of all sample points in the cluster, and the standard deviation of all points in the cluster.
[0058] The mean of all sample points in each cluster obtained through the above operation can be calculated. If the difference between the mean of the cluster and the adjacent cluster is less than the threshold eps_c (S450: Yes), the two clusters can be merged (S451). Otherwise (S450: No), the cluster remains unchanged (S452). The result of the above pre-clustering operation based on mean difference is shown in Figure 7.
[0059] The clustering result of the above pre-clustering operation may be updated according to the result of the cluster merging operation.
[0060] Based on the results of the pre-clustering operation, all clusters can be clustered using the DBSCAN algorithm based on the mean of all sample points in each cluster (S460). According to an embodiment, clustering can be further performed based on the standard deviation of all sample points in each cluster. The clustering results of the above clustering operation can be updated. The results of the above clustering operation are shown in Figure 8.
[0061] In the above example, eps_p controls the difference between adjacent sample points, MinPts is the minimum number of points in a cluster, and eps_c controls the difference between the average values of adjacent clusters in pre-clustering. eps_p, MinPts, and eps_c are adjustable parameters. For example, the floating-point data stored in the memory can be converted into an array, and the optimal parameters can be determined using the nearest neighbor search. The nearest neighbor search is an optimization problem of finding the point closest to (or most similar to) a given point in a given set. The nearest neighbor search can be implemented by establishing a KD tree or a ball tree, etc. The present disclosure applies the above method to cluster the real-time time series data based on the determined optimal parameters and obtains the clustering results.
[0062] FIG. 9 is a diagram showing clustering results of simulation data using the DBSCAN algorithm.
[0063] FIG10 is a diagram showing clustering results of simulation data by a method for clustering time series data according to an embodiment of the present disclosure.
[0064] By comparing FIG9 and FIG10, it can be seen that the clustering method of the present invention has a better clustering effect on the sample points in the simulated data sample set than the clustering effect using the DBSCAN algorithm, and has good robustness to the noise in the simulated data sample set.
[0065] FIG11 is a diagram showing the clustering results of the experimental data using the DBSCAN algorithm.
[0066] FIG12 is a diagram showing clustering results of experimental data by a method for clustering time series data according to an embodiment of the present disclosure.
[0067] Similarly, by comparing Figure 11 and Figure 12, it can be seen that the clustering effect of the clustering method disclosed in the present invention on the sample points in the experimental data sample set is better than the clustering effect using the DBSCAN algorithm, and has good robustness to the noise in the experimental data sample set.
[0068] FIG13 is a table showing a comparison of processing time using the DBSCAN algorithm and a method for clustering time series data according to an embodiment of the present disclosure.
[0069] As can be seen from the table in Figure 13, when the number of sample points in the sample set is 1000, 30000 and 50000, the clustering time using the DBSCAN algorithm is 0.01s, 6.238s and 29.395s respectively, while the clustering time using the clustering method of the present invention is 0.0038s, 0.0409s and 0.0719s respectively. The clustering speed using the clustering method of the present invention is 2.63 times, 152.53 times and 408.80 times higher than that using the DBSCAN algorithm. Therefore, the clustering time using the clustering method of the present invention is significantly reduced compared to the clustering time using the DBSCAN algorithm, and the efficiency is significantly improved.
[0070] In summary, the present invention creatively pre-clusters one-dimensional time series data, which is conducive to improving the accuracy and efficiency of subsequent DBSCAN-based clustering and can solve the problem that current time series data clustering is sensitive to noise and inefficient.
[0071] The present invention can be implemented based on the development environment shown in Table 1. Although the present invention can run stably on machines with standard CPU configurations, the algorithm module involves reading, writing, and storing large amounts of data, as well as model training and testing, requiring extensive data computation and manipulation. Running the system in a high-performance CPU or GPU-based hardware and software environment significantly improves efficiency and stability. Data acquisition is performed using a single-channel nanopore sequencer and a PC. The target library data is collected using the nanopore sequencer, and the collected signal data and other information are saved as a .dat file structure and stored in real time on the PC hard drive.
[0072]
[0073] Table 1. Development environment
[0074] FIG14 schematically illustrates a block diagram of a device 1400 for clustering time series data according to an embodiment of the present disclosure. The device shown in FIG14 can be any device with processing capabilities. It should be noted that the device shown in FIG14 is merely an example and should not limit the functionality or scope of use of the embodiments of the present application.
[0075] As shown in FIG. 14 , the device 1400 according to this embodiment includes a central processing unit (CPU) 1401 (e.g., the CPU in Table 1: 1403) can perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 1402 or the programs loaded from the storage unit 1408 into the random access memory (RAM) 1403. The RAM 1403 also stores various programs and data required for the operation of the device 1400. The CPU or GPU 1401, the ROM 1402, and the RAM 1403 are connected to each other via a bus 1404. An input / output (I / O) interface 1405 is also connected to the bus 1404.
[0076] Device 1400 may also include one or more of the following components connected to I / O interface 1405: an input section 1406 including a keyboard or mouse, etc.; an output section 1407 including a cathode ray tube (CRT) or liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1408 including a hard disk, etc.; and a communication section 1409 including a network interface card, such as a LAN card or modem. Communication section 1409 performs communication processing via a network, such as the Internet. A drive 1410 is also connected to I / O interface 1405 as needed. Removable media 1411, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, etc., is installed in drive 1410 as needed, so that computer programs read from the removable media can be installed in storage section 1408 as needed.
[0077] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1409, and / or installed from a removable medium 1411. When the computer program is executed by the central processing unit (CPU) 1401, the above-mentioned functions defined in the device of the embodiment of the present application are executed.
[0078] It should be noted that the computer-readable medium described in this disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code embodied on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical cable, RF, or any suitable combination thereof.
[0079] The method of the present invention and the devices involved have been described above in conjunction with the embodiments. Those skilled in the art will appreciate that the methods shown above are merely exemplary. The method of the present invention is not limited to the steps and sequence shown above. The device shown above may include more modules, for example, modules that can be developed or developed in the future and can be used for the device, etc. The various identifiers shown above are merely exemplary and not restrictive, and the present disclosure is not limited to the specific information elements that serve as examples of these identifiers. Those skilled in the art can make many improvements, modifications, and substitutions based on the teachings of the illustrated embodiments, and can extend the method to other fields where time series clustering is required.
[0080] The program running on the device according to the present disclosure may be a program that controls a central processing unit (CPU) to enable a computer to implement the functions of the embodiments of the present disclosure. The program or the information processed by the program may be temporarily stored in a volatile memory (such as a random access memory RAM), a hard disk drive (HDD), a non-volatile memory (such as a flash memory), or other memory systems.
[0081] The program for realizing each embodiment function of the present disclosure can be recorded on a computer-readable recording medium. The corresponding function can be realized by making a computer system read the program recorded on the recording medium and executing these programs. The so-called "computer system" herein can be a computer system embedded in the device, and can include an operating system or hardware (such as a peripheral device). "Computer-readable recording medium" can be a semiconductor recording medium, an optical recording medium, a magnetic recording medium, a short-term dynamic storage program recording medium or any other recording medium that is computer-readable.
[0082] The various features or functional modules of the devices used in the above embodiments can be implemented or executed by circuits (e.g., single-chip or multi-chip integrated circuits). The circuits designed to perform the functions described in this specification may include a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of the above devices. The general-purpose processor may be a microprocessor, or any existing processor, controller, microcontroller, or state machine. The above circuits may be digital circuits or analog circuits. In the case where new integrated circuit technologies have emerged to replace existing integrated circuits due to advances in semiconductor technology, one or more embodiments of the present disclosure may also be implemented using these new integrated circuit technologies.
[0083] As described above, the embodiments of the present disclosure have been described in detail with reference to the accompanying drawings. However, the specific structure is not limited to the above-mentioned embodiments, and the present disclosure also includes any design changes that do not deviate from the main purpose of the present disclosure. In addition, various modifications can be made to the present disclosure within the scope of the claims, and embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present disclosure. In addition, components with the same effect described in the above-mentioned embodiments can be replaced with each other.
Claims
1. A method for clustering time series data, comprising: Pre-clustering the sample points in the time series data, wherein the pre-clustering includes: Traversing all sample points among the sample points, clustering the sample points into corresponding clusters based on the distance between the current sample point and the previous sample point of the current sample point; and Calculating a mean value for all sample points in each of all generated clusters, and determining whether to merge the current cluster with its adjacent clusters based on a mean difference between the current cluster and its adjacent clusters; and All the pre-clustered clusters are clustered based on their means.
2. The method according to claim 1, wherein The time series data includes nanopore sequence data and is stored in a memory in a floating point format. The nanopore sequence data includes a sequencing signal and a pore current signal.
3. The method according to claim 2, further comprising: The time series data stored in the memory is converted into an array.
4. The method according to claim 1, wherein The distance is obtained by calculating the L1 norm value between the current sample point and the sample point before the current sample point.
5. The method according to claim 4, wherein Clustering the sample points into corresponding clusters based on the distance between the current sample point and a previous sample point of the current sample point includes: When the L1 norm of the current sample point and the previous sample point of the current sample point is less than the initial minimum distance ε, the current sample point and the previous sample point are marked as the same cluster; When the L1 norm of the current sample point and the previous sample point of the current sample point is greater than or equal to the ε, if the total number of sample points in the current cluster is less than the minimum number of sample points MinPts, the current sample point and the previous sample point are marked as the same cluster; otherwise, the current sample point is set to a new cluster.
6. The method according to claim 5, wherein: Determining whether to merge the current cluster and the adjacent clusters based on the mean difference between the current cluster and the adjacent clusters includes: If the mean difference between the current cluster and its adjacent clusters is less than the threshold eps_c, the current cluster and its adjacent clusters are merged; Otherwise, the current cluster remains unchanged.
7. The method according to claim 6, further comprising: The initial minimum distance ε, the minimum number of sample points MinPts and the threshold eps_c are determined by nearest neighbor search.
8. The method according to claim 7, wherein: The initial minimum distance ε, the minimum number of sample points MinPts, and the threshold eps_c are adjustable parameters.
9. The method according to claim 8, further comprising: The nearest neighbor search is implemented by establishing a KD tree or a ball tree.
10. The method according to claim 1, wherein Based on the mean, all clusters after the pre-clustering are clustered using the DBSCAN algorithm and also based on the standard deviation.
11. A device for clustering time series data, comprising: one or more processors; as well as One or more memories storing computer-executable instructions which, when executed by the one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 10.
12. A computer storage medium having computer executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 10.