Analysis and Governance Method and Product for Periodic Time-Series Data

By cutting, encoding, clustering and Kalman filter processing of periodic timing data, the problems of noise and deviation in timing data are solved, and data quality improvement and event moments are achieved.

CN119248760BActive Publication Date: 2025-06-17OPEN UNIVERSITY OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411305632.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2025-06-17
Estimated Expiration
2044-09-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively manage noise and deviations in timing data, resulting in a decline in data quality and it is difficult to accurately reflect the moment when events or actions occur.

Method used

By cutting and encoding periodic time series data, data clustering and Kalman filters are used to reorganize and reduce data noise, correct data deviations, and improve data quality.

Benefits of technology

It realizes effective governance of periodic timing data, corrects data deviations, improves data quality, and makes the data more accurately reflect the moment when events or actions occur.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119248760B_ABST
    Figure CN119248760B_ABST
Patent Text Reader

Abstract

The present invention provides a method and product for analyzing and governing periodic time-series data, relating to the technical field of data analysis and governance. In an embodiment of the present invention, time-series data is cut according to the periodicity of the data, and the data is encoded according to the occurrence time of the periodic time-series data, and the encoded time-series data is clustered. In the clustering result obtained by clustering based on the encoded time-series data, abnormal data will form a separate cluster because it cannot correspond to the time slots of data in other periods. Thus, abnormal data can be discovered and excluded based on the clustering result. For the remaining data after excluding abnormal data, the data in each time slot can be denoised separately to correct the deviation of the data corresponding to the same time slot, so as to achieve the purpose of improving data quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of data analysis and governance, and in particular, to a method and product for analyzing and governing periodic time series data. Background Art

[0002] Time series data refers to a sequence composed of the moments when a certain event or action continuously occurs. For example, the time series of the departure moments of subway lines, or the time series of the regular startup moments of devices. Periodic time series data refers to time series data with a certain periodicity. For example, the subway departs at 5:30 every morning, or the lawn sprinkler device is turned on regularly in the morning and afternoon every day. Time series data is obtained by a collection device or a receiving device and transmitted through a network to the data warehouse. During this process, there are often noises, resulting in data deviation and making it difficult to accurately reflect the moment when an event or action occurs.

[0003] Therefore, there is an urgent need for a method to govern time series data currently. Summary of the Invention

[0004] The embodiments of the present invention provide a method for analyzing and governing periodic time series data, which cuts and encodes the time series data according to the periodicity of the data, reorganizes and denoises the data through data clustering and a Kalman filter, thereby correcting data deviation and achieving the purpose of improving data quality.

[0005] In a first aspect of the embodiments of the present invention, a method for analyzing and governing periodic time series data is provided, and the method includes:

[0006] Cut the input periodic time series data according to the period, and encode the time series data within each period according to the occurrence time of the periodic time series data to obtain encoded time series data;

[0007] Cluster the encoded time series data, exclude abnormal data, and locate the time slots where the periodic time series data occurs;

[0008] Denoise the encoded time series data in each time slot;

[0009] Re - splice and reverse - encode the denoised data to obtain output data.

[0010] Optionally, encoding the time series data within each period includes:

[0011] Encoding the time series data of each period respectively based on the offset of the time corresponding to the time series data within the period relative to the starting point of the period time.

[0012] Optionally, clustering the encoded time series data, excluding abnormal data, and locating the time slots where the periodic time series data occurs includes:

[0013] Cluster the encoded timing data for multiple cycles to obtain a clustering result; one cluster in the clustering result corresponds to one time slot.

[0014] Based on the clustering result, determine the clusters with the number of data less than the first preset threshold as abnormal clusters, and remove the abnormal data in the corresponding time slots.

[0015] Optionally, the method further includes:

[0016] When the cycle with the most encoded timing data among all cycles contains N encoded timing data, perform M + 2 times of clustering on the encoded timing data, where M ≥ 1, and specify the target number of clusters as {N - 1, N, N + 1, …, N + M} during the clustering process;

[0017] After each clustering is completed, calculate the sum of variances of all encoded timing data to the cluster centers to obtain a variance set {S N-1 , S N , S N+1 , …, S N+M};

[0018] Determine a value P between N - 1 and N + M such that the value of |S P-1 - S P | is greater than the second preset threshold, and the value of |S P - S P+1 | is less than the second preset threshold, and the second preset threshold ranges from 1.5 to 3.5;

[0019] Take P as the optimal number of time slots;

[0020] Clustering the encoded timing data for multiple cycles includes:

[0021] Take the optimal number of time slots as the number of clusters and perform clustering.

[0022] Optionally, the method further includes:

[0023] When it is impossible to determine a value P such that the value of |S P-1 - S P | is greater than the second preset threshold, and the value of |S P - S P+1 | is less than the second preset threshold, take the value N as the optimal number of time slots.

[0024] Optionally, denoising the encoded timing data in each time slot includes:

[0025] For the encoded time-series data falling into the same time slot and having similar occurrence times, the Kalman filter is used to correct the encoded time-series data in the same time slot one by one to obtain the denoised data.

[0026] In the second aspect of the embodiments of the present invention, an analysis and governance device for periodic time-series data is provided. The device includes:

[0027] An encoding module, configured to cut the input periodic time-series data by period and encode the time-series data within each period to obtain encoded time-series data;

[0028] A clustering module, configured to cluster the encoded time-series data, exclude abnormal data, and locate the time slots where the periodic time-series data occurs;

[0029] A denoising module, configured to perform denoising on the encoded time-series data in each time slot;

[0030] A reverse encoding module, configured to splice and reverse-encode the denoised data again to obtain output data.

[0031] Optionally, the encoding module is specifically configured to encode the time-series data of each period respectively based on the offset of the time corresponding to the time-series data within the period relative to the starting point of the period time.

[0032] Optionally, the clustering module is specifically configured to:

[0033] Cluster the encoded time-series data of multiple periods to obtain a clustering result; one cluster in the clustering result corresponds to one time slot;

[0034] Based on the clustering result, determine the clusters with the number of data less than the first preset threshold in the clusters as abnormal clusters, and remove the abnormal data in the corresponding time slots.

[0035] Optionally, the device further includes:

[0036] A first determination module, configured to perform M + 2 times of clustering on the encoded time-series data when the period with the most encoded time-series data among all periods contains N encoded time-series data, M≥1. When clustering, the target number of clusters is specified as {N - 1, N, N + 1, …, N + M} respectively; after each clustering is completed, calculate the sum of variances of all encoded time-series data to the cluster centers to obtain a variance set {S N-1 , S N , S N+1 , …, S N+M}; determine a value P between N - 1 and N + M such that |S P-1 - S PThe value of | is greater than the second preset threshold, and |S P -S P+1 The value of | is less than the second preset threshold, and the second preset threshold ranges from 1.5 to 3.5; Take P as the optimal number of time slots;

[0037] The clustering module is specifically configured to: Take the optimal number of time slots as the number of clusters and perform clustering.

[0038] Optionally, the device further includes:

[0039] A second determination module, configured to, when it is impossible to determine a value P such that |S P-1 -S P The value of | is greater than the second preset threshold, and |S P -S P+1 The value of | is less than the second preset threshold, take the value N as the optimal number of time slots.

[0040] Optionally, the noise reduction module is specifically configured to:

[0041] For the encoded time series data that fall into the same time slot with adjacent occurrence times, use the Kalman filter to correct the encoded time series data in the same time slot one by one to obtain the noise-reduced data.

[0042] A third aspect of the embodiments of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the method for analyzing and managing periodic time series data as described in the first aspect of the present invention.

[0043] A fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for analyzing and managing periodic time series data as described in the first aspect of the present invention.

[0044] A fifth aspect of the embodiments of the present invention provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, it implements the steps in the method for analyzing and managing periodic time series data as described in the first aspect of the present invention.

[0045] In the embodiments of the present invention, the time-series data is cut according to the periodicity of the data, and the data is encoded according to the occurrence time of the periodic time-series data. After clustering the encoded time-series data, in the clustering results obtained based on the encoded time-series data, the abnormal data will form a separate cluster because it cannot correspond to the time slots of the data in other cycles. Therefore, the abnormal data can be discovered and excluded based on the clustering results. For the remaining data after excluding the abnormal data, the data in each time slot can be denoised separately to correct the deviation of the data corresponding to the same time slot, so as to achieve the purpose of improving the data quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 is a flowchart of the steps of the method for analyzing and managing periodic time-series data provided by the embodiments of the present invention;

[0048] Figure 2 is a schematic diagram of the clustering results obtained in the method for analyzing and managing periodic time-series data provided by the embodiments of the present invention;

[0049] Figure 3 is a schematic diagram of the results of using the elbow method to find the optimal number of time slots in the method for analyzing and managing periodic time-series data provided by the embodiments of the present invention;

[0050] Figure 4 is a schematic diagram of the data denoising results in the method for analyzing and managing periodic time-series data provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0052] The embodiments of the present invention provide a method for analyzing and managing periodic time-series data, as Figure 1 shown, the method may include the following steps:

[0053] S101, cut the input periodic time-series data according to the period, and encode the time-series data in each period according to the occurrence time of the periodic time-series data to obtain the encoded time-series data.

[0054] S102. Cluster the encoded time-series data, exclude abnormal data, and locate the time slots where periodic time-series data occurs.

[0055] S103. Denoise the encoded time-series data in each time slot.

[0056] S104. Re - splice and reverse - encode the denoised data to obtain the output data.

[0057] The method provided by the embodiments of the present invention can be applied to a data processing device, which can be configured on any electronic device to receive input periodic time - series data, process the received data, and output the processed data. The input periodic time - series data can be uploaded through the external interface of the electronic device, or stored locally in the electronic device, or obtained through a network device communicatively connected to the electronic device.

[0058] In the embodiments of the present invention, in step S101, first, the periodicity of the data can be determined, that is, the time length included in each period. It can be determined based on the knowledge of business logic or the statistical period obtained by analyzing the data. Secondly, the starting point of each period can be determined. It can be a fixed time interval (such as the beginning of each day, week, or month), or determined according to certain feature points in the data. Once the starting point of the period is determined, the data set can be cut into multiple subsets according to these points. Each subset contains the data of a complete period. During the process of cutting the data, if the start or end of the data set does not completely contain a complete period, these data can be discarded, or these data can be processed separately, or these data can participate in the subsequent clustering steps as a complete period.

[0059] In the embodiments of the present invention, the input periodic time - series data is cut by period. Specifically, the input time - series data is divided by period. The period length can be hours, days, weeks, etc., which is determined by the periodic characteristics of the time - series data. Thus, according to the periodicity of the time - series data, the periodic time - series data can be cut. For example, if the subway departs at 5:30 every morning and departs every 5 minutes, one day can be used as a period, and the time - series data of the subway departures within one day can be used as the time - series data within one period, and the periodic time - series data can be cut with one day as a period. In the cut data, the time - series data generated within one day is a time - series data group.

[0060] In the embodiments of the present invention, encoding the time - series data can include: encoding the time - series data of each period respectively based on the offset of the time corresponding to the time - series data within the period relative to the starting point of the period time, so as to represent each time - series data based on the occurrence time of each time - series data.

[0061] Specifically, in the embodiments of the present invention, the time series data within a period can be encoded based on the accuracy of the time series data (usually seconds or minutes), and the corresponding encoding can be calculated according to the offset of each time series data (moment) within the period relative to the starting point of the period. For example, taking a day as a period, starting at 0:00 and ending at 24:00, with an accuracy of minutes, there are 1440 encoding positions within a period. An event or behavior occurring at 12:00 noon is encoded as 720. Each moment within the period when an event or behavior occurs can be converted into the corresponding encoding. After encoding, the time series of events or behaviors occurring within each period is represented by a set of integer encodings.

[0062] In the embodiments of the present invention, the k-means algorithm can be used to cluster the encoded time series data to determine the time slots for finding the moments when events or behaviors occur, and each cluster corresponds to a time slot.

[0063] The k-means algorithm can divide the samples in the encoded time series data into K clusters, making the samples within the clusters as similar as possible and the samples between the clusters as different as possible. Specifically, the k-means algorithm can include the following steps:

[0064] Step 1: Randomly select K data points as the initial cluster centers.

[0065] Step 2: For each sample in the dataset, calculate its distances from the K cluster centers and assign it to the cluster represented by the nearest cluster center.

[0066] Step 3: For each cluster, calculate the mean of all samples within the cluster and use this mean as the new cluster center.

[0067] Repeat steps 2 and 3 until the cluster centers no longer change or reach the preset number of iterations. At this time, each sample is assigned to a cluster, and these clustering results are output.

[0068] In the embodiments of the present invention, for the time series data of events or behaviors occurring within each period, after encoding according to the occurrence moment, and then clustering the encoded time series data corresponding to multiple periods, it can be understood that the encoded time series data at the same moment in multiple periods will form a cluster, and the moment when the data in this cluster occurs is a time slot that will periodically generate time series data.

[0069] For example: The subway departs at 5:30 every morning and departs every 5 minutes. Then, one day can be taken as a period, and the time series data of the subway departures within one day can be regarded as the time series data within a period. After clustering, it can be determined that 5:30 am every day is a time slot, and a time slot is determined every 5 minutes.

[0070] In the embodiments of the present invention, during the clustering process, the encoded time-series data at the same moment in multiple cycles will form a cluster, and the moment when the data corresponding to this cluster occurs is a time slot that periodically generates time-series data. For the abnormal data generated at a certain moment within a certain cycle, since there is no corresponding data at its occurrence moment in other cycles, this abnormal data will form a separate cluster. Based on this, abnormal data can be discovered and excluded.

[0071] Specifically, in the embodiments of the present invention, step S102 may include the following sub-steps:

[0072] S1021, perform clustering on the encoded time-series data of multiple cycles to obtain a clustering result; one cluster in the clustering result corresponds to a time slot.

[0073] S1022, based on the clustering result, determine the clusters with the number of data in the cluster less than the first preset threshold as abnormal clusters, and remove the abnormal data in the corresponding time slots.

[0074] Due to reasons such as equipment failures or program errors, there will be abnormal values in the time-series data. Considering that abnormal values often deviate from the normal data range, after clustering, the abnormal values will form a separate category. In the embodiments of the present invention, these abnormal values can be discovered by clustering the encoded time-series data, and then the abnormal values can be excluded. Specifically, all the clusters obtained after clustering can be checked one by one, and the clusters with the number of moments within the cluster less than the preset threshold are determined as abnormal clusters, and the values in the abnormal clusters are determined as abnormal values and excluded.

[0075] In the embodiments of the present invention, the first preset threshold can be set according to the actual situation, specifically as Max(total number of cycles × 0.05, 1). If the product of the number of cycles and 0.05 is greater than 1, the first preset threshold is the product of the number of cycles and 0.05; if the result is less than 1, the first preset threshold is 1.

[0076] As Figure 2 shown, it shows a schematic diagram of the clustering result, where the data from cycle 1 to cycle n is obtained by cutting the periodic time-series data. As Figure 2 can be seen, the events that occur periodically at normal moments will be clustered into one cluster, while the abnormal values will form a separate cluster, from which the abnormal values can be intuitively and accurately discovered.

[0077] In the embodiments of the present invention, after excluding the abnormal data, the number of the remaining clusters is the number of time slots when the periodic time-series data occurs, and each cluster corresponds to a time slot.

[0078] In the embodiments of the present invention, incomplete cycles with missing data can also be filled based on the located time slots to further improve data quality. Specifically, after determining the time slots, the data of each cycle after cutting can be searched according to the time slots. In the case where data is missing in a certain time slot, corresponding data can be filled in the corresponding time slot.

[0079] In the embodiments of the present invention, during the clustering process, it is also necessary to determine the optimal number of time slots. Specifically, the optimal number of time slots can be determined based on the cycle with the largest number of time points among all cycles.

[0080] In the embodiments of the present invention, the elbow method can be used to find the optimal number of time slots. The specific method includes: when the cycle with the most encoded time series data among all cycles contains N encoded time series data, perform M + 2 clusterings on the encoded time series data, where M ≥ 1. During the clustering process, the target number of clusters is specified as {N - 1, N, N + 1, …, N + M} respectively; after each clustering is completed, calculate the sum of variances of all encoded time series data to the cluster centers to obtain a variance set {S N-1 , S N , S N+1 , …, S N+M}; determine a value P between N - 1 and N + M such that the value of |S P-1 - S P | is greater than the second preset threshold, and the value of |S P - S P+1 | is less than the second preset threshold. The second preset threshold ranges from 1.5 to 3.5; take P as the optimal number of time slots. Take this optimal number of time slots as the number of clusters and perform clustering.

[0081] Specifically, as Figure 3 shown, it shows a schematic diagram of the result of using the elbow method to find the optimal number of time slots. Figure 3 In N-1 , S N , S N+1 , …, S N+M}, where the abscissa represents the target number of clusters {N - 1, N, N + 1, …, N + M}, and the ordinate represents the variance set {S N-1 , S N , S N+1 , …, S N+M}. When the target number of clusters is 3, the sum of variances of all encoded time series data to the cluster centers obtained based on the clustering result meets the requirements of the optimal number of time slots, then take the target number of clusters 3 as the optimal number of time slots.

[0082] In the embodiments of the present invention, when it is impossible to determine a value P such that |S P-1 - S P| The value is greater than the second preset threshold, and |S P -S P+1 When the value of | is less than the second preset threshold, the numerical value N is used as the optimal number of time slots.

[0083] In the embodiments of the present invention, it is also considered that since the input data is periodic, that is, events or behaviors occur at similar times across cycles. After the foregoing clustering process, similar times will fall into the same cluster. Due to measurement errors, the times at which events or behaviors occur in the same cluster will fluctuate.

[0084] Based on this, in the embodiments of the present invention, in step S103, the encoded time-series data in each time slot is denoised. Specifically, for the encoded time-series data that falls into the same time slot with similar occurrence times, the Kalman filter is used to correct the encoded time-series data in the same time slot one by one to obtain the denoised data. Specifically, the Kalman filter is a linear filter that requires a linear relationship between the state at the current time and the state at the previous time. The occurrence times of events or behaviors between different cycles can be described as X t =X t-1 , that is, in different cycles, the occurrence times of events or behaviors are theoretically the same, and this linear relationship exactly meets the precondition for using the Kalman filter. Therefore, the Kalman filter can be used to correct the values in the same cluster one by one. As Figure 4 shown, it shows a schematic diagram of the data denoising result. Among them, the red dotted line is the reference true value, the green solid line is the measured value with noise (that is, the occurrence time of the input data in the embodiments of the present invention), and the blue broken line is the corrected value (that is, the occurrence time of the output data in the embodiments of the present invention). It can be seen that the corrected value is overall closer to the true value. The Kalman filter can process noisy signals through system input and output observation data to obtain the optimal estimate of the system state or the true signal.

[0085] In the embodiments of the present invention, in step S104, the time-series data after the above steps S101 to S103 is re-spliced in the order of cycles, and the data is reverse-encoded to convert the encoding into a time value, and the final data processing result is output.

[0086] In the embodiments of the present invention, the data processing device first cuts the input time-series data by cycle and encodes the time-series data within the cycle. Secondly, the encoded time-series data is clustered to locate the time slots where periodic events or actions occur. Then, based on the clustering result, abnormal time-series data is checked and excluded (such as Figure 2The clusters in it 4). After judging the normal clusters, the Kalman filter is used to denoise the data within the clusters. Finally, the data is re-spliced and reverse-encoded to obtain the output data. The purpose of improving the data quality can be achieved.

[0087] Based on the same inventive concept, an embodiment of the present invention further provides an analysis and governance device for periodic time-series data, and the device includes:

[0088] An encoding module, configured to cut the input periodic time-series data by period, and encode the time-series data within each period to obtain encoded time-series data;

[0089] A clustering module, configured to cluster the encoded time-series data, exclude abnormal data, and locate the time slots where the periodic time-series data occurs;

[0090] A denoising module, configured to denoise the encoded time-series data in each time slot;

[0091] A reverse encoding module, configured to re-splice and reverse-encode the denoised data to obtain output data.

[0092] Optionally, the encoding module is specifically configured to encode the time-series data of each period respectively based on the offset of the time corresponding to the time-series data within the period relative to the starting point of the period time.

[0093] Optionally, the clustering module is specifically configured to:

[0094] Cluster the encoded time-series data of multiple periods to obtain a clustering result; one cluster in the clustering result corresponds to one time slot;

[0095] Based on the clustering result, determine the clusters with the number of data in the cluster less than the first preset threshold as abnormal clusters, and remove the abnormal data in the corresponding time slots.

[0096] Optionally, the device further includes:

[0097] A first determination module, configured to perform M + 2 clusterings on the encoded time-series data when the period with the most encoded time-series data among all periods contains N encoded time-series data, M≥1, and specify the target number of clusters as {N - 1, N, N + 1, …, N + M} during the clustering process; after each clustering is completed, calculate the sum of the variances of all the encoded time-series data to the cluster centers to obtain a variance set {S N-1 , S N , S N+1 , …, S N+M}; determine a value P between N - 1 and N + M such that |S P-1 - S PThe value of | is greater than the second preset threshold, and |S P -S P+1 The value of | is less than the second preset threshold, and the second preset threshold ranges from 1.5 to 3.5; Take P as the optimal number of time slots;

[0098] The clustering module is specifically configured to: Take the optimal number of time slots as the number of clusters for clustering.

[0099] Optionally, the device further includes:

[0100] The second determination module is used to, when it is impossible to determine a value P such that |S P-1 -S P The value of | is greater than the second preset threshold, and |S P -S P+1 The value of | is less than the second preset threshold, take the value N as the optimal number of time slots.

[0101] Optionally, the noise reduction module is specifically configured to:

[0102] For the encoded timing data that fall into the same time slot with close occurrence times, use the Kalman filter to correct the encoded timing data in the same time slot one by one to obtain the noise-reduced data.

[0103] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the steps in the analysis and governance method of periodic timing data as described in any one of the above embodiments.

[0104] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by the processor, it implements the steps in the analysis and governance method of periodic timing data as described in any one of the above embodiments.

[0105] Based on the same inventive concept, an embodiment of the present invention provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by the processor, they implement the steps in the analysis and governance method of periodic timing data as described in any one of the above embodiments.

[0106] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same and similar parts among the embodiments can be referred to each other.

[0107] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the embodiments of the present invention can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0108] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (apparatuses), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable terminal devices generate a device for realizing the functions specified in Figure 1 one or more of the processes Figure 1 or blocks or the combination of multiple processes and / or blocks.

[0109] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable terminal devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one or more of the processes Figure 1 or blocks or the combination of multiple processes and / or blocks.

[0110] These computer program instructions can also be loaded onto a computer or other programmable terminal devices, so that a series of operation steps are executed on the computer or other programmable terminal devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal devices provide steps for realizing the functions specified in Figure 1 one or more of the processes Figure 1 or blocks or the combination of multiple processes and / or blocks.

[0111] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0112] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising said element.

[0113] The above has introduced in detail a method and product for analyzing and governing periodic time-series data provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for analyzing and managing periodic time series data, characterized in that: The method comprises: The input periodic time series data is cut into periods, and the time series data in each period is encoded according to the occurrence time of the periodic time series data to obtain the encoded time series data; Cluster the encoded time series data, exclude abnormal data, and locate the time slot where periodic time series data occurs; De-noising the encoded time series data in each time slot; Rejoin and reverse encode the denoised data to obtain output data; According to the occurrence time of periodic time series data, the time series data in each period is encoded, including: Based on the offset of the time corresponding to the time series data in the period relative to the starting point of the period time, the time series data of each period are respectively encoded to represent each time series data based on the time when each time series data occurs; Cluster the encoded time series data, exclude abnormal data, and locate the time slots where periodic time series data occurs, including: Clustering the encoded time series data of multiple periods to obtain a clustering result; a cluster in the clustering result corresponds to a time slot; Based on the clustering result, a cluster whose data quantity is less than a first preset threshold is determined as an abnormal cluster, and the abnormal data in the corresponding time slot is removed; The method further comprises: When the period with the most encoded time series data among all periods contains N encoded time series data, the encoded time series data are clustered M+2 times, M≥1, and the number of target clusters is specified as {N-1, N, N+1,…, N+M} during the clustering process; After each clustering is completed, the sum of the variances of all encoded time series data to the cluster center is calculated to obtain the variance set {S N-1 , S N , S N+1 ,…, S N+M }; Determine a value P between N-1 and N+M such that |S P-1 -S P The value of | is greater than the second preset threshold, and |S P -S P+1 The value of | is less than a second preset threshold value, and the second preset threshold value is between 1.5 and 3.5; Let P be the optimal number of slots at the moment; Clustering of encoded time series data of multiple periods, including: The optimal number of time slots is used as the number of clusters for clustering.

2. The method for analyzing and managing periodic time series data according to claim 1, characterized in that: The method further comprises: When it is impossible to determine a value P such that |S P-1 -S P The value of | is greater than the second preset threshold, and |S P -S P+1 When the value of | is less than the second preset threshold, the value N is used as the optimal number of time slots.

3. The method for analyzing and managing periodic time series data according to claim 1, characterized in that: De-noise the encoded time series data in each slot at each moment, including: For the encoded time series data that fall into the same time slot with similar occurrence time, the Kalman filter is used to correct the encoded time series data in the same time slot one by one to obtain the denoised data.

4. A device for analyzing and managing periodic time series data, characterized in that: The device comprises: The encoding module is used to cut the input periodic time series data into periods, encode the time series data in each period according to the occurrence time of the periodic time series data, and obtain the encoded time series data; The clustering module is used to cluster the encoded time series data, exclude abnormal data, and locate the time slot where the periodic time series data occurs; A noise reduction module, used to reduce noise on the encoded time series data in each slot at each moment; A reverse encoding module is used to reassemble and reverse encode the denoised data to obtain output data; The encoding module is specifically used to encode the time series data of each cycle respectively based on the offset of the time corresponding to the time series data in the cycle relative to the starting point of the cycle time, so as to represent each time series data based on the occurrence time of each time series data; The clustering module is specifically used for: Clustering the encoded time series data of multiple periods to obtain a clustering result; a cluster in the clustering result corresponds to a time slot; Based on the clustering result, a cluster whose data quantity is less than a first preset threshold is determined as an abnormal cluster, and the abnormal data in the corresponding time slot is removed; The device also includes: The first determination module is used to cluster the encoded time series data M+2 times when the period with the most encoded time series data in all periods contains N encoded time series data, M≥1, and the number of target clusters is specified as {N-1, N, N+1,…, N+M} during the clustering process; after each clustering is completed, the sum of the variances of all encoded time series data to the cluster center is calculated to obtain the variance set {S N-1 , S N , S N+1 ,…, S N+M }; Determine a value P between N-1 and N+M such that |S P-1 -S P The value of | is greater than the second preset threshold, and |S P -S P+1 | is less than a second preset threshold value, and the second preset threshold value is between 1.5 and 3.5; P is used as the optimal number of slots at the moment; The clustering module is specifically used to: perform clustering by taking the optimal number of time slots as the number of clusters.

5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for analyzing and managing periodic time series data described in any one of claims 1-3 is implemented.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method for analyzing and managing periodic time series data described in any one of claims 1 to 3.

7. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps in the method for analyzing and managing periodic time series data described in any one of claims 1-3 are implemented.

Citation Information

Patent Citations

  • Livestock blood product processing process monitoring method based on Internet of Things

    CN117762106A