Time-Series Data Compression and Graphical Signature Analysis
By generating compressed data sets and graphical signatures to represent time series data, the complexity of visualization and analysis of modern systems when processing large amounts of time series data is solved, and simplified storage of data and identification of exception events are achieved.
Patent Information
- Application Number
- CN202080051183.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-15
- Filing Date
- 2020-05-17
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2040-05-17
AI Technical Summary
Modern systems and services are difficult to effectively visualize in charts and dashboards when processing large amounts of time series data, resulting in complex data storage and analysis, and difficult to distinguish between reduced reliability and normal operational deviations.
By generating a compressed dataset, the number of occurrences of each unique data value pair is recorded, and the time series data is represented using graphical signatures to reduce noise and highlight abnormal events.
It significantly simplifies data storage and analysis, provides simplified interpretation and visualization of time series data, and can distinguish between normal and abnormal operations and reduce specific deviations of individual operations.
Smart Images

Figure CN114096959B_ABST
Abstract
Description
Technical Field
[0001] Embodiments described herein relate to systems and methods for storing, viewing, and analyzing sequential data, including, for example, time-series data. Summary of the Invention
[0002] Modern systems and services continue to become more complex with an ever-increasing number of features and functionality. Modern monitoring systems can be configured to detect problems, but the increasing number of data points collected in time series (e.g., in telemetry environments) can make it difficult to visualize the data in meaningful ways in charts and "dashboards." In some examples described herein, information from large amounts of time series data is compressed into a much smaller number of data points, which can significantly simplify the storage and analysis of the data, including, for example, diagnostic analysis of fluctuations in the behavior of telemetry services. The systems and methods described herein are particularly capable of distinguishing between significant degradations in reliability and deviations from normal operations. Furthermore, the compressed graphical representations of time series data (i.e., "graphic signatures") provide a new, simplified language that simplifies the interpretation of a set of time series data over any given time period, while also providing visualization of the time dimension as a directional plane. Certain systems and methods described herein also reduce the noise of deviations specific to individual operations and focus on isolating and identifying anomalous events in time series data.
[0003] One embodiment provides a method for compressing a sequential data set on a computer system. The computer system receives the sequential data set in the form of a plurality of data values in a serial sequence. The computer system analyzes the sequential data set to identify the number of times each of a plurality of unique data value pairs occurs in the sequential data set. Each unique data value pair includes a first data value and a second data value different from the first data value. The computer system then generates a compressed data set based on the sequential data set. The compressed data set includes data elements for each of the plurality of unique data value pairs. Each data element includes: an identification of the first data value of the unique data value pair, an identification of the second data value of the unique data value pair, and a count indicating the number of times the second data value of the unique data value pair occurs immediately after the first data value of the unique data value pair in the sequential data set.
[0004] Some embodiments generate a graphical signature indicating the contents of a sequential data set by a computer system. The graphical signature includes a plurality of nodes and a plurality of vectors extending between different nodes. Each node corresponds to a different data value in a compressed data set. Each vector corresponds to a different data element in the compressed data set and is displayed in the graphical signature as a line starting from a node corresponding to a first data value of the data element and extending to a node corresponding to a second data value of the data element.
[0005] These and other features, aspects and advantages will be apparent from a reading of the following detailed description and a review of the associated drawings.It is to be understood that both the foregoing general description and the following detailed description are illustrative and are not restrictive of the aspects, as claimed. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 is a block diagram of a system for capturing time series data, compressing the captured time series data, and displaying a graphical signature indicative of the time series data, according to one embodiment.
[0007] Figure 2 is used Figure 1 Flowchart of a method for compressing time series data in a system.
[0008] Figure 3 is listed by Figure 1 A table showing examples of time series data captured by the system.
[0009] Figure 4 is listed by Figure 2 The method generates a table of compressed data.
[0010] Figure 5 Is to use dynamic rounding to further compress Figure 3 The method outputs data as a flow chart of the method.
[0011] Figure 6 is used Figure 1 A flow chart of a method for a system to generate and display a graphical signature indicative of time series data.
[0012] Figure 7 is Figure 1 A diagram of the first collection of time series data captured by the system.
[0013] Figure 8 is Figure 1 A diagram of a second collection of time series data captured by the system.
[0014] Figure 9 is Figure 6 Generated by the method Figure 7 Example of a graphical signature for time series data.
[0015] Figure 10 is Figure 6 Generated by the method Figure 8 Example of a graphical signature for time series data.
[0016] Figure 11 It is through application Figure 5 Dynamic rounding and Figure 6 Generated by the method Figure 7Example of further compressed graphical signatures for time series data.
[0017] Figure 12 It is through application Figure 5 Dynamic rounding and Figure 6 Generated by the method Figure 8 Example of further compressed graphical signatures for time series data.
[0018] Figure 13 is a flow chart of a method for testing new software builds for a telemetry service using data compression and graph signature analysis.
[0019] Figure 14 It is an instruction Figure 13 Example of a graphical signature of a "healthy build" in the method.
[0020] Figure 15 It is an instruction Figure 13 Example of a graphical signature of an "unhealthy build" in the method. DETAILED DESCRIPTION
[0021] One or more embodiments are described and illustrated in the following description and accompanying drawings. These embodiments are not limited to the specific details provided herein and may be modified in various ways. Furthermore, other embodiments not described herein may exist. Furthermore, functions described herein as being performed by one component may be performed by multiple components in a distributed manner. Similarly, functions performed by multiple components may be combined and performed by a single component. Similarly, components described as performing a particular function may also perform additional functions not described herein. For example, a device or structure "configured" in a certain manner is configured in at least that manner, but may also be configured in other ways not listed. Furthermore, some embodiments described herein may include one or more electronic processors configured to perform the described functions by executing instructions stored on a non-transitory computer-readable medium. Similarly, the embodiments described herein may be implemented as a non-transitory computer-readable medium storing instructions executable by one or more electronic processors to perform the described functions. As used herein, "non-transitory computer-readable medium" includes all computer-readable media but does not include transitory, propagating signals. Thus, non-transitory computer-readable media may include, for example, a hard disk, a CD-ROM, an optical storage device, a magnetic storage device, ROM (read-only memory), RAM (random access memory), register memory, processor cache, or any combination thereof.
[0022] In addition, the words and terms used herein are for descriptive purposes and should not be considered limiting. For example, the use of "including," "containing," "comprising," "having" and variations thereof in this document is intended to cover the items listed thereafter and their equivalents as well as additional items. The terms "connected" and "coupled" are used broadly and include direct and indirect connections and couplings. In addition, "connected" and "coupled" are not limited to physical or mechanical connections or couplings and may include direct or indirect electrical connections or couplings. In addition, electronic communications and notifications may be performed using wired connections, wireless connections, or a combination thereof and may be transmitted directly or through one or more intermediate devices over various types of networks, communication channels, and connections. In addition, relational terms such as first and second, top and bottom, etc. may be used herein solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between these entities or actions.
[0023] Figure 1 An example of a system configured to collect, analyze, compress, and visually present time series data in a graphical format is illustrated. Figure 1 The system includes a controller 101 having an electronic processor 103 and non-transitory computer-readable memory 105. The memory 105 stores data and instructions that are executed by the electronic processor 103 to provide the functionality of the controller 101, including the functionality described in the following examples. In some implementations, the controller 101 can be implemented as a computer (e.g., a desktop computer or a server), while in other implementations, the controller can be implemented as an embedded system. One or more sensors 107 are communicatively coupled to the controller 101 and are configured to collect time series data and provide it to the controller 101. The collected time series data is then stored in the memory 105 and / or analyzed by operations performed by the electronic processor 103.
[0024] The sensor 107 may include any type of sensor configured to monitor and collect time series data. For example, in some implementations, the sensor 107 may include electrodes configured to monitor the heart rate or ECG of a human patient. In other implementations, the sensor 107 may include a component configured to monitor the performance of the CPU (e.g., execution speed, etc.) or the data transfer rate. In some implementations, the sensor 107 is directly coupled to the controller 101 via a wired or wireless communication interface. However, in other implementations, the sensor 107 is provided as a "client" computer or device (or part of a client computer or device), and the controller 101 is provided as a remote computer server or cloud computing environment.
[0025] The display 109 is also communicatively coupled to the controller 101 and is configured to display, for example, a graphical user interface and data in a numeric, textual, and / or graphical format based on output received from the controller 101. In some implementations, the controller 101 may also be communicatively coupled to one or more actuators 111 and configured to operate the one or more actuators 111 to perform operations based on the analysis of the time series data. For example, in some implementations, the actuator 111 may include a patient alarm, and the controller 101 may be configured to automatically activate the patient alarm in response to determining that the collected time series data indicates an emergency cardiac condition (e.g., a heart attack). In other implementations, the controller 101 may be configured to activate or utilize additional computing resources in response to determining that the collected time series data indicates an anomaly or defect in the computing environment.
[0026] Sequential data is generated by Figure 1 107 as a series of values (e.g., integer values) output by the sensor 107. Time series data is an example of sequential data in which new data values are measured / recorded at a defined sampling frequency. Depending on the duration over which the time series data was collected and the sampling frequency used, the number of data points in a time series data set can be very large. Therefore, the storage and analysis of sequential data, such as time series data, can be improved by mechanisms for compressing (or condensing) the collected data while retaining sufficient information for analyzing the original sequential data.
[0027] Figure 2 An example of a method for compressing sequential data is illustrated. Although the examples described herein are specifically directed to "time series data," in other implementations, the systems and methods described herein can be extended to other types of data sequences. Figure 2 The method loops through the time series data by comparing each pair of adjacent data points in the time series and counting the number of times each unique combination of data values occurs as a consecutive pair of data points.
[0028] like Figure 2 As shown, time series data is collected by sensor 107 and provided to controller 101 (step 201). As additional data is received from sensor 107, it is stored in memory 105 for later analysis. The collected time series data is then processed sequentially to analyze each pair of adjacent data elements in the time series. For example, counter (i) can be set to zero (step 203) to start the analysis at the beginning of the data set. The value of the first data point (f(x i )) and the value of the next consecutive data point (f(x i+1 )) are compared (step 205). If adjacent data points have different values, the system checks to see if the data value pair (f(x i), f(x i+1 )) is already contained in the count table of the data set (step 207). If not, the data value pair is added to the count table as a new unique data value pair (step 209) and the "count" value of the new data value pair is incremented (i.e., to "1") (step 211). If the analysis has not reached the end of the time series data (step 213), the loop counter i is incremented (step 215) and the controller 101 analyzes the next pair of data points in the time series data (step 205).
[0029] As described above, if the adjacent data values are not equal and the same data value pair is not already included in the count table, the controller 101 includes the new data value pair as a new entry in the count table (step 209). However, when the controller 101 subsequently encounters the same sequential data value pair, the controller 101 simply increments the "count" value for that data value pair in the count table (step 211). Additionally, if the controller 101 discovers that two consecutive data points in the time series have the same value, the controller 101 simply moves on to the next consecutive pair of data entries without modifying the count table. Once the controller 101 has analyzed each consecutive pair of adjacent data points in the time series (step 213), the count table is output as a compressed representation of the time series dataset (step 217).
[0030] To further explain Figure 2 method, Figure 3 An example of a very small time series data set is provided, where the x value indicates the time at which the data point was collected, and the y value is the measurement of the sensor output at time x (ie, y = f(x)). Figure 4 is Figure 2 Generated by the method Figure 3 An example of a compressed representation of a dataset.
[0031] Figure 3 The first pair of adjacent data points in the time series are (1,0) and (2,0). i )=0 and f(x i+1 )=0) is compared by the controller 101 (step 205, Figure 2 ) and are determined to be the same value. Thus, Figure 2 The method moves to the next consecutive data point pair without modifying the count table. Figure 3 The second pair of adjacent data points in the time series is (2,0) and (3,2). The controller 101 determines that the measurement values (y) of these two consecutive data points are different (i.e., "0" and "2") (step 205). Therefore, the controller 101 adds the data value pair as a new entry to Figure 4The count table of the new data value pair is stored in step 209 and the “count” column of the count table is incremented (step 211). Figure 3 The third pair of adjacent data points in the time series is (3, 2) and (4, 3). This is again a new combination of unequal values, so the data value pair is added as a new entry to the count table (step 209).
[0032] Figure 2 The method continues until the Y values of each pair of consecutive data items in the time series have been analyzed and the count table has been updated accordingly. Figure 4 The count table indicates that there are two different data values y=2 (i.e., i =2 and x i+1 =3 and at x i =7 and x i+1 =8) follows the data value y = 0 in sequence. Furthermore, the data value y = 3 follows the data value y = 2 twice in sequence, the data value y = 1 follows the data value y = 3 once in sequence, the data value y = 0 follows the data value y = 1 once in sequence, and the data value y = 2 follows the data value y = 3 once in sequence.
[0033] In this way, a time series data set of any number of data points can be compressed and represented by a "count" that indicates the number of times each unique pair of different data values occurs sequentially in the time series data set. Although the original time series data set cannot be fully reconstructed from the compressed data set (i.e., the "count table"), the compressed data still provides important information about the time series data set. For example, the compressed data set provides an indication / confirmation of how many "events" occurred in which the signal measured by the sensor was below (or above) a certain threshold, an indication of the maximum / minimum measured sensor values, and an indication of high deviations of consecutive data points (indicating rapid changes in the time series data). As discussed in further detail below, the time series data set represented by Figure 2 The compressed datasets generated by our approach also encode useful and important information about the overall variability of the time series data signal.
[0034] The data size of the compressed dataset can be further reduced by applying rounding to the values of the time series data to reduce the number of unique pairs of data values that occur immediately after each other in the time series. For example, the data values in the time series can be Figure 2The method is applied to the data before or after rounding (e.g., rounding to the nearest integer, rounding to the nearest multiple of 10, rounding to the nearest multiple of 5, etc.). In some specific implementations, "dynamic" rounding is applied to one or more data sets to achieve a target number of unique data value pairs (or a target number of "nodes" in a graphical representation as discussed further below). As shown in the figure below, applying the same dynamic rounding to multiple different time series data sets can help illustrate different degrees of variation between the time series data sets.
[0035] Figure 5 Illustration of a method for applying dynamic rounding to a Figure 2 An example of a method for compressing a data set is shown in FIG. First, a compressed data set is received (e.g., accessed from a memory or by Figure 2 (step 501). The controller then determines a target or "maximum" number of nodes for further compressing the data (step 503) and applies dynamic rounding to the data values to reach the target number of nodes (step 505). After adjusting the values in the compressed data set (e.g., a count table) through dynamic rounding, the controller updates "counts" to sum the counts of any data value pairs that became duplicates after applying dynamic rounding (step 507) and outputs the further reduced / rounded data set (step 509).
[0036] In some implementations, the controller can be configured to apply dynamic rounding to reduce the data set to the total number of unique data values (or "nodes") that appear in one or more data value pairs in the count table. Alternatively, in some implementations, the controller can be configured to determine and apply appropriate dynamic rounding to reach a target or "maximum" number of unique data value pairs.
[0037] For example, Figure 4 The compressed data set shown in the "Count Table" includes five unique data value pair combinations ((0,2); (2,3); (3,1); (1,0); and (3,2)) and four unique data values / nodes (0,1,2,3). A controller configured to apply "dynamic rounding" to reduce the total number of unique data values / nodes in the compressed data set to three unique values / nodes may determine (at step 505) to round the data values in the compressed data set to the nearest multiple of two. Apply Figure 5 The values of the count table before dynamic rounding are shown in Table 1 below. After dynamic rounding (i.e. Figure 5 The values of the count table after step 505 of ) are shown in Table 2 below. Figure 5The values of the count table after "counting" the data value pairs that have become redundant (after step 507) are shown in Table 3 below. As shown in this example, if the data values in the count table are rounded to the nearest multiple of two to reduce the total number of "nodes" in the compressed data set to three, the third and fifth entries in the original count table become redundant (both equal to (4, 2) after dynamic rounding). Therefore, the counts associated with the now redundant entries are added (now totaling two) and the redundant data value pair entries are deleted from the compressed data set.
[0038]
[0039]
[0040] Table 1: Original compressed dataset ( Figure 5 Step 501)
[0041]
[0042] Table 2: Compressed dataset after dynamic rounding ( Figure 5 Step 505)
[0043]
[0044] Table 3: Compressed dataset after merging redundant data value pairs ( Figure 5 Step 505)
[0045] As an additional example, if the dynamic rounding mechanism determines Figure 4 The data values in the count table will be rounded to the nearest multiple of five, so the total number of unique values / nodes in the compressed data set is reduced to two (0,5), and the number of unique data value pairs is also reduced to two ((0,5) and (5,0)). As a result of rounding, data value pairs whose values are both adjusted to the same value (e.g., the first and fourth entries in the count table) are deleted from the updated count table, and data value pairs that become duplicates after rounding (e.g., the third and fifth entries in the count table) are merged. This example is further illustrated in Tables 4-6 below, where Table 4 shows the data before dynamic rounding (i.e., after Figure 5 Table 5 shows the value of the count table after dynamic rounding (i.e., after step 501 in Figure 5 Table 6 shows the values of the count table after removing redundant data value pairs and data value pairs whose values are no longer different from the count table (i.e., after step 507 of FIG. 5 ). Figure 5 After step 509 ), the value of the counting table.
[0046]
[0047]
[0048] Table 4: Original compressed dataset ( Figure 5 Step 501)
[0049]
[0050] Table 5: Compressed dataset after dynamic rounding ( Figure 5 Step 505)
[0051]
[0052] Table 6: Compressed dataset after merging redundant data value pairs ( Figure 5 Step 505)
[0053] As mentioned above, according to Figure 2 The time series data compressed by Figure 5 The dynamic rounding further compresses the time series data and can be stored in memory for future analysis. The amount of memory required to store the compressed data set (e.g., in a "count table" format) is in many cases significantly less than the amount of memory required to store the original time series data. However, the compressed data set can still be analyzed to provide useful and important information about the original time series data, including, for example, the maximum / minimum data values, the maximum change between adjacent data values, and information about the variability of the time series data.
[0054] In some implementations, the controller is configured to store or display the compressed data as a graphical representation of the compressed data set to provide a "graphic signature" of the original time series data. For example, the graphical signature can provide a visual indication of variability in the original time series data and can be used to distinguish between normal time series signal conditions and abnormal / abnormal conditions. Figure 6 The diagram shows the Figure 2 An example of a method for generating a graphical signature for a time series dataset using a compressed dataset (i.e., a "count table" format) generated by a method. First, a compressed dataset is received (e.g., accessed from a memory as Figure 2 The output of the method is received, or sent to the controller from a remote server or cloud storage / computing environment) (step 601).
[0055] The compressed data set is then analyzed to select a suitable layout / topology template (step 603). In some implementations, the controller can be configured to store a plurality of predefined layout / topology templates that the controller can use to generate a graphical signature for the time series data. In various different implementations, the controller can be configured to select a suitable template based on, for example, the number of different data value pairs in the compressed data set, the number of different nodes / values in the compressed data set, and the percentage of times the same value appears in different unique data value pairs. For example, if the vast majority of data value pairs include the number zero as one of their values, this can indicate that the original time series data was primarily based on zero values and had a relatively large amount of short-term deviations. To account for this in the graphical signature, the controller can be configured to select a "star" layout (e.g., as Figure 10 ), where the node corresponding to the value that appears in the largest number of distinct data value pairs is located at the center of the graphical representation, and where the nodes corresponding to other values that appear in the compressed data set are arranged in an elliptical pattern around the central node.
[0056] Once a suitable template has been determined for the time series data set, the positions of the various nodes are plotted according to the selected template. Each individual node represents a different value in one or more data value pairs that appear in the compressed data set. Each unique data value pair in the compressed data set is illustrated in the graphical signature as a vector extending from one node to another. Each vector begins at the node corresponding to the first value in the data value pair and ends at the node corresponding to the second value in the data value pair. The "count" of each data value pair (i.e., the number of times the data value pair appears sequentially in the time series data) is represented by the thickness of the vector, as shown in the graphical signature. For example, data value pairs with relatively high "counts" will be represented in the graphical signature by relatively thick vectors, while data value pairs with relatively low "counts" will be represented in the graphical signature by relatively thin vectors.
[0057] Back to Figure 6After drawing the node positions based on the selected template (step 605), the controller determines the "thickness" of the vector corresponding to the first data value pair in the compressed data set based on the size of the "count" of the first data value pair (step 607), and then adds the vector of the determined thickness to the graphic signature between the two nodes corresponding to the two values of the data value pair (step 609). If the compressed data set includes more data value pairs that have not yet been represented by vectors in the graphic signature (step 611), the controller proceeds to the next data value pair in the compressed data set (step 613) and repeats steps 607 and 609 for each data value pair in the compressed data set. When each data value pair in the compressed data set is represented as a vector in the graphic signature, the graphic signature is displayed on the display screen (step 615).
[0058] After a graphical signature is generated for a particular time series dataset, the graphical signature can be visually inspected by a user, analyzed by an automated process implemented by a controller, and / or stored in memory. In some implementations, graphical signatures can be generated for two different time series datasets and compared. Because the node layout templates and vectors are determined based on analysis of the compressed dataset, the graphical signature for an abnormal / anomalous event will have a significantly different appearance than the graphical signature for a normal time series.
[0059] Figure 7 and Figure 8 It is a time series plot of two different events of the same process. Figure 8 Indicates that the operation was successful, and Figure 7 The diagram shows an abnormal (or "less successful") occurrence of the same operation. In particular, Figure 7 The time series illustrates some degree of functional loss in the operation of the telemetry system. Figure 8 The time series consists of approximately 500 data points collected over a period of approximately 18 days. Figure 7 The time series consists of approximately 700 data points collected over approximately 26 days. Figure 7 and Figure 8 There are certainly clear differences between the plots for , but it is difficult to estimate the differences due to the complexity of the time series profiles. Figure 7 The impact of abnormal events that occur during the time series. Figure 2 and for each of the two time series datasets, Figure 6 Generating graphical signatures, the system effectively compresses the data, provides a greatly simplified visual representation of these complex time series data sets, and provides effective noise suppression. In some implementations, compressing time series data into graphical signatures can also enable faster and simplified interpretation and analysis of the original time series data.
[0060] Figure 9 and Figure 10 Use Figure 6 The method targets Figure 7 and Figure 8 Example of a graphical signature generated from time series data. Graphical signature of an abnormal event ( Figure 9 ) and the graphical signature of normal events ( Figure 10 ) provides distinct compressed data summary visualizations, including, for example, different numbers of nodes in the graph signature and different layouts / topologies of the graph signature.
[0061] You can use, for example Figure 5 Dynamic rounding further compresses the datasets for each of these time series. Each individual value in each dataset is rounded to the nearest multiple of 10. The controller selects a multiple of 10 for the dynamic rounding operation based on the expected maximum number of nodes in the resulting compressed dataset. Figure 11 and Figure 12 After applying dynamic rounding to further compress the dataset, Figure 7 and Figure 8 Example of a graphical signature for time series data of . Note that because the same dynamic rounding is applied to both time series and because Figure 7 The time series shows the Figure 8 The time series has more variability, so Figure 11 The further compressed graph signature nodes (7 nodes) are Figure 12 The further compressed graph signature (only 3 nodes) is more. It should also be noted that by Figure 6 Before generating the graphic signature, the method Figure 2 Compression technology and Figure 5 Dynamic rounding, the system can Figure 8 The 500 data points of the original time series data are compressed into Figure 12 The graphical signature of the coin represents only three nodes, while still providing important information about the operational behavior over the 18-day period.
[0062] Figure 9 and Figure 11 The graphical signatures of both correspond to Figure 7 time series datasets (i.e., time series data corresponding to some degree of functional loss in the telemetry system). Figure 10 and Figure 12 The graphical signatures of both correspond to Figure 8 time series data (i.e., time series data corresponding to the normal operation of the telemetry system). Figure 11 The further compression of the graphical signature highlights the Figure 7The time series corresponds to some operationally specific dynamics of loss of functionality in the telemetry system over a period of time. Figure 9 and Figure 10 Comparison of graphic signatures (and Figure 11 and Figure 12 Comparison of further compressed graphic signatures of Figure 8 Time series data ratio Figure 7 The time series data are more stable.
[0063] Figure 13 The diagram illustrates a specific example of a system for updating, validating, and deploying software updates for a telemetry service that is configured to utilize the aforementioned data compression and graph signature analysis techniques. Due to the scale and complexity of modern distributed cloud systems, software engineers and service reliability engineers face numerous challenges. Distributed services may follow a continuous deployment model, where application code is continuously modified and updates are then delivered / deployed to customers on a scheduled basis (e.g., weekly or daily). For large systems that may have tens of thousands of instrument operations and, therefore, may generate tens or hundreds of thousands of time series, reviewing the complete telemetry dataset can be challenging. For example, due to the sheer size of the data objects and the limited ability of the human eye to process a specific, finite number of data points at a given time, it is impossible to observe this data simultaneously. However, by utilizing the aforementioned data compression and graph signature mechanisms, the telemetry data for each new build can be compressed into a specific graph signature, making it easier to categorize it as a "healthy build" or an "unhealthy build."
[0064] like Figure 13 As shown, the application code can be modified at any time during the continuous delivery cycle by any of a number of different developers (step 1301). Telemetry instrumentation for all operations is enabled and a large amount of telemetry data is emitted from the application. This telemetry data includes, for example, a large number of individual quality of service (QoS) time series data sets (step 1303). This telemetry data is routed to a data repository in real time (step 1305). Before "releasing" (e.g., deploying or delivering) a new "build" of the application software, the system applies the above-mentioned data compression technique to the QoS time series data set in the data repository and generates a "graphic signature" indicating the status of the current build / update (e.g., whether the current build is a "healthy build" or an "unhealthy build") (step 1307).
[0065] Figure 14 An example of a graphical signature for a "healthy build" indicating fluctuations within an allowed range of Quality of Service (QoS) value variations is illustrated. Figure 15 An example of a graph signature for an "unhealthy build" is shown. An unhealthy signature includes more vectors and more nodes, indicating a greater variation in QoS value changes.
[0066] Back to Figure 13 In the example of , the system is configured to automatically analyze the graphical signature to determine whether the signature indicates a "healthy build" or an "unhealthy build" (step 1309). In other implementations, such analysis of the graphical signature may be performed manually or with input from a "build release" engineer. If it is determined that the graphical signature indicates a "healthy build" (step 1311), the updated application code is advanced to a new build release (step 1313) and the new build is deployed (step 1315). However, if the graphical signature is determined to indicate an "unhealthy build" (step 1317), the release / deployment of the unhealthy build is stopped (step 1319) and negative customer impact is avoided. An error log is generated (step 1321) and the development team is notified that the planned build release has been blocked. The development team reviews recent changes and error logs to locate the cause of the error / regression and updates the application code again to fix the issue. The process then returns to step 1301, where the system collects new QoS data for the "fixed" application code and generates a new graphical signature. This process is repeated until the application code is again identified as a "healthy build" (step 1311) and deployed to the customer (step 1315).
[0067] It should be apparent from the above description that in some implementations, time series data of any size range can be compressed by identifying adjacent pairs of values in the time series data, rounding the values to integers (or other dynamically or statically defined multiples), and counting the number of occurrences of each unique value pair combination. This counting operation is used to convert the time series data into a three-column dataset, where each row (or "entry") in the new dataset includes the first value in the data value pair (f(x i )), the adjacent values in the data value pair (f(x i+i )), as well as a "count" indicating the number of times this pair of values appears immediately consecutively in the time series data. A graph structure (i.e., a graph signature) is then generated, in which a node represents each unique value in the three-column dataset, and vectors connect the nodes to represent each unique pair of values in the three-column dataset. The structure can be plotted using a graphics library such as GGPlot R. By generating a graph signature in this way, complex time series (e.g., up to 5,000 or more consecutive data points) can be represented in a simplified manner, which can then be classified as more stable behavior or less stable behavior that requires attention / intervention.
[0068] The various systems and methods described in the examples above provide significant advantages and can be used to automatically classify complex time series of extremely large sizes using data compression algorithms. The compressed visual data representations can also be classified and used for anomaly detection or evaluation of services in a variety of different systems in which time series or other sequential data are collected / monitored. For example, while the specific examples presented above describe the use of these techniques to monitor loss of function in telemetry systems, these techniques can also be applied to health and patient monitoring systems. For example, the data compression techniques and graphical signatures described herein can be applied to time series data of an electrocardiogram. More specifically, abnormal sequences that may indicate a heart attack can be detected by analysis of the graphical signatures. Similarly, a graphical signature can be created on a personal device (e.g., a wrist-worn device with a heart rate monitoring feature) and displayed to the user to provide the user with a graphical representation of their current heart function (e.g., at rest or while exercising).
[0069] In some implementations, the controller can be configured to apply a trained artificial intelligence (AI) mechanism, such as a trained artificial neural network, in generating the compressed data set and / or the graphical signature. For example, in some implementations, the artificial neural network can be configured to receive some or all of the compressed data set (e.g., Figure 2 output of the time series data set) and output an identification of an appropriate layout / topology template to be used for the graphical signature of the data series. Artificial neural networks may also be used in some implementations of the dynamic rounding process. For example, the artificial neural network may be configured to receive as input the time series data and / or a compressed data set (i.e., an initial "count" table) and produce as output an identification of appropriate multiples to which the data values will be rounded in the dynamic rounding step. Alternatively, in some implementations, the artificial neural network may be configured to receive as input the time series data set or the initial compressed data set and produce as output a further compressed data set - such that the artificial neural network performs the entire dynamic rounding process, or in some cases, the entire compression process.
[0070] Thus, embodiments provide, among other things, systems and methods for compressing sequential data into a data structure indicating the number of sequential occurrences of each of a plurality of unique data value pairs in the time series data and generating a graphical signature from the compressed data indicating various aspects of the original time series data. Various features and advantages are set forth in the following claims.
Claims
1. A system for compressing a sequential data set on a computer system, the system comprising a controller configured to: receiving a sequential data set comprising a plurality of data values in a serial sequence; analyzing the sequential data set to identify a number of times each unique data value pair in a plurality of unique data value pairs occurs in the sequential data set, wherein Each unique data value pair includes a first data value and a second data value, the second data value being different from the first data value; as well as generating a compressed data set based on the sequential data set, wherein the compressed data set includes a data element for each unique data value pair in the plurality of unique data value pairs, wherein each data element includes an identification of the first data value of the unique data value pair, an identification of the second data value of the unique data value pair, and a count indicating the number of times the second data value of the unique data value pair occurs immediately after the first data value of the unique data value pair in the sequential data set, The controller is further configured to generate a graphical signature indicating the content of the sequential data set by: locating a plurality of nodes in the graphical signature, wherein each node in the plurality of nodes corresponds to a different data value in the compressed data set, and A plurality of vectors are located in the graphical signature, wherein each vector corresponds to a different data element of the compressed data set, and wherein each vector appears in the graphical signature as a line that begins at a node corresponding to a first data value of the data element and extends to a node corresponding to a second data value of the data element.
2. The system of claim 1, wherein: The controller is configured to generate the compressed data set by: rounding the first data value and the second data value of each unique data value pair to the nearest multiple of a rounding value, and The first data value and the second data value of each unique data value pair in the compressed data set are replaced with a rounded data value pair comprising the rounded first data value and the rounded second data value of the unique data value pair.
3. The system of claim 2, wherein: The controller is further configured to determine a rounding value that will result in rounding of the first data value and the second data value of each unique data value pair to reduce the number of unique data values in the compressed data set below a defined threshold.
4. The system of claim 1, wherein: The controller is further configured to: storing the compressed data set in a non-transitory computer-readable memory; and Determine whether a threshold is exceeded in the sequential data set by: accessing the compressed data set from the memory, and The data values of the compressed data set are analyzed to determine whether one or more data values exceed the threshold.
5. The system of claim 1, wherein: The controller is further configured to: storing the compressed data set in a non-transitory computer-readable memory; and The stability condition in the sequential dataset was quantified by the following operation: accessing the compressed data set from the memory, and A value indicative of the stability condition is calculated based at least in part on a number of unique data value pairs in the compressed data set.
6. The system of claim 1, wherein: Each vector of the plurality of vectors includes a coarseness corresponding to a count of data elements corresponding to the vector, wherein data elements having higher counts in the graphical signature are shown as vectors having a greater coarseness than vectors corresponding to data elements having lower counts.
7. The system of claim 1, wherein: The controller is configured to generate the graphical signature indicating the contents of the sequential data set by further selecting a node layout template from a plurality of stored node layout templates based on the analysis of the compressed data set, and wherein the controller is configured to position the plurality of nodes in the graphical signature by positioning each of the plurality of nodes according to the selected node layout template.
8. The system of claim 1, wherein: The controller is further configured to detect the occurrence of an abnormal event by performing the following operations: comparing the graphical signature generated for the sequential data set with a second graphical signature of a sequential data set indicative of normal operation, and A difference between the graphical signature and the second graphical signature is detected.
9. The system of claim 8, wherein: The controller is further configured to activate at least one actuator to modify an operation corresponding to the sequential data set in response to detecting the occurrence of the abnormal event.
10. The system of claim 1, wherein: The controller is configured to receive the sequential data set by receiving a time series data set, wherein each data value in the time series data set indicates a value of at least one condition at different consecutive times within a defined time period based on a sampling frequency.
11. The system of claim 10, wherein: The controller is configured to receive the sequential data set by receiving a sequence of output values from a sensor at the sampling frequency, wherein each output value of the sequence of output values is indicative of a condition measured by the sensor.
12. The system of claim 1, wherein: The controller is configured to receive the sequential data sets by receiving sequential data sets indicative of operation of the telemetry system.
13. A method for compressing a sequential data set on a computer system, the method comprising: receiving, by the computer system, a sequential data set comprising a plurality of data values in a serial sequence; analyzing, by the computer system, the sequential data set to identify a number of times each unique data value pair in a plurality of unique data value pairs occurs in the sequential data set, wherein each unique data value pair includes a first data value and a second data value, the second data value being different from the first data value; and generating, by the computer system, a compressed data set based on the sequential data set, wherein the compressed data set includes a data element for each unique data value pair in the plurality of unique data value pairs, wherein each data element includes an identification of the first data value of the unique data value pair, an identification of the second data value of the unique data value pair, and a count indicating the number of times the second data value of the unique data value pair occurs immediately after the first data value of the unique data value pair in the sequential data set, The method further comprises generating, by the computer system, a graphical signature indicating the content of the sequential data set by: locating a plurality of nodes in the graphical signature, wherein each node in the plurality of nodes corresponds to a different data value in the compressed data set, and A plurality of vectors are located in the graphical signature, wherein each vector corresponds to a different data element of the compressed data set, and wherein each vector appears in the graphical signature as a line that begins at a node corresponding to a first data value of the data element and extends to a node corresponding to a second data value of the data element.
14. A non-transitory computer-readable memory having stored thereon computer-executable instructions which, when executed by an electronic processor, cause a computer-based system to: receiving a sequential data set comprising a plurality of data values in a serial sequence; analyzing the sequential data set to identify a number of times each unique data value pair in a plurality of unique data value pairs occurs in the sequential data set, wherein Each unique data value pair includes a first data value and a second data value, the second data value being different from the first data value; as well as generating a compressed data set based on the sequential data set, wherein the compressed data set includes a data element for each unique data value pair in the plurality of unique data value pairs, wherein each data element includes an identification of the first data value of the unique data value pair, an identification of the second data value of the unique data value pair, and a count indicating the number of times the second data value of the unique data value pair occurs immediately after the first data value of the unique data value pair in the sequential data set, wherein the computer-executable instructions, when executed by the electronic processor, further cause the computer-based system to generate a graphical signature indicative of the contents of the sequential data set by: locating a plurality of nodes in the graphical signature, wherein each node in the plurality of nodes corresponds to a different data value in the compressed data set, and A plurality of vectors are located in the graphical signature, wherein each vector corresponds to a different data element of the compressed data set, and wherein each vector appears in the graphical signature as a line that begins at a node corresponding to a first data value of the data element and extends to a node corresponding to a second data value of the data element.
Citation Information
Patent Citations
Data compression
CN106452450A