Time series data analysis method, device, equipment, medium and product
By generating the heatmap of the timing data set and calculating the similarity between the heatmap, the problem in the prior art is solved that it is difficult to quickly and accurately obtain the similarity between high-dimensional timing control data with high density sampling, and the rapid and accurate similarity analysis of high-dimensional timing data is achieved.
Patent Information
- Application Number
- CN202411747312.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to quickly and accurately obtain the similarity between high-dimensional timing control data of high-density sampling, and the calculation method is too complex and lacks sufficient consideration of timing characteristics.
By obtaining at least two time series data sets, a heat map for each data set is generated and the similarity is calculated using the average grayscale, standard deviation and covariance of the heat map, combined with weighting processing to obtain similarity between data sets.
This method can quickly and accurately obtain the similarity between high-dimensional timing control data of high-density sampling, and fully consider the timing characteristics of the data set.
Smart Images

Figure CN119939263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of industrial control technology, and in particular to a time series data analysis method, device, equipment, medium and product. Background Art
[0002] With the rapid development of digital information technology, data is showing an exponential growth trend and gradually showing the characteristics of high dimension, multivariate and time series.
[0003] High-dimensional time series data in the field of industrial control is collected by many sensors and has the characteristics of high-density sampling and multiple time periods. When analyzing the similarity between multiple high-dimensional data sets, correlation metrics are usually used for quantification. In order to realize data mining of high-dimensional feature space, the problem of similarity measurement of high-dimensional data sets is introduced. At present, most studies mainly focus on the improvement of metric functions or the division of high-dimensional space regions, so as to realize the metric analysis of high-dimensional data under a single data set, which cannot effectively solve the metric problem between high-dimensional multi-time period data sets. In addition, the calculation method is too complicated and lacks sufficient consideration of time series characteristics, making it difficult to quickly and accurately obtain the similarity between high-dimensional time series control data with high-density sampling. Summary of the invention
[0004] In view of this, the present application provides a time series data analysis method, device, equipment, medium and product, which can accurately obtain the similarity between high-dimensional time series control data with high density sampling. The technical solution is as follows.
[0005] In a first aspect, a time series data analysis method is provided, the method comprising:
[0006] Obtain at least two time series data sets;
[0007] For each time series data set, a heat map corresponding to the time series data set is generated with the time series as the horizontal axis and the time series data type as the vertical axis;
[0008] The similarity between the at least two time series data sets is obtained according to the heat maps respectively corresponding to the at least two time series data sets.
[0009] In a possible implementation, the heat maps corresponding to the at least two time series data sets respectively include a first heat map corresponding to the first time series data set and a second heat map corresponding to the second time series data set;
[0010] Acquiring the similarity between the at least two time series data sets according to the heat maps respectively corresponding to the at least two time series data sets includes:
[0011] Obtaining a brightness contrast function value according to an average grayscale of the first heat map and an average grayscale of the second heat map;
[0012] Obtaining a standard deviation comparison function value according to a standard deviation of the first thermogram and a standard deviation of the second thermogram;
[0013] Obtaining a structure contrast function value according to the covariance between the first thermodynamic map and the second thermodynamic map;
[0014] The brightness contrast function value, the standard deviation contrast function value, and the structure contrast function value are weighted to obtain a similarity between the first time series data set and the second time series data set.
[0015] In a possible implementation, before generating a heat map corresponding to each time series data set with the time series as the horizontal axis and the time series data type as the vertical axis, the method further includes:
[0016] Obtain the maximum value and the minimum value of each time series data type in the time series data set;
[0017] Based on the maximum value and the minimum value of each time series data type, the data of each time series data type in the time series data set is normalized.
[0018] In a possible implementation, the method further includes:
[0019] Sort the data of each time series data type in the time series data set by time, and determine the position of missing data in the time series data set;
[0020] Fill the missing data position with data according to the time series data adjacent to the missing data position in the time series data set.
[0021] In a possible implementation, acquiring the similarity between the at least two time series data sets according to the heat maps respectively corresponding to the at least two time series data sets includes:
[0022] According to the heat maps corresponding to the at least two time series data sets, a diagonal matrix data set corresponding to the at least two time series data sets is obtained; the arrangement order of each time series data type of different time series data sets corresponding to the horizontal coordinates and the vertical coordinates in the diagonal matrix data set; the elements in the diagonal matrix data set are used to indicate the similarity of the time series data sets;
[0023] The method further comprises:
[0024] Clustering is performed on the time series data types in the diagonal matrix data set to obtain a time series data type set, and the time series data type set is used to analyze the correlation between the time series data types.
[0025] In a possible implementation, the method further includes:
[0026] According to the element values in the diagonal matrix data set, the corresponding matrix heat map is displayed on the display device.
[0027] In a second aspect, a time series data analysis device is provided, the device comprising:
[0028] A data set acquisition module, used to acquire at least two time series data sets;
[0029] A heat map generation module is used to generate a heat map corresponding to each time series data set, with the time series as the horizontal axis and the time series data type as the vertical axis;
[0030] The similarity acquisition module is used to acquire the similarity between the at least two time series data sets according to the heat maps respectively corresponding to the at least two time series data sets.
[0031] In a possible implementation, the heat maps corresponding to the at least two time series data sets respectively include a first heat map corresponding to the first time series data set and a second heat map corresponding to the second time series data set;
[0032] The similarity acquisition module is used to acquire a brightness contrast function value according to the average grayscale of the first heat map and the average grayscale of the second heat map;
[0033] Obtaining a standard deviation comparison function value according to a standard deviation of the first thermogram and a standard deviation of the second thermogram;
[0034] Obtaining a structure contrast function value according to the covariance between the first thermodynamic map and the second thermodynamic map;
[0035] The brightness contrast function value, the standard deviation contrast function value, and the structure contrast function value are weighted to obtain a similarity between the first time series data set and the second time series data set.
[0036] In a possible implementation manner, the device further includes:
[0037] A maximum value acquisition module, used to obtain the maximum value and the minimum value of each time series data type in the time series data set;
[0038] The normalization module is used to perform normalization processing on the data of each time series data type in the time series data set based on the maximum value and the minimum value of each time series data type.
[0039] In a possible implementation manner, the device further includes:
[0040] A missing data detection module is used to sort the data of each time series data type in the time series data set according to time, and determine the position of missing data in the time series data set;
[0041] The data filling module is used to fill data at the missing data position according to the time series data adjacent to the missing data position in the time series data set.
[0042] In a possible implementation, the similarity acquisition module is further used to obtain a diagonal matrix data set corresponding to the at least two time series data sets according to the heat maps respectively corresponding to the at least two time series data sets; the arrangement order of each time series data type of different time series data sets corresponding to the horizontal coordinates and the vertical coordinates in the diagonal matrix data set; the elements in the diagonal matrix data set are used to indicate the similarity of the time series data sets;
[0043] The device also includes:
[0044] A clustering module is used to perform clustering processing on the time series data types in the diagonal matrix data set to obtain a time series data type set, and the time series data type set is used to analyze the correlation between the time series data types.
[0045] In a possible implementation manner, the device further includes:
[0046] The heat map display module is used to display the corresponding matrix heat map on the display device according to the element values in the diagonal matrix data set.
[0047] In a third aspect, a computer device is provided, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the above-mentioned time series data analysis method by executing the computer instructions.
[0048] In a fourth aspect, a computer-readable storage medium is provided, on which computer instructions are stored, and the computer instructions are used to enable a computer to execute the above-mentioned time series data analysis method.
[0049] In a fifth aspect, a computer program product or a computer program is provided, comprising computer instructions, wherein the computer instructions are used to enable a computer to execute the above-mentioned time series data analysis method.
[0050] The technical solution provided by this application may have the following beneficial effects:
[0051] When processing time series data in the industrial field, at least two time series data sets can be obtained; for each time series data set, a heat map corresponding to the time series data set is generated with the time series as the horizontal axis and the time series data type as the vertical axis; according to the heat maps corresponding to the at least two time series data sets, the similarity between the at least two time series data sets is obtained. The above scheme converts the time series data into a heat map, and then calculates the similarity between the heat maps, thereby fully considering the time series characteristics of the data set, and can quickly and accurately obtain the similarity between high-dimensional time series control data with high-density sampling. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0053] Figure 1 The present invention is a flow chart of a method for analyzing time series data according to an exemplary embodiment.
[0054] Figure 2 A heat map corresponding to a time series data set involved in an embodiment of the present application is shown.
[0055] Figure 3 The present invention is a flow chart of a method for analyzing time series data according to an exemplary embodiment.
[0056] Figure 4 The result matrix diagram of the high-dimensional parameter heat map after similarity analysis by the SSIM algorithm is shown.
[0057] Figure 5 It is a structural schematic diagram of a time series data analysis device provided in an embodiment of the present application.
[0058] Figure 6 It is a structural schematic diagram of a computer device provided by an optional embodiment of the present invention. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0060] In the description of the embodiments of the present application, the term "corresponding" may indicate a direct or indirect correspondence between two items, or an association relationship between the two items, or a relationship between indication and being indicated, configuration and being configured, and the like.
[0061] With the rapid development of digital information technology, data is growing exponentially and gradually showing high-dimensional, multivariate and time-series characteristics. In the field of high-dimensional time series data, data not only has high-dimensional multivariate characteristics, but also has multi-period characteristics, which is of great significance to the research of high-dimensional data mining. When analyzing the similarity between multiple high-dimensional data sets, correlation metrics are usually used to quantify. In order to realize data mining in high-dimensional feature space, the problem of similarity measurement of high-dimensional data sets is introduced.
[0062] At present, most of the research focuses on the improvement of metric functions or the division of high-dimensional space regions, so as to achieve metric analysis of high-dimensional data under a single data set, which cannot effectively solve the metric problem between high-dimensional multi-period data sets. In addition, the calculation method is too complicated and lacks sufficient consideration of time series characteristics, making it difficult to quickly and accurately obtain the similarity between high-dimensional time series control data with high density sampling. Therefore, this application proposes a time series data analysis method based on image similarity from the perspective of the data set itself.
[0063] Figure 1 FIG. 1 is a flow chart of a method for analyzing time series data according to an exemplary embodiment. The method is executed by a computer device. Figure 1 As shown, the time series data analysis method may include the following steps:
[0064] Step 101: Obtain at least two time series data sets.
[0065] The actual data set may be a time series data set obtained from an actual industrial scenario and requiring comparative analysis. In an embodiment of the present application, the time series data set may be collected from different devices in the same field, or collected from different fields in the same time period. By analyzing the time series data set, the correlation between data sets from different devices in the same field, or the correlation between data sets in different fields in terms of time series, may be determined.
[0066] Step 102 : for each time series data set, a heat map corresponding to the time series data set is generated with the time series as the horizontal axis and the time series data type as the vertical axis.
[0067] A heat map is a two-dimensional image that maps certain attributes of data to colors, and can intuitively show the distribution characteristics of high-dimensional data. The generation of a heat map is equivalent to reducing the dimension of high-dimensional time series data and visualizing it, making the original complex data easier to analyze and compare. The horizontal axis (time series) represents the time dimension and shows the dynamic changes of data over time. The vertical axis (time series data type) represents the different feature dimensions of the data, such as sensor number, indicator category, or variable sequence. Color value: represents the size or state of the corresponding moment and feature value. High-dimensional numerical data can be converted into image format (such as grayscale or color) through normalization or standardization methods.
[0068] Please refer to Figure 2 , which shows a heat map corresponding to a time series data set involved in an embodiment of the present application.
[0069] Step 103: Obtain the similarity between the at least two time series data sets according to the heat maps respectively corresponding to the at least two time series data sets.
[0070] The generated heat map is used as input, and the similarity between images is calculated using image processing algorithms (such as cosine similarity, structural similarity index (SSIM), or feature matching algorithms). That is, the similarity between two data sets is quantified based on the color distribution, shape characteristics, and texture differences of the heat map.
[0071] The image similarity method makes it easier to capture the global and local characteristics of time series data, and the heat map representation can better consider the interaction between the time series characteristics and multidimensional features of the data.
[0072] In summary, when processing time series data in the industrial field, at least two time series data sets can be obtained; for each time series data set, a heat map corresponding to the time series data set is generated with the time series as the horizontal axis and the time series data type as the vertical axis; according to the heat maps corresponding to at least two time series data sets, the similarity between the at least two time series data sets is obtained. The above scheme converts the time series data into a heat map, and then calculates the similarity between the heat maps, thereby fully considering the time series characteristics of the data set, and can quickly and accurately obtain the similarity between high-dimensional time series control data with high-density sampling.
[0073] Figure 3 FIG. 1 is a flow chart of a method for analyzing time series data according to an exemplary embodiment. The method is executed by a computer device. Figure 3 As shown, the time series data analysis method may include the following steps:
[0074] Step 301: Obtain at least two time series data sets.
[0075] Step 302: Obtain the maximum value and the minimum value of each time series data type in the time series data set.
[0076] The time series data type is the dimension of the time series data, and different time series data types represent different dimensions, such as sensor number, indicator category, etc. After obtaining the time series data set, the maximum and minimum values of each time series data set over the entire time period can be calculated as a reference for subsequent normalization.
[0077] Step 303 : normalize the data of each time series data type in the time series data set based on the maximum value and the minimum value of each time series data type.
[0078] After obtaining each time series data set, the data of each time series data type in the time series data set can be normalized according to the minimum and maximum values of each time series data type, so that data of different dimensions can be easily compared on the same scale, thereby improving the image quality of subsequent heat map generation and making the color distribution more uniform.
[0079] Step 304 , sort the data of each time series data type in the time series data set according to time, and determine the location of missing data in the time series data set.
[0080] Sort the data of each feature dimension by timestamp to ensure the continuity of the data on the time axis, and mark the positions without corresponding values in the time axis as missing data points. Missing data will reduce the accuracy of heat map generation and similarity calculation, so it needs to be identified and processed.
[0081] Therefore, before obtaining the time series data set, the data needs to be preprocessed. First, due to the complex characteristics of the data, directly analyzing the original data is not conducive to the output and conversion of experimental results. In order to improve the computational efficiency and results of the algorithm, this patent uses normalization to preprocess the data. The initial data is converted to the [0, 1] interval through the min-max method. The specific formula is as follows:
[0082]
[0083] Since time series data has the characteristics of high-density sampling, it is inevitable that there will be some missing values in the collected time series data. The methods for processing missing values include backward filling method, mean interpolation method, modeling prediction method, etc. This application mainly uses mean interpolation method to process missing values. Mean interpolation is to use the average value of the effective value of the measured data to interpolate the missing value. The missing value is x, and the normal value is x1 to x n , where n is the total length of the data set. The specific formula is as follows:
[0084]
[0085] Step 305 , filling data at the missing data position according to the time series data adjacent to the missing data position in the time series data set.
[0086] When there are missing data locations in the time series data set, data needs to be filled in at that location to ensure the integrity of the time series data and make the visualized heat map more accurate.
[0087] Optionally, a linear interpolation method may be used in an embodiment of the present application, that is, the mean of the adjacent data before and after the missing position is used to fill in the missing position; optionally, a forward / backward filling method may also be used in an embodiment of the present application, that is, the value of the previous data or the next data of the missing position is used to fill in the missing data position.
[0088] Step 306 , for each time series data set, a heat map corresponding to the time series data set is generated with the time series as the horizontal axis and the time series data type as the vertical axis.
[0089] In an embodiment of the present application, for each time series data set, the time axis is used as the horizontal axis and the time series data type is used as the vertical axis, so that the normalized data is mapped to the color value of the corresponding position, thereby generating a heat map corresponding to the time series data set.
[0090] Step 307 : acquiring the similarity between the at least two time series data sets according to the heat maps respectively corresponding to the at least two time series data sets.
[0091] In the embodiment of the present application, the similarity between the at least two time series data sets is the similarity of the heat maps corresponding to the at least two time series data sets respectively.
[0092] Generally, image similarity measurement technology can be used to obtain the similarity between heat maps. Image similarity measurement technology describes and extracts the characteristic rules in the image, compares the similarity of the content of two images, and thus measures the similarity between the images. It has a wide range of applications in face recognition, image retrieval, and correlation determination.
[0093] At present, the simplest technology is the similarity measurement technology based on grayscale histogram, which uses the grayscale pixel index of the image for discrimination. This method converts all the color channels of the original image into gray, so that the pixel range of the image is between 0 and 255, and the grayscale value of the image is between 0 and 1. Then, the correlation between images is obtained by calculating the probability distribution of the histogram after statistical conversion. Although the algorithm is simple to implement and the amount of calculation is relatively small, it does not take into account the structural distribution characteristics of the image, which is easy to cause large errors.
[0094] The cosine similarity measurement technology measures the similarity by measuring the cosine value of the inner product space of two vectors, and is suitable for feature similarity measurement in high-dimensional space. Images often contain many feature codes and belong to high-dimensional feature space. Therefore, the image measurement technology based on cosine similarity converts the feature group of each image into a vector in high-dimensional space and calculates the cosine value between the vectors to represent the similarity of the two images. The principle of this algorithm is simple and intuitive, but the amount of calculation is large. The calculation formula is as follows:
[0095]
[0096] Among them, X i and Y i It represents the transformed vector of the image feature group. The similarity is measured by calculating the cosine angle value of the vector inner product space. The value range is -1 to 1.
[0097] Structural Similarity (SSIM) is a full-reference image similarity evaluation index based on the three elements of image brightness, contrast and structure. First, brightness is estimated by the average grayscale of the image, and the brightness contrast function is recorded as L(x,y), and the formula is as follows:
[0098]
[0099] The contrast is estimated using the standard deviation of the image, denoted as C(x,y), and the formula is as follows:
[0100]
[0101] The structure is the standard deviation of the image divided by itself, denoted as S(x,y), and the formula is as follows:
[0102]
[0103] Among them, μ x represents the mean of the image, σ x represents the standard deviation, σ xy It is expressed as covariance. C1, C2, and C3 are constants to avoid zero in the denominator. The SSIM algorithm is obtained by the weighted product of these three factors:
[0104] SSIM(x,y)=[L(x,y)] α [C(x,y)] β [S(x,y)] γ
[0105] Usually, α, β, and γ are all set to 1, and C2=2C3. Substituting the above three formulas into the simplified formula, we get:
[0106]
[0107] In practical applications, the idea of sliding windows is usually used to divide the image into different blocks, and then the mean, standard deviation and covariance of each window block are calculated, and then the corresponding block structural similarity is calculated, and finally the global structural similarity of the two images is obtained by cumulative average.
[0108] The similarity measurement between data is generally based on the distance between the data. Common methods include Euclidean distance, Manhattan distance, etc. Although this method is simple and intuitive, it is only applicable to low-dimensional feature space and has large errors in high-dimensional space. Based on this, similarity measurement methods based on grids and subspaces are introduced. This type of method divides the high-dimensional space of data into subspaces, and measures the similarity between data in the subspace to measure the similarity of the original data. When the data under study not only has high-dimensional characteristics, but also has multi-time periods and parallel characteristics between data sets. Traditional algorithms based on distance and subspace division generally focus on the metric analysis of a single high-dimensional data set, and cannot effectively analyze high-dimensional multi-time period data sets.
[0109] Since the above-mentioned metric analysis methods are applicable to different fields, efficient analysis techniques are selected based on the data set itself. This patent uses high-dimensional time series data in the field of industrial control, which is collected by many sensors and has the characteristics of high-density sampling and multiple time periods. Based on the characteristics of the data, the commonly used image similarity analysis techniques are experimentally analyzed, and the specific information of the results is shown in the table:
[0110] Table 1. Comparative analysis of image similarity algorithms
[0111]
[0112] As can be seen from Table 1 above, the similarity algorithm based on grayscale histogram only judges based on the similarity of color information, with a low pixel conversion rate, and as long as the images have similar color distributions, the similarity will be judged as high, lacking consideration of the image structure distribution, so the error is large. Although the discrimination algorithm based on cosine similarity has a high pixel conversion rate, the amount of computation required to convert the image into a vector is large, and it has a high time complexity. SSIM measures the color distribution law of different regional blocks based on the image structure distribution characteristics, and has good contrast and algorithm efficiency. Therefore, this application intends to use the SSIM algorithm to calculate the similarity between thermal map images.
[0113] Feature analysis of high-dimensional time series data based on image similarity
[0114] Finally, since it is difficult to directly perform feature analysis on high-dimensional multi-period data, this patent needs to use visualization methods to first encode different parameter data into images, visualize them as heat maps, map the changing laws and patterns of the parameters, and then use the image similarity algorithm to measure the heat map, so as to effectively extract the parameter similarity features in the original data, convert the original high-dimensional multi-period data set into a single high-dimensional feature data set, and then analyze the extracted feature data.
[0115] Heatmaps can map the changing rules and patterns of parameters. First, according to the extreme values of different parameters in the data set, normalization is achieved to unify the mapping range of each column color in the heatmap. Figure 2 As shown in the figure, the time series is the horizontal axis and the time series data type is the vertical axis. Each column of color is mapped from blue to red, which represents the normalized value of the current parameter. The heat map maps the parameter value to the color space, thereby representing the parameter change characteristics based on color, thereby realizing the measurement of image similarity.
[0116] The SSIM algorithm is used to calculate the similarity between the data sets encoded as heat maps, extract the parameter similarity features in the original data set, convert the original high-dimensional multi-period parameter data set into a single high-dimensional feature data set, and then analyze the parameter feature data. How to quickly obtain regular features from the results of metric analysis is crucial for feature analysis. Therefore, it is necessary to use reasonable visualization methods to display the metric results. This application designs a matrix diagram to visually analyze the similarity calculation results;
[0117] In a possible implementation, the heat maps corresponding to the at least two time series data sets respectively include a first heat map corresponding to the first time series data set and a second heat map corresponding to the second time series data set;
[0118] According to the average grayscale of the first heat map and the average grayscale of the second heat map, a brightness contrast function value is obtained; according to the standard deviation of the first heat map and the standard deviation of the second heat map, a standard deviation contrast function value is obtained; according to the covariance of the first heat map and the second heat map, a structure contrast function value is obtained; the brightness contrast function value, the standard deviation contrast function value and the structure contrast function value are weighted to obtain the similarity between the first time series data set and the second time series data set.
[0119] The calculation formula for the similarity between the first time series data set and the second time series data set refers to the calculation formula for the structural similarity measurement, which will not be repeated here.
[0120] In a possible implementation, a diagonal matrix data set corresponding to at least two time series data sets is obtained based on the heat maps corresponding to the at least two time series data sets; the arrangement order of each time series data type of different time series data sets corresponding to the horizontal coordinates and the vertical coordinates in the diagonal matrix data set; the elements in the matrix are used to indicate the similarity of the time series data set. Clustering the time series data types in the matrix can obtain different sets of time series data types, so as to quickly obtain feature rules from a global perspective.
[0121] Furthermore, according to the element values in the diagonal matrix data set, a corresponding matrix heat map is displayed on a display device.
[0122] The SSIM algorithm is used to calculate the similarity between the data sets encoded as heat maps, extract the parameter similarity features in the original data set, convert the original high-dimensional multi-period parameter data set into a single high-dimensional feature data set, and then analyze the parameter feature data. How to quickly obtain regular features from the results of metric analysis is crucial for feature analysis. Therefore, it is necessary to use reasonable visualization methods to display the metric results. This application designs a matrix diagram to perform visual analysis on the similarity calculation results, such as Figure 4 As shown in the figure, the result matrix of the high-dimensional parameter heat map after similarity analysis by the SSIM algorithm is shown. The horizontal axis of the matrix is arranged from left to right according to the arrangement order of the parameters of the data set, and the vertical axis is arranged from top to bottom. The square blocks represent the similarity of the two heat maps calculated by the SSIM algorithm, where the color of the block is mapped by the similarity value, and the mapping range is from light blue to dark blue. Light blue indicates low similarity and dark blue indicates high similarity. This view can clearly show the similarity between the heat maps that characterize the influencing parameters.
[0123] Therefore, when analyzing a data set with high-dimensional time series characteristics, the data is first preprocessed using normalization and missing value processing methods. Based on the extreme values of different parameters in the data set, the min-max method is used to achieve normalization, and the data is normalized to [0, 1] to unify the mapping range of each column color in the heat map. Then the mean interpolation method is used to process missing values in the data set.
[0124] Secondly, the preprocessed data set is converted and encoded using visualization technologies such as Echarts.js or D3.js, and visualized as a thermal image. Since different data sets after preprocessing have non-equal length characteristics, isometric processing is performed based on the shortest parameter time series before view conversion to characterize the initial change mode of the data feature state. Then, based on the normalized data set, the color mapping in the matrix thermal map is determined. After visual encoding, the original data set is converted into a matrix thermal map that represents the feature change state.
[0125] Finally, the image similarity measurement technology is used to measure and analyze the encoded heat map to achieve feature analysis and efficient extraction. This step uses the SSIM algorithm to calculate the similarity between the heat maps, and the result is a diagonal matrix data set. Then a clustering algorithm, such as the K-Means algorithm, is used to perform cluster analysis on the extracted parameter similarity features to mine the characteristic laws between high-dimensional time series data sets. In addition, in order to accelerate the exploration process of feature analysis, the diagonal matrix data is mapped with the help of visualization methods to quickly obtain the characteristic laws;
[0126] In summary, when processing time series data in the industrial field, at least two time series data sets can be obtained; for each time series data set, a heat map mapping the time series data set is generated with the time series as the horizontal axis and the time series data type as the vertical axis; according to the heat maps corresponding to the at least two time series data sets, the similarity between the at least two time series data sets is obtained. The above scheme converts the time series data into a heat map, and then calculates the similarity between the heat maps, thereby fully considering the time series characteristics of the data set, and can quickly and accurately obtain the similarity between high-dimensional time series control data with high-density sampling.
[0127] In the embodiments of the present application, a time series data analysis device is also provided, which is used to implement the above embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0128] The present application provides a time series data analysis device. Figure 5 : is a schematic diagram of the structure of a time series data analysis device provided in an embodiment of the present application, the device comprising:
[0129] A data set acquisition module 501 is used to acquire at least two time series data sets;
[0130] A heat map generation module 502 is used to generate a heat map corresponding to each time series data set, with the time series as the horizontal axis and the time series data type as the vertical axis;
[0131] The similarity acquisition module 503 is used to acquire the similarity between the at least two time series data sets according to the heat maps respectively corresponding to the at least two time series data sets.
[0132] In a possible implementation, the heat maps corresponding to the at least two time series data sets respectively include a first heat map corresponding to the first time series data set and a second heat map corresponding to the second time series data set;
[0133] The similarity acquisition module is used to acquire a brightness contrast function value according to the average grayscale of the first heat map and the average grayscale of the second heat map;
[0134] Obtaining a standard deviation comparison function value according to a standard deviation of the first thermogram and a standard deviation of the second thermogram;
[0135] Obtaining a structure contrast function value according to the covariance between the first thermodynamic map and the second thermodynamic map;
[0136] The brightness contrast function value, the standard deviation contrast function value, and the structure contrast function value are weighted to obtain a similarity between the first time series data set and the second time series data set.
[0137] In a possible implementation manner, the device further includes:
[0138] A maximum value acquisition module, used to obtain the maximum value and the minimum value of each time series data type in the time series data set;
[0139] The normalization module is used to perform normalization processing on the data of each time series data type in the time series data set based on the maximum value and the minimum value of each time series data type.
[0140] In a possible implementation manner, the device further includes:
[0141] A missing data detection module is used to sort the data of each time series data type in the time series data set according to time, and determine the position of missing data in the time series data set;
[0142] The data filling module is used to fill data at the missing data position according to the time series data adjacent to the missing data position in the time series data set.
[0143] In a possible implementation, the similarity acquisition module is further used to obtain a diagonal matrix data set corresponding to the at least two time series data sets according to the heat maps respectively corresponding to the at least two time series data sets; the arrangement order of each time series data type of different time series data sets corresponding to the horizontal coordinates and the vertical coordinates in the diagonal matrix data set; the elements in the diagonal matrix data set are used to indicate the similarity of the time series data sets;
[0144] The device also includes:
[0145] A clustering module is used to perform clustering processing on the time series data types in the diagonal matrix data set to obtain a time series data type set, and the time series data type set is used to analyze the correlation between the time series data types.
[0146] In a possible implementation manner, the device further includes:
[0147] The heat map display module is used to display the corresponding matrix heat map on the display device according to the element values in the diagonal matrix data set.
[0148] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0149] The timing data analysis device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0150] The embodiment of the present invention also provides a computer device having the above Figure 5 The timing data analysis device shown.
[0151] See also Figure 6 , Figure 6 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display graphic information in a graphical user interface on an external input / output device (such as a display device coupled to an interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.
[0152] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0153] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0154] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0155] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0156] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0157] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0158] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.
[0159] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the present invention.
Claims
1. A time series data analysis method, characterized in that: The method comprises: Obtain at least two time series data sets; For each time series data set, a heat map corresponding to the time series data set is generated with the time series as the horizontal axis and the time series data type as the vertical axis; The similarity between the at least two time series data sets is obtained according to the heat maps respectively corresponding to the at least two time series data sets.
2. The method according to claim 1, characterized in that The heat maps corresponding to the at least two time series data sets respectively include a first heat map corresponding to the first time series data set and a second heat map corresponding to the second time series data set; Acquiring the similarity between the at least two time series data sets according to the heat maps respectively corresponding to the at least two time series data sets includes: Obtaining a brightness contrast function value according to an average grayscale of the first heat map and an average grayscale of the second heat map; Obtaining a standard deviation comparison function value according to a standard deviation of the first thermogram and a standard deviation of the second thermogram; Obtaining a structure contrast function value according to a covariance between the first thermodynamic map and the second thermodynamic map; The brightness contrast function value, the standard deviation contrast function value, and the structure contrast function value are weighted to obtain a similarity between the first time series data set and the second time series data set.
3. The method according to claim 2, characterized in that Before generating a heat map corresponding to each time series data set with the time series as the horizontal axis and the time series data type as the vertical axis, the method further includes: Obtain the maximum value and the minimum value of each time series data type in the time series data set; Based on the maximum value and the minimum value of each time series data type, the data of each time series data type in the time series data set is normalized.
4. The method according to claim 3, characterized in that The method further comprises: Sort the data of each time series data type in the time series data set by time, and determine the position of missing data in the time series data set; Fill the missing data position with data according to the time series data adjacent to the missing data position in the time series data set.
5. The method according to any one of claims 1 to 4, characterized in that: The obtaining the similarity between the at least two time series data sets according to the heat maps respectively corresponding to the at least two time series data sets includes: According to the heat maps corresponding to the at least two time series data sets, a diagonal matrix data set corresponding to the at least two time series data sets is obtained; the arrangement order of each time series data type of different time series data sets corresponding to the horizontal coordinates and the vertical coordinates in the diagonal matrix data set; the elements in the diagonal matrix data set are used to indicate the similarity of the time series data sets; The method further comprises: Clustering is performed on the time series data types in the diagonal matrix data set to obtain a time series data type set, and the time series data type set is used to analyze the correlation between the time series data types.
6. The method according to claim 5, characterized in that The method further comprises: According to the element values in the diagonal matrix data set, the corresponding matrix heat map is displayed on the display device.
7. A time series data analysis device, characterized in that: The device comprises: A data set acquisition module, used to acquire at least two time series data sets; A heat map generation module is used to generate a heat map corresponding to each time series data set, with the time series as the horizontal axis and the time series data type as the vertical axis; The similarity acquisition module is used to acquire the similarity between the at least two time series data sets according to the heat maps respectively corresponding to the at least two time series data sets.
8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the time series data analysis method according to any one of claims 1 to 6 by executing the computer instructions.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the time series data analysis method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the time series data analysis method according to any one of claims 1 to 6.