Test data inflection point identification method based on multivariate time sequence analysis

By adjusting the neighborhood distance of the DBSCAN algorithm and combining the distribution uniformity and high-dimensional redundant correlation of multivariate time series data, the problem of insufficient adaptability of the neighborhood distance of the traditional DBSCAN algorithm in multivariate time series test data is solved, and more accurate inflection point identification is achieved.

CN120724142APending Publication Date: 2025-09-30BEIJING XTEST TESTING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510851810.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

In the inflection point identification of multivariate time series experimental data, the neighborhood distance parameter of the traditional DBSCAN density clustering algorithm cannot adapt to the high-density distribution of local multivariate time series data, resulting in over-segmentation of data and inaccurate inflection point identification.

Method used

By collecting multivariate time series test data, forming each test data sequence, obtaining the peak-valley subsequences of adjacent peaks and troughs, calculating the distribution uniformity and linear trend correlation, combining high-dimensional redundant correlation, and adjusting the neighborhood distance of the DBSCAN algorithm to identify inflection points.

Benefits of technology

It improves the accuracy and stability of inflection point identification in multivariate time series data, enhances the adaptability to high-dimensional complex scenarios, reduces the impact of noise interference and outliers, and ensures the reliability of inflection point identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120724142A_ABST
    Figure CN120724142A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a test data inflection point identification method based on multivariate time sequence analysis, and the method comprises the steps: collecting multivariate time sequence test data, and forming each test data sequence; based on adjacent wave crests and wave troughs in each test data sequence, obtaining each peak-to-millet sequence in each test data sequence; determining the distribution uniformity of each test data sequence; obtaining the linear trend correlation degree of each test data sequence; obtaining the high-density tendency of each test data sequence, obtaining the multivariate data correlation degree of each peak millet sequence in each test data sequence, and determining the high-dimensional redundancy correlation degree of each test data sequence; and obtaining a neighborhood distance adjustment factor of each test data sequence, and adjusting the neighborhood distance when the DBSCAN density clustering algorithm identifies the inflection point of the multivariate time sequence test data. According to the invention, the accuracy of test data inflection point identification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method for identifying inflection points of experimental data based on multivariate time series analysis. Background Art

[0002] With the development of science and technology and advancements in data acquisition techniques, vast amounts of time series data have accumulated across various fields. In many disciplines and practical applications, inflection points in test data often signal performance degradation, failure occurrence, or the onset of a critical state. Accurately identifying inflection points in multivariate time series test data can provide a more comprehensive analysis and processing of information relationships within this data. However, modern multivariate time series data often exhibits characteristics such as high dimensionality, nonlinearity, and strong noise, posing significant challenges to accurately identifying inflection points in this data.

[0003] Traditional techniques for studying inflection points in multivariate experimental time series data often employ unsupervised learning clustering algorithms for inflection point identification. These algorithms eliminate the need for complex pre-labeling and offer the advantages of rapid convergence and high interpretability for large-scale multivariate time series data. However, when using the DBSCAN density clustering algorithm to identify inflection points in multivariate time series experimental data, the global parameter neighborhood distance cannot adapt to the high-density distribution of local multivariate time series experimental data. Excessively high neighborhood distances can lead to over-segmentation of the experimental data, and the high-dimensionality of the multivariate experimental data sequence itself can result in a high uniformity in the data density distribution, making it difficult to identify inflection points through effective clustering results. This results in inaccurate analysis of inflection points in multivariate time series experimental data. Summary of the Invention

[0004] In order to solve the above technical problems, the present application provides a test data inflection point identification method based on multivariate time series analysis to solve the existing problems.

[0005] The test data inflection point identification method based on multivariate time series analysis in this application adopts the following technical solutions:

[0006] One embodiment of the present application provides a method for identifying inflection points in experimental data based on multivariate time series analysis, the method comprising the following steps:

[0007] Collect multivariate time series test data and form each test data sequence;

[0008] Based on adjacent peaks and valleys in each test data sequence, each peak-valley subsequence in each test data sequence is obtained; based on the degree of dispersion of each peak-valley subsequence in each test data sequence and the difference between the peak-valley subsequences, the distribution uniformity of each test data sequence is determined;

[0009] The linear trend correlation of each test data sequence is determined by using the distance relationship between the data in each test data sequence and the value range of the data in each peak and valley subsequence in each test data sequence;

[0010] Combining the distribution uniformity with the linear trend correlation, a high-density trend of each test data sequence is obtained; for any test data sequence, each associated peak-valley subsequence of each peak-valley subsequence in the any test data sequence is obtained in each test data sequence other than the any test data sequence;

[0011] The time difference and distribution difference between each peak-valley subsequence and its associated peak-valley subsequence in each test data sequence are analyzed to determine the multivariate data correlation degree of each peak-valley subsequence in each test data sequence. The high-dimensional redundant correlation degree of each test data sequence is obtained by combining the number of principal components of each peak-valley subsequence in each test data sequence and the similarity between the principal components.

[0012] The linear trend correlation and the high-dimensional redundancy association are integrated to obtain the neighborhood distance adjustment factor of each test data sequence, and adjust the neighborhood distance when the DBSCAN density clustering algorithm identifies the inflection point of the multivariate time series test data.

[0013] In one embodiment, obtaining each peak-valley subsequence includes:

[0014] For any peak in each test data sequence, the trough with the shortest time interval with the peak is obtained and recorded as the associated trough of the peak. The time series test data between the peak and its associated trough are combined into a peak-trough subsequence.

[0015] In one embodiment, the process of determining the distribution uniformity is:

[0016] The cumulative sum of the discrete degrees of all peak-valley subsequences in each test data sequence is calculated, recorded as the first sum value, and the cumulative sum of the differences between all arbitrary two peak-valley subsequences in each test data sequence is calculated, recorded as the second sum value. The distribution uniformity is negatively correlated with the first sum value and the second sum value.

[0017] In one embodiment, determining the linear trend correlation includes:

[0018] Based on the correlation between the data in each test data sequence, the data in each test data sequence is divided into clusters; the cumulative sum of the metric distances between all data in each cluster and the corresponding cluster center is calculated, recorded as the third sum value, the discrete degree of the third sum value of all clusters in each test data sequence is calculated, and multiplied by the cumulative sum of the extreme values ​​of all peak and valley subsequences in each test data sequence. The smaller the result of the multiplication, the greater the linear trend correlation.

[0019] In one embodiment, the process of determining each associated peak-valley subsequence is as follows:

[0020] The mean of the moments at which the peaks and troughs in each peak-valley subsequence are located is taken as the central moment of each peak-valley subsequence. In each test data sequence other than any one of the test data sequences, the peak-valley subsequences closest to the central moment of any peak-valley subsequence in any one of the test data sequences are obtained as the associated peak-valley subsequences of any one of the peak-valley subsequences in the any one of the test data sequences.

[0021] In one embodiment, the process of determining the multivariate data association degree is as follows:

[0022] The cumulative sum of the differences between the central moments of each peak-valley subsequence and all its associated peak-valley subsequences in each test data sequence is calculated, and recorded as the fourth sum value. The cumulative sum of the differences between each peak-valley subsequence and all its associated peak-valley subsequences in each test data sequence is calculated, and recorded as the fifth sum value. The multivariate data correlation degree is the inverse of the product of the fourth sum value and the fifth sum value.

[0023] In one embodiment, determining the high-dimensional redundancy correlation degree includes:

[0024] The principal component analysis algorithm is used to obtain the principal component of each peak-valley subsequence, and the cumulative sum of the similarities between all arbitrary two principal components of each peak-valley subsequence is calculated, recorded as the sixth sum value. The ratio of the sixth sum value of each peak-valley subsequence to the number of its principal components is calculated. Combined with the multivariate data association degree, the high-dimensional redundant association degree of each test data sequence is determined.

[0025] In one embodiment, the multiplication result of the ratio of each peak-valley subsequence in each test data sequence and the multivariate data correlation is calculated, and the sum of the multiplication results of all peak-valley subsequences in each test data sequence is calculated as the high-dimensional redundant correlation of each test data sequence.

[0026] In one embodiment, adjusting the neighborhood distance when the DBSCAN density clustering algorithm identifies the inflection point of multivariate time series test data includes:

[0027] Each experimental data sequence of the multivariate time series experimental data is used as the input of the DBSCAN density clustering algorithm to obtain each clustering cluster of each experimental data sequence. During the clustering process, when the neighborhood distance adjustment factor of the experimental data sequence is less than or equal to the preset threshold, the initial neighborhood distance of the DBSCAN density clustering algorithm is not adjusted. Otherwise, the neighborhood distance of the DBSCAN density clustering algorithm is reduced based on the neighborhood distance adjustment factor of the experimental data sequence.

[0028] In one embodiment, the mean of all data in each cluster of each experimental data sequence is calculated as the cluster center of each cluster, and a segmentation threshold of the absolute value of the difference between all data in each cluster and the corresponding cluster center is obtained, and the experimental data corresponding to the absolute value of the difference in each cluster that is greater than the segmentation threshold is used as the inflection point.

[0029] This application has at least the following beneficial effects:

[0030] The present application collects multivariate time series test data to form each test data sequence; based on the adjacent peaks and troughs in each test data sequence, obtains each peak and valley subsequence in each test data sequence; determines the distribution uniformity of each test data sequence according to the degree of dispersion of each peak and valley subsequence in each test data sequence and the difference between the peak and valley subsequences; the determination of distribution uniformity improves the ability to perceive the data fluctuation pattern in a refined manner, helps to reflect the data density of the test data sequence, and improves the accuracy of the advance prediction of the data distribution characteristics in the test data sequence; utilizes the distance relationship between the data in each test data sequence and the value range of the data in each peak and valley subsequence in each test data sequence to determine the linear trend correlation of each test data sequence; the linear trend correlation reflects the linear correlation of the data in each cluster in the test data sequence and the density change condition near the data fluctuation point, enhances the collaborative trend capture ability of the multivariate time series data, and improves the reliability of the subsequent inflection point detection; combines the distribution uniformity with the linear trend correlation to obtain the high-density trend of each test data sequence; the high-density trend reflects the data distribution density condition and the linear trend correlation degree of each test data sequence in the multivariate test data. The accurate identification of the distribution of test data improves the accuracy of the subsequent adjustment of the neighborhood distance of the DBSCAN density clustering algorithm; further, the time difference and distribution difference of each peak-valley subsequence and its associated peak-valley subsequence in each test data sequence are analyzed, and the multivariate data correlation of each peak-valley subsequence in each test data sequence is determined. Combined with the number of principal components of each peak-valley subsequence in each test data sequence and the similarity between the principal components, the high-dimensional redundant correlation of each test data sequence is obtained; the high-dimensional redundant correlation reflects the data association status between multivariate test data and the degree of interference by high-dimensional redundant information, which is helpful The aim is to break through the limitations of traditional density clustering in high-dimensional complex scenarios, ensure the appropriateness of neighborhood distance, improve the identification accuracy of high-dimensional data cluster structure, and enhance the robustness against noise interference and outliers; the linear trend correlation and the high-dimensional redundant correlation are integrated to obtain the neighborhood distance adjustment factor of each test data sequence, and adjust the neighborhood distance when the DBSCAN density clustering algorithm identifies the inflection point of multivariate time series test data, thereby improving the adaptability of the DBSCAN algorithm to complex density distribution, reducing the negative impact of local density fluctuations of test data on clustering results, and ensuring the stability and accuracy of inflection point identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0032] Figure 1 A flowchart of the steps of the method for identifying inflection points of experimental data based on multivariate time series analysis provided in this application;

[0033] Figure 2 Flowchart for determining the neighborhood distance adjustment factor. DETAILED DESCRIPTION

[0034] In order to further illustrate the technical means and effects adopted by this application to achieve the predetermined invention objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features and effects of the test data inflection point identification method based on multivariate time series analysis proposed in this application. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics of one or more embodiments may be combined in any suitable form.

[0035] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0036] The specific scheme of the test data inflection point identification method based on multivariate time series analysis provided by the present application is described in detail below with reference to the accompanying drawings.

[0037] An embodiment of the present application provides a method for identifying inflection points of test data based on multivariate time series analysis. Specifically, the following method for identifying inflection points of test data based on multivariate time series analysis is provided. Figure 1 , the method comprises the following steps:

[0038] Step S001: Collect multivariate time series test data to form each test data sequence.

[0039] Sensor instruments are used to obtain multivariate test data in the test scenario, wherein all types of test data are collected synchronously, and the data sampling frequency is set to 1 Hz. All test data obtained in this embodiment are numerical data. Z-Score standardization is used to standardize and normalize the collected multivariate test data to avoid the subsequent analysis results being affected by different data dimensions. The collected multivariate test data are filled with missing values ​​and denoised by linear interpolation and Kalman filtering to avoid data loss and serious data noise interference caused by transmission fluctuations and external environmental interference during data transmission. Since Z-Score standardization, linear interpolation and Kalman filtering are all well-known technologies, the specific process will not be repeated. The implementer can use other existing methods to standardize, fill in values ​​and filter the multivariate test data to achieve preprocessing of multivariate time series test data.

[0040] Each type of preprocessed test data is organized into a test data sequence in a time series order, thereby obtaining each test data sequence.

[0041] Step S002: Based on adjacent peaks and troughs in each test data sequence, obtain each peak-valley subsequence in each test data sequence; and determine the distribution uniformity of each test data sequence according to the degree of dispersion of each peak-valley subsequence in each test data sequence and the difference between the peak-valley subsequences.

[0042] In the process of identifying inflection points of multivariate experimental data using the DBSCAN density clustering algorithm, the neighborhood distance parameter in the DBSCAN density clustering algorithm is highly affected by the local data density. The data density in different time series ranges is different, and the fixed neighborhood distance cannot effectively adapt to this data density change; the nonlinear trend in the experimental data changes will also make the local data density more complex, and the uneven distribution of data points will further aggravate the local data density difference, making the boundaries of the clusters unclear and affecting the accuracy of inflection point identification.

[0043] Specifically, in each experimental data sequence, the stronger the relationship between data changes and time dependence in the experimental data sequence, the smaller the difference between data fluctuation intervals, and the more blurred the difference in probability distribution of data fluctuation intervals, the higher the data density of the experimental data sequence. At this time, when using the DBSCAN clustering algorithm to identify the inflection points of the experimental data, the neighborhood distance value should be further reduced to avoid ignoring subtle changes in local density, resulting in inaccurate inflection point identification; at the same time, the nonlinear trend in the experimental data may cause the alternation of inflection point clustering areas and low-density clustering areas, resulting in significant changes in data density near the data fluctuation points and unclear boundaries of the clustering results, thereby concealing the true inflection points.

[0044] Based on the above analysis, this embodiment analyzes any test data sequence in the multivariate time series test data, and uses each test data sequence as the input of the AMPD (Automatic Multiscale-based Peak Detection) multi-scale peak automatic search algorithm to obtain all peaks and troughs in each test data sequence. The trough with the shortest time interval with each peak in each test data sequence is recorded as the associated trough of each peak, and the sequence composed of all data between each peak and its associated trough in each test data sequence is used as each peak and trough subsequence of each test data sequence.

[0045] The distribution uniformity of each test data sequence is determined based on the degree of dispersion of each peak-valley subsequence in each test data sequence and the differences between the peak-valley subsequences.

[0046] It should be noted that the degree of dispersion can be specifically calculated using variance, standard deviation, coefficient of variation, etc. The difference between peak and valley subsequences represents the degree of difference between the two sequences, which can be specifically calculated using DTW distance, Euclidean distance, KL divergence, JS divergence, etc. This embodiment does not impose any restrictions on this.

[0047] In this embodiment, the coefficient of variation of each peak-valley subsequence in each test data sequence is calculated. A larger coefficient of variation for each peak-valley subsequence indicates a more discrete data distribution and a lower data density within the peak-valley subsequence. Furthermore, the KL divergence is calculated between any two peak-valley subsequences in each test data sequence. The cumulative sum of the coefficients of variation of all peak-valley subsequences in each test data sequence is calculated, recorded as a first sum, and the cumulative sum of the KL divergences between all any two peak-valley subsequences in each test data sequence is calculated, recorded as a second sum. The distribution uniformity is negatively correlated with both the first sum and the second sum. In this embodiment, the reciprocal of the product of the first sum and the second sum for each test data sequence is used as the distribution uniformity of each test data sequence.

[0048] It should be noted that when there is only one peak-valley subsequence in the test data sequence, the reciprocal of the product of the coefficient of variation of the test data sequence and the information entropy is used as the distribution uniformity of the test data sequence.

[0049] The distribution uniformity of each experimental data sequence reflects the data distribution discreteness in the experimental data sequence and the difference between the data fluctuation ranges. Among them, the data fluctuation range is the peak-valley subsequence.

[0050] Step S003: Determine the linear trend correlation of each test data sequence using the distance relationship between the data in each test data sequence and the value range of the data in each peak and valley subsequence in each test data sequence. Combined with the distribution uniformity and the linear trend correlation, the high-density trend of each test data sequence is obtained.

[0051] Furthermore, based on the correlation between the data in each test data sequence, the data in each test data sequence is divided into clusters. In this embodiment, each test data sequence is used as input to the K-mediods clustering algorithm to obtain all clusters and cluster centers corresponding to each test data sequence. The number of cluster centers in the K-mediods clustering algorithm is set to K, and the absolute value of the data difference between data points is used as the metric distance between data points. The cumulative metric distance between all data within each cluster in each test data sequence and the cluster center is recorded as the third sum of each cluster, representing the intra-cluster distance of each cluster. Since the K-mediods clustering algorithm is a well-known technique, the specific process will not be described in detail. Using the K-mediods clustering algorithm facilitates obtaining the actual cluster centers in each test data sequence. The number of cluster centers K is obtained using the elbow method, which is a well-known technique. In this embodiment, the metric distance is calculated using the Euclidean distance. Implementers can choose other feasible metric distance calculation methods.

[0052] Next, the degree of dispersion of the third sum of all clusters in each test data sequence is calculated. In this embodiment, the variance of the third sum of all clusters in each test data sequence is calculated, the range of the data within each peak-valley subsequence in each test data sequence is determined, and the product of the variance of the third sum of all clusters in each test data sequence and the cumulative sum of the ranges of all peak-valley subsequences in each test data sequence is calculated, recorded as a first product. Based on the first product, the linear trend correlation of each test data sequence is determined, wherein the smaller the first product, the greater the linear trend correlation. In this embodiment, the reciprocal of the first product is used as the linear trend correlation of each test data sequence.

[0053] The linear trend correlation of each experimental data sequence reflects the linear correlation of the data within each cluster in the experimental data sequence and the density change near the data fluctuation point.

[0054] This embodiment combines the distribution uniformity and the linear trend correlation to determine the high-density trend of each test data sequence, which is used to characterize the data time series correlation and unclear clustering results in each test data sequence. The specific calculation method is:

[0055] Where, is the high density trend of the i-th experimental data sequence; is the distribution uniformity of the i-th test data sequence, is the linear trend correlation of the i-th experimental data sequence.

[0056] In another embodiment, the sum of the distribution uniformity of the i-th test data sequence and the linear trend correlation of the i-th test data sequence is used as the high-density trend of the i-th test data sequence.

[0057] It should be understood that high-density trend reflects the data distribution density of each test data sequence in the multivariate test data and the degree of correlation of the linear trend; when the time series data density in the test data sequence is higher and the linear trend is more obvious, the probability distribution difference between the data fluctuation intervals in the test data sequence is more vague, the data dispersion within the data fluctuation interval is smaller, and the calculation of distribution uniformity is more accurate. At the same time, the clustering result boundary of the test data sequence is clearer due to the high linear trend of the test data, the data distribution of each cluster in the test data sequence is tighter, and the data distribution in each data fluctuation range is denser, and the linear trend correlation is calculated. Get bigger.

[0058] When the high-density trend is stronger, it means that the data distribution density in the test data sequence is higher and the linear trend is stronger. At this time, when using the DBSCAN density clustering algorithm to obtain the inflection point of the multivariate test data, the neighborhood distance of the DBSCAN density clustering algorithm should be reduced to avoid slight changes in the test data sequence being ignored, resulting in inaccurate inflection point identification or the true inflection point being concealed.

[0059] Step S004: For any test data sequence, obtain, from each test data sequence other than the test data sequence, each associated peak-valley subsequence of each peak-valley subsequence within the test data sequence. Analyze the temporal and distributional differences between each peak-valley subsequence and its associated peak-valley subsequence in each test data sequence to determine the multivariate data correlation of each peak-valley subsequence in each test data sequence. Combined with the number of principal components of each peak-valley subsequence in each test data sequence and the similarity between the principal components, a high-dimensional redundancy correlation of each test data sequence is obtained.

[0060] However, if we only rely on the data distribution density and linear trend in the test data as the basis for adjusting the neighborhood distance of the DBSCAN density clustering algorithm, it is easy to lack the analysis of data correlation between multivariate time series test data and the consideration of the high-dimensional information redundancy of the test data itself, which may cause the inflection points identified by the clustering results to lack interpretability, that is, they cannot reflect the essential characteristics of the test data; at the same time, when the test data sequence is more seriously disturbed by its own high-dimensional redundant information, the test data will contain a large amount of meaningless noise data, diluting the effective information in the test data. At this time, it is more important to limit the neighborhood distance of the DBSCAN density clustering algorithm to offset the high-dimensional information redundancy dilution effect, enhance the sensitivity of the DBSCAN density clustering algorithm to local density, and thus accurately identify inflection points.

[0061] Specifically, in the process of using the DBSCAN density clustering algorithm to identify inflection points in each experimental data sequence in multivariate time series experimental data, the more similar the data differences and data change trends in the data fluctuation ranges between the experimental data sequences are, the higher the correlation between the experimental data sequences is, and a smaller DBSCAN algorithm neighborhood distance is needed to ensure that the close relationship between the data points can be captured more accurately; and the weaker the orthogonality between the dimensional components of the data fluctuation range in each experimental data sequence is, the data information in the data fluctuation range in the experimental data sequence cannot be well dispersed into multiple orthogonal dimensional components, the more seriously the data sequence is affected by high-dimensional redundancy, and a smaller DBSCAN algorithm neighborhood distance should be used to avoid the true inflection point being concealed.

[0062] Based on the above analysis, this embodiment determines the high-dimensional redundancy correlation of each test data sequence, which is used to characterize the similarity of the data fluctuation range between each type of test data and the rest of the test data in the multivariate time series test data, as well as the interference status of high-dimensional redundant information. Specifically:

[0063] First, the mean of the peak moments and trough moments of each peak-valley subsequence in each test data sequence is calculated as the center moment of each peak-valley subsequence in each test data sequence. The peaks and troughs in each peak-valley subsequence are the peaks and troughs used to divide each peak-valley subsequence when performing peak detection on each test data sequence. Taking the jth peak-valley subsequence in the i-th test data sequence as an example, for each test data sequence except the i-th test data sequence in the multivariate time series test data, calculate the absolute value of the difference between the central moment of each peak-valley subsequence and the j-th peak-valley subsequence, and record it as the absolute value of the central moment difference. In each test data sequence except the i-th test data sequence, obtain the peak-valley subsequence with the smallest absolute value of the central moment difference of the j-th peak-valley subsequence, as the associated peak-valley subsequence of the j-th peak-valley subsequence in the i-th test data sequence, that is, the number of associated peak-valley subsequences of the j-th peak-valley subsequence is the number of test data sequences except the i-th test data sequence in the multivariate time series test data.

[0064] Next, the time and distribution differences between each peak-valley subsequence and its associated peak-valley subsequences in each test data sequence are analyzed to determine the multivariate data correlation of each peak-valley subsequence in each test data sequence. Specifically, the cumulative absolute value of the center time difference between each peak-valley subsequence in each test data sequence and all its associated peak-valley subsequences is calculated, recorded as the fourth sum, and the cumulative JS divergence between each peak-valley subsequence in each test data sequence and all its associated peak-valley subsequences is calculated, recorded as the fifth sum. The multivariate data correlation of each peak-valley subsequence in each test data sequence is the inverse of the product of the fourth sum and the fifth sum.

[0065] The multivariate data correlation reflects the time difference of the data fluctuation range and the data difference status between each experimental data sequence and the other experimental data sequences.

[0066] Furthermore, each peak-valley subsequence within each test data sequence in the multivariate test data is used as input, and a principal component analysis (PCA) algorithm is used to obtain all principal dimensional components and the corresponding eigenvectors of each principal dimensional component for each peak-valley subsequence in each test data sequence. Since the PCA principal component analysis method is a well-known technique, the specific acquisition process is not described in detail here.

[0067] In this embodiment, the high-dimensional redundant correlation degree of each test data sequence is calculated as follows:

[0068] Where, is the high-dimensional redundant correlation degree of the i-th test data sequence; is the multivariate data correlation degree of the j-th peak-valley subsequence of the i-th experimental data sequence, is the total number of peak-valley subsequences in the i-th test data sequence, is the total number of dimensional components corresponding to the j-th peak-valley subsequence in the i-th test data sequence, It is the cumulative result of the similarity between all arbitrary two dimensional component feature vectors of the j-th peak-valley subsequence in the i-th experimental data sequence, recorded as the sixth sum value.

[0069] It should be noted that the similarity described in this embodiment is calculated using cosine similarity, and implementers can choose other existing feasible similarity calculation methods, such as Pearson correlation coefficient.

[0070] It should be understood that the high-dimensional redundant correlation reflects the data association status between multivariate test data and the degree of interference by high-dimensional redundant information; in multivariate test data, the smaller the difference in the generation time of the data fluctuation interval between any test data sequence and the rest of the test data sequences, and the smaller the data difference between the fluctuation intervals of each associated data, the higher the data association between the test data sequence and the rest of the test data sequences. becomes larger; and when the high-dimensional redundant correlation condition is more serious, the independent information sources that affect the data changes within the range of each data fluctuation interval in the test data sequence are fewer, and the total number of main dimension components in each data fluctuation interval in the test data sequence is smaller. The smaller it is, the less the data information in the data fluctuation range in the test data sequence can be dispersed into multiple orthogonal principal dimension components, and the higher the similarity of the eigenvectors corresponding to the principal dimension components in the data fluctuation range. Get bigger.

[0071] Step S005 , fusing the linear trend correlation and the high-dimensional redundancy correlation to obtain a neighborhood distance adjustment factor for each test data sequence, and adjusting the neighborhood distance when the DBSCAN density clustering algorithm identifies the inflection point of the multivariate time series test data.

[0072] In multivariate time series test data, when the data distribution density of all test data sequences is higher and the linear trend correlation is stronger, it is easier to ignore subtle changes in the high-density data distribution when using the DBSCAN density clustering algorithm to identify inflection points; at the same time, when the correlation between the test data sequence and the other test data sequences is stronger and the interference from its own high-dimensional redundant information is more serious, when using the DBSCAN density clustering algorithm to identify inflection points, it is more necessary to use a smaller DBSCAN algorithm neighborhood distance to ensure that the close relationship between data points can be captured more accurately, avoiding the true inflection points from being obscured.

[0073] Based on the above analysis, this embodiment constructs a neighborhood distance adjustment factor for each experimental data sequence, which is used to characterize the extent to which the neighborhood distance of the DBSCAN density clustering algorithm needs to be reduced when using the DBSCAN density clustering algorithm to identify inflection points in multivariate experimental time series data.

[0074] Where, is the neighborhood distance adjustment factor of the i-th experimental data sequence in the multivariate time series experimental data; is the high density trend of the i-th experimental data sequence; is the high-dimensional redundant correlation of the i-th test data sequence; norm() is the normalization function, so that The value range is between [0,1]. The flowchart for determining the neighborhood distance adjustment factor is as follows: Figure 2 shown.

[0075] In the process of identifying inflection points in multivariate test data using the DBSCAN density clustering algorithm, when the data distribution density in each test data sequence is higher and the linear trend correlation is stronger, it is more necessary to promptly lower the neighborhood distance of the DBSCAN density clustering algorithm to capture subtle changes in the inflection points in the high-density data distribution; at the same time, when the correlation between each test data sequence is stronger and the interference from its own high-dimensional redundant information is more serious, the neighborhood distance in the DBSCAN density clustering algorithm should be lowered to avoid the effective information of the test data being masked by the high-dimensional redundant information, resulting in inaccurate inflection point identification.

[0076] Therefore, this embodiment sets a neighborhood distance adjustment threshold Q and an initial value W of the neighborhood distance in the DBSCAN density clustering algorithm. When the neighborhood distance adjustment factor of the test data sequence is less than or equal to the neighborhood distance adjustment threshold Q, it is determined that the DBSCAN density clustering algorithm is used to obtain the inflection point in the current test data sequence. It is able to accurately capture subtle changes in data under high-density distribution, and is slightly interfered with by its own high-dimensional redundant information. There is no need to adjust the initial value of the neighborhood distance in the DBSCAN density clustering algorithm. Otherwise, when the neighborhood distance adjustment factor of the test data sequence is greater than the neighborhood distance adjustment threshold Q, it is determined that the DBSCAN density clustering algorithm is used to obtain the inflection point in the current test data sequence. It is unable to accurately capture subtle changes in data under high-density distribution, and is seriously interfered with by its own high-dimensional redundant information. It is necessary to promptly reduce the neighborhood distance in the DBSCAN density clustering algorithm to ensure accurate identification of inflection points in subsequent test data.

[0077] It should be noted that when it is necessary to reduce the neighborhood distance in the DBSCAN density clustering algorithm, this embodiment obtains the neighborhood distance of the DBSCAN density clustering algorithm through the neighborhood distance adjustment factor of the corresponding test data sequence. Specifically, a nonlinear mapping method is used to map the neighborhood distance adjustment factor of the corresponding test data sequence to the interval [b, W), where W is the neighborhood distance initial value and b is the neighborhood distance lower limit, to prevent the clusters from being over-segmented and unable to form an effective clustering structure. The nonlinear mapping described in this embodiment uses the Sigmoid function for mapping. The implementer can choose other existing feasible nonlinear mapping methods at will, and this embodiment does not limit this. In this embodiment, the neighborhood distance adjustment threshold Q=0.7, the neighborhood distance initial value W=0.4, and the neighborhood distance lower limit b=0.25. The implementer can adjust them according to the actual situation, and this embodiment does not limit this.

[0078] The specific method of using the DBSCAN density clustering algorithm to obtain the inflection points in each test data sequence is as follows:

[0079] Each test data sequence in the multivariate time series test data is used as input, and the DBSCAN density clustering algorithm is used to obtain all clusters in each test data sequence. The mean of all data in each cluster is used as the cluster center of each cluster. The minimum sample size is set to 200. The sequence consisting of the absolute values ​​of the difference between all test data and the cluster center in each cluster in ascending order is recorded as the center distance sequence of each cluster. The center distance sequence corresponding to each cluster in each test data sequence is used as the input of the OTSU method. The segmentation threshold of the center distance sequence of each cluster in the test data sequence is obtained. The test data corresponding to the absolute value of the difference in the center distance sequence of each cluster that is greater than the segmentation threshold is used as the inflection point in the corresponding test data sequence to complete the inflection point identification of the multivariate time series test data. Among them, the DBSCAN density clustering algorithm and the OTSU method are both existing well-known technologies, and the specific process is not described in detail.

[0080] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0081] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0082] The above-described embodiments are only used to illustrate the technical solutions of the present application, and not to limit them. Modifications to the technical solutions described in the aforementioned embodiments, or equivalent replacements of some of the technical features therein, do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for identifying inflection points of experimental data based on multivariate time series analysis, characterized in that: The method comprises the following steps: Collect multivariate time series test data and form each test data sequence; Based on adjacent peaks and valleys in each test data sequence, each peak-valley subsequence in each test data sequence is obtained; based on the degree of dispersion of each peak-valley subsequence in each test data sequence and the difference between the peak-valley subsequences, the distribution uniformity of each test data sequence is determined; The linear trend correlation of each test data sequence is determined by using the distance relationship between the data in each test data sequence and the value range of the data in each peak and valley subsequence in each test data sequence; Combining the distribution uniformity with the linear trend correlation, a high-density trend of each test data sequence is obtained; for any test data sequence, each associated peak-valley subsequence of each peak-valley subsequence in the any test data sequence is obtained in each test data sequence other than the any test data sequence; The time difference and distribution difference between each peak-valley subsequence and its associated peak-valley subsequence in each test data sequence are analyzed to determine the multivariate data correlation degree of each peak-valley subsequence in each test data sequence. The high-dimensional redundant correlation degree of each test data sequence is obtained by combining the number of principal components of each peak-valley subsequence in each test data sequence and the similarity between the principal components. The linear trend correlation and the high-dimensional redundancy association are integrated to obtain the neighborhood distance adjustment factor of each test data sequence, and adjust the neighborhood distance when the DBSCAN density clustering algorithm identifies the inflection point of the multivariate time series test data.

2. The method for identifying inflection points of experimental data based on multivariate time series analysis according to claim 1, characterized in that: The acquisition of each peak-valley subsequence includes: For any peak in each test data sequence, the trough with the shortest time interval with the peak is obtained and recorded as the associated trough of the peak. The time series test data between the peak and its associated trough are combined into a peak-trough subsequence.

3. The method for identifying inflection points of experimental data based on multivariate time series analysis according to claim 1, wherein: The process of determining the distribution uniformity is as follows: The cumulative sum of the discrete degrees of all peak-valley subsequences in each test data sequence is calculated, recorded as the first sum value, and the cumulative sum of the differences between all arbitrary two peak-valley subsequences in each test data sequence is calculated, recorded as the second sum value. The distribution uniformity is negatively correlated with the first sum value and the second sum value.

4. The method for identifying inflection points of experimental data based on multivariate time series analysis according to claim 1, wherein: The determination of the linear trend correlation includes: Based on the correlation between the data in each test data sequence, the data in each test data sequence is divided into clusters; the cumulative sum of the metric distances between all data in each cluster and the corresponding cluster center is calculated, recorded as the third sum value, the discrete degree of the third sum value of all clusters in each test data sequence is calculated, and multiplied by the cumulative sum of the extreme values ​​of all peak and valley subsequences in each test data sequence. The smaller the result of the multiplication, the greater the linear trend correlation.

5. The method for identifying inflection points of test data based on multivariate time series analysis according to claim 1, wherein: The process of determining each associated peak-valley subsequence is as follows: The mean of the moments at which the peaks and troughs in each peak-valley subsequence are located is taken as the central moment of each peak-valley subsequence. In each test data sequence other than any one of the test data sequences, the peak-valley subsequences closest to the central moment of any peak-valley subsequence in any one of the test data sequences are obtained as the associated peak-valley subsequences of any one of the peak-valley subsequences in the any one of the test data sequences.

6. The method for identifying inflection points of experimental data based on multivariate time series analysis according to claim 5, characterized in that: The process of determining the multivariate data association degree is as follows: The cumulative sum of the differences between the central moments of each peak-valley subsequence and all its associated peak-valley subsequences in each test data sequence is calculated, and recorded as the fourth sum value. The cumulative sum of the differences between each peak-valley subsequence and all its associated peak-valley subsequences in each test data sequence is calculated, and recorded as the fifth sum value. The multivariate data correlation degree is the inverse of the product of the fourth sum value and the fifth sum value.

7. The method for identifying inflection points of experimental data based on multivariate time series analysis according to claim 1, wherein: The determination of the high-dimensional redundancy correlation degree includes: The principal component analysis algorithm is used to obtain the principal component of each peak-valley subsequence, and the cumulative sum of the similarities between all arbitrary two principal components of each peak-valley subsequence is calculated, recorded as the sixth sum value. The ratio of the sixth sum value of each peak-valley subsequence to the number of its principal components is calculated. Combined with the multivariate data association degree, the high-dimensional redundant association degree of each test data sequence is determined.

8. The method for identifying inflection points of test data based on multivariate time series analysis according to claim 7, wherein: The multiplication result of the ratio of each peak-valley subsequence in each test data sequence and the multivariate data correlation degree is calculated, and the sum of the multiplication results of all peak-valley subsequences in each test data sequence is calculated as the high-dimensional redundant correlation degree of each test data sequence.

9. The method for identifying inflection points of test data based on multivariate time series analysis according to claim 1, wherein: The adjustment of the neighborhood distance when the DBSCAN density clustering algorithm identifies the inflection point of the multivariate time series test data includes: Each experimental data sequence of the multivariate time series experimental data is used as the input of the DBSCAN density clustering algorithm to obtain each clustering cluster of each experimental data sequence. During the clustering process, when the neighborhood distance adjustment factor of the experimental data sequence is less than or equal to the preset threshold, the initial neighborhood distance of the DBSCAN density clustering algorithm is not adjusted. Otherwise, the neighborhood distance of the DBSCAN density clustering algorithm is reduced based on the neighborhood distance adjustment factor of the experimental data sequence.

10. The method for identifying inflection points of test data based on multivariate time series analysis according to claim 9, characterized in that: Calculate the mean of all data in each cluster of each test data sequence as the cluster center of each cluster, obtain the segmentation threshold of the absolute value of the difference between all data in each cluster and the corresponding cluster center, and take the test data corresponding to the absolute value of the difference in each cluster that is greater than the segmentation threshold as the inflection point.