Method and device for detecting change point position and category
By baseline extraction, screening, grouping voting and tree structure split clustering of time sequence data of multi-dimensional physiological variable graphs, combined with GLR algorithm, the problem of variable point detection under label-free conditions is solved, efficient and accurate variable point position and category detection is achieved, and automated health monitoring in the medical field is supported.
Patent Information
- Application Number
- CN202210520719.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-13
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-05-13
AI Technical Summary
In the field of medical technology, it is difficult for the prior art to effectively detect variable point positions and categories in physiological variable data such as EEG and ECG without variable point positions and category labels.
By baseline extraction, screening candidate variable points, grouping voting and tree structure split clustering, combining generalized likelihood ratio (GLR) algorithm and top-down split hierarchical clustering, the detection of variable point positions and categories is achieved.
It realizes the accuracy and efficiency of variable point detection without relying on artificial parameters, and can automatically monitor the health status of medical patients.
Smart Images

Figure CN114818969B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical technology, and in particular, to a method for detecting the position and category of change points. Background Art
[0002] In the field of medical technology, it is often necessary to monitor the health status of medical patients by detecting the change trends of physiological variables such as heart rate, electroencephalogram (EEG), and electrocardiogram (ECG). Therefore, the position and category of change points in data such as EEG and ECG play extremely important roles.
[0003] Currently, in many practical application scenarios, due to the constraints of the scale and complexity of time-series data such as EEG and ECG, it is difficult to obtain the occurrence time and category labels of change points in the above data.
[0004] Without relying on the position and category labels of change points, realizing the detection of the position and category of change points in different application scenarios is the main challenge faced by the current related fields. Summary of the Invention
[0005] Embodiments of the present invention provide a method and device for detecting the position and category of change points, so as to solve the problem of detecting the position and category of change points in the time-series data of a physiological variable graph in a certain dimension when there is a lack of annotation of the position and category of change points in the time-series data of a multi-dimensional physiological variable graph.
[0006] In a first aspect, embodiments of the present invention provide a method for detecting the position and category of change points, including:
[0007] Performing baseline extraction on the input time-series data of a multi-dimensional physiological variable graph to obtain a baseline sequence of the time-series data of each dimension of the physiological variable graph;
[0008] Screening all time points in the baseline sequence of the time-series data of each dimension of the physiological variable graph to obtain candidate change points corresponding to the time-series data of each dimension of the physiological variable graph;
[0009] Grouping the candidate change points corresponding to the time-series data of each dimension of the physiological variable graph according to the time gap between the candidate change points to obtain a plurality of candidate change point groups;
[0010] Voting on each time point in the plurality of candidate change point groups by using a generalized likelihood ratio GLR algorithm based on voting, and determining change points according to the voting results;
[0011] Extracting change point sequence data from the time-series data of each dimension of the physiological variable graph according to the position of the change points, where the change point sequence data includes the change points;
[0012] Cluster the extracted change point sequence data in the form of a tree - shaped structure split to obtain the change point classification result.
[0013] In a possible implementation, the baseline extraction of the input multi - dimensional physiological variable graph time - series data to obtain the baseline sequence of each dimension of the physiological variable graph time - series data includes:
[0014] For each time point, calculate the mean value of the time point and the previous N time points, where N is an integer greater than 1, and use the mean value as the baseline corresponding value of the time point.
[0015] Arrange the baseline corresponding values of each time point in time series to form the baseline sequence of each dimension of the physiological variable graph time - series data.
[0016] In a possible implementation, the screening of all time points in the baseline sequence of each dimension of the physiological variable graph time - series data to obtain the candidate change points corresponding to each dimension of the physiological variable graph time - series data includes:
[0017] Perform the following screening process on the baseline sequence of each dimension of the physiological variable graph time - series data in the order of dimensions:
[0018] Compare the ratio of the data difference between each time point and the next time point in the baseline sequence corresponding to the current dimension to the maximum data difference within the GLR window corresponding to the time point with a preset first threshold. Add the time points with a difference greater than the first threshold to the first candidate set corresponding to the current dimension, and delete the time points with a difference not greater than the first threshold. The GLR window corresponding to the time point includes at least the time point and the next time point.
[0019] Calculate the absolute value of the difference between two consecutive points in the first candidate set corresponding to the current dimension. If there are time points in the first candidate set corresponding to the current dimension whose absolute value of the difference is greater than the second threshold, then use the time points in the first candidate set corresponding to the current dimension whose absolute value of the difference is greater than the second threshold as the candidate change points and end the screening process. If there are no time points in the first candidate set corresponding to the current dimension whose absolute value of the difference is greater than the second threshold, then screen the baseline sequence of the physiological variable graph time - series data of the next dimension. The second threshold is the absolute value of the difference corresponding to the M - th position after sorting the absolute values of the differences of all time points in the baseline sequence corresponding to the current dimension from large to small, where M = B×(1 / W), B is the number of differences of all time points in the baseline sequence corresponding to the current dimension, and W is the size of the GLR window.
[0020] Alternatively, when there is no time point in the corresponding first candidate set of the multiple dimensions where the absolute value of the difference is greater than the second threshold, calculate the absolute value of the difference between two consecutive points in the first candidate set, and then calculate the absolute value of the difference of the absolute value of the difference between the two consecutive points to obtain the absolute value of the second-order difference. If there is a time point in the first candidate set of the current dimension where the absolute value of the second-order difference is greater than the third threshold, then use the time point in the current dimension where the absolute value of the second-order difference is greater than the third threshold as the candidate change point, and end the screening process. If there is no time point in the first candidate set corresponding to the current dimension where the absolute value of the second-order difference is greater than the third threshold, then screen the baseline sequence of the time series data of the physiological variable diagram of the next dimension. Wherein, the third threshold is the absolute value of the second-order difference corresponding to the D-th position after sorting the absolute values of the second-order differences of all time points in the baseline sequence corresponding to the current dimension from large to small, and D = T×(1 / W), where T is the number of second-order differences of all time points in the baseline sequence corresponding to the current dimension, and W is the size of the GLR window.
[0021] In a possible implementation manner, grouping the candidate change points corresponding to the time series data of the physiological variable diagrams of each dimension according to the time gap between the candidate change points, to obtain a plurality of candidate change point groups, including:
[0022] Calculate the time difference between two consecutive candidate change points among the candidate change points corresponding to the time series data of the physiological variable diagrams of each dimension;
[0023] When the time difference is less than the length of the smaller part of the front part and the rear part of the change point within the GLR window, add the candidate change point with a later time among the two candidate change points to the candidate change point group corresponding to the candidate change point with an earlier time;
[0024] When the time difference is not less than the length of the smaller part of the front part and the rear part of the change point within the GLR window, create a new candidate change point group for the candidate change point with a later time among the two candidate change points.
[0025] In a possible implementation manner, using the generalized likelihood ratio GLR algorithm based on voting to vote on each time point within the plurality of candidate change point groups, and determining the change point according to the voting result, including:
[0026] Delete the single-point groups and the change point groups with a time gap greater than the preset time threshold within the plurality of candidate change point groups to obtain available candidate change point groups;
[0027] Calculate the GLR value of each time point within the available candidate change point groups according to the GLR window;
[0028] Traverse each time point of the available candidate change point groups by sliding the outer voting window, and vote for each time point within the candidate change point groups. Among them, when voting for each time point within the outer voting window, the voting value of each time point within the outer voting window is the sum of the GLR value of each time point and the GLR values of all time points after this time point within the GLR window. The time point with the largest voting value within the outer voting window gets one vote;
[0029] When the number of votes of the time point with the most votes within the available candidate change point groups exceeds half of the voting opportunity, determine the time point with the most votes within the available candidate change point groups as the change point;
[0030] When the number of votes of the time point with the most votes within the available candidate change point groups does not exceed half of the voting opportunity, determine that there is no change point within the available candidate change point groups.
[0031] In a possible implementation manner, clustering the extracted change point sequence data in the form of splitting in a tree structure to obtain a change point classification result, including:
[0032] Use the extracted change point sequence data as the root node and split it in the form of splitting in a tree structure. Among them,
[0033] Split all cluster nodes at the current depth at each layer;
[0034] Before splitting the current node, obtain a similarity matrix by taking the weighted average of the soft-DTW distance and the SBD distance and then taking the negative value, and estimate the optimal k value of the current node. Use the optimal k value and the similarity matrix as the input of spectral clustering to complete the clustering of the data within the current node;
[0035] If the optimal k value of the current node is 1, the current node is not split and the current node is used as a leaf node;
[0036] Complete the splitting in a layer-by-layer and node-by-node manner. When the k value of all nodes is 1, the clustering ends and a clustering result is obtained;
[0037] Obtain the change point classification result according to the clustering result.
[0038] In a possible implementation manner, the obtaining the change point classification result according to the clustering result includes:
[0039] Calculate the internal temporal mean of each cluster of the clustering result according to the soft-DTW distance, and use the mean as the representation of the cluster;
[0040] Calculate the difference between each piece of data within the cluster and the cluster representation, calculate the mean and standard deviation of all the differences, and add the mean to the standard deviation to obtain the internal statistical feature of the cluster;
[0041] Calculate the Fréchet distance between the representation of each cluster and the representations of other clusters. When the Fréchet distance between the representations of two clusters is less than the mean of the internal statistical features of the two clusters, merge the two clusters.
[0042] In a possible implementation manner, it further includes:
[0043] Calculate the weight coefficients of the time-series data of the physiological variable diagrams for each dimension according to the baseline sequences of the time-series data of the physiological variable diagrams for each dimension;
[0044] Calculate the soft-DTW distance and the SBD distance according to the weight coefficients of the time-series data of the physiological variable diagrams for each dimension.
[0045] In a possible implementation manner, the calculating the weight coefficients of the time-series data of the physiological variable diagrams for each dimension according to the baseline sequences of the time-series data of the physiological variable diagrams for each dimension includes:
[0046] Calculate the basic weight coefficients of the time-series data of the physiological variable diagrams for each dimension, where the basic weight coefficient is the difference between the maximum value and the minimum value in the time-series data of the physiological variable diagrams for each dimension;
[0047] For the time-series data of the physiological variable diagrams for each dimension, calculate the ratio of the basic weight coefficient of the time-series data of the physiological variable diagrams for the dimension to a first parameter to obtain a comprehensive weight, where the first parameter = E × F, where E is the mean of the absolute values of the differences between two consecutive points in the time-series data of the physiological variable diagrams for the dimension, and F is the sum of the absolute values of the Pearson correlation coefficients between the dimension and the time-series data of the physiological variable diagrams for all other dimensions;
[0048] Use the sum of the comprehensive weights of the time-series data of the physiological variable diagrams for each dimension to regularize the comprehensive weights of the time-series data of the physiological variable diagrams for each dimension to obtain the weight coefficients of the time-series data of the physiological variable diagrams for each dimension.
[0049] In a second aspect, an apparatus for detecting the change point position and category provided by an embodiment of the present invention includes:
[0050] A first extraction module, configured to perform baseline extraction on the input multi-dimensional time-series data of physiological variable diagrams to obtain the baseline sequences of the time-series data of the physiological variable diagrams for each dimension;
[0051] A screening module, configured to screen all time points in the baseline sequences of the time-series data of the physiological variable diagrams for each dimension to obtain candidate change points corresponding to the time-series data of the physiological variable diagrams for each dimension;
[0052] A grouping module, configured to group the candidate change points corresponding to the time series data of the physiological variable diagrams of each dimension according to the time gap of the candidate change points, so as to obtain a plurality of candidate change point groups;
[0053] A voting module, configured to vote on each time point within the plurality of candidate change point groups by using a generalized likelihood ratio GLR algorithm based on voting, and determine a change point according to the voting result;
[0054] A second extraction module, configured to extract change point sequence data from the time series data of the physiological variable diagrams of each dimension according to the position of the change point, where the change point sequence data includes the change point;
[0055] A clustering module, configured to cluster the extracted change point sequence data in a form of splitting a tree structure to obtain a change point classification result.
[0056] In a possible implementation manner, the first extraction module is specifically configured to:
[0057] For each time point, calculate the mean value of the time point and the N time points before the time point, and use the mean value as the baseline corresponding value of the time point, where N is an integer greater than 1;
[0058] Arrange the baseline corresponding values of each time point in time sequence to form a baseline sequence of the time series data of the physiological variable diagrams of each dimension.
[0059] In a possible implementation manner, the screening module is specifically configured to:
[0060] Perform the following screening process on the baseline sequences of the time series data of the physiological variable diagrams of each dimension in the dimension order:
[0061] Compare the ratio of the data difference between each time point and the next time point in the baseline sequence corresponding to the current dimension to the maximum data difference within the GLR window corresponding to the time point with a preset first threshold, add the time points with the difference greater than the first threshold to the first candidate set corresponding to the current dimension, and delete the time points with the difference not greater than the first threshold, where the GLR window corresponding to the time point includes at least the time point and the next time point;
[0062] Calculate the absolute value of the difference between two consecutive points in the first candidate set corresponding to the current dimension. If there is a time point in the first candidate set corresponding to the current dimension where the absolute value of the difference is greater than the second threshold, then use the time point in the first candidate set corresponding to the current dimension where the absolute value of the difference is greater than the second threshold as the candidate change point, and end the screening process. If there is no time point in the first candidate set corresponding to the current dimension where the absolute value of the difference is greater than the second threshold, then screen the baseline sequence of the time series data of the physiological variable graph for the next dimension. Wherein, the second threshold is the absolute value of the difference corresponding to the Mth position after sorting the absolute values of the differences between all time points in the baseline sequence corresponding to the current dimension from large to small, and M = B×(1 / W), where B is the number of differences between all time points in the baseline sequence corresponding to the current dimension, and W is the size of the GLR window;
[0063] Alternatively, when there is no time point in the first candidate sets corresponding to the multiple dimensions where the absolute value of the difference is greater than the second threshold, calculate the absolute value of the difference between two consecutive points in the first candidate set, and take the absolute value of the difference of the absolute value of the difference between two consecutive points to obtain the absolute value of the second-order difference. If there is a time point in the first candidate set corresponding to the current dimension where the absolute value of the second-order difference is greater than the third threshold, then use the time point in the current dimension where the absolute value of the second-order difference is greater than the third threshold as the candidate change point, and end the screening process. If there is no time point in the first candidate set corresponding to the current dimension where the absolute value of the second-order difference is greater than the third threshold, then screen the baseline sequence of the time series data of the physiological variable graph for the next dimension. Wherein, the third threshold is the absolute value of the second-order difference corresponding to the Dth position after sorting the absolute values of the second-order differences between all time points in the baseline sequence corresponding to the current dimension from large to small, and D = T×(1 / W), where T is the number of second-order differences between all time points in the baseline sequence corresponding to the current dimension, and W is the size of the GLR window.
[0064] In a possible implementation manner, the grouping module is specifically configured to:
[0065] Calculate the time difference between two consecutive candidate change points among the candidate change points corresponding to the time series data of the physiological variable graph of each dimension;
[0066] When the time difference is less than the length of the smaller part of the front part and the rear part of the change point within the GLR window, add the candidate change point with a later time among the two candidate change points to the candidate change point group corresponding to the candidate change point with an earlier time;
[0067] When the time difference is not less than the length of the smaller part among the front part and the rear part of the change point within the GLR window, a new candidate change point group is established for the candidate change point that is later among the two candidate change points.
[0068] In a possible implementation manner, the voting module is specifically configured to:
[0069] Delete the single-point groups and the change point groups with a time gap within the group greater than a preset time threshold in the multiple candidate change point groups, and obtain available candidate change point groups;
[0070] Calculate the GLR value of each time point within the available candidate change point groups according to the GLR window;
[0071] Traverse each time point of the available candidate change point groups by sliding an outer voting window, and vote for each time point within the candidate change point groups. Among them, when voting for each time point within the outer voting window, the voting value of each time point within the outer voting window is the sum of the GLR value of each time point and the GLR values of all time points after this time point within the GLR window, and the time point with the largest voting value within the outer voting window gets one vote;
[0072] When the number of votes of the time point with the most votes within the available candidate change point groups exceeds half of the winning chance, determine the time point with the most votes within the available candidate change point groups as the change point;
[0073] When the number of votes of the time point with the most votes within the available candidate change point groups does not exceed half of the winning chance, determine that there is no change point within the available candidate change point groups.
[0074] In a possible implementation manner, the clustering module is specifically configured to:
[0075] Use the extracted change point sequence data as the root node and perform splitting in the form of a tree structure split, where
[0076] Each layer splits all cluster nodes at the current depth;
[0077] Before splitting the current node, obtain a similarity matrix by taking the weighted average of the soft-DTW distance and the SBD distance and then taking the negative value, and estimate the best k value of the current node. Use the best k value and the similarity matrix as the input of spectral clustering to complete the clustering of the data within the current node;
[0078] If the best k value of the current node is 1, the current node is not split and the current node is used as a leaf node;
[0079] The splitting is completed in a layer-by-layer and node-by-node manner. When the k value of all nodes is 1, the clustering ends and the clustering result is obtained;
[0080] The change point classification result is obtained according to the clustering result.
[0081] In a possible implementation manner, the clustering module is further configured to:
[0082] Calculate the within-cluster temporal mean for each cluster of the clustering result according to the soft-DTW distance, and use the mean as the representation of the cluster;
[0083] Calculate the difference between each piece of data in the cluster and the cluster representation, calculate the mean and standard deviation of all the differences, and add the mean and the standard deviation to obtain the internal statistical feature of the cluster;
[0084] Calculate the Fréchet distance between the representations of each cluster and the representations of other clusters. When the Fréchet distance between the representations of two clusters is less than the mean of the internal statistical features of the two clusters, the two clusters are merged.
[0085] In a possible implementation manner, it further includes a calculation module, and the calculation module is configured to:
[0086] Calculate the weight coefficient of the time series data of each dimension physiological variable graph according to the baseline sequence of the time series data of each dimension physiological variable graph;
[0087] Calculate the soft-DTW distance and the SBD distance according to the weight coefficients of the time series data of each dimension physiological variable graph.
[0088] In a possible implementation manner, the calculation module is specifically configured to:
[0089] Calculate the basic weight coefficient of the time series data of each dimension physiological variable graph, and the basic weight coefficient is the difference between the maximum value and the minimum value in the time series data of each dimension physiological variable graph;
[0090] For the time series data of each dimension physiological variable graph, calculate the ratio of the basic weight coefficient of the time series data of the dimension physiological variable graph to the first parameter to obtain the comprehensive weight, where the first parameter = E × F, and E is the mean of the absolute values of the differences between two consecutive points in the time series data of the dimension physiological variable graph, and F is the sum of the absolute values of the Pearson correlation coefficients between the dimension and the time series data of all other dimension physiological variable graphs;
[0091] Use the sum of the comprehensive weights of the time series data of each dimension physiological variable graph to regularize the comprehensive weights of the time series data of each dimension physiological variable graph to obtain the weight coefficients of the time series data of each dimension physiological variable graph.
[0092] Third aspect, an embodiment of the present invention provides a detection device for change point positions and categories, including:
[0093] At least one processor;
[0094] And a memory communicatively connected to the at least one processor;
[0095] Wherein, the memory is used to store instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a detection method for change point positions and categories provided in the first aspect.
[0096] Fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which computer-executable instructions are stored, and the computer-executable instructions are executed by a processor to implement a detection method for change point positions and categories provided in the first aspect.
[0097] The detection method for change point positions and categories provided by the present invention obtains the change point positions and categories existing in the time series data of physiological variable diagrams in each dimension by processing the time series data of physiological variable diagrams in each dimension through two stages of unsupervised change point extraction and change point clustering. This method gets rid of the dependence on artificially set parameters. At the same time, the change point extraction algorithm provided by the present invention combines a probability model with a pre-screening and group voting mechanism to ensure the accuracy of change point extraction. The clustering algorithm provided by the present invention combines top-down split hierarchical clustering with the NME algorithm to improve the accuracy of the clustering result. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.
[0099] Figure 1 It is a schematic flowchart of a detection method for change point positions and categories provided in Embodiment 1 of the present invention;
[0100] Figure 2 It is a schematic flowchart of a detection method for change point positions and categories provided in Embodiment 2 of the present invention;
[0101] Figure 3 It is a schematic flowchart of a detection method for change point positions and categories provided in Embodiment 3 of the present invention;
[0102] Figure 4 It is a schematic diagram for grouping candidate change points;
[0103] Figure 5 It is a schematic flowchart of a detection method for change point positions and categories provided in Embodiment 4 of the present invention;
[0104] Figure 6 Schematic diagram for voting on time points within a candidate change point group;
[0105] Figure 7 Schematic flowchart of a method for detecting change point positions and categories provided in Embodiment 5 of the present invention;
[0106] Figure 8 Schematic flowchart of a method for detecting change point positions and categories provided in Embodiment 6 of the present invention;
[0107] Figure 9 Schematic flowchart of a method for detecting change point positions and categories provided in Embodiment 7 of the present invention;
[0108] Figure 10 Schematic structural diagram of a device for detecting change point positions and categories provided in Embodiment 8 of the present invention;
[0109] Figure 11 Schematic structural diagram of a device for detecting change point positions and categories provided in Embodiment 9 of the present invention.
[0110] Through the above-mentioned drawings, specific embodiments of the present invention have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the inventive concept in any way, but to illustrate the concept of the present invention to those skilled in the art by referring to specific embodiments. Detailed Embodiments
[0111] Here, exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0112] In the medical field, the detection of change point positions and categories can play an extremely important role. For physiological variables such as EEG and ECG, by detecting the positions and categories of their change points and judging the occurrence of various events based on the change points, real-time automated health monitoring can be achieved. Starting from the actual application scenario, the present invention processes the time series data of multi-dimensional physiological variable graphs to obtain change point positions and categories that conform to the actual situation, thereby assisting in information extraction of relevant data and automated health monitoring. Currently, traditional methods for change point detection rely on change point position and category labels, and this method is more efficient than traditional methods and reduces the dependence on manually set parameters.
[0113] Figure 1This is a schematic flowchart of a method for detecting the position and category of change points provided in the first embodiment of the present invention. As Figure 1 shown, the method for detecting the position and category of change points includes the following steps.
[0114] S101. Extract the baseline of the input multi-dimensional physiological variable graph time series data to obtain the baseline sequence of the time series data of each dimension of the physiological variable graph.
[0115] The multi-dimensional physiological variable graph time series data is a physiological variable data graph used in the medical field to monitor the health status of medical patients. Specifically, it can be an electroencephalogram, an electrocardiogram, etc. Among them, the dimension of the physiological variable graph time series data corresponds to the number of data groups in the physiological variable graph. Exemplarily, for an electrocardiogram, it can be divided into three-lead electrocardiogram, six-lead electrocardiogram, twelve-lead electrocardiogram, and eighteen-lead electrocardiogram according to the number of its leads. When the electrocardiogram is a twelve-lead electrocardiogram, the dimension of the corresponding electrocardiogram time series data is twelve dimensions.
[0116] To extract the baseline of the input multi-dimensional physiological variable graph time series data, first calculate the mean value of each time point and the N time points before this time point for each time point in the input multi-dimensional physiological variable graph time series data, and use this mean value as the baseline corresponding value of this time point. Among them, N is an integer greater than 1. Arrange the obtained baseline corresponding values in time sequence to form the baseline sequence of the time series data of each dimension of the physiological variable graph. By performing baseline extraction processing on the time series data of each dimension of the physiological variable graph, the influence of noise on the time series data of each dimension of the physiological variable graph can be reduced.
[0117] S102. Screen all time points in the baseline sequence of the time series data of each dimension of the physiological variable graph to obtain the candidate change points corresponding to the time series data of each dimension of the physiological variable graph.
[0118] The change points existing in the time series data of each dimension of the physiological variable graph are key points for monitoring the health status of medical patients. Exemplarily, when the time series data of each dimension of the physiological variable graph is the time series data of each dimension of the electrocardiogram, its change points can be key points such as P wave, Q wave, U wave, QSR wave group, ST segment, T wave, etc. By screening all time points in the baseline sequence of the time series data of each dimension of the physiological variable graph, the corresponding candidate change points in this time series data can be obtained, and the candidate change points can be used for subsequent voting processing.
[0119] S103. Group the candidate change points corresponding to the time series data of each dimension of the physiological variable graph according to the time gap of the candidate change points to obtain multiple candidate change point groups.
[0120] For candidate change points, group the candidate change points according to the time gap between the change points. Group the candidate change points with a smaller time gap into one group. Exemplarily, this time gap can be 30 time points.
[0121] S104. Use the Generalized Likelihood Ratio (GLR) algorithm based on voting to vote on the candidate change points within multiple candidate change point groups, and determine the change points according to the voting results.
[0122] The Generalized Likelihood Ratio (GLR) algorithm is used to estimate the difference in the distribution and characteristics of the two subsequences before and after the change point using the generalized likelihood ratio. Among them, for a point with a larger generalized likelihood ratio, it can be considered that the distribution characteristics of the two subsequences before and after the change point have a larger gap. When using the GLR algorithm based on voting to vote on the candidate change points within multiple candidate change point groups, in order to detect the change points, this method uses an outer voting sliding window to traverse the entire time series, and takes the sum of the GLR values of the current point and all subsequent points within the outer window as the measurement value for voting. The time point with the largest value within the current window gets one vote, and the change points are determined according to the final voting results.
[0123] S105. Extract the change point sequence data from the time series data of the physiological variable graphs in each dimension according to the positions of the change points. The change point sequence data includes the change points.
[0124] Extract the change point sequence data from the time series data of the physiological variable graphs in each dimension according to the positions of the change points. The selection of the change point sequence data is determined by the size of the GLR window. Specifically, the change points are generally located in the middle of the GLR window, and the change point sequence data includes the time points before and after the change point within the GLR window.
[0125] S106. Cluster the extracted change point sequence data in the form of a tree structure split to obtain the change point classification result.
[0126] Cluster the extracted change point sequence data in the form of a tree structure split. This clustering method is a top-down divisive hierarchical clustering. That is, first consider all data samples as one cluster, and then recursively split the current cluster into smaller clusters according to a certain partitioning strategy. Using this method for clustering, the tree structure formed during the entire clustering process is convenient for visualization.
[0127] In this embodiment, by extracting change points and clustering change points from the time series data of the multi-dimensional physiological variable map, the positions and categories of the change points existing in the time series data of the multi-dimensional physiological variable map are obtained. Specifically, the change point extraction algorithm extracts the baseline of the time series data of each dimension of the physiological variable map to obtain the baseline sequence of the time series data of each dimension of the physiological variable map. Further, a method combining pre-screening and group voting is used to extract all time points of the baseline sequence to obtain the change point result. This method not only ensures the accuracy of the extracted change points but also gets rid of the dependence of the algorithm on artificially set thresholds and the total number of change points. The change point clustering adopts a top-down splitting hierarchical clustering, and the tree structure formed during the entire clustering process is convenient for visualization.
[0128] Figure 2 FIG. is a schematic flow chart of a method for detecting the position and category of change points provided in the second embodiment of the present invention. This embodiment is a detailed description of step S102 in the first embodiment. As Figure 2 shown, the method for detecting the position and category of change points includes the following steps.
[0129] S201, calculate the ratio of the data difference between each time point and the next time point in the baseline sequence corresponding to the current dimension to the maximum data difference within the GLR window corresponding to this time point.
[0130] In this embodiment, by way of example, when the time series data of the multi-dimensional physiological variable map is the electrocardiogram data of twelve leads, then the time series data of the physiological variable map has a total of twelve dimensions, and the baseline sequences of the time series data of the multi-dimensional physiological variable map are screened in the order of dimensions.
[0131] Among them, the GLR window is a pre-set window, and its size is pre-set according to the density of the input time series data of the multi-dimensional physiological variable map. When the time series data of the multi-dimensional physiological variable map is relatively dense, the pre-set GLR window value is larger; conversely, the GLR window value is smaller. By way of example, the window size can be 100 time points or 50 time points.
[0132] S202, determine whether this ratio is greater than the first threshold.
[0133] The size of the pre-set first threshold is determined according to the sampling frequency of the time series data of the multi-dimensional physiological variable map. By way of example, it can be 0.1. When this ratio is greater than the first threshold, step S203 is executed; when this ratio is not greater than the first threshold, step S204 is executed.
[0134] S203, add the time points with a difference greater than the first threshold to the first candidate set corresponding to the current dimension.
[0135] A time point with a difference greater than the first threshold may be a change point, and it is added to the first candidate set corresponding to the current dimension for further screening.
[0136] S204. Delete the time points with differences not greater than the first threshold.
[0137] For a time point with a difference not greater than the first threshold, it is considered that the point is relatively smooth and does not belong to a change point. This time point is deleted and does not enter the subsequent screening. By performing steps S203 and S204 to screen the time points in the baseline sequence corresponding to the current dimension, some smooth points in the baseline sequence can be removed, which can improve the accuracy of change point extraction and shorten the time for change point extraction.
[0138] S205. Calculate the absolute value of the difference between two consecutive points in the first candidate set corresponding to the current dimension.
[0139] S206. Determine whether there are time points in the first candidate set corresponding to the current dimension with an absolute value of the difference greater than the second threshold.
[0140] If there are time points in the first candidate set corresponding to the current dimension with an absolute value of the difference greater than the second threshold, then execute step S207. If there are no time points in the first candidate set corresponding to the current dimension with an absolute value of the difference greater than the second threshold, then execute step S208. Among them, the second threshold is the absolute value of the difference corresponding to the Mth position after sorting the absolute values of the differences between all time points in the baseline sequence corresponding to the current dimension from large to small. M = B×(1 / W), where B is the number of differences between all time points in the baseline sequence corresponding to the current dimension, and W is the size of the GLR window. Exemplarily, when the number of differences between all consecutive two time points in the baseline sequence corresponding to the current dimension is 1000 and the size of the GLR window is 100, then the second threshold is the absolute value of the difference corresponding to the 10th position after sorting the absolute values of the differences between all consecutive two time points in the corresponding baseline sequence from large to small.
[0141] S207. Take the time points in the first candidate set corresponding to the current dimension with an absolute value of the difference greater than the second threshold as candidate change points and end the screening process.
[0142] S208. Screen the baseline sequence of the time series data of the physiological variable graph for the next dimension.
[0143] If the absolute values of the differences in the first candidate set corresponding to the current dimension are all not greater than the second threshold, then screen the baseline sequence of the time series data of the physiological variable graph for the next dimension, and its screening method is the same as the screening methods in steps S201 - S206.
[0144] After screening the baseline sequence of the multi-dimensional physiological variable diagram time series data according to the above screening method, when there is no time point in the corresponding first candidate sets of multiple dimensions where the absolute value of the difference is greater than the second threshold, it is necessary to continue screening the time points in the corresponding first candidate sets of multiple dimensions. A possible implementation is to calculate the absolute value of the difference between two consecutive points in the first candidate set, calculate the absolute value of the difference of the absolute value of the difference between two consecutive points to obtain the absolute value of the second-order difference, and sort the absolute value of the second-order difference from largest to smallest. If there is a time point in the first candidate set of the current dimension where the absolute value of the second-order difference is greater than the third threshold, then use the time point in the current dimension where the absolute value of the second-order difference is greater than the third threshold as the candidate change point and end the screening process. If there is no time point in the corresponding first candidate set of the current dimension where the absolute value of the second-order difference is greater than the third threshold, then screen the baseline sequence of the time series data of the physiological variable diagram of the next dimension. Among them, the third threshold is the absolute value of the second-order difference corresponding to the D-th position after sorting the absolute values of the second-order differences of all time points in the baseline sequence corresponding to the current dimension from largest to smallest. D = T×(1 / W), where T is the number of second-order differences of all time points in the baseline sequence corresponding to the current dimension, and W is the size of the GLR window. The method for determining its threshold is the same as the method in step S206 and will not be elaborated here.
[0145] In this embodiment, according to the dimension order, the baseline sequences of the time series data of the physiological variable diagrams of each dimension are first screened for change points according to the preset first threshold and the ratio of the difference between the data of two consecutive time points in the baseline sequence to the maximum data difference within the GLR window corresponding to the time point. Then, the result obtained from the first screening is secondarily screened according to the second threshold and the absolute value of the difference between two consecutive points in the first candidate set. Or, according to the third threshold and the absolute value of the second-order difference between two consecutive points in the first candidate set, the candidate change points in the time series data of the physiological variable diagrams of each dimension are obtained. This pre-screening method improves the accuracy of change point extraction.
[0146] Figure 3 This is a method for detecting the position and category of change points provided in Embodiment 3 of the present invention. This embodiment is a detailed description of step S103 in Embodiment 1. As Figure 3 shown, the method for detecting the position and category of change points includes the following steps.
[0147] S301, calculate the time difference between two consecutive candidate change points among the candidate change points corresponding to the time series data of the physiological variable diagrams of each dimension.
[0148] According to the candidate change points corresponding to the time series data of the physiological variable diagrams of each dimension obtained in Embodiment 2, calculate the time difference between two consecutive candidate change points among the candidate change points, where the candidate change points are arranged according to time sequence.
[0149] S302. Determine whether the time difference is less than the length of the smaller part of the pre-change part and the post-change part within the GLR window.
[0150] When the time difference between two consecutive candidate change points is less than the length of the smaller part of the pre-change part and the post-change part within the GLR window, execute step S303; when the time difference between two consecutive candidate change points is not less than the length of the smaller part of the pre-change part and the post-change part within the GLR window, execute step S304.
[0151] The positions of the candidate change points in the GLR window are preset. Specifically, the candidate change points may be at any position within the GLR window. According to the positions of the candidate change points, the part in front of the change point within the GLR window is taken as the pre-change part, and the part behind the change point within the GLR window is taken as the post-change part. Calculate the difference between the change point and the first time point of the pre-change part as the length of the pre-change part, and calculate the difference between the change point and the last time point of the post-change part as the length of the post-change part. When the change point is in the middle of the GLR window, the lengths of the pre-change part and the post-change part are the same. Figure 4 It is a schematic diagram for grouping candidate change points. As shown in Figure 4, the rectangular frame is the GLR window, the black dots are the candidate change points, the black square dots are the first time points of the pre-change parts, and the black triangle dots are the last time points of the post-change parts. Calculate the differences between two consecutive candidate change points respectively, and select the smaller length between the length of the pre-change part and the length of the post-change part within the GLR window to compare with the difference between the two consecutive candidate change points. In Figure 41, it shows the case where the time difference between two consecutive candidate change points is less than the smaller length of the pre-change part and the post-change part within the GLR window, and in Figure 42, it shows the case where the time difference between two consecutive candidate change points is greater than the smaller length of the pre-change part and the post-change part within the GLR window.
[0152] S303. Add the candidate change point with a later time among the two candidate change points to the candidate change point group corresponding to the candidate change point with an earlier time.
[0153] When the time difference between two consecutive candidate change points is less than the length of the smaller part of the pre-change part and the post-change part within the GLR window, add the candidate change point with a later time among the two candidate change points to the candidate change point group corresponding to the candidate change point with an earlier time.
[0154] S304. Create a new candidate change point group for the candidate change point with a later time among the two candidate change points.
[0155] When the time difference between two consecutive candidate change points is not less than the length of the smaller part of the pre-change part and the post-change part within the GLR window, create a new candidate change point group for the candidate change point with a later time among the two candidate change points. Group all the candidate change points corresponding to the time series data of the physiological variable diagrams in each dimension according to the time gap.
[0156] Exemplarily, starting from the first candidate change point for grouping, calculate the time difference between the first candidate change point and the second candidate change point, the length of the front part of the change point within the GLR window of the first candidate change point, and the length of the rear part of the change point. When the time difference between the first candidate change point and the second candidate change point is less than the smaller length of the front part of the change point and the rear part of the change point within the GLR window of the first candidate change point, establish a candidate change point group for the first candidate change point and the second candidate change point. When the time difference between the first candidate change point and the second candidate change point is not less than the length of the smaller part of the front part of the change point and the rear part of the change point within the GLR window, establish a candidate change point group for the first candidate change point. Then, calculate the time difference between the second candidate change point and the third candidate change point, the length of the front part of the change point within the GLR window of the second candidate change point, and the length of the rear part of the change point. When the time difference between the second candidate change point and the third candidate change point is less than the smaller length of the front part of the change point and the rear part of the change point within the GLR window of the second candidate change point, if the second candidate change point and the first candidate change point are in the same candidate change point group, put the third candidate change point into the same candidate change point group as the first candidate change point and the second candidate change point. If the second candidate change point and the first candidate change point are not in the same candidate change point group, establish a new candidate change point group for the second candidate change point and the third candidate change point. When the time difference between the second candidate change point and the third candidate change point is not less than the length of the smaller part of the front part of the change point and the rear part of the change point within the GLR window of the second candidate change point, if the second candidate change point and the first candidate change point are not in the same candidate change point group, establish a new candidate change point group for the second candidate change point. And so on, complete the grouping for all candidate points.
[0157] In this embodiment, sort the candidate change points according to the time sequence of the physiological variable diagrams in each dimension, and group all the candidate change points in turn according to the time difference between two consecutive candidate change points and the length of the smaller part of the front part of the change point and the rear part of the change point within the GLR window. Group the candidate change points with the time difference between two candidate change points less than the length of the smaller part of the front part of the change point and the rear part of the change point within the GLR window into one group. By grouping the candidate change points, it is beneficial to select the best change point from each change point group subsequently, thereby improving the accuracy of change point extraction.
[0158] Figure 5 It is a schematic flowchart of a method for detecting the position and category of change points provided by an embodiment of the present invention. This embodiment is a detailed description of step S104 in Embodiment 1. As Figure 5 shown, the method for detecting the position and category of change points includes the following steps.
[0159] S501, delete the single-point groups and the change point groups with a time gap greater than a preset time threshold in multiple candidate change point groups to obtain available candidate change point groups.
[0160] For a single-point group in the candidate change point group and a change point group with an intra-group time gap greater than the preset time threshold, it is considered that the change points within the group are mutational, and then they are deleted without subsequent voting processing.
[0161] S502. Calculate the GLR value of each time point within the available candidate change point group according to the GLR window.
[0162] The candidate change points within the candidate change point group are arranged in chronological order. Calculate the GLR value of each time point within the available candidate change point group according to the GLR window. Specifically, the GLR value can be obtained by the following formula:
[0163]
[0164] where x j represents the current candidate change point, and u1 and σ1 2 respectively represent the mean and variance of the data in the rear window of x within the GLR window, and u0 and σ0 j respectively represent the mean and variance of the data in the front window of x within the GLR window. 2 respectively represent the mean and variance of the data in the front window of x within the GLR window. j
[0165] S503. Traverse each time point of the available candidate change point group by sliding the outer voting window, and vote on each time point within the candidate change point group.
[0166] The value of the outer voting window is not less than the GLR window. When the size of the outer voting window is twice the size of the GLR window, the voting effect on the candidate change points is better. Of course, the present invention does not limit the size of the outer voting window, and this size can be any size not less than the GLR window.
[0167] Traverse each time point of the available candidate change point group by sliding the outer voting window, and vote on each time point within the candidate change point group. Figure 6 is a schematic diagram of voting on the time points within the candidate change point group. As Figure 6 shown, when voting on each time point within the outer voting window, for any one time point, the voting value of this time point is the sum of the GLR value of this time point and the GLR values of all time points after this time point within the GLR window. This voting value can be S0, S k and S n-1 etc. in the figure. Among them, the time point with the largest voting value within the outer voting window gets one vote.
[0168] S504. Determine whether the number of votes of the time point with the most votes within the available candidate change point group exceeds half of the voting opportunity.
[0169] When the number of votes at the time point with the most votes in the available candidate change point group exceeds half of the voting opportunity, step S505 is executed. When the number of votes at the time point with the most votes in the available candidate change point group does not exceed half of the voting opportunity, step S506 is executed.
[0170] When voting for the time points in the candidate change point group by sliding the outer voting window, the voting opportunity for each time point in each group is the same. Specifically, the outer voting window slides with a step size of one time point. Therefore, each time point participating in the voting has the opportunity to obtain multiple votes, and finally the point with the most votes is selected according to the voting situation of each time point.
[0171] S505, determine the time point with the most votes in the available candidate change point group as the change point.
[0172] When the number of votes at the time point with the most votes in the available candidate change point group exceeds half of the voting opportunity, determine the time point with the most votes in the available candidate change point group as the change point.
[0173] S506, determine that there is no change point in the available candidate change point group.
[0174] When the number of votes at the time point with the most votes in the available candidate change point group does not exceed half of the voting opportunity, determine that there is no change point in the available candidate change point group.
[0175] In this embodiment, the available candidate change point group is obtained by deleting the single point group and the change point group with the time gap within the group greater than the preset time threshold in the candidate change point group. The voting-based generalized likelihood ratio GLR algorithm is used to vote for the time points in multiple available candidate change point groups. When the number of votes at the time point with the most votes in the available candidate change point group exceeds half of the voting opportunity, the time point with the most votes in the available candidate change point group is determined as the change point, which ensures the accuracy of the obtained change point.
[0176] Figure 7 It is a schematic flowchart of a method for detecting the position and category of change points provided in Embodiment 5 of the present invention. This embodiment is a detailed description of step S106 in Embodiment 1. As Figure 7 shown, the method for detecting the position and category of change points includes the following steps.
[0177] S701, take the extracted change point sequence data as the root node and split it in the form of a tree structure split.
[0178] Take the extracted change point sequence data as the root node and split it in the form of a tree structure split, where each layer splits all the cluster nodes at the current depth.
[0179] S702. Before splitting the current node, a similarity matrix is obtained by taking the negative value of the weighted average of the soft-DTW distance and the SBD distance, and the optimal value of the current node is estimated.
[0180] Optionally, the soft-DTW distance and the SBD distance are calculated according to the extracted change point sequence and the weight coefficients of the time series data of the physiological variable diagrams in each dimension, and a similarity matrix is obtained by taking the negative value of the weighted average of the soft-DTW distance and the SBD distance. The optimal k value of the current node is estimated through this similarity matrix. Among them, the weight coefficients of the time series data of the physiological variable diagrams in each dimension can be calculated according to the baseline sequences of the time series data of the physiological variable diagrams in each dimension, and the specific calculation method can refer to Embodiment 7.
[0181] Specifically, the optimal k value of the current node corresponding to this similarity matrix is obtained by using the NME (Normalized Maximum Eigengap) algorithm. Assuming this similarity matrix is A, a threshold P is selected according to the similarity matrix A, the largest P values in each row of the similarity matrix A are converted to 1, and the remaining similarity values are converted to 0, and is used as the similarity matrix corresponding to the current threshold P. Traverse from 1 to N as the threshold P, and calculate the list of eigenvalue differences e for the current situation p and the maximum eigenvalue λ p,N , and further calculate g p and r (p) and, take the p corresponding to the smallest r (p) value as the finally used threshold, and obtain the similarity matrix and node k corresponding to this threshold. Among them, the calculation formulas of g p and r (p) are as follows:
[0182]
[0183] S703. Determine whether the optimal k value of the current node is 1.
[0184] When the optimal k value of the current node is 1, execute step S705. When the optimal k value of the current node is not 1, execute step S704.
[0185] S704. Complete the splitting in a layer-by-layer and node-by-node manner. When the k value of all nodes is 1, the clustering ends and the clustering result is obtained.
[0186] Complete the splitting of the extracted change point sequence data in a layer-by-layer and node-by-node manner in the form of a tree structure splitting. When the k value of all nodes is 1, the clustering ends and the clustering result is obtained.
[0187] S705. The current node does not split and is used as a leaf node.
[0188] When the best k value of the current node is 1, the current node does not split and is used as a leaf node.
[0189] In this embodiment, a top - down splitting clustering method is combined with the NME algorithm to cluster the extracted change - point sequence data. By using top - down splitting hierarchical clustering, the tree - shaped structure formed during the entire clustering process is convenient for visualization. Using the NME algorithm can reduce the influence of similarity noise and improve the accuracy of the clustering result.
[0190] Figure 8 It is a schematic flowchart of a method for detecting the change - point position and category provided in Embodiment VI of the present invention. This embodiment is a detailed description of step S704 in Embodiment V. As Figure 8 shown, the method for detecting the change - point position and category includes the following steps.
[0191] S801. Calculate the within - cluster temporal mean for each cluster of the clustering result according to the soft - DTW distance, and use this mean as the representation of the cluster.
[0192] S802. Calculate the difference between each piece of data within the cluster and the cluster representation, calculate the mean and standard deviation of all differences, and add the mean and the standard deviation to obtain the internal statistical characteristics of the cluster.
[0193] S803. Calculate the Fréchet distance between the representations of each cluster and the representations of other clusters. When the Fréchet distance between the representations of two clusters is less than the mean of the internal statistical characteristics of the two clusters, merge the two clusters.
[0194] The Fréchet distance is a method for solving the similarity of spatial paths, used to calculate the maximum distance between corresponding points of two curves. Calculate the Fréchet distance between the representations of each cluster and the representations of other clusters. When the Fréchet distance between the representations of two clusters is less than the mean of the internal statistical characteristics of the two clusters, it indicates that the similarity between the two clusters is relatively high, so the two clusters are merged. To avoid over - merging, the merging step is only carried out once, that is, calculate whether each pair of clusters needs to be merged respectively, merge all the interconnected clusters simultaneously, and re - extract the representation to obtain the final category division result.
[0195] In this embodiment, by post - processing the clustering result, that is, by comparing the internal statistical characteristics of the clusters and the Fréchet distance between the representation of each cluster and the representations of other clusters, the clusters with relatively high similarity are merged, and this method further improves the accuracy of the clustering result.
[0196] Figure 9The flowchart shows a method for detecting the change point position and category provided in the seventh embodiment of the present invention. This embodiment details the calculation method of the weight coefficients of each dimension in step S702 of the fifth embodiment. As Figure 9 shown, the method for detecting the change point position and category includes the following steps.
[0197] S901, calculate the basic weight coefficient of the time series data of the physiological variable diagram of each dimension. The basic weight coefficient is the difference between the maximum value and the minimum value in the time series data of the physiological variable diagram of each dimension.
[0198] According to the baseline sequence of the time series data of the physiological variable diagram of each dimension, calculate the difference between the maximum value and the minimum value of the time series data of the physiological variable diagram of each dimension, and use this difference as the basic weight coefficient of the time series data of the physiological variable diagram of each dimension.
[0199] S902, for the time series data of the physiological variable diagram of each dimension, calculate the ratio of the basic weight coefficient of the time series data of the physiological variable diagram of the dimension to the first parameter to obtain the comprehensive weight.
[0200] The first parameter = E × F. Specifically, E is the mean value of the absolute values of the differences between two consecutive points in the time series data of the physiological variable diagram of the current dimension, and F is the sum of the absolute values of the Pearson correlation coefficients between the time series data of the current dimension and the time series data of all other dimensions. The Pearson correlation coefficient is used to measure the linear correlation between two variables, and its value ranges from -1 to 1. The larger the absolute value of the Pearson coefficient, the higher the linear correlation between the two variables.
[0201] Calculate the ratio of the basic weight coefficient of the time series data of the physiological variable diagram of each dimension to the first parameter, and use this ratio as the comprehensive weight of the time series data of the physiological variable diagram of each dimension.
[0202] S903, use the sum of the comprehensive weights of the time series data of the physiological variable diagram of each dimension to regularize the comprehensive weights of the time series data of the physiological variable diagram of each dimension, and obtain the weight coefficients of the time series data of the physiological variable diagram of each dimension.
[0203] Use the sum of the comprehensive weights of the time series data of the physiological variable diagram of each dimension to regularize the comprehensive weights of the time series data of the physiological variable diagram of each dimension, and obtain the weight coefficients of the time series data of the physiological variable diagram of each dimension. Through the regularization process, overfitting of the time series data of the physiological variable diagram of each dimension can be avoided.
[0204] In this embodiment, the basic weight coefficients of each dimension are obtained according to the baseline sequence of the time series data of the physiological variable diagrams of each dimension. The comprehensive weight of the time series data of the physiological variable diagrams of each dimension is obtained according to the basic weight coefficients of each dimension. The weight coefficients of the time series data of the physiological variable diagrams of each dimension are obtained by normalizing the comprehensive weight according to the sum of the comprehensive weights. The weight coefficients are applied to the clustering of change points, further improving the accuracy of clustering.
[0205] Figure 10 FIG. 10 is a schematic structural diagram of a change point position and category detection device provided in Embodiment 8 of the present invention. As shown in FIG. 10, the change point position and category detection device 10 includes: a first extraction module 110, a screening module 120, a grouping module 130, a voting module 140, a second extraction module 150, and a clustering module 160. Specifically:
[0206] The first extraction module 110 is configured to perform baseline extraction on the input multi-dimensional physiological variable diagram time series data to obtain the baseline sequence of the time series data of the physiological variable diagrams of each dimension;
[0207] The screening module 120 is configured to screen all time points in the baseline sequence of the time series data of the physiological variable diagrams of each dimension to obtain candidate change points corresponding to the time series data of the physiological variable diagrams of each dimension;
[0208] The grouping module 130 is configured to group the candidate change points corresponding to the time series data of the physiological variable diagrams of each dimension according to the time gap of the candidate change points to obtain a plurality of candidate change point groups;
[0209] The voting module 140 is configured to vote on each time point candidate change point in a plurality of candidate change point groups by using a generalized likelihood ratio GLR algorithm based on voting, and determine the change point according to the voting result;
[0210] The second extraction module 150 is configured to extract change point sequence data including change points from the time series data of the physiological variable diagrams of each dimension according to the position of the change point;
[0211] The clustering module 160 is configured to cluster the extracted change point sequence data in the form of a tree structure split to obtain a change point classification result.
[0212] In a possible implementation manner, the first extraction module 110 is specifically configured to:
[0213] For each time point, calculate the mean value of the time point and the N time points before the time point, and use the mean value as the baseline corresponding value of the time point, where N is an integer greater than 1;
[0214] Arrange the baseline corresponding values of each time point in time sequence to form the baseline sequence of the time series data of the physiological variable diagrams of each dimension.
[0215] In a possible implementation, the screening module 120 is specifically configured to:
[0216] Perform the following screening process on the baseline sequence of the time-series data of the physiological variables of each dimension in the order of dimensions:
[0217] Compare the ratio of the data difference between each time point and the next time point in the baseline sequence corresponding to the current dimension to the maximum data difference within the GLR window corresponding to the time point with a preset first threshold. Add the time points with a difference greater than the first threshold to the first candidate set corresponding to the current dimension, and delete the time points with a difference not greater than the first threshold. Among them, the GLR window corresponding to the time point includes at least the time point and the next time point;
[0218] Calculate the absolute value of the difference between two consecutive points in the first candidate set corresponding to the current dimension. If there is a time point in the first candidate set corresponding to the current dimension where the absolute value of the difference is greater than the second threshold, then use the time point in the first candidate set corresponding to the current dimension where the absolute value of the difference is greater than the second threshold as a candidate change point, and end the screening process. If there is no time point in the first candidate set corresponding to the current dimension where the absolute value of the difference is greater than the second threshold, then perform screening on the baseline sequence of the time-series data of the physiological variables of the next dimension. Among them, the second threshold is the absolute value of the difference corresponding to the Mth position after sorting the absolute values of the differences between all time points in the baseline sequence corresponding to the current dimension from largest to smallest. M = B × (1 / W), where B is the number of differences between all time points in the baseline sequence corresponding to the current dimension, and W is the size of the GLR window;
[0219] Alternatively, when there is no time point in the first candidate sets corresponding to multiple dimensions where the absolute value of the difference is greater than the second threshold, calculate the absolute value of the difference between two consecutive points in the first candidate set, and take the absolute value of the difference of the absolute value of the difference between two consecutive points to obtain the absolute value of the second-order difference. If there is a time point in the first candidate set of the current dimension where the absolute value of the second-order difference is greater than the third threshold, then use the time point in the current dimension where the absolute value of the second-order difference is greater than the third threshold as a candidate change point, and end the screening process. If there is no time point in the first candidate set corresponding to the current dimension where the absolute value of the second-order difference is greater than the third threshold, then perform screening on the baseline sequence of the time-series data of the physiological variables of the next dimension. Among them, the third threshold is the absolute value of the second-order difference corresponding to the Dth position after sorting the absolute values of the second-order differences between all time points in the baseline sequence corresponding to the current dimension from largest to smallest. D = T × (1 / W), where T is the number of second-order differences between all time points in the baseline sequence corresponding to the current dimension, and W is the size of the GLR window.
[0220] In a possible implementation, the grouping module 130 is specifically configured to:
[0221] Calculate the time difference between two consecutive candidate change points among the candidate change points corresponding to the time series data of the physiological variable diagrams in each dimension;
[0222] When the time difference is less than the length of the smaller part of the front and rear of the change point within the GLR window, add the candidate change point with a later time among the two candidate change points to the candidate change point group corresponding to the candidate change point with an earlier time;
[0223] When the time difference is not less than the length of the smaller part of the front and rear of the change point within the GLR window, create a new candidate change point group for the candidate change point with a later time among the two candidate change points.
[0224] In a possible implementation, the voting module 140 is specifically configured to:
[0225] Delete the single-point groups and the change point groups with a time gap greater than the preset time threshold in the multiple candidate change point groups to obtain available candidate change point groups;
[0226] Calculate the GLR value of each time point within the available candidate change point group according to the GLR window;
[0227] Traverse each time point of the available candidate change point group by sliding the outer voting window and vote for each time point within the candidate change point group. Among them, when voting for each time point within the outer voting window, the voting value of each time point within the outer voting window is the sum of the GLR value of each time point and the GLR values of all time points after this time point within the GLR window, and the time point with the largest voting value within the outer voting window gets one vote;
[0228] When the number of votes of the time point with the most votes within the available candidate change point group exceeds half of the voting opportunities, determine the time point with the most votes within the available candidate change point group as the change point;
[0229] When the number of votes of the time point with the most votes within the available candidate change point group does not exceed half of the voting opportunities, determine that there is no change point within the available candidate change point group.
[0230] In a possible implementation, the clustering module 160 is specifically configured to:
[0231] Use the extracted change point sequence data as the root node and perform splitting in the form of a tree structure split, where
[0232] At each layer, split all the cluster nodes at the current depth;
[0233] Before splitting the current node, a similarity matrix is obtained by taking the negative value of the weighted average of the soft-DTW distance and the SBD distance, and the optimal value of the current node is estimated. The optimal k value and the similarity matrix are used as the input for spectral clustering to complete the clustering of the data within the current node;
[0234] If the optimal k value of the current node is 1, the current node is not split and is used as a leaf node;
[0235] The splitting is completed in a layer-by-layer and node-by-node manner. When the k value of all nodes is 1, the clustering ends and the clustering result is obtained;
[0236] The change point classification result is obtained according to the clustering result.
[0237] In a possible implementation, the clustering module 160 is further configured to:
[0238] Calculate the within-cluster temporal mean for each cluster of the clustering result according to the soft-DTW distance, and use this mean as the representation of the cluster;
[0239] Calculate the difference between each piece of data within the cluster and the cluster representation, calculate the mean and standard deviation of all differences, and add the mean and the standard deviation to obtain the internal statistical characteristics of the cluster;
[0240] Calculate the Fréchet distance between the representations of each cluster and the representations of other clusters. When the Fréchet distance between the representations of two clusters is less than the mean of the internal statistical characteristics of the two clusters, the two clusters are merged.
[0241] In a possible implementation, it further includes a calculation module, and this calculation module is used for:
[0242] Calculate the weight coefficients of the temporal sequence data of each dimension physiological variable graph according to the baseline sequence of the temporal sequence data of each dimension physiological variable graph;
[0243] Calculate the soft-DTW distance and the SBD distance according to the weight coefficients of the temporal sequence data of each dimension physiological variable graph.
[0244] In a possible implementation, the calculation module is further used for:
[0245] Calculate the basic weight coefficients of the temporal sequence data of each dimension physiological variable graph, and the basic weight coefficient is the difference between the maximum value and the minimum value in the temporal sequence data of each dimension physiological variable graph;
[0246] For the time-series data of the physiological variable diagram for each dimension, calculate the ratio of the basic weight coefficient of the time-series data of the physiological variable diagram for the dimension to the first parameter to obtain the comprehensive weight. The first parameter = E × F, where E is the mean of the absolute values of the differences between two consecutive points in the time-series data of the physiological variable diagram for the current dimension, and F is the sum of the absolute values of the Pearson correlation coefficients between the time-series data of the physiological variable diagram for the current dimension and the time-series data of the physiological variable diagrams for all other dimensions;
[0247] Use the sum of the comprehensive weights of the time-series data of the physiological variable diagrams for each dimension to regularize the comprehensive weights of the time-series data of the physiological variable diagrams for each dimension, and obtain the weight coefficients of the time-series data of the physiological variable diagrams for each dimension.
[0248] The device provided in this embodiment can be used to execute the method steps in the above-mentioned method embodiment. The specific implementation manners and technical effects are similar and will not be elaborated here.
[0249] Figure 11 A detection device for the change point position and category provided in Embodiment IX of the present invention. As Figure 11 shown, the detection device 11 for the change point position and category includes:
[0250] At least one processor 111; and
[0251] A memory 112 communicatively connected to at least one processor 111; wherein,
[0252] The memory 112 stores instructions executable by at least one processor 111. The instructions are executed by at least one processor 111 so that at least one processor 111 can execute the method for detecting the change point position and category as described above.
[0253] The specific implementation process of the processor 111 can be referred to the above method embodiment. The specific implementation manners and technical effects are similar and will not be elaborated here.
[0254] Embodiment X of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed by a processor, they are used to implement the method steps in the above-mentioned method embodiment. The specific implementation manners and technical effects are similar and will not be elaborated here.
[0255] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the invention. The present invention is intended to cover any variations, uses, or adaptations of the invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present invention. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.
[0256] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A method for detecting variable point positions and categories, characterized in that Including: Performing baseline extraction on the input multi-dimensional physiological variable graph time series data to obtain the baseline sequence of the time series data of each dimension of physiological variable graphs; The multi-dimensional physiological variable graph time series data is a physiological variable data graph in the medical field for monitoring the health status of medical patients; Screening all time points in the baseline sequence of the time series data of each dimension of physiological variable graphs to obtain the candidate change points corresponding to the time series data of each dimension of physiological variable graphs; Grouping the candidate change points corresponding to the time series data of each dimension of physiological variable graphs according to the time gap between the candidate change points to obtain multiple candidate change point groups; Using the voting-based Generalized Likelihood Ratio (GLR) algorithm to vote on each time point within the multiple candidate change point groups, and determining the change points according to the voting results; Extracting the change point sequence data from the time series data of each dimension of physiological variable graphs according to the positions of the change points, where the change point sequence data includes the change points; Clustering the extracted change point sequence data in the form of a tree structure split to obtain the change point classification result.
2. The method according to claim 1, wherein The screening of all time points in the baseline sequence of the time series data of each dimension of physiological variable graphs to obtain the candidate change points corresponding to the time series data of each dimension of physiological variable graphs includes: Performing the following screening process on the baseline sequence of the time series data of each dimension of physiological variable graphs in the dimension order: Comparing the ratio of the data difference between each time point and the next time point in the baseline sequence corresponding to the current dimension to the maximum data difference within the GLR window corresponding to the time point with a preset first threshold, adding the time points with the difference greater than the first threshold to the first candidate set corresponding to the current dimension, and deleting the time points with the difference not greater than the first threshold, where the GLR window corresponding to the time point includes at least the time point and the next time point; Calculating the absolute value of the difference between two consecutive points in the first candidate set corresponding to the current dimension. If there are time points in the first candidate set corresponding to the current dimension with the absolute value of the difference greater than the second threshold, then taking the time points in the first candidate set corresponding to the current dimension with the absolute value of the difference greater than the second threshold as the candidate change points and ending the screening process. If there are no time points in the first candidate set corresponding to the current dimension with the absolute value of the difference greater than the second threshold, then screening the baseline sequence of the time series data of the physiological variable graph of the next dimension, where the second threshold is the absolute value of the difference corresponding to the Mth position after sorting the absolute values of the differences between all time points in the baseline sequence corresponding to the current dimension from large to small, and M = B×(1 / W), where B is the number of differences between all time points in the baseline sequence corresponding to the current dimension, and W is the size of the GLR window; Alternatively, when there is no time point in the corresponding first candidate set of the multiple dimensions where the absolute value of the difference is greater than the second threshold, calculate the absolute value of the difference between two consecutive points in the first candidate set, and then calculate the absolute value of the difference of the absolute value of the difference between the two consecutive points to obtain the absolute value of the second-order difference. If there is a time point in the first candidate set of the current dimension where the absolute value of the second-order difference is greater than the third threshold, then use the time point in the current dimension where the absolute value of the second-order difference is greater than the third threshold as the candidate change point and end the screening process. If there is no time point in the first candidate set corresponding to the current dimension where the absolute value of the second-order difference is greater than the third threshold, then screen the baseline sequence of the time series data of the physiological variable graph of the next dimension. Wherein, the third threshold is the absolute value of the second-order difference corresponding to the D-th position after sorting the absolute values of the second-order differences of all time points in the baseline sequence corresponding to the current dimension from large to small, and D = T×(1 / W), where T is the number of second-order differences of all time points in the baseline sequence corresponding to the current dimension, and W is the size of the GLR window.
3. The method according to claim 2, wherein Grouping the candidate change points corresponding to the time series data of the physiological variable graph of each dimension according to the time gap between the candidate change points, to obtain a plurality of candidate change point groups, including: Calculating the time difference between two consecutive candidate change points among the candidate change points corresponding to the time series data of the physiological variable graph of each dimension; When the time difference is less than the length of the smaller part of the front part and the rear part of the change point within the GLR window, add the candidate change point with a later time among the two candidate change points to the candidate change point group corresponding to the candidate change point with an earlier time; When the time difference is not less than the length of the smaller part of the front part and the rear part of the change point within the GLR window, create a new candidate change point group for the candidate change point with a later time among the two candidate change points.
4. The method according to claim 3, wherein Using the generalized likelihood ratio GLR algorithm based on voting to vote on each time point within the multiple candidate change point groups, and determining the change point according to the voting results, including: Deleting the single-point groups and the change point groups with a time gap greater than the preset time threshold within the multiple candidate change point groups to obtain available candidate change point groups; Calculating the GLR value of each time point within the available candidate change point groups according to the GLR window; Traversing each time point of the available candidate change point groups by sliding the outer voting window and voting on each time point within the candidate change point groups. When voting on each time point within the outer voting window, the voting value of each time point within the outer voting window is the sum of the GLR value of each time point and the GLR values of all time points after this time point within the GLR window, and the time point with the largest voting value within the outer voting window gets one vote; When the number of votes of the time point with the most votes within the available candidate change point group exceeds half of the chance of getting votes, determine the time point with the most votes within the available candidate change point group as the change point; When the number of votes at the time point with the most votes in the available candidate change point group does not exceed half of the voting opportunity, it is determined that there is no change point in the available candidate change point group.
5. The method according to any one of claims 1-4, characterized in that, Clustering the extracted change point sequence data in the form of splitting in a tree structure to obtain a change point classification result, including: Taking the extracted change point sequence data as the root node and splitting it in the form of splitting in a tree structure, where each layer splits all cluster nodes at the current depth; Before splitting the current node, obtaining a similarity matrix by taking the negative value of the weighted average of the soft-DTW distance and the SBD distance, estimating the optimal k value of the current node, and using the optimal k value and the similarity matrix as the input of spectral clustering to complete the clustering of the data within the current node; If the optimal k value of the current node is 1, the current node is not split and the current node is used as a leaf node; Completing the splitting in a layer-by-layer and node-by-node manner. When the k value of all nodes is 1, the clustering ends and a clustering result is obtained; Obtaining the change point classification result according to the clustering result.
6. The method according to claim 5, characterized in that, The obtaining the change point classification result according to the clustering result includes: Calculating the internal time series mean of each cluster of the clustering result according to the soft-DTW distance, and using the mean as the representation of the cluster; Calculating the difference between each piece of data in the cluster and the cluster representation, calculating the mean and standard deviation of all differences, and adding the mean and the standard deviation to obtain the internal statistical characteristics of the cluster; Calculating the Fréchet distance between the representations of each cluster and the representations of other clusters. When the Fréchet distance between the representations of two clusters is less than the mean of the internal statistical characteristics of the two clusters, the two clusters are merged.
7. The method according to claim 6, wherein It also includes: Calculating the weight coefficient of the time series data of each dimension physiological variable graph according to the baseline sequence of the time series data of each dimension physiological variable graph; Calculating the soft-DTW distance and the SBD distance according to the weight coefficients of the time series data of each dimension physiological variable graph.
8. The method according to claim 7, wherein The calculating the weight coefficient of the time series data of each dimension physiological variable graph according to the baseline sequence of the time series data of each dimension physiological variable graph includes: Calculating the basic weight coefficient of the time series data of each dimension physiological variable graph, where the basic weight coefficient is the difference between the maximum value and the minimum value in the time series data of each dimension physiological variable graph; For the time series data of each dimension physiological variable graph, calculating the ratio of the basic weight coefficient of the time series data of the dimension physiological variable graph to the first parameter to obtain a comprehensive weight, where the first parameter = E × F, where E is the mean of the absolute values of the differences between two consecutive points in the time series data of the dimension physiological variable graph, and F is the sum of the absolute values of the Pearson correlation coefficients between the dimension and the time series data of all other dimension physiological variable graphs; Using the sum of the comprehensive weights of the time series data of each dimension physiological variable graph to regularize the comprehensive weights of the time series data of each dimension physiological variable graph to obtain the weight coefficients of the time series data of each dimension physiological variable graph.
9. A detection device for variable point positions and categories, characterized in that, It includes: A first extraction module for extracting a baseline from the input time-series data of multi-dimensional physiological variable graphs to obtain a baseline sequence of the time-series data of each dimension of physiological variable graphs; The multi-dimensional physiological variable graph time-series data is a physiological variable data graph used in the medical field to monitor the health status of medical patients; A screening module for screening all time points in the baseline sequence of the time-series data of each dimension of physiological variable graphs to obtain candidate change points corresponding to the time-series data of each dimension of physiological variable graphs; A grouping module for grouping the candidate change points corresponding to the time-series data of each dimension of physiological variable graphs according to the time gap between the candidate change points to obtain a plurality of candidate change point groups; A voting module for voting on each time point in the plurality of candidate change point groups by using a generalized likelihood ratio GLR algorithm based on voting, and determining a change point according to the voting result; A second extraction module for extracting change point sequence data from the time-series data of each dimension of physiological variable graphs according to the position of the change point, where the change point sequence data includes the change point; A clustering module for clustering the extracted change point sequence data in the form of a tree structure split to obtain a change point classification result.
10. A detection device for variable point positions and categories, characterized in that, Including: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory is used to store instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Harmonic anomaly identification method based on variable point segmentation and sequence clustering
CN111611961A
Time series data classification method based on multi-level shape
CN111814897A