A Multivariate Time Series Classification Method Based on Deep Learning
Through a multivariate time series classification method based on deep learning, feature extraction and classification network training of knee six-degree-of-freedom movement data is solved, and misdiagnosis of the diagnosis of anterior cruciate ligament injury in the existing technology is solved, achieving more efficient, accurate and safe diagnostic effects.
Patent Information
- Application Number
- CN202111217841.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-10-19
AI Technical Summary
The prior art has missed or misdiagnosed in the diagnosis of anterior cruciate ligament injury of knee joints. Traditional imaging technology cannot explore kinematic characteristics, and there are problems such as inflexible use, slow examination speed, high cost and radiation hazards.
A multivariate time series classification method based on deep learning is adopted to transform feature, global and local feature extraction of knee joint six-degree of freedom movement data, and build a classification network for multi-objective learning training to effectively capture the correlation characteristics between variables and the timing characteristics within univariate.
This method can more accurately diagnose anterior cruciate ligament injury of the knee joint, improves the accuracy and efficiency of diagnosis, reduces costs, and avoids radiation hazards.
Smart Images

Figure CN114154551B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and more particularly, to a multi-variate time series classification method based on deep learning. Background Art
[0002] As one of the most complex weight-bearing joints in the human body, the knee joint bears more load during walking and exercise, playing an important shock-absorbing role. At the same time, knee joint injuries are very common. Among them, anterior cruciate ligament deficiency (ACL-D) is one of the most common sports injuries. With the development of the economy and the improvement of living standards, people's sports needs have gradually increased, followed by an increase in sports injuries. The anterior cruciate ligament (ACL) not only restricts the anterior translation of the tibia, but also controls the axial rotation and varicose vein movement of the knee joint, playing a high role value in controlling the stability of the human knee joint. It is an important static structure for stabilizing the knee joint. After injury, it can cause knee joint instability and cannot heal on its own, which will seriously affect the quality of life of patients. If not treated surgically in time, it is very likely to cause serious articular cartilage damage, seriously affecting the quality of life and sports level of patients, and ultimately artificial joint replacement is required.
[0003] The survey results of medical institutions in developed foreign regions show that in the more than 20 years from the late 1990s to 2016, the initial diagnosis accuracy rate of anterior cruciate ligament (ACL) injury has hardly improved, and only 14.4% of patients were correctly diagnosed at the initial diagnosis. At present, the detection and evaluation diagnosis of such injuries in clinical practice mainly rely on subjective interrogation, manual examination and imaging methods, so there are still missed diagnoses or misdiagnoses in the diagnosis of anterior cruciate ligament injury or rupture of the knee joint. Common traditional imaging techniques for anterior cruciate ligament injury include CT and MRI, etc., which obtain the anatomical structure information of the injury site for diagnosis. However, these techniques cannot explore the kinematic characteristics of patients, and have the disadvantages of inflexible use, slow examination speed, high cost and radiation hazards. In addition, arthroscopic examination is the gold standard for diagnosing anterior cruciate ligament injury, but it is expensive, invasive, and causes greater harm to patients.
[0004] With the update and progress of biomechanical techniques, recent studies have found that the analysis of knee joint kinematics can provide important theoretical basis for the evaluation of knee joint injuries in the dynamic activity state, and may become one of the important diagnostic methods for knee joint injuries, laying an important fundamental role for the effective diagnosis of knee joint injuries in clinical practice. Structure determines function, so to a certain extent, the change of motor function can reflect the change of anterior cruciate ligament structure. Such as Figure 1The 6-degree-of-freedom (6DoF) kinematic data shown can be obtained from the Opti-knee 3D motion analysis system. This system has the advantages of flexible use, fast detection speed, low cost, and low risk. As far as is known, although 6-degree-of-freedom (6DoF) kinematic data can optimize the ACL-D diagnosis process, deep neural networks are rarely used in the study of ACL-D diagnosis. In this work, an attempt is made to transform the diagnosis of ACL-D into a multivariate time series classification problem based on 6DoF data.
[0005] A time series is a set of real-valued observations arranged in chronological order. A multivariate time series (MTS) is a set of co-evolving time series, usually recorded simultaneously over time by a set of sensors. With the advancement of sensor technology, the multivariate time series classification (MTSC) problem may be one of the most important problems in the field of time series data mining and has received a great deal of attention in recent decades. However, there is not much work specifically studying anterior cruciate ligament injuries of the knee, and even less research on the Opti-Knee six-degree-of-freedom dataset.
[0006] Existing methods for multivariate time series classification (MTSC) can be divided into methods based on subsequences of time series, bags of patterns, and methods based on neural networks. In the past decade, classifiers based on self-sequences and bags of patterns have been well studied, such as WeaselMuse, ShapeNet, etc. These methods require converting time series into a set of subsequences as candidate features. However, the size of the feature space and the number of subsequences make feature selection difficult. Recently, deep learning-based methods have achieved good performance in MTSC, such as MLSTM-FCN, TapNet, etc. However, these methods require sufficient data to train a large number of parameters.
[0007] In addition, the features of the 6-degree-of-freedom dataset in the ACL are also different from the most common datasets in classical MTSC problems. The observations are summarized as follows:
[0008] 1) Complex correlations among 6-degree-of-freedom variables.
[0009] The 6 degrees of freedom of the knee joint are affected by the knee joint, which will generate complex spatio-temporal dynamic correlations in the sequence data. In classical MTSC models, the features of each time series are extracted independently and then concatenated for subsequent learning processes. It is believed that the simple extraction and concatenation processes may lose the feature relationships between variables.
[0010] 2) High variability and high volatility. Due to the high variability of human body postures and gaits, even if the data come from the same category, there may be significant fluctuations in the 6-degree-of-freedom data. This is different from typical MTSC data. The six degrees of freedom of the knee joint are affected by individual differences and spatio-temporal differences. Individuals with different knee health levels have different six degrees of freedom at different times, and the movement patterns are complex and variable, with complex change rules.
[0011] 3) Limited data. Medical data must be annotated by professional doctors before it can be used. Compared with common datasets in MTSC, there are only dozens to hundreds of samples in ACL. Therefore, the amount of data is not sufficient to support overly complex models. When constructing a model, the complexity of the model must be considered. Summary of the Invention
[0012] The present invention provides a multivariate time series classification method based on deep learning, which can effectively capture the local and global correlation features between variables, as well as the local temporal features and global features within a single variable.
[0013] In order to achieve the above technical effects, the technical solution of the present invention is as follows:
[0014] A multivariate time series classification method based on deep learning, comprising the following steps:
[0015] S1: Perform feature transformation on the variables;
[0016] S2: Extract global features and local features from the variables after feature transformation;
[0017] S3: Build a classification network;
[0018] S4: Perform multi-objective learning training on the built classification network, input the features obtained in step S2 into the trained classification network, and finally make the feature distances of similar samples closer and the distances of dissimilar samples farther.
[0019] Further, the specific process of step S1 is as follows:
[0020] The time series corresponding to the i-th variable can be expressed as a time series vector where l is the sequence length. After performing phase space reconstruction on the original time series vector using delay embedding, the j-th term of the reconstructed vector corresponding to the i-th variable is expressed as:
[0021]
[0022] where j = 1, 2, …, l - (d t - 1)·t, d t is the embedding dimension of the phase space, t represents the time delay, and the s-th sample after delay embedding is expressed as follows:
[0023]
[0024] where m = l - (d t - 1)·t, and each point in the new phase space represents a possible state. The delay embedding dimension d t is 3, and the time delay t is 1.
[0025] Furthermore, the process of global feature extraction in step S2 is as follows:
[0026] Each multivariate time series sample learns a low-dimensional embedding where d e is the final embedding dimension, and a neural network-based method is always used where θ is a set of function parameters, is the time series data after delay embedding;
[0027] First, a pooling operation is performed on the third dimension of the time series data and a permute operation is carried out to convert the 3D to 2D Then, where d l is the size of the hidden size parameter of the LSTM layer, that is, the global feature extraction is completed.
[0028] Furthermore, the process of local feature extraction in step S2 is as follows:
[0029] 1), Using two layers of convolution, regarding 96 time points as 96 channels, and focusing on learning the correlation features among 18 variables. At this time, no convolution is performed on three adjacent time dimensions of the delay embedding. The average pooling after two layers of convolution strengthens the role of the delay embedding, realizing the extraction of local features among different variables at adjacent time points; at the same time, pooling is performed in the feature direction to reduce the subsequent calculation amount. Here, average pooling is used because it is in the initial feature mining stage. Compared with max pooling, average pooling can retain more information;
[0030] 2), First, convert the features of all channels learned in step 1) into one channel. After completing the dimension conversion, it is convenient for the subsequent convolutional layer to deeply mine the relationship among the features learned in the first part. Finally, use squeeze to reduce the dimension, which is convenient for subsequent use of 1DCNN. Compared with always using 2DCNN, 1DCNN reduces the number of parameters of the subsequent convolution. After convolution, max pooling is used because in the deep mining stage, some unobvious features can be ignored and the attention can be focused on the most significant features;
[0031] 3) Use 1D-CNN and pooling to reduce the data dimension while capturing deep interactions to reduce the number of parameters in the subsequent fully connected layer.
[0032] Further, after passing through the fully connected layer, the final output of a single local feature extraction sub-network is The final output after concatenating the outputs of a sub-networks is The final output of the entire feature extraction part is Next, concatenate the outputs of the two parts of the sub-network to obtain the final embedding
[0033] Further, the process of step S3 is as follows:
[0034] The classification network consists of two fully connected layers. Between the two fully connected layers, the non-linear activation function ReLu is used to obtain non-linear features, and a batch normalization layer is used to stabilize the training and prevent overfitting. The classification network can be expressed as:
[0035]
[0036] Further, the process of step S4 is as follows:
[0037] In the first step, use supervised training CenterLoss to make the inter-class distance of similar samples smaller. While CenterLoss keeps the features of different classes separable, it minimizes the intra-class variation. The formula is as follows:
[0038]
[0039] represents the feature center of the s-th class. Since it is necessary to consider the entire training set and average the features of each class in each iteration, the efficiency is extremely low. Therefore, in each iteration, instead of updating the center for the entire training set, the update is performed based on mini-batch training, and the center is calculated by averaging the features of the corresponding class;
[0040] And in order to improve the training effect of the subsequent classification network, in this stage, the pre-training of the classification network will also be completed. In this stage, the joint supervision of BCELoss and CenterLoss is used to train the network for feature learning. The formula for the first step is as follows:
[0041]
[0042] After the similar samples are sufficiently aggregated, in the second step, TripletLoss is used to improve the generated features to increase the differences between different classes, thereby enhancing the quality of the generated features and reducing the classification burden on the classification network. Since in the first step of training, CenterLoss has already made the similar samples as close as possible, there is no need to select multiple positive and negative samples as in related works such as ShapeNet. Instead, select the most similar sample farthest from the target sample as the positive sample, and the most dissimilar sample closest to the target sample as the negative sample. The training formula for the second step is as follows:
[0043]
[0044] After the final effective features are generated, in the third step, BCELoss is used alone to formally train the classification network module to improve the performance of the classification network.
[0045] Furthermore, the classification network consists of 2 fully connected layers, 1 activation function layer, 1 batch normalization layer, and 1 sigmoid classification function layer, and finally converts the features into probability values between 0 and 1.
[0046] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:
[0047] The present invention carefully designs 4 sets of local feature extraction sub-networks that can fully explore the correlation features between variables and the internal features of a single variable. Different from general convolutional networks, the designed sub-networks enable two-dimensional convolutional networks of two dimensions to be friendly integrated into one network. Compared with a pure 1DCNN network, this network can capture deeper features; compared with a pure 2DCNN network, this network can reduce the number of parameters. While using BaseCNN to capture the interactions between variables, an LSTM with long-range dependence ability is also used to construct a sub-network, and its long-term memory characteristics are utilized to strengthen the global temporal characteristics of the model and help the model better capture global temporal features. In addition, a three-step training mode is cleverly designed to effectively play the roles of CenterLoss and TripletLoss to process the volatile and specific features of this dataset, providing good feature embeddings for the final classification network. The present invention greatly promotes the existing research on artificial intelligence-assisted diagnosis of anterior cruciate ligament and has great clinical significance and practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Six-degree-of-freedom data obtained by Opti-Knee;
[0049] Figure 2 Overall framework structure diagram of the present invention;
[0050] Figure 3Local feature extraction sub-network structure diagram. Specific implementation mode
[0051] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the present patent;
[0052] To better illustrate this embodiment, some components in the accompanying drawings are omitted, enlarged or reduced, and do not represent the dimensions of the actual product;
[0053] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.
[0054] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0055] For the diagnosis of anterior cruciate ligament injury, in order to fully explore the individual kinematic features contained in the six-degree-of-freedom dataset of both legs obtained from the Opti-Knee test, the present invention designs a deep learning-based multivariate time series classification method. This method can effectively capture the local and global correlation features between variables, as well as the local temporal features and global features within a single variable, so as to obtain the spatio-temporal gait features embedded in the six-degree-of-freedom motion data of the joints measured by the joint three-dimensional motion measurement system.
[0056] In this section, we introduce the proposed multi-view multi-objective neural network to fully explore the individual kinematic features contained in the six-degree-of-freedom data. The overall framework of this method is as Figure 2 shown. To solve the problem that most MTSC models lose the correlation features between variables during the feature extraction stage. Our model focuses on exploring the correlation features between variables, including local features and global features. In addition, our model can also mine the internal temporal features of each variable. This method is mainly divided into the following parts: feature transformation, feature extraction, and multi-objective learning.
[0057] The first part: Feature transformation
[0058] To retain the features of the gait system dynamics, we use the phase space reconstruction (PSR) technique to transform the features of the time series into the topological properties of geometric objects in the embedding space. Delay embedding is the most common PSR technique. All possible states of the original dynamic system are represented as unique points in the reconstructed space. After using delay embedding, our research object is no longer multiple one-dimensional time series, but a phase space with the same topological meaning as the original time series. After delay embedding, the dimension of the data at a single moment increases, making the neighborhood correlation of the data stronger. The state represented by each point can reflect the information of multiple moments, enhancing the expressiveness of the data and facilitating the capture of local features.
[0059] The time series corresponding to the i-th variable can be represented as a time series vector where \(l\) is the sequence length. After performing phase space reconstruction on the original time series vector using delay embedding, the \(j\)-th term of the reconstructed vector corresponding to the \(i\)-th variable is expressed as:
[0060]
[0061] where \(j = 1, 2, \cdots, l-(d t -1)\cdot t\), \(d t is the embedding dimension of the phase space, and \(t\) represents the time delay. The \(s\)-th sample after delay embedding is represented as follows:
[0062]
[0063] where \(m = l-(d t -1)\cdot t\). Each point in the new phase space represents a possible state. The delay embedding dimension \(d t is 3, and the time delay \(t\) is 1.
[0064] Part Two: Feature Transformation
[0065] The present invention can learn a low-dimensional embedding for each multivariate time series sample where \(d e is the final embedding dimension. Generally speaking, the present invention uses a neural network-based method where \(\theta\) is a set of function parameters, and is the time series data after delay embedding.
[0066] To solve the problem that most MTSC models lose the correlation features between variables during the feature extraction stage, we design a local feature extraction part that focuses on exploring the local correlation features between variables. While completing the local feature extraction, we also use LSTM with long-range dependence ability to construct a global feature extraction part, leveraging its long-term memory characteristics to strengthen the global time series characteristics of the model and help the model better capture the global time series features. First, we set two groups of inputs arranged by leg and arranged by degree of freedom: 1) Arranged by leg: We put the 6 degrees of freedom of the same leg together, and then connect the variables in the order of the leg to be predicted, the auxiliary leg, and the difference between the two legs; 2) Arranged by 6 degrees of freedom: We put the variables with the same degree of freedom together, and then connect 6 groups of variables.
[0067] In the global feature extraction part, since the dimension of the reconstructed data is 3D and not suitable for LSTM, we first perform a pooling operation on the third dimension of the time series data and a permute operation to convert the 3D to 2D Then we obtain where \(d lIt is the size of the hidden size parameter of the LSTM layer.
[0068] The local feature extraction part consists of four sub-networks. The first sub-network is used to capture the features among all variables; the second sub-network focuses on the features within a single leg; the third sub-network focuses on the features within a single degree of freedom; and the fourth sub-network is used to capture the features within a single variable. Different from general pure one-dimensional or pure two-dimensional convolutional neural networks (CNNs), we enable 1D-CNN and 2D-CNN to be friendly integrated into a network by compressing and expanding channels and exchanging dimensions among corresponding layers, so as to give full play to their roles in obtaining features. The differences among the four sub-networks are mainly reflected in the first stage, which is achieved by adjusting the convolution stride and convolution kernel size parameters of the 1D-CNN in stage 1. Since these sub-networks are similar in structure, we choose one of them as an example to introduce. The specific structural parameters of the four sub-networks are shown in Tables 1-4.
[0069] As Figure 3 shown, a local feature extraction sub-network can be roughly divided into three stages.
[0070] Stage 1: Using two layers of convolution, regarding 96 time points as 96 channels, and focusing on learning the correlation features among 18 variables. At this time, no convolution is performed on the three adjacent time dimensions of the delay embedding. The average pooling after the two layers of convolution strengthens the role of the delay embedding and realizes the extraction of local features among different variables at adjacent time points; at the same time, pooling is performed in the feature direction to reduce the subsequent computational amount. Here, average pooling is used because it is in the initial feature mining stage, and compared with max pooling, average pooling can retain more information.
[0071] Stage 2: First, transform the features of all channels learned in the first stage into one channel. After completing the dimension transformation, it is convenient for the subsequent convolutional layer to deeply mine the relationship among the features learned in the first part. Finally, use squeeze to reduce the dimension, which is convenient for subsequent use of 1DCNN. Compared with always using 2DCNN, 1DCNN reduces the number of parameters in the subsequent convolution. After convolution, max pooling is used because in the deep mining stage, we can ignore some unobvious features and focus on the most significant features.
[0072] Stage 3: Using 1D-CNN and pooling to capture deep interactions while reducing the data dimension to reduce the number of parameters in the subsequent fully connected layer.
[0073] Different from other models, as the issue of small data volume needs to be fully considered, instead of directly concatenating the outputs of the two parts and then connecting them to the fully-connected layer, in this paper, max pooling is performed separately on each part, and the most prominent features are retained to reduce the dimension before connecting to the fully-connected layer, and then the features are concatenated after the dimension is reduced. After each part passes through the fully-connected layer, the final output of a single local feature extraction sub-network is The final output after concatenating the outputs of a sub-networks is The final output of the entire feature extraction part is Next, the outputs of the two sub-networks are concatenated to obtain the final embedding
[0074] Table 1: Parameter settings of sub-network one
[0075]
[0076]
[0077] Table 2: Parameter settings of sub-network two
[0078]
[0079]
[0080] Table 3: Parameter settings of sub-network three
[0081]
[0082]
[0083] Table 4: Parameter settings of sub-network four
[0084] Number of layers Name Stride Kernel size Number of kernels Input dimension Output dimension 1 permute - - - (94,18,3) (3,94,18) 2 Conv1d (2,1) (24,1) 10 (3,94,18) (10,36,18) 3 relu - - - (10,36,18) (10,36,18) 4 avg_pool2d (3,1) (3,1) - (10,36,18) (10,12,18) 5 Conv1d (3,1) (3,1) 50 (10,12,18) (50,4,18) 6 relu - - - (50,4,18) (50,4,18) 7 avg_pool2d (4,1) (4,1) - (50,4,18) (50,1,18) 8 permute - - - (50,1,18) (18,50,1) 9 max_pool2d (5,1) (5,1) - (18,50,1) (18,10,1) 10 reshape - - - (18,10,1) (180) 11 Linear - - - (180) (2)
[0085] Part Three: Classification network
[0086] The classification network is composed of two fully-connected layers. Between the two fully-connected layers, we use the non-linear activation function ReLu to obtain non-linear features, and use a batch normalization layer to make the training stable and prevent overfitting. This classification network can be expressed as:
[0087]
[0088] Part Four: Multi-objective learning
[0089] The ultimate goal of this invention is to accurately judge whether the anterior cruciate ligament is ruptured. Without a doubt, the quality of the features generated before the classification network will affect the classification task. Therefore, in order to improve the quality of the generated features, this invention proposes a three-step training mode.
[0090] The goal of learning / training in the feature generation part is to ensure that time series with the same label obtain similar representations, and vice versa. However, the six degrees of freedom of the original knee joint data in this paper are affected by individual differences and spatio-temporal differences. The six degrees of freedom of individuals with different knee health levels are different at different times, and the motion patterns are complex and variable, and the change rules are complex. Even professional doctors cannot accurately identify them. Therefore, the generated features of the same category also have differences. The commonly used loss for feature generation is TripletLoss. The idea is that for each target sample, a positive sample of the same category and a negative sample of a different category are selected, so that the distance between the target sample and the positive sample is shortened, and the distance from the negative sample is lengthened. On this basis, there are many improved works, mostly improving the selection of positive and negative samples. For example, for each target sample, one positive sample and multiple negative samples are selected; or multiple positive samples and multiple negative samples. However, when the features of the same category have large differences, directly lengthening the distance from the negative sample and shortening the distance from the positive sample will cause serious training oscillation and non-convergence in the feature generation stage. In this case, the sample distribution of the trained features will be more chaotic, that is, ineffective training.
[0091] To solve this problem, in the first step of the training stage, this paper first uses supervised training CenterLoss to make the inter-class distance of samples of the same category smaller. While keeping the features of different classes separable, CenterLoss minimizes the intra-class variation. The formula is as follows:
[0092]
[0093] represents the feature center of the s-th class. Since we need to consider the entire training set and average the features of each class in each iteration, the efficiency is extremely low. Therefore, in each iteration, instead of updating the center for the entire training set, the update is performed based on mini-batch training, and the center is calculated by averaging the features of the corresponding class.
[0094] And to improve the training effect of the subsequent classification network, in this stage, we will also complete the pre-training of the classification network. In this stage, we use the joint supervision of BCELoss and CenterLoss to train the network for feature learning. The formula of the first step is as follows:
[0095]
[0096] After the similar samples are sufficiently aggregated, in the second step, TripletLoss is used to improve the generated features to increase the differences between different classes, thereby improving the quality of the generated features and reducing the classification burden on the classification network. Since in the first step of training, CenterLoss has made the similar samples as close as possible, there is no need to select multiple positive and negative samples as in related works such as ShapeNet. We select the most distant similar sample from the target sample as the positive sample and the closest dissimilar sample to the target sample as the negative sample. The training formula for the second step is as follows:
[0097]
[0098] After the final effective features are generated, in the third step, we use BCELoss alone to formally train the classification network module to improve the performance of the classification network.
[0099] Experiment
[0100] All the experimental data come from the real world. Through the joint three-dimensional motion measurement system, the six-degree-of-freedom motion data of the individual's single-leg joints can be captured, including: abduction / adduction (unit: degree), flexion / extension (unit: degree), internal rotation / external rotation (unit: degree), anterior-posterior displacement (unit: centimeter), medial-lateral displacement (unit: centimeter), and vertical displacement (unit: centimeter). A total of 209 individuals participated in the test. By using the OptiKnee joint three-dimensional motion measurement system, the six-degree-of-freedom joint data of both legs of each individual were obtained. Among the 209 subjects, due to the missing gait signals of 4 individuals, the total number of individual samples in the experiment was 205. Among the 205 samples: the number of samples with both legs normal was 60, the number of samples with the left leg broken was 77, the number of samples with the right leg broken was 50, and the number of samples with both legs broken was 18.
[0101] Based on the above analysis, in order to ensure a relatively balanced distribution of positive and negative samples, the left leg was selected as the leg to be predicted and the right leg as the auxiliary leg, and the label value for the anterior cruciate ligament injury of the left leg was set to 1 and 0 for normal. After this division, the number of positive samples (left leg broken) was 95, and the number of negative samples (right leg normal) was 110. To ensure sample balance, we adopted an oversampling strategy during training, randomly selecting a certain number of positive samples from the positive samples in the training set to ensure the same ratio of positive and negative samples in the training set.
[0102] For each model, we calculated the accuracy (acc), specificity (Spe), recall, the area under the ROC curve (AUC), F1, and precision (Pre) for the classification task (ACL-D diagnosis). We compared the method of the present invention with 10 other different methods. Among traditional machine learning methods, we selected SVM, Random Forest, and XGBoost; among neural networks, we selected VGG, GRU, and LSTM. Among the methods specifically for dealing with multivariate time series classification problems, we selected three representative methods, MLSTM-FCN, TapNet, and ShapeNet. Among them, ShapeNet is currently the best method based on time series subsequences, and TapNet is currently the best method based on neural networks.
[0103] We used five-fold cross-validation to test all methods and took the average of ten five-fold experiments as the final result. The results are as follows:
[0104]
[0105] From the above experimental analysis, it can be seen that the multi-view multi-object neural network method of the present invention has significant advantages in the anterior cruciate ligament injury diagnosis task, can better capture the local temporal features and global association features within and between single variables, and is more conducive to mining the feature patterns of single-leg joint six-degree-of-freedom motion data.
[0106] The same or similar reference numerals correspond to the same or similar components;
[0107] The positional relationships described in the drawings are for illustrative purposes only and should not be construed as limitations on this patent;
[0108] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A multi-variate time series classification method based on deep learning, characterized in that, it includes the following steps: S1: Through a joint three-dimensional motion measurement system, capture the six-degree-of-freedom motion data of an individual's single-leg joints. The variables include: abduction / adduction, flexion / extension, internal rotation / external rotation, anterior-posterior displacement, medial-lateral displacement, and vertical displacement; two groups of inputs are set, arranged by leg and by degree of freedom: 1) Arranged by leg: Place the 6 degrees of freedom of the same leg together, and then connect the variables in the order of the leg to be predicted, the assisting leg, and the difference between the two legs; 2) Arranged by 6 degrees of freedom: Place the variables of the same degree of freedom together, and then connect 6 groups of variables; perform feature transformation on the variables; The specific process of step S1 is: The time series corresponding to the i-th variable is represented as a time series vector , where l is the sequence length. After performing phase space reconstruction on the original time series vector using delay embedding, the j-th term of the reconstructed vector corresponding to the i-th variable is expressed as: Among them , is the embedding dimension of the phase space, t represents the time delay, and the s-th sample after time-delay embedding is expressed as follows: Among them , each point in the new phase space represents a state, the delay embedding dimension is 3, and the time delay t is 1; S2: Perform global feature extraction and local feature extraction on the variables after feature transformation; The process of performing global feature extraction in step S2 is: Each multivariate time series sample learns a low-dimensional embedding , where is the final embedding dimension, and a neural network-based method is always used , where is the set of function parameters, is the time series data after delay embedding; First, perform a pooling operation on the third dimension of the time series data and a permute operation to convert the 3D to 2D , and then obtain , where is the size of the hidden size parameter of the LSTM layer, that is, global feature extraction is completed; In step S2, the local feature extraction part consists of four sub-networks; The first sub-network is used to capture the features between all variables; The second sub-network focuses on the features within a single leg; The third sub-network focuses on the features within a single degree of freedom; The fourth sub-network is used to capture the features within a single variable; By compressing and expanding channels and exchanging dimensions between corresponding layers, 1D-CNN and 2D-CNN are friendly integrated into a network, so as to fully play their roles to obtain features; The differences between the 4 sub-networks are reflected in the first stage, and are achieved by adjusting the convolution stride and convolution kernel size parameters of 1D-CNN in stage 1; The process of a local feature extraction sub-network performing local feature extraction is: 1), Use two layers of convolution, regard 96 time points as 96 channels, and focus on learning the correlation features between 18 variables. At this time, no convolution is performed on the three adjacent time dimensions of the delay embedding. The average pooling after two layers of convolution strengthens the role of the delay embedding, and realizes the extraction of local features between different variables at adjacent time points; At the same time, pooling is performed in the feature direction to reduce the subsequent calculation amount. Here, average pooling is used because it is in the preliminary feature mining stage. Compared with max pooling, average pooling can retain more information; 2), First, transform the features of all channels learned in step 1) into one channel. After completing the dimension transformation, it is convenient for the subsequent convolution layer to deeply mine the relationship between the features learned in the first part. Finally, use squeeze to reduce the dimension, which is convenient for subsequent use of 1DCNN. Compared with always using 2DCNN, 1DCNN reduces the number of parameters in the subsequent convolution. After convolution, max pooling is used because in the deep mining stage, some unobvious features can be ignored and the attention can be focused on the most significant features; 3), Use 1D-CNN and pooling to reduce the data dimension while capturing deep interactions to reduce the number of parameters in the subsequent fully connected layer; S3: Build a classification network; S4: Conduct multi-objective learning training on the constructed classification network. Input the features obtained in step S2 into the trained classification network, ultimately making the distances between the features of the same-class samples closer and the distances between different-class samples farther apart. Classify the newly obtained individual six-degree-of-freedom motion data of the leg joints through the classification network to determine whether the anterior cruciate ligament is ruptured.
2. The deep learning-based multi-variate time series classification method according to claim 1, wherein, After passing through the fully connected layer, the final output of a single local feature extraction sub-network is , After concatenating the final outputs of the sub-networks, we get . The final output of the entire feature extraction part is . Next, we concatenate the outputs of the two parts of the sub-networks to obtain the final embedding .
3. The deep learning-based multi-variate time series classification method according to claim 2, wherein, the process of step S3 is as follows: The classification network is composed of two fully connected layers. Between the two fully connected layers, the non-linear activation function ReLu is used to obtain non-linear features, and a batch normalization layer is used to stabilize the training and prevent overfitting. This classification network can be expressed as: 。 4. The deep learning-based multi-variate time series classification method according to claim 3, wherein, the process of step S4 is as follows: In the first step, first use supervised training CenterLoss to make the inter-class distance of the same-class samples smaller. While CenterLoss keeps the features of different classes separable, it minimizes the intra-class variation. The formula is as follows: Denote the feature center of the s-th class. Since it is necessary to consider the entire training set and average the features of each class in each iteration, the efficiency is extremely low. Therefore, in each iteration, instead of updating the center for the entire training set, the update is performed based on mini-batch training, and the center is calculated by averaging the features of the corresponding class; And in order to improve the training effect of the subsequent classification network, at this stage, the pre-training of the classification network is also completed. At this stage, the joint supervision of BCELoss and CenterLoss is used to train the network for feature learning. The formula for the first step is as follows: After the same-class samples are aggregated, in the second step, TripletLoss is used to improve the generated features to increase the difference between different classes, thereby improving the quality of the generated features and reducing the classification burden on the classification network. Since in the first step of training, CenterLoss has already made the same-class samples closer, there is no need to select multiple positive and negative samples as in the related work of ShapeNet. Select the same-class sample farthest from the target sample as the positive sample and the different-class sample closest to the target sample as the negative sample. The formula for the second step of training is as follows: After the final effective features are generated, in the third step, BCELoss is used alone to formally train the classification network module to improve the performance of the classification network.
5. The deep learning-based multi-variate time series classification method according to claim 4, wherein, The classification network consists of 2 fully connected layers, 1 activation function layer, 1 batch normalization layer, and 1 sigmoid classification function layer to finally convert the features into probability values between 0 and 1.