A nanopore protein sequencing data processing method based on machine learning and its application

By applying machine learning methods in nanopore protein sequencing, combining DTW and K-Means algorithms to judge the directionality of protein single molecules, and using convolutional neural networks and recurrent neural networks to identify the identities of different molecules, the problem of directionality judgment and molecular identity recognition in nanopore protein sequencing is solved, and efficient and accurate sequencing results are achieved.

CN116741265BActive Publication Date: 2025-05-16CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310705437.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-14
Publication Date
2025-05-16
Estimated Expiration
2043-06-14

AI Technical Summary

Technical Problem

In nanopore protein sequencing, the lack of effective signal processing methods makes it difficult to judge the directionality of the single molecule of labelless protein when transmembrane translocations within the pores, and it is difficult to accurately identify the identities of different molecules.

Method used

Machine learning-based method is used to combine dynamic time regularization algorithm (DTW) and K-Means clustering algorithm to judge the directionality of protein single molecules, and the identities of different molecules are identified through the classification algorithm of convolutional neural network and recurrent neural network.

Benefits of technology

The directional judgment of single molecules of label-free proteins and the identification of different molecules are achieved, the accuracy and reliability of protein sequencing data are improved, and the high robustness and generalization ability are high.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116741265B_ABST
    Figure CN116741265B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of nanopore protein sequencing, and specifically discloses a nanopore protein sequencing data processing method based on machine learning and an application thereof; the present invention discloses the use of machine learning for processing and analysis of nanopore sequencing data, constructs a clustering algorithm based on a dynamic time warping algorithm (Dynamic Time Warping, DTW) plus K-Means, and a classification algorithm based on a convolutional neural network and a recurrent neural network, which solves the problem of self-directional judgment of unlabeled molecules when they are translocated through a nanopore in nanopore protein de novo sequencing experimental data, and the problem of identity recognition of different molecules in the translocation process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of nanopore protein sequencing, and specifically discloses a nanopore protein sequencing data processing method based on machine learning and an application thereof. Background Art

[0002] Protein analysis provides key information for most biological processes, cell phenotypes, and diseases. Studies have shown that many diseases are caused by protein dysfunction, and the primary structure of a protein, i.e., the order of amino acids (AA), determines its higher-order structure and ultimately the function of the protein. Therefore, de novo protein sequencing is one of the core technologies that promote the further development of proteomics.

[0003] Since nanopore sensors have the advantages of single-molecule sensitivity, long read length, and high throughput, in addition to being successfully applied in DNA sequencing (such as MinION from Oxford Nanopore Technologies in the UK and QNome-3841 from Qi Carbon Technology in China), they also show great potential in protein analysis and protein sequencing. Studies have shown that the primary structure of proteins can be read using solid-state nanopores (G.Timp et al. 2016, Nature Nanotechnology 11,968–976), and it has been proven that 13 of the 20 protein amino acids can be distinguished using aerolysin nanopores based on the characteristics of their respective current blockade signals, and that the level of residual current is closely related to the size of the detected amino acid (A.Oukhaled et al. 2020, Nature Nanotechnology 38,176–181). The latest research results in 2021 show that the application of nanopores can achieve repeated reading of single peptides, thereby distinguishing amino acid substitutions at a single residue accuracy (C. Dekker et al. 2021, Science 374 (6574): 1509-1513). The amino acid sequence information of a single protein molecule that is translocated across the membrane through a nanopore is hidden in the corresponding blocking current fluctuations. Therefore, after implementing scientific and effective signal processing methods, we can read the primary structure of the protein by obtaining the subtle residual level fluctuation pattern in the blocking current signal.

[0004] The essence of protein sequencing is to read the primary structure of proteins, that is, to measure the order of amino acids on the peptide chain. Compared with gene sequencing, protein sequencing lacks a reference genome, so there is a lack of reference benchmarks in the data analysis process; the volume of amino acids that make up the peptide chain unit is small, only about one-tenth of the nucleotide, and there is a lot of noise interference, which will lead to a low signal-to-noise ratio, thus affecting the accuracy and reliability of the data; and there are as many as 20 amino acids, far more than the four bases that make up the DNA chain. The increase in sequence diversity will cause the complexity of the signal to increase exponentially, thereby greatly increasing the difficulty of signal processing and analysis, so that the existing data processing and analysis methods for nanopore DNA sequencing are difficult to apply to the field of protein sequencing. Traditional proteomic experimental analysis methods, such as mass spectrometry, are difficult to distinguish different analytes at the resolution of a single amino acid. At present, proteomics research urgently needs a protein sequencing technology with single-site residue specificity. It is worth noting that nanopore single-molecule sequencing technology has been successful in the field of genomics and has begun to advance into proteomics. However, there is still a lack of research on the problem of determining the self-direction of unlabeled protein molecules when they translocate through nanopores, that is, how to accurately determine whether each polypeptide chain is amino-terminal first or carboxyl-terminal first when it translocates across the membrane in the pore after being captured by the electric field. This is a prerequisite for the nanopore to accurately read the primary structure of the protein. Summary of the invention

[0005] To solve the above problems, the present invention discloses a nanopore protein sequencing data processing method based on machine learning and its application, which solves the problem of judging the directionality of a single unlabeled protein molecule when it translocates across the membrane in the pore, improves the calculation of the consensus current trajectory, and solves the problem of identifying the identities of different molecules during the translocation process.

[0006] To achieve the above object, the present invention adopts the following technical solution:

[0007] A method for processing nanopore protein sequencing data based on machine learning, comprising the following steps:

[0008] 1) Use the dynamic time warping algorithm DTW and the K-Means clustering algorithm to determine the directionality of the unlabeled molecules when they translocate through the nanopore;

[0009] 2) Use classification algorithms based on convolutional neural networks and recurrent neural networks to determine the identities of different molecules in the translocation process;

[0010] The DTW is a nonlinear warping technique for comparing the similarity of two time series, by aligning the coordinates in the two time series to find the optimal alignment path between them;

[0011] The K-means algorithm is an unsupervised learning algorithm that measures the similarity between samples by Euclidean distance and assigns samples with high similarity to the same category through iterative optimization.

[0012] The label-free molecule is a linear protein polypeptide molecule.

[0013] Further, in the above-mentioned method for processing nanopore protein sequencing data based on machine learning, the step 1) comprises the following steps:

[0014] First, the DTW distance is used to measure the similarity between time series and the distance matrix of the corresponding data set is constructed;

[0015] Then, the K-Means clustering principle was applied, the cluster category number K was set to 2, and the linear amino acid crystallographic volume model corresponding to the primary structure of the protein to be tested and its sequence after flipping along the time axis were used as the two initial cluster centers. The distance between the time domain current sequence in each blocking event and the two cluster centers was calculated, and it was assigned to the cluster center closest to it.

[0016] After all events are assigned, the means of all events in the two clusters are calculated respectively, and the cluster centers are updated. The above process is then repeated to continuously update and iterate the cluster centers to improve the accuracy of the clustering results. The iteration is stopped when the cluster centers no longer change.

[0017] Finally, the clustering results are analyzed, the calculation method of the consensus current trajectory is improved, and the clustering results are evaluated based on similarity indicators, including but not limited to the Pearson correlation coefficient PCC.

[0018] Furthermore, in the above-mentioned method for processing nanopore protein sequencing data based on machine learning, the step 1) comprises the following steps:

[0019] a. Collect and obtain the blockade current time series data generated by the translocation of unlabeled molecules in the nanopore;

[0020] b. Preprocess the data by using Matlab's built-in function interp1 to uniformly interpolate current blocking events of varying lengths to 500 points, and perform Z-Score normalization to eliminate differences in data dimensions;

[0021] c. Use K-Means clustering algorithm to process the data set. According to the data characteristics, set the cluster category K value to 2, select the DTW algorithm that can be scaled on the time axis as the similarity measure between current time series, and define two initial cluster centers, which are the volume model corresponding to the primary structure of the unlabeled molecule and the volume model after flipping along the time axis. Add the two volume models to the data set and participate in the iterative process.

[0022] d. In the clustering process, the distance between each event and the two cluster centers is calculated respectively, and the event is assigned to the cluster center with the closest distance. The mean of the two clusters is calculated based on the newly assigned result to replace the original cluster center. The above process is repeated until the cluster center is no longer updated. The iteration is stopped and the current cluster center and the assignment result of each event in the data set are output;

[0023] e. All events in the data set are assigned to two clusters, and the volume models before and after flipping are assigned to two clusters respectively. The PCC value between the two cluster centers is calculated. The correctness of the clustering result is judged by comparing whether the two cluster centers have a correlation of medium strength or above before and after flipping, that is, |PCC| ≥ 0.3; (For example, the PCC value between the two cluster centers is -0.6, that is, the two are negatively correlated, and after flipping one of the cluster centers, the PCC value between it and the other cluster center is 0.58, that is, the two are positively correlated, and the correctness of the clustering result can be preliminarily judged. Generally speaking, if the absolute value of the above two PCC values ​​is ≥ 0.3, it can be said that there is a medium-strength correlation between the two variables, and the correctness of the clustering result can be preliminarily judged.)

[0024] f. Determine the event directionality based on the clustering results, select several typical events for averaging, obtain the consensus current trajectory of the unlabeled molecule, and determine the PCC value between it and the volume model of the unlabeled molecule;

[0025] g. Choose a suitable W value so that after aligning the consensus current trajectory to the amino acid volume model in DTW, the PCC value between the consensus current trajectory and the volume model reaches the highest value.

[0026] Furthermore, in the above-mentioned method for processing nanopore protein sequencing data based on machine learning, the label-free molecule is β-amyloid protein Native Aβ 1-42 .

[0027] Furthermore, in the above-mentioned method for processing nanopore protein sequencing data based on machine learning, the label-free molecule is Scrambled Aβ 1-42 .

[0028] Furthermore, in the above-mentioned method for processing nanopore protein sequencing data based on machine learning, in step 2), an encoder-decoder framework is used, and the encoder part first uses a convolutional neural network to perform multiple convolution operations on the input current signals of different categories to extract the spatial features of the current signals, and then uses a recurrent neural network to learn its temporal features to capture the temporal correlation in the sequence;

[0029] The decoder part uses a multi-layer perceptron to perform nonlinear dimensionality reduction on the signal to predict the final current signal category. By continuously iteratively comparing the differences between the predicted results and the true labels, the parameter values ​​of each intermediate layer of the neural network are reversely updated to improve the accuracy of the prediction results.

[0030] Furthermore, in the above-mentioned method for processing nanopore protein sequencing data based on machine learning, in step 2), the encoder-decoder method is used to identify two different label-free molecules, comprising the following steps:

[0031] I. The encoder part includes a convolutional neural network and a recurrent neural network. The convolutional neural network includes two convolutional layers. Each convolutional layer contains 256 convolution kernels. Each convolution kernel performs a convolution operation on the input current signal, multiplies the elements in the current sequence with the weights in the convolution kernel and sums them up to extract the frequency components and spatial features of the sequence. Different numbers of convolution kernels are used to extract features at different levels of the sequence to obtain the final output feature sequence of the convolutional layer.

[0032] II. Use the recurrent neural network in the encoder to learn the time-step features of the current signal. While processing the input sequence, remember the previously processed information and model the sequence data by taking the hidden state of the previous moment as the input of the current moment, thereby capturing the time correlation in the sequence.

[0033] III. In the decoder part, a multi-layer perceptron is used to map the feature sequence output by the encoder to the final current signal category. After the output sequence representation is normalized by the sigmoid function, the probability of the final category of the current signal is obtained;

[0034] IV. Update the weights of each layer of the model by comparing the error between the model output and the true label so that the model can more accurately predict the category belonging probability.

[0035] Furthermore, in the above-mentioned nanopore protein sequencing data processing method based on machine learning, the recurrent neural network structure in step II is a long short-term memory LSTM structure, and the LSTM structure includes three gating mechanisms: an input gate, a forget gate, and an output gate, which respectively control the input, output, and forgetting of previous information in the current sequence, thereby realizing effective long-term information storage and control; each LSTM gate contains 64 neurons, and each neuron learns different feature representations of the signal.

[0036] Furthermore, in the above-mentioned nanopore protein sequencing data processing method based on machine learning, the multilayer perceptron in step III is composed of three fully connected layers, including two hidden layers and one output layer, and the two hidden layers contain 100 and 50 neurons respectively. Each neuron is connected to all neurons in the previous layer, and each layer combines and abstracts different features through nonlinear transformation to generate a low-dimensional representation of the sequence. Finally, after the transformation of the hidden layer, the signal reaches the output layer, and the output sequence representation is normalized by the sigmoid function to obtain the probability of the final category attribution of the current signal.

[0037] On the other hand, the present invention discloses the application of the above-mentioned nanopore protein sequencing data processing method based on machine learning in large-scale protein sequencing.

[0038] The present invention has the following beneficial effects:

[0039] Native Aβ 1-42 and Scrambled Aβ 1-42 When processing nanopore sequencing data, before calculating the consensus current trajectory composed of hundreds of single-molecule transmembrane translocation events, the problem that needs to be solved is how to determine the self-directivity of the unlabeled protein single molecule during translocation. In the present invention, based on the data characteristics, that is, there are only two possible directions, the K-Means clustering algorithm with a K value of 2 is selected, combined with DTW as a distance metric, and the cluster center initialization method is changed. The volume model corresponding to the primary structure of the single molecule of the protein to be tested and the volume model after flipping along the time axis are used as two initialized clustering centers, added to the data set, and the directionality of the unlabeled time series is judged, and the data set is finally divided into two clusters. After obtaining the directionality of each translocation event based on the clustering results, the consensus current trajectory of the protein to be tested is calculated, and it is found that it is highly correlated with the corresponding linear amino acid volume model, and the PCC value is as high as 0.8.

[0040] By using a neural network-based classification algorithm to classify different target protein sequencing signals, efficient and accurate classification results were achieved, and the classification AUC value (AUC∈[0.5 1]) reached 0.97±0.01. (AUC measures the model's ability to distinguish between positive and negative samples at different thresholds, and can be used to evaluate the overall performance of the model; while accuracy refers to the ratio of the number of samples correctly classified by the classification model to the total number of samples. AUC is more suitable for the case of unbalanced samples, and can more comprehensively evaluate the performance of the classification model. The value range of AUC is between 0 and 1. The closer the value is to 1, the better the classification performance of the model is, and the more accurately it can distinguish between positive and negative samples.) Compared with traditional methods, this algorithm has a strong diversity classification ability and can accurately distinguish the differences and similarities between different protein sequencing signals. It can capture the abstract features in the signal and make full use of the sequence information and structural features of different proteins to achieve accurate classification. At the same time, it has high robustness and generalization ability, can handle protein sequencing signals of various types and lengths, and has good adaptability and promotion ability. In addition, through efficient algorithm design and continuous optimization, it achieves a faster classification speed while ensuring accuracy. It can classify a large number of protein sequencing signals in a shorter time, improving processing efficiency and production benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 : Flowchart for judging the directionality of single-molecule unlabeled protein translocation through nanopores based on DTW and K-Means clustering algorithms. (a) Original time current series of different lengths generated when a single-molecule protein translocates through a nanopore. (b) Each blocking event is interpolated to 500 points and Z-Score processed. (c) Settings of the K-Means clustering algorithm, setting the K value to 2, initializing the cluster center to the volume model before and after flipping, the distance metric is DTW, A: blocking event, B: cluster center, W: the size of the DTW warp window. (d) Clustering results, all events in the data set are divided into two clusters. Gray: blocking event, black: cluster center;

[0042] Figure 2 : (a) and (d) are Native Aβ 1-42 Datasets and Scrambled Aβ 1-42Characteristic distribution of events in the dataset in terms of blockade duration and current blockade ratio. Each square in the figure represents a blockade event. Black squares: events with blockade duration between 100-500μs. Red squares: typical blockade events. Contour: Black contour lines, formed by connecting those points with 50%×PDFmax value in the normalized heat map of the probability density function (PDF) of the dataset. (b) Red solid line: Native Aβ averaged from 50 typical events in Figure (a) 1-42 Consensus current traces. Black solid line: Native Aβ 1-42 Linear amino acid volume model (kmer = 1). Gray solid line: The correspondence between the two sequences when the two sequences are aligned using the DTW algorithm. (c) Red solid line, when W = 12, the Native consensus current trajectory is aligned to the Native Aβ using the DTW method 1-42 (e) Red solid line: Scrambled Aβ averaged from 50 typical events in Figure (d) 1-42 Consensus current trace. Black solid line: Scrambled Aβ 1-42 Linear amino acid volume model (kmer = 1). Gray solid line: The correspondence between the two sequences when the two sequences are aligned using the DTW algorithm. (f) Red solid line, when W = 12, the DTW method is used to align the Scrambled consensus current trajectory to Scrambled Aβ 1-42 The results after alignment with the linear amino acid volume model (kmer=1).

[0043] Figure 3 The classification algorithm based on convolutional neural network and recurrent neural network solves the problem of identity recognition of different translocated molecules.

[0044] (a) Framework diagram of classification algorithm based on neural network model (b) Receiver operating characteristic (ROC) curve of classification algorithm: a group of Scrambled Aβ 1-42 The data sets are respectively and three different groups of Native Aβ 1-42 The data set was input into the neural network model for training and evaluation, and the average AUC was 0.97. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present invention are described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] The reagents or instruments used in the examples of the present invention without indicating the manufacturer are all conventional reagent products that can be obtained through commercial purchase.

[0047] Example 1

[0048] Native Aβ 1-42 Determined by the nanopore's own directionality during translocation.

[0049] In obtaining β-amyloid protein (Native Aβ 1-42 ) Based on the blockade current time series dataset generated by translocation in the nanopore, Native Aβ 1-42 The process of blocking event direction is as follows Figure 1 As shown. First, the data was preprocessed. The interp1 function provided by Matlab was used to uniformly interpolate the current blocking events of different lengths to 500 points, and the Z-Score was normalized to eliminate the difference in data dimensions. After that, the K-Means clustering algorithm was used to process the data set. According to the data characteristics, the number of cluster categories K was set to 2, and the DTW algorithm that can be scalable on the time axis was selected as the similarity measure between current time series. The two customized initialization cluster centers were Native Aβ and 1-42 The volume model corresponding to the primary structure and the volume model after flipping along the time axis are calculated, and the two volume models are added to the data set to participate in the iteration process. During the clustering process, the distance between each event and the two cluster centers is calculated respectively, and the event is assigned to the cluster center with the nearest distance. The mean of the two clusters is calculated based on the newly assigned results to replace the original cluster center. The above process is repeated until the cluster center is no longer updated, and the iteration is stopped. The current cluster center and the allocation result of each event in the data set are output. Finally, all events in the data set are assigned to two clusters. The volume models before and after flipping are assigned to two clusters respectively, and the PCC value between the two cluster centers is -0.63, and the PCC value between cluster center 1 and cluster center 2 after flipping is 0.58, indicating that the clustering results are initially correct. Afterwards, the directionality of the event is judged based on the clustering results, and 50 typical events (such as Figure 2 (a)) and average to get Native Aβ 1-42 The consensus current traces, with Native Aβ 1-42The PCC value between the volume models (k-mer = 1) is 0.84, such as Figure 2 (b), the two are highly correlated. When W = 12, DTW is used to align the consensus current trajectory to the amino acid volume model, as shown in Figure 2 As shown in (c), the PCC value between the consensus current trajectory and the volume model is as high as 0.96.

[0050] Example 2

[0051] Scrambled Aβ 1-42 Determined by the nanopore's own directionality during translocation.

[0052] In obtaining Scrambled Aβ 1-42 Scrambled Aβ was identified based on the blockade current time series dataset generated by translocation in the nanopore 1-42 The process of blocking event direction is as follows Figure 1 As shown. First, the data was preprocessed. The current blocking events of different lengths were uniformly interpolated to 500 points using Matlab's built-in function interp1, and the Z-Score was normalized to eliminate the difference in data dimensions. Then, the K-Means clustering algorithm was used to process the data set. According to the data characteristics, the number of cluster categories K was set to 2, and the DTW algorithm that can be scalable on the time axis was selected as the similarity measure between current time series. The two customized initialization cluster centers were Scrambled Aβ and Scrambled Aβ. 1-42 The volume model corresponding to the primary structure and the volume model after flipping along the time axis are calculated, and the two volume models are added to the data set to participate in the iteration process. During the clustering process, the distance between each event and the two cluster centers is calculated respectively, and the event is assigned to the cluster center with the nearest distance. The mean of the two clusters is calculated based on the newly assigned results to replace the original cluster center. The above process is repeated until the cluster center is no longer updated, and the iteration is stopped. The current cluster center and the allocation result of each event in the data set are output. Finally, all events in the data set are assigned to two clusters. The volume models before and after flipping are assigned to two clusters respectively, and the PCC value between the two cluster centers is -0.65, and the PCC value between cluster center 1 and cluster center 2 after flipping is 0.56, indicating that the clustering results are initially correct. Afterwards, the directionality of the event is judged based on the clustering results, and 50 typical events (such as Figure 2 (d)) were averaged to obtain Scrambled Aβ 1-42 Consensus current traces, with Scrambled Aβ 1-42 The PCC value between the amino acid volume model (k-mer = 1) is 0.76, and the two are highly correlated, such as Figure 2(e) When W = 12, the consensus current trajectory is aligned to the amino acid volume model using DTW. Figure 2 As shown in (f), the PCC value between the consensus current trajectory and the volume model is as high as 0.96.

[0053] Example 3

[0054] Identification of the different molecules involved in the translocation.

[0055] For Native Aβ 1-42 and Scrambled Aβ 1-42 The blockade current time series data set generated by translocation in the nanopore, the present invention uses an encoder-decoder method to identify two different target proteins. The neural network model of the present invention is implemented by using the Python programming language and the TensorFlow.Keras framework.

[0056] Specifically, the convolutional neural network of the encoder part includes two convolutional layers. Each convolutional layer contains 256 convolution kernels, each of which performs convolution operation on the input current signal, multiplies the elements in the current sequence with the weights in the convolution kernel and sums them, extracts the frequency components, spatial features, etc. of the sequence, and extracts the features of different levels of the sequence through different numbers of convolution kernels to obtain the output feature sequence of the final convolutional layer. By stacking two convolutional layers, the expression ability and prediction accuracy of the model are improved. Next, the recurrent neural network is used to learn the time step characteristics of the current signal. The recurrent neural network can memorize the previously processed information while processing the input sequence, and model the sequence data by taking the hidden state of the previous moment as the input of the current moment, thereby capturing the time correlation in the sequence. The long short-term memory (LSTM) used in the present invention is a special recurrent neural network structure that can effectively avoid the gradient vanishing and gradient explosion problems when learning long sequences, and improve the learning ability and generalization ability of the model. The LSTM structure includes three gating mechanisms: input gate, forget gate and output gate. These gating mechanisms can control the input, output, and forgetting of previous information in the current sequence, thereby achieving effective long-term information storage and control. Each LSTM layer contains 64 neurons, each of which learns different feature representations of the signal. By stacking three LSTM layers, long-term dependencies in the sequence can be captured more effectively and the expressiveness of the model can be improved. The pooling layer downsamples the feature sequence output by the LSTM layer to reduce the dimension of the feature sequence. It aggregates the elements in the feature sequence, takes the average value of each feature channel in the sequence, and uses this value as the representative value to reduce the size of the feature sequence, reduce the number of model parameters and computational complexity, and improve the robustness of the model.

[0057] In the decoder part, a multilayer perceptron is used to map the feature sequence output by the encoder to the final current signal category. The multilayer perceptron consists of three fully connected layers, including two hidden layers and one output layer. Both hidden layers contain 100 and 50 neurons, and each neuron is connected to all neurons in the previous layer. Each layer combines and abstracts different features through nonlinear transformation to generate a low-dimensional representation of the sequence. Finally, after the transformation of the hidden layer, the signal reaches the output layer. After the output sequence representation is normalized by the sigmoid function, the probability of the final category of the current signal is obtained.

[0058] The weights of each layer of the model are updated by comparing the error between the model output and the true label, so that the model can more accurately predict the probability of category belonging. During the training process, the neural network model is updated for Native Aβ through 100 rounds of iterative weight updates. 1-42 and Scrambled Aβ 1-42 The classification effect reached AUC = 0.97 ± 0.01.

[0059] According to the above embodiments, the present invention selects the K-Means clustering algorithm with a K value of 2 based on the data characteristics, that is, there are only two possible directions, combines DTW as the distance metric, and changes the cluster center initialization method, using the volume model corresponding to the single-molecule primary structure of the protein to be tested and the volume model after flipping along the time axis as two initialized cluster centers, adding the data set, judging the directionality of the unlabeled time series, and finally dividing the data set into two clusters. After obtaining the directionality of each translocation event based on the clustering results, the consensus current trajectory of the protein to be tested is calculated, and it is found that it is highly correlated with the corresponding linear amino acid volume model, and the PCC value is as high as 0.8.

[0060] By using a neural network-based classification algorithm to classify different target protein sequencing signals, efficient and accurate classification results were achieved, and the classification AUC value reached 0.97±0.01. Compared with traditional methods, this algorithm has a strong diversity classification ability and can accurately distinguish the differences and similarities between different protein sequencing signals. It can capture the abstract features in the signal and make full use of the sequence information and structural features of different proteins to achieve accurate classification. At the same time, it has high robustness and generalization ability, can handle protein sequencing signals of various types and lengths, and has good adaptability and promotion ability. In addition, through efficient algorithm design and continuous optimization, a faster classification speed is achieved while ensuring accuracy. It can classify a large number of protein sequencing signals in a short time, improving processing efficiency and production benefits.

[0061] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form, and any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.

Claims

1. A method for processing nanopore protein sequencing data based on machine learning, characterized in that: The following steps are involved: 1) Use the dynamic time warping algorithm DTW and the K-Means clustering algorithm to determine the directionality of the unlabeled molecules when they translocate through the nanopore; 2) Use classification algorithms based on convolutional neural networks and recurrent neural networks to determine the identities of different molecules in the translocation process; The DTW is a nonlinear warping technique for comparing the similarity of two time series, by aligning the coordinates in the two time series to find the optimal alignment path between them; The K-means algorithm is an unsupervised learning algorithm that measures the similarity between samples by Euclidean distance and assigns samples with high similarity to the same category through iterative optimization. The label-free molecule is a linear protein polypeptide molecule; The step 1) comprises the following steps: First, the DTW distance is used to measure the similarity between time series and the distance matrix of the corresponding data set is constructed; Then, the K-Means clustering principle was applied, the cluster category number K was set to 2, and the linear amino acid crystallographic volume model corresponding to the primary structure of the protein to be tested and its sequence after flipping along the time axis were used as the two initial cluster centers. The distance between the time domain current sequence in each blocking event and the two cluster centers was calculated, and it was assigned to the cluster center closest to it. After all events are assigned, the means of all events in the two clusters are calculated respectively, and the cluster centers are updated. The above process is then repeated to continuously update and iterate the cluster centers to improve the accuracy of the clustering results. The iteration is stopped when the cluster centers no longer change. Finally, the clustering results are analyzed, the calculation method of the consensus current trajectory is improved, and the clustering results are evaluated according to similarity indicators, including but not limited to the Pearson correlation coefficient PCC∈[-1 1], and the experimental results of strong positive correlation are obtained, that is, PCC≥0.

5.

2. The method for processing nanopore protein sequencing data based on machine learning according to claim 1, characterized in that: The step 1) comprises the following steps: a. Collect and obtain the blockade current time series data generated by the translocation of unlabeled molecules in the nanopore; b. Preprocess the data by using Matlab's built-in function interp1 to uniformly interpolate current blocking events of varying lengths to 500 points, and perform Z-Score normalization to eliminate differences in data dimensions; c. Use K-Means clustering algorithm to process the data set. According to the data characteristics, set the cluster category K value to 2, select the DTW algorithm that can be scaled on the time axis as the similarity measure between current time series, and define two initial cluster centers, which are the volume model corresponding to the primary structure of the unlabeled molecule and the volume model after flipping along the time axis. Add the two volume models to the data set and participate in the iterative process. d. In the clustering process, the distance between each event and the two cluster centers is calculated respectively, and the event is assigned to the cluster center with the closest distance. The mean of the two clusters is calculated based on the newly assigned result to replace the original cluster center. The above process is repeated until the cluster center is no longer updated. The iteration is stopped and the current cluster center and the assignment result of each event in the data set are output; e. All events in the data set are assigned to two clusters, and the volume models before and after flipping are assigned to two clusters respectively. The PCC value between the two cluster centers is calculated. The correctness of the clustering result is judged by comparing whether the two cluster centers have a correlation of medium intensity or above before and after flipping, that is, |PCC|≥0.3; f. Determine the event directionality based on the clustering results, select several typical events for averaging, obtain the consensus current trajectory of the unlabeled molecule, and determine the PCC value between it and the volume model of the unlabeled molecule; g. Choose a suitable W value so that after aligning the consensus current trajectory to the amino acid volume model in DTW, the PCC value between the consensus current trajectory and the volume model reaches the highest value.

3. The method for processing nanopore protein sequencing data based on machine learning according to claim 1, characterized in that: The unlabeled molecule is β-amyloid protein Native Aβ 1-42 .

4. The method for processing nanopore protein sequencing data based on machine learning according to claim 1, characterized in that: The unlabeled molecule is Scrambled Aβ 1-42 .

5. The method for processing nanopore protein sequencing data based on machine learning according to claim 1, characterized in that: In the step 2), an encoder-decoder framework is used, where the encoder part first uses a convolutional neural network to perform multiple convolution operations on the input current signals of different categories to extract the spatial features of the current signals, and then uses a recurrent neural network to learn its temporal features to capture the temporal correlation in the sequence; The decoder part uses a multi-layer perceptron to perform nonlinear dimensionality reduction on the signal to predict the final current signal category. By continuously iteratively comparing the differences between the predicted results and the true labels, the parameter values ​​of each intermediate layer of the neural network are reversely updated to improve the accuracy of the prediction results.

6. The method for processing nanopore protein sequencing data based on machine learning according to claim 1, characterized in that: In step 2), the encoder-decoder method is used to identify two different unlabeled molecules, comprising the following steps: I. The encoder part includes a convolutional neural network and a recurrent neural network. The convolutional neural network includes two convolutional layers. Each convolutional layer contains 256 convolution kernels. Each convolution kernel performs a convolution operation on the input current signal, multiplies the elements in the current sequence with the weights in the convolution kernel and sums them up to extract the frequency components and spatial features of the sequence. Different numbers of convolution kernels are used to extract features at different levels of the sequence to obtain the final output feature sequence of the convolutional layer. II. Use the recurrent neural network in the encoder to learn the time-step features of the current signal. While processing the input sequence, remember the previously processed information and model the sequence data by taking the hidden state of the previous moment as the input of the current moment, thereby capturing the time correlation in the sequence. III. In the decoder part, a multi-layer perceptron is used to map the feature sequence output by the encoder to the final current signal category. After the output sequence representation is normalized by the sigmoid function, the probability of the final category of the current signal is obtained; IV. Update the weights of each layer of the model by comparing the error between the model output and the true label so that the model can more accurately predict the category belonging probability.

7. The method for processing nanopore protein sequencing data based on machine learning according to claim 6, characterized in that: The recurrent neural network structure in step II is a long short-term memory LSTM structure. The LSTM structure includes three gating mechanisms: an input gate, a forget gate, and an output gate, which respectively control the input, output, and forgetting of previous information in the current sequence, thereby achieving effective long-term information storage and control; each LSTM gate contains 64 neurons, and each neuron learns different feature representations of the signal.

8. The method for processing nanopore protein sequencing data based on machine learning according to claim 6, characterized in that: The multilayer perceptron in step III is composed of three fully connected layers, including two hidden layers and one output layer. The two hidden layers contain 100 and 50 neurons respectively. Each neuron is connected to all neurons in the previous layer. Each layer combines and abstracts different features through nonlinear transformation to generate a low-dimensional representation of the sequence. Finally, after the transformation of the hidden layer, the signal reaches the output layer. After the output sequence representation is normalized by the sigmoid function, the probability of the final category attribution of the current signal is obtained.

9. Application of the nanopore protein sequencing data processing method based on machine learning according to any one of claims 1 to 8 in large-scale protein sequencing.

Citation Information

Patent Citations

  • Nanopore sequencing data base identification method based on deep learning

    CN113870949A

  • Continuous wavelet-based dynamic time warping method and system

    US20200035325A1