Construction and application of a polypeptide classifier
Patent Information
- Application Number
- CN202380104751.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2026-07-17
AI Technical Summary
The prior art is difficult to accurately classify more categories of polypeptide nanopore sequencing signals, resulting in limited classification accuracy and category types.
By establishing a polypeptide classification model, model training is performed using the feature set of known polypeptides, the feature set of the polypeptides to be tested is detected and input into the model to obtain the category of the polypeptides to be tested. The feature extraction step includes extracting 13 feature values from the sequencing signal and the spectrum signal to form a feature set.
The classification accuracy of the six polypeptides was achieved to reach 95.52%, 14% higher than the prior art, and the type of polypeptide that the classifier should be distinguished was increased.
Smart Images

Figure CN122422941A_ABST
Abstract
Description
Construction and application of a peptide classifier Technical Field
[0001] The present disclosure relates to a method for detecting polypeptides, and more particularly to extracting effective features from polypeptide signal fragments to achieve classification of different polypeptides. Background Art
[0002] The function or dysfunction of proteins depends largely on their primary structure, specifically their amino acid sequence. Identifying and sequencing proteins can provide crucial information for disease diagnosis and treatment. Nanopore detection technology is a technique that enables the transport of molecules through ion channels within nanopores. Nanopores provide nanoscale confinement, allowing peptide molecules to interact with the nanopore sensing interface, thereby causing changes in the ion current flowing through them. By detecting changes in the ion current induced by peptides, the peptides passing through can be analyzed. Classifying these peptides is crucial for their identification and detection. The continuous advancement of nanopore technology has significantly improved nanopore fabrication, control, and signal detection, laying the foundation for signal classification in peptide nanopore sequencing. Signal classification in peptide nanopore sequencing requires complex signal data analysis and processing, supported by advances in artificial intelligence. Machine learning and deep learning technologies enable intelligent analysis and processing of peptide nanopore sequencing signal data, improving the accuracy and efficiency of signal classification and accelerating the detection of peptide classes.
[0003] One existing peptide detection method relies on the direct translocation of a single peptide molecule under the influence of an electric field. Because different peptides require varying amounts of translocation time, the magnitude of the resulting current change also varies. Therefore, simple peptide classification often relies on analyzing the translocation time and blockade current (DOI: 10.1002 / celc.201800288; 10.1038 / s41587-019-0345-2). To increase detection resolution, Luning Yu et al. added ten negatively charged asparagine residues to the end of the protein and denatured the full-length protein with guanidine hydrochloride, resulting in a relatively long pore translocation time and thus providing more detection information (DOI: 0.1038 / s41587-022-01598-3). In this approach, seven features, including the mean, median, and maximum current values, were extracted from the peptide nanopore sequencing signal. A gradient boosting decision tree was then used to classify the peptides based on these features, resulting in the classification of three proteins. Patent WO2021111125A1 provides a sequencing solution based on helicase rate control, which ensures that the polypeptide has a longer translocation time, which is beneficial for obtaining richer sequencing signals. However, this method has not yet reported how to classify its polypeptide signals.
[0004] There are still many defects in the prior art. For example, in the case of direct translocation of polypeptides, the translocation efficiency is affected by the electrical properties of the polypeptide, and neutral polypeptides are more difficult to achieve translocation. For charged polypeptides, due to their excessive speed and short residence time in the pore, only very little information can be obtained, such as translocation time, blocking current, etc. Therefore, the classification of polypeptides can only be based on these two signals, and the accuracy and category types of the classification are limited. For the method of classification based on current characteristic values reported by Luning Yu et al. (10.1038 / s41587-022-01598-3), the classification object is the overall translocation signal of the full-length protein. The data set used is the nanopore signal of three polypeptides at most, and the accuracy of its classifier is about 80.8%. It has fewer classification categories and a lower accuracy rate.
[0005] Summary of the Invention
[0006] The technical problem to be solved by the present invention is how to more accurately classify nanopore sequencing signals of more types of polypeptides.
[0007] In the present invention, a polypeptide classification model is first established using a feature set of known polypeptides and their corresponding categories, and then the feature set of the polypeptide to be tested is detected and input into the polypeptide classification model, thereby obtaining the category of the polypeptide to be tested.
[0008] Specifically, the present invention provides the following aspects:
[0009] 1. A method for polypeptide classification, comprising the following steps:
[0010] A sequencing step, which includes sequencing the polypeptide to be tested or a linker containing the polypeptide to be tested, obtaining a sequencing signal of the polypeptide, and obtaining a spectrum signal of the polypeptide from the sequencing signal;
[0011] a feature extraction step, comprising extracting a plurality of feature values from the sequencing signal and the spectrum signal respectively through a feature set function to form a feature set;
[0012] In the prediction step, the feature set is input into a polypeptide classification model for classification, and the category with the highest frequency of output is the category of the polypeptide to be tested.
[0013] 2. The method according to item 1, wherein the polypeptide classification model is determined based on a feature set of known polypeptides;
[0014] Optionally, the training process of the polypeptide classification model includes:
[0015] Performing the sequencing and feature extraction steps on the polypeptides of known categories to obtain a feature set of the polypeptides of known categories;
[0016] The feature set of polypeptides of known categories and their corresponding polypeptide categories are used for model training to obtain the polypeptide classification model.
[0017] 3. The method according to item 1 or 2, wherein the feature values extracted in the feature extraction step include: the standard error of the linear trend slope of the sequencing signal and spectrum signal blocks, the second coefficient of the 10th-order autoregressive model, the change of the sequencing signal and spectrum signal from the 0 quantile to the 0.4 quantile, the change of the sequencing signal and spectrum signal from the 0.4 quantile to the 0.8 quantile, the change of the sequencing signal and spectrum signal from the 0.6 quantile to the 0.8 quantile, the coefficient of the continuous wavelet transform, the first coefficient angle of the FFT, the quality quantile of the index of the sequencing signal and spectrum signal, the permutation entropy of the sequencing signal and spectrum signal, the 0.3 quantile, skewness, and variance of the sequencing signal and spectrum signal.
[0018] 4. The method of item 1 or 2, wherein the method further comprises a ligation step before the sequencing step, wherein the ligation step comprises ligating the polypeptide to be tested or the polypeptide of the known type to a non-polypeptide portion (NPP) to obtain a linker, wherein the structure of the linker is NPP1-polypeptide-NPP2;
[0019] Optionally, wherein the non-polypeptide portion (NPP) comprises a nucleic acid portion and / or a non-basic portion, for example, the non-basic portion may be a non-basic spacer; optionally, the NPP may comprise the same or different nucleic acids, the same or different non-basic portions, or a combination of nucleic acids and non-basic portions;
[0020] Optionally, the NNP1 and NNP2 are nucleic acid 1 and nucleic acid 2, respectively. Preferably, the sequences of nucleic acid 1 and nucleic acid 2 are SEQ ID NOs: 5 and 6, respectively.
[0021] 5. The method according to item 4, wherein the sequencing step comprises:
[0022] Obtaining a nanopore sequencing signal of the connector NPP1-polypeptide-NPP2 by nanopore sequencing technology;
[0023] intercepting the sequencing signal of the polypeptide from the nanopore sequencing signal, for example, intercepting the sequencing signal of the polypeptide using a YOLO recognition model;
[0024] A spectrum signal is extracted from the intercepted polypeptide sequencing signal, for example, by extracting the spectrum signal through a fast Fourier transform function.
[0025] 6. The method according to any one of items 1-5, wherein the polypeptide classification model is a decision tree classification model, such as a random forest model, a GBDT model, or an Adaboost model.
[0026] 7. A system for polypeptide classification, comprising:
[0027] (1) a sequencing module, which is used to obtain a sequencing signal sequence and a spectrum signal sequence of a polypeptide;
[0028] (2) a feature extraction module, which is used to extract multiple feature values from the sequencing signal sequence and the spectrum signal sequence to form a feature set;
[0029] (3) A prediction module, which is used to input the feature set into a polypeptide classification model for processing to obtain the category of the polypeptide to be tested.
[0030] 8. A computer program product comprising instructions which, when executed by a processor, cause the processor to perform the method according to any one of items 1-6.
[0031] 9. A computer-readable storage medium storing instructions which, when executed by a processor, cause the processor to perform the method according to any one of items 1-6.
[0032] 10. A system comprising:
[0033] processor;
[0034] A memory having instructions stored thereon, which, when executed by the processor, cause the processor to perform the method according to any one of items 1-6.
[0035] 11. A kit for polypeptide classification, comprising reagents for performing the method according to any one of items 1 to 6, and / or the system according to item 10.
[0036] 12. The kit of claim 11, comprising reagents for sequencing and feature extraction, optionally comprising a polypeptide portion and a non-polypeptide portion (NPP), and optionally comprising instructions for use. Beneficial effects
[0037] In an example application, for classifying six peptides, the peptide classification scheme implemented by the present invention achieved an accuracy of 95.52%. This is 14% higher than the accuracy of the existing classification scheme proposed in the article doi:10.1038 / s41587-022-01598-3. The classifier also distinguished three more peptide types than the example described in that article. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1: Mass spectrometry characterization results of OPO linkers, where (a) is DNA1-peptide1-DNA2, (b) is DNA1-peptide2-DNA2, (c) is DNA1-peptide3-DNA2, (d) is DNA1-peptide4-DNA2, (e) is DNA1-peptide5-DNA2, and (f) is DNA1-peptide6-DNA2.
[0039] Figure 2: Nanopore sequencing signals for OPO adapters, where (a) is DNA1-peptide1-DNA2, (b) is DNA1-peptide2-DNA2, (c) is DNA1-peptide3-DNA2, (d) is DNA1-peptide4-DNA2, (e) is DNA1-peptide5-DNA2, and (f) is DNA1-peptide6-DNA2. Sequencing signals for peptide fragments are shown within the dashed box.
[0040] Figure 3: Data generation process.
[0041] Figure 4: Peptide classifier training process.
[0042] Figure 5: Peptide class prediction pipeline.
[0043] Figure 6: Classification results of the classifier for six peptides. DETAILED DESCRIPTION
[0044] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0045] (I) Polypeptide Classification Method
[0046] In the present invention, unless otherwise specified or limited by the context, the terms peptide, polypeptide, and protein used are not particularly limited and can be replaced with each other under appropriate circumstances. For example, any one of the peptides, polypeptides, or proteins mentioned alone may include more than 2 amino acid residues, such as about 5-9 amino acid residues (peptide), 10-100 amino acid residues (polypeptide), or more than 100 amino acid residues (protein).
[0047] In the present invention, the term "polypeptide classification method" refers to a method for classifying polypeptides based on their amino acid sequence; that is, polypeptides with the same amino acid sequence are classified into the same class, and polypeptides with different amino acid sequences are classified into different classes. In specific embodiments of the present invention, the polypeptides of known class used in model training refer to polypeptides whose amino acid sequences are known.
[0048] In a specific embodiment of the present invention, the polypeptide classification method disclosed herein first includes a scheme for model training, which mainly includes: obtaining a spectral signal from the sequencing signal of the polypeptide fragment, and extracting 13 eigenvalues from the sequencing signal and the spectral signal, respectively, for a total of 26 eigenvalues, and performing model training on the eigenvalue set and the category corresponding to the polypeptide to obtain a polypeptide classification model.
[0049] The sequencing signal of the polypeptide fragment can be obtained by any sequencing technology known in the art, as long as the current sequencing signal of the polypeptide can be obtained, for example, by nanopore sequencing technology.
[0050] In a specific embodiment of the present invention, model training may include the following steps (1) to (9).
[0051] (1) Use non-polypeptide parts (NPP1 and NPP2) to connect the N-terminus and C-terminus of a polypeptide fragment to form an "NPP1-polypeptide-NPP2" connector (wherein, in the model training scheme, the category of the polypeptide fragment is known).
[0052] In specific embodiments of the present invention, in order to address issues such as uneven charge distribution and transmembrane movement of polypeptides, it is contemplated to link a non-polypeptide portion (NPP) to the C-terminus and / or N-terminus of the polypeptide. Therefore, the term non-polypeptide portion (NPP) as used herein is not particularly limited and may include any portion that is linked to the polypeptide of interest in order to facilitate polypeptide sequencing and / or generate a signal distinct from the polypeptide sequencing signal, thereby facilitating the identification and extraction of polypeptide signals. For example, such an NPP may include a nucleic acid (oligonucleotide or polynucleotide), a non-basic portion such as a non-basic spacer (AP site) or a non-basic linker, or any other suitable compound.
[0053] The sequencing signal obtained from the NPP-containing linker can include a mixed sequencing signal of a polypeptide signal and a non-polypeptide signal (NPS). Therefore, the term non-polypeptide signal as used herein includes, for example, a signal generated from the NPP portion during sequencing.
[0054] As used herein, the NPP can be linked to the polypeptide of interest (the polypeptide to be tested) in any suitable manner, for example, via a chemical bond. In some embodiments, the NPP can be any structural moiety capable of generating a signal distinct from the polypeptide sequencing signal, thereby facilitating the identification and extraction of the polypeptide signal. Therefore, in some embodiments, any suitable NPP can be selected based on whether the generated NPS signal is distinguishable from the polypeptide sequencing signal.
[0055] In some embodiments, the polypeptide of interest (test molecule) may include two or more non-polypeptide moieties (NPPs), wherein the two or more NPPs may be the same or different. For example, the two or more NPPs may include the same or different nucleic acids, the same or different non-basic moieties, or a combination of nucleic acids and non-basic moieties. In some embodiments, the two or more NPPs may include a combination of one or more nucleic acids and one or more non-basic moieties, wherein the nucleic acids and non-basic moieties may be the same or different, respectively. In some embodiments, the NPPs may be located on one or both sides of the polypeptide moiety. In some embodiments, the test molecule may include NPP1-polypeptide-NPP2, for example, the test molecule may include nucleic acid 1-polypeptide-nucleic acid 2, wherein nucleic acid 1 and nucleic acid 2 may be the same or different. In some embodiments, the NPP may change or remain unchanged during the sequencing process, preferably remaining unchanged (i.e., the NPPs attached to different polypeptide sequences are the same), thereby maintaining the NPS generated therefrom. In an exemplary embodiment of the present invention, the sequences of nucleic acid 1 and nucleic acid 2 are SEQ ID NOs: 5 and 6, respectively.
[0056] (2) Obtain k nanopore sequencing signals (e.g., current signals) of the connector described in (1) by sequencing technology such as nanopore sequencing technology. For example, a line graph of each sequencing signal can be drawn with time as the x-axis and signal as the y-axis.
[0057] Among them, the nanopore technology is a simple and efficient single-molecule detection technology that has been widely studied in DNA sequencing. By applying voltage, negatively charged DNA translocates through the nanopore, interacts with the pore, causes current changes, and sequence information is obtained by analyzing the current signal. Similarly, proteins can also be analyzed using this method. However, unlike nucleic acids, proteins are generally not uniformly charged, so that the probability of different electrical polypeptides translocating through the nanopore under the action of the electric field force is different. For nanopore sequencing of polypeptides, the polypeptide can be formed into a composite connector with NPP, and then passed through the nanopore under the control of the helicase or polymerase. The electrical signal obtained contains mixed sequencing signals of both the polypeptide and NPP.
[0058] (3) Using the YOLO recognition model, the sequencing signal of the polypeptide fragment is intercepted from the k nanopore sequencing signals described in (2).
[0059] In a specific embodiment, the polypeptide signal extraction method of this step may include a data generation process, a model fine-tuning process, and a polypeptide signal extraction process.
[0060] In a specific embodiment, the method for extracting the polypeptide signal in this step may include obtaining an image containing a sequencing signal from a sequencing sample, wherein the sequencing sample contains a molecule to be tested having a polypeptide portion and a non-polypeptide portion (NPP), and the sequencing signal contains a mixed sequencing signal having a polypeptide signal and a non-polypeptide signal, using a neural network model to identify the coordinates of a bounding box for the non-polypeptide signal in the image, and converting the coordinates into a time index point in the original sequencing signal, and extracting the polypeptide signal from the mixed sequencing signal through the time index point.
[0061] The method for extracting polypeptide signals in this step includes using a deep learning neural network model to identify the coordinates of a bounding box of non-polypeptide signals within a sequencing image. In some embodiments, the deep learning neural network can be trained using an expanded dataset, such that the trained deep learning neural network is capable of identifying the coordinates of a bounding box of non-polypeptide signals from an image containing a polypeptide signal and one or more non-polypeptide signals. In some embodiments, the trained deep learning neural network takes as input an image containing a polypeptide signal and one or more non-polypeptide signals and outputs the coordinates of a bounding box of non-polypeptide signals. In some embodiments, each bounding box has an associated probability score that an NPP is present within the bounding box. The method for extracting polypeptide signals disclosed herein also includes converting the coordinates into time index points in the original sequencing signal and extracting the polypeptide signal from the mixed sequencing signal using the time index points. The probability score for identifying the NPP and the extracted polypeptide signal are then output for storage and use by an end user. In some embodiments, the neural network model can be constructed based on an object detection model architecture (such as an SSD architecture, a YOLO architecture, or a RetinaNet architecture).
[0062] In a specific embodiment, the method of this step may include a model training phase. In some embodiments, the model can be a machine learning model, such as a convolutional neural network (CNN), for example, a starting neural network, Resnet, RetinaNet, SSD network, YOLO network, or RNN, such as an LSTM model or a GRU model, or any combination thereof. The model can also be any other suitable ML model trained in object detection based on an image, such as a three-dimensional 3DCNN, DTW technology, HMM, etc., or a combination of one or more such technologies, such as CNN-HMM or MCNN. The same type of model or different types of models can be used to identify the coordinates of the bounding box of the non-polypeptide signal.
[0063] In a specific embodiment, the model training scheme may include, for example, obtaining multiple line graphs containing mixed sequencing signals from a sequencing sample containing a test molecule having a polypeptide portion and a non-polypeptide portion, marking a segment of the non-polypeptide signal on each line graph as a label, repeating the above steps with samples containing different polypeptides, obtaining an appropriate number of line graphs (for example, 5,000 or more) and position labels of the non-polypeptide signal, and using the above line graphs and labels as training data to fine-tune the YOLO model. For example, a pre-trained model can be obtained from the GitHub page of YOLOv7, and fine-tuned using the prepared line graphs and labels as training data.
[0064] In some embodiments, the method for extracting polypeptide signals in this step may include the following model training scheme.
[0065] (a) A polypeptide sequence is coupled to two or more fixed NPPs (eg, nucleic acid sequences) by coupling (eg, chemical bonding) to form a linker of NPP1-polypeptide-NPP2 (eg, "nucleic acid 1-polypeptide-nucleic acid 2").
[0066] (b) Obtaining images of multiple sequencing signals (e.g., current signals) of the mixture described in (a) by sequencing technology (e.g., nanopore sequencing). For example, a line graph of each sequencing signal can be drawn with time as the x-axis and the signal as the y-axis.
[0067] (c) Mark (eg, mark with a box) the sequencing signal fragments of NPP1 and NPP2 as labels on each image (eg, line graph) described in (b).
[0068] (d) Repeat steps (a) to (c) with different polypeptide chain samples to obtain 5,000 or more sequencing signal images and position labels of NPP1 and NPP2.
[0069] (e) Select a pre-trained object detection model (such as the YOLO model) and fine-tune it using the images and labels described in (d) as training data.
[0070] In some embodiments, the methods of the present disclosure may include a polypeptide chain fragment extraction scheme for nanopore sequencing signals in the practical application of steps (f) to (h).
[0071] (f) During the application process, a new "NPP1-polypeptide-NPP2" conjugate is prepared using the method of step (a), and the sequencing signal of the mixture is obtained using the method of step (b) and an image is drawn.
[0072] (g) Use the fine-tuned target detection model described in step (e) to perform target detection on the image described in step (f). If the NPP1 and NPP2 segments are recognized and the recognition confidence is above the set threshold, the target detection is successful and proceeds to the next step; otherwise, the detection fails.
[0073] (h). Extract the x coordinates of NPP1 and NPP2 from the target detection coordinates of NPP1 and NPP2 in step (g), and intercept the NPP1 and NPP2 fragments of the corresponding linker sequencing signal. The remaining portion after interception is the polypeptide chain fragment in the linker sequencing signal.
[0074] This step can automate the polypeptide fragment cutting process of the mixed structure (eg, "NPP1-polypeptide-NPP2") sequencing signal, thereby accelerating the data analysis process of the polypeptide fragment signal.
[0075] (4) Using a fast Fourier transform function, a spectrum signal is extracted from the sequencing signal of the polypeptide fragment obtained by intercepting (3).
[0076] Among them, in the present invention, the fast Fourier transform function is abbreviated as FFT, which is an algorithm commonly used in the prior art for efficiently calculating Fourier transform. It can convert the representation of the signal in the time domain into a representation in the frequency domain. By analyzing the amplitude and phase information of different frequency components, the characteristics and structure of the signal are revealed, thereby obtaining the distribution of different frequencies in the signal and analyzing and processing the signal.
[0077] (5) Using a feature set function, extract the following feature values from the nanopore sequencing signals of the k polypeptide fragments obtained in (3) and the spectrum signals obtained in (4) to obtain a feature set.
[0078] Specifically, let x=(x1,x2,…,x n ) is the signal sequence data, and the following 13 eigenvalues are calculated, where x i represents the i-th data point, n is the length of the signal sequence, and there are k polypeptide fragments in total. The "signal sequence" and the "signal sequence" involved in the following 13 characteristic values refer to the sequence of the sequencing signal in step (3) or the sequence of the spectrum signal obtained in step (4); that is, when x is a sequencing signal sequence, the "signal sequence" refers to the sequencing signal sequence; when x is a spectrum signal sequence, the "signal sequence" refers to the spectrum signal sequence.
[0079] Feature 1: Standard error of the slope of the linear trend of the signal sequence block.
[0080] Here m=5 is the block length, (a k ,b k) is the linear regression x i =a k +b k *i is the parameter, n is the length of the signal sequence, if n / m is not divisible, the integer part before the decimal point is retained.
[0081] Feature 2: The second coefficient of the 10th-order autoregressive model. f2(x) = AR2(x)
[0082] Here, AR2(x) is the second coefficient of the AR2(10) model fitted to the data.
[0083] Features 3, 4, and 5: Quartile variation of the signal sequence. Three features are calculated here:
[0084] The x here i Satisfies the 0th quantile to the 0.4th quantile of the signal sequence. f4(x)=Var([x i+1 -x i ])
[0085] The x here i Satisfies the 0.4 quantile to the 0.8 quantile of the signal sequence. f5(x)=Var([x i+1 -x i ])
[0086] The x here i Satisfy the 0.6 quantile to the 0.8 quantile of the signal sequence
[0087] Features 6 and 7: Continuous wavelet transform coefficients. f6(x) = CWT 1,2 (x) f7(x)=CWT 14,10 (x)
[0088] CWT here i,j (x) is the i-th coefficient of the continuous wavelet transform of width j on x. Feature 8: The angle of the first coefficient of the FFT. f8(x) = angle(FFT(x))
[0089] The FFT here is the Fast Fourier Transform function, and the angle function calculates the angle of the first coefficient.
[0090] Feature 9: Quality quantile of signal sequence index.
[0091] Feature 10: Permutation entropy of signal sequence.
[0092] Here p i is the probability that the i-th arrangement appears in the dimensional 3 arrangement with time delay τ=1.
[0093] Feature 11: Calculate the 0.3 quantile. f 11 (x) = Q 0.3 (x)
[0094] Q here 0.3 (x) is the 0.3 quantile of the signal sequence.
[0095] Feature 12: Skewness
[0096] Here is the mean and σ is the standard deviation.
[0097] Feature 13: Variance
[0098] Here is the mean of the signal sequence.
[0099] First, let x=(x1,x2,…,x n ) is the sequencing signal sequence data, where x i represents the i-th data point, and n is the length of the signal sequence. The 13 characteristic values in the sequencing signal are obtained through the above formula.
[0100] Then, let x be the spectrum signal sequence, and obtain the 13 eigenvalues in the spectrum signal through the above method and formula.
[0101] The 13 eigenvalues obtained from the sequencing signal and the 13 eigenvalues obtained from the spectrum signal are combined to obtain a feature set of 26 values, and the feature set is normalized.
[0102] In this step, the feature set function is known in the prior art, but there are a large number of eigenvalues therein. The inventors selected specific 13 eigenvalues from the numerous eigenvalues as a feature set, and obtained these 13 eigenvalues from the sequencing signal and the spectrum signal, respectively, a total of 26 eigenvalues as a feature set. The results showed that the selected feature set has good accuracy.
[0103] (6) Based on the normalized feature set described in (5), the average value of each feature is calculated, and all feature average values are formed into an array as the center point of the feature set.
[0104] (7) removing samples with poor uniformity in the normalized feature set described in (5), specifically by calculating the Euclidean distance between each sample in the normalized feature set described in (5) and the center point obtained in (6), and removing thirty percent of the samples with the farthest distance, thereby obtaining a calibrated feature set.
[0105] (8) Repeat steps (1)-(7) using different known classes of peptides and obtain a calibrated feature set for each peptide.
[0106] (9) The calibration feature sets of all categories of polypeptides described in (8) and their corresponding polypeptide categories are trained to form a polypeptide classifier such as a random forest classifier; wherein the random forest classifier is only exemplary, and those skilled in the art can use any decision tree classification model known in the prior art (such as a random forest model, a GBDT model, an Adaboost model) for training; in the present invention, the classification model obtained by the training is also referred to as a polypeptide classification model or a polypeptide classifier.
[0107] In an embodiment of the present invention, for a classification scheme of a polypeptide to be tested whose category is unknown, the method disclosed herein may include the following steps.
[0108] First, multiple polypeptides of known categories are selected and trained according to the above steps (1)-(9) to obtain a polypeptide classification model such as a random forest classifier; or a polypeptide classification model that has been built based on the feature values can be directly selected for the following polypeptide classification.
[0109] (10) During the application process, for a peptide to be tested whose class is unknown, the above steps (1) to (7) are performed to generate a calibration feature set of the peptide whose class is unknown.
[0110] (11) The calibration feature set of the peptide to be tested obtained in (10) is input into the random forest classifier obtained in (9) or previously established and classified. The class with the highest frequency of output is the classifier's prediction of the class of the unknown peptide.
[0111] (II) Other aspects
[0112] In some embodiments, the present disclosure provides a computer program product comprising instructions that, when executed by a processor, cause the processor to perform part or all of one or more methods and / or part or all of one or more processes of the present disclosure.
[0113] In some embodiments, the present disclosure provides a computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform part or all of one or more methods and / or part or all of one or more processes of the present disclosure.
[0114] In some embodiments, the present disclosure provides a system, apparatus, or device comprising one or more processors and a memory having stored thereon instructions that, when executed by the processor, cause the processor to perform part or all of one or more methods and / or part or all of one or more processes of the present disclosure.
[0115] In some embodiments, the present disclosure provides a kit for polypeptide classification, which includes one or more reagents for performing the method for polypeptide classification of the present disclosure, and / or the system described in the present disclosure. For example, in some embodiments, the kit for polypeptide classification may include one or more reagents for preparing a molecule to be tested having a polypeptide portion and a non-polypeptide portion (NPP). In some embodiments, the kit may include a polypeptide portion and / or a nucleic acid (oligonucleotide or polynucleotide) as an NPP, a non-basic portion (e.g., a non-basic spacer or a non-basic linker), or any other appropriate compound. In some embodiments, the kit may also include one or more reagents for sequencing (e.g., nanopore sequencing), such as primers, enzymes, buffers, etc., and may optionally include instructions for use. In some embodiments, the components in the kit may be appropriately placed in one or more containers.
[0116] Experimental Materials
[0117] The DNA used in the following examples was synthesized by Sangon Biotech (Shanghai) Co., Ltd. (sequences are expressed as 5'→3'), wherein the spacer C3Spacer is represented by iSpC3 and has the structural formula:
[0118] Spacer 18 is represented by iSp18 and has the structural formula:
[0119] All peptides were synthesized by GenScript Biotech Co., Ltd. (sequences are expressed as N-terminus → C-terminus). The C-terminal lysine side chain amino group of the peptide sequence was modified with an azide, represented by LYS(N3). The sequence information is as follows:
[0120] Y-top:
[0121] XXXXXXXXXXXXXXXXXXXXXXXXXXXXXX-SEQ ID NO:1-YYYY-SEQ ID NO:2, X=iSpC3; Y=iSp18
[0122] Among them, SEQ ID NO: 1 is TTTTTTTTTT, and SEQ ID NO: 2 is GGTTGTTTCTGTTGGTGCTGATATTGCT
[0123] Y-bottom (SEQ ID NO.3):
[0124] Phosphorylation (i.e., pho)
[0125] -GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA
[0126] Y-tether:
[0127] Cholesterol-triethylene glycol (ie, Chol-TEG) / TTYYYY-SEQ ID NO: 4, Y=iSp18, wherein SEQ ID NO: 4 is -TTGACCGCTCGCCTC
[0128] DNA1-top (SEQ ID NO.5):
[0129] pho-GCTTCTCGTG-dibenzocyclooctyne (i.e., DBCO)
[0130] DNA2-top (SEQ ID NO.6):
[0131] Maleimide (ie, EMCS)
[0132] -GCTGTCTTCTGTCGTCGTTTCCTTCTCTGC
[0133] DNA1-bottom (SEQ ID NO.7):
[0134] CACGAGAAGCA
[0135] DNA2-bottom (SEQ ID NO.8):
[0136] GCAGAGAAGGAAACGACGACAGAAGACAGC
[0137] peptide1 (SEQ ID NO.9):
[0138] CMEASSEPPLDA{Lys(N3)}
[0139] peptide2 (SEQ ID NO.10):
[0140] CSDVTNQLVDFQW{Lys(N3)}
[0141] peptide3 (SEQ ID NO.11):
[0142] CYPYVAVML{Lys(N3)}
[0143] peptide4 (SEQ ID NO.12):
[0144] CVADHSGQV{Lys(N3)}
[0145] peptide5:
[0146] CT{Lys(N3)}
[0147] peptide6 (SEQ ID NO.13):
[0148] CFEMTIPQFQNFYRQF{Lys(N3)}
[0149] Example 1: Preparation and sequencing of OPO conjugates (O: ssDNA, P: peptide)
[0150] 1. Experimental methods
[0151] Prepare OP linker (where O represents oligonucleotide and P represents polypeptide). Take 12 nmol of peptide (i.e., peptide 1 to peptide 6, SEQ ID NO. 9-13) in an EP tube, add an equal molar amount of TCEP solution, and react in 0.1M HEPES buffer (5mM EDTA, pH = 7.2) at 25°C for 1 hour; then add 0.6 nmol of DNA2-top (SEQ ID NO. 6) solution (DNA2-top final concentration: 100μM), vortex and shake to mix, and react in a metal bath at 25°C for 4 hours. After the reaction is completed, use NEB's PCR & DNA Cleanup Kit (#T1030L) was used for purification to remove unreacted peptides and TCEP.
[0152] Prepare OPO conjugates. Dissolve DNA1-top (SEQ ID NO.5) in 0.1M HEPES buffer (5mM EDTA, pH=7.2) and quantify the concentration using the Qubit ssDNA Detection Kit (ThermoFisher). Add 5 times the molar amount of DNA1-top (SEQ ID NO.5) solution to the above OP conjugate product, vortex and mix, and react in a metal bath at 25°C overnight. After the reaction, use NEB's Unreacted DNA 1-top (SEQ ID NO. 5) was purified using a PCR & DNA Cleanup Kit (#T1030L) to obtain a crude product. The crude product was further purified using an Agilent 1260 Infinity high-performance liquid chromatography (HPLC). The target product fraction was collected, lyophilized, and a portion was analyzed by mass spectrometry. The remaining sample was stored at -80°C until further use.
[0153] Sequencing on the machine. Y-type linker: After annealing the top chain Y-top of the linker with the bottom chain Y-bottom of the linker to obtain an annealing product, incubate it with DNA helicase to obtain a linker containing DNA helicase, which is a Y-type linker complex. The freeze-dried OPO linker is annealed with the complementary sequences DNA1-bottom (SEQ ID NO.7) and DNA2-bottom (SEQ ID NO.8) to form a double-stranded part, which is the OPO annealing product. Use NEB's T4 ligase (#M2200L) kit to connect the OPO annealing product and the Y-type linker complex. The resulting mixture is the library on the machine, and the final concentration of the OPO linker in the mixture is 0.4μM. Use nanopore sequencing to collect nanopore signals at a frequency of 5kHZ.
[0154] 2. Experimental results
[0155] Figure 1 shows the mass spectrometry results for six OPO conjugates. The experimental molecular weights obtained by mass spectrometry match the theoretical molecular weights of the six OPO conjugates, demonstrating that six OPO conjugates were obtained in this example. Figure 2 shows the nanopore sequencing signals for six peptides, with the dashed boxes representing the sequencing signals for the peptide fragments.
[0156] Example 2: Peptide classification model training
[0157] Peptide signal feature extraction. The YOLO recognition model is used to remove the nucleic acid signal from the sequencing signal, leaving only the peptide sequence signal. The spectrum of the peptide fragment sequencing signal is extracted using the Fast Fourier Transform function. The feature extraction package tsfresh is used to extract the feature set of the nanopore sequencing signals and their spectrum signals for k peptide fragments, and the feature set is normalized. For details on the extracted feature value parameters, see the Summary of the Invention section.
[0158] Homogeneity calculation. Calculate the mean value of each feature in the extracted feature set. Compute all the mean values in an array as the center point of the feature set. Calculate the Euclidean distance between each sample in the feature set and the center point. Remove the 30% of samples with the greatest distance to obtain a calibrated feature set.
[0159] Classification model training. After performing the "Peptide Signal Feature Extraction" and "Homogeneity Calculation" operations on the peptides, a calibrated feature set is output for each peptide. All feature sets and their corresponding peptide class labels are combined into a dataset, which is used to train a random forest classifier with parameters n_estimators = 50 and min_samples_split = 10.
[0160] Example 3: Classification of polypeptides to be tested
[0161] The trained classifier was used to perform classification tests on the test set containing the six polypeptide mixtures. The test results are shown in FIG6 . The classification accuracy of the classifier for each polypeptide was greater than 92%, and the overall classification accuracy was 95.52%.
Claims
1. A method for classifying polypeptides, comprising the following steps: A sequencing step, which includes sequencing a polypeptide to be tested or a conjugate containing the polypeptide to be tested, obtaining a sequencing signal of the polypeptide, and obtaining a spectral signal of the polypeptide from the sequencing signal; A feature extraction step, which includes extracting a plurality of feature values from the sequencing signal and the spectral signal respectively through a feature set function to form a feature set; A prediction step, inputting the feature set into a polypeptide classification model for classification, and the category with the highest occurrence frequency in the output categories is the category of the polypeptide to be tested.
2. The method according to claim 1, wherein, The polypeptide classification model is determined according to the feature sets of known polypeptides; Optionally, the training process of the polypeptide classification model includes: Performing the sequencing step and the feature extraction step on polypeptides with known categories to obtain the feature sets of the polypeptides with known categories; Performing model training on the feature sets of the polypeptides with known categories and their corresponding polypeptide categories to obtain the polypeptide classification model.
3. The method according to claim 1 or 2, wherein, The feature values extracted in the feature extraction step include: the standard error of the linear trend slope of the sequencing signal and spectral signal blocks, the second coefficient of the 10th-order autoregressive model, the changes of the sequencing signal and spectral signal from the 0 quantile to the 0.4 quantile, the changes of the sequencing signal and spectral signal from the 0.4 quantile to the 0.8 quantile, the changes of the sequencing signal and spectral signal from the 0.6 quantile to the 0.8 quantile, the coefficients of continuous wavelet transform, the angle of the first coefficient of FFT, the quality quantiles of the indices of the sequencing signal and spectral signal, the permutation entropy of the sequencing signal and spectral signal, the 0.3 quantile of the sequencing signal and spectral signal, skewness, and variance.
4. The method according to claim 1 or 2, wherein, Before the sequencing step, there is also a ligation step, which includes ligating the polypeptide to be tested or the polypeptide with a known category with a non-polypeptide part (NPP) to obtain a conjugate, and the structure of the conjugate is NPP1-polypeptide-NPP2; Optionally, the non-polypeptide part (NPP) includes a nucleic acid part and / or a non-base part. For example, the non-base part can be a non-base spacer; optionally, the NPP can include the same or different nucleic acids, the same or different non-base parts, or a combination of nucleic acid and non-base parts; Optionally, the NNP1 and NNP2 are nucleic acid 1 and nucleic acid 2 respectively. Preferably, the sequences of the nucleic acid 1 and nucleic acid 2 are SEQ ID NO:5 and 6 respectively.
5. The method according to claim 4, wherein, The sequencing step includes: Obtaining a nanopore sequencing signal of the conjugate NPP1-polypeptide-NPP2 through nanopore sequencing technology; Intercepting the sequencing signal of the polypeptide from the nanopore sequencing signal, for example, intercepting the sequencing signal of the polypeptide through a YOLO recognition model; Extracting and obtaining a spectral signal from the intercepted polypeptide sequencing signal, for example, extracting and obtaining a spectral signal through a fast Fourier transform function.
6. The method according to any one of claims 1-5, wherein, The polypeptide classification model is a decision tree classification model, such as a random forest model, a GBDT model, or an Adaboost model.
7. A system for classifying polypeptides, comprising: (1) A sequencing module, which is used to obtain a sequencing signal sequence and a spectral signal sequence of a polypeptide; (2) A feature extraction module, which is used to extract a plurality of feature values from the sequencing signal sequence and the spectral signal sequence to form a feature set; (3) A prediction module, which is configured to input the feature set into a polypeptide classification model for processing to obtain the category of the polypeptide to be tested.
8. A computer program product comprising instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1-6.
9. A computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1-6.
10. A system, comprising: A processor; A memory storing instructions thereon, the instructions, when executed by the processor, causing the processor to execute the method according to any one of claims 1-6.
11. A kit for classifying polypeptides, comprising reagents for performing the method according to any one of claims 1-6, and / or the system according to claim 10.
12. The kit according to claim 11, which comprises reagents for sequencing and feature extraction, optionally further comprising a polypeptide part and a non-polypeptide part (NPP), and optionally comprising an instruction manual.