Method and apparatus for classifying nanopore sequencing time series electrical signals
Patent Information
- Application Number
- CN202280101611.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-16
- Publication Date
- 2025-06-24
AI Technical Summary
The existing technology has high computational complexity, long training time, and relies on base interpretation in nanopore sequencing time series electrical signal classification, resulting in low efficiency and large errors.
The MultiRocket algorithm is used to directly extract feature values from the nanopore sequencing time series electrical signals, and a linear classification model is used for classification, omitting the base reading process, improving efficiency and accuracy.
Significantly improves the accuracy and efficiency of nanopore sequencing signal classification, reduces computational complexity and training time, and adapts to scenarios where features change.
Smart Images

Figure CN120202468A_ABST
Abstract
Description
Method and apparatus for classifying nanopore sequencing time series electrical signals Technical Field
[0001] The present disclosure relates to the field of time series classification, and more particularly, to a method and apparatus for classifying nanopore sequencing time series electrical signals and a method and apparatus for training a classification model. Background Art
[0002] Most time series classification methods that achieve state-of-the-art accuracy have high computational complexity, requiring extensive training time even for smaller datasets and prohibitive computation time for larger datasets. Furthermore, many existing methods focus on a single type of feature, such as shape or frequency.
[0003] Therefore, the present disclosure proposes to perform sequence classification tasks directly based on time series signals, which can omit many intermediate processes and greatly improve efficiency.
[0004] Summary of the Invention
[0005] According to the embodiments of the present disclosure, the classification problem of sequencing results can be solved by moving it forward to the stage of processing time series signals, and the application of the advanced time series classification algorithm MultiRocket achieves high efficiency and extremely high accuracy.
[0006] Specifically, embodiments of the present disclosure provide a method and apparatus for classifying nanopore sequencing time-series electrical signals.
[0007] One aspect of the present disclosure provides a method for classifying nanopore sequencing time series electrical signals, the method comprising: performing feature extraction on the nanopore sequencing time series electrical signals to obtain feature values, the feature values comprising at least one of the proportion of positive values (PPV), the mean value of positive values (MPV), the mean value of positive value indexes (MIPV), and the maximum number of consecutive positive values (LSPV); and inputting the obtained feature values into a linear classification model to classify the nanopore sequencing time series electrical signals.
[0008] According to an embodiment of the present disclosure, the method further includes: preprocessing the nanopore sequencing time series electrical signal to convert it into a sequence of a predetermined length.
[0009] According to an embodiment of the present disclosure, the preprocessing includes: using truncation or zero padding, or using a resampling method.
[0010] According to an embodiment of the present disclosure, the method further includes: performing feature extraction on the nanopore sequencing time series electrical signal by using a MultiRocket algorithm.
[0011] According to an embodiment of the present disclosure, the MultiRocket algorithm uses multiple convolution kernels of fixed length.
[0012] According to an embodiment of the present disclosure, the linear classification model is a ridge regression model.
[0013] According to an embodiment of the present disclosure, the method further includes training the feature extraction and the linear classification model, wherein the training includes: labeling the sequencing signals to generate training data; iteratively extracting features from the training data and inputting the extracted feature values into the linear classification model for classification so that the classification results match the labeling results; and storing the linear classification model and the parameters used in the feature extraction.
[0014] According to an embodiment of the present disclosure, the annotation includes: converting the sequencing signal into a base sequence; for each base sequence, truncating the first M bases and the last N bases, aligning them with the corresponding barcode based on predetermined rules and calculating scores, and taking the highest score of the two alignments as the score, wherein M and N are appropriate values estimated based on the sequencing speed and the size of the adapter, so that the truncated bases completely cover the connected barcode sequence; and marking the sequencing signal with a barcode alignment score exceeding the predetermined score with a corresponding barcode label.
[0015] According to an embodiment of the present disclosure, the labeling further includes: converting the sequencing signal into a base sequence; for each base sequence, truncating the first M bases and the last N bases to compare and score with the corresponding barcode based on a predetermined rule, where M and N are natural numbers greater than 0; and marking the sequencing signal whose barcode comparison score exceeds the predetermined score with a corresponding barcode label.
[0016] According to an embodiment of the present disclosure, the MultiRocket algorithm is directly used to obtain the characteristic value of the nanopore sequencing time series electrical signal without converting it into a base sequence.
[0017] One aspect of the present disclosure provides a method for training a classification model, the method comprising: converting sequencing signals into base sequences; for each base sequence, extracting the first M bases and the last N bases, and performing alignment and scoring with corresponding barcodes based on predetermined rules, where M and N are natural numbers greater than 0; labeling sequencing signals with barcode alignment scores exceeding a predetermined score with corresponding barcode labels; preprocessing the barcode-labeled sequencing signals to obtain training data; and inputting the training data into a classification model to train the classification model.
[0018] One aspect of the present disclosure provides a method for training a classification model, the method comprising: converting sequencing signals into base sequences; for each base sequence, truncating the first M bases and the last N bases, aligning them with corresponding barcodes based on predetermined rules and calculating scores, and taking the highest score of the two alignments as the score, wherein M and N are appropriate values estimated based on sequencing speed and adapter size, so that the truncated bases completely cover the connected barcode sequences; labeling sequencing signals with barcode alignment scores exceeding a predetermined score with corresponding barcode tags; preprocessing the sequencing signals with barcode tags to obtain training data; and inputting the training data into a classification model to train the classification model.
[0019] According to an embodiment of the present disclosure, the classification model includes: MultiRocket feature extraction and linear classification model.
[0020] According to an embodiment of the present disclosure, the preprocessing includes: intercepting the data of the first m seconds and the last n seconds of all sequencing signals with barcode labels, and using the data as the training data, where m>0 and n>0.
[0021] According to an embodiment of the present disclosure, the predetermined rules include: adding 2 points for matching a base; subtracting 1 point for mismatching a base; an initial gap penalty of 0.5 points; and an extended gap penalty of 0.1 points.
[0022] According to an embodiment of the present disclosure, the predetermined score is 45 points.
[0023] According to an embodiment of the present disclosure, the linear classification model is a ridge regression model.
[0024] One aspect of the present disclosure provides a device for classifying nanopore sequencing time-series electrical signals, the device comprising a processing circuit and a memory, the memory storing instructions that, when executed by the processing circuit, cause the device to perform the above-mentioned method for classifying nanopore sequencing time-series electrical signals.
[0025] One aspect of the present disclosure provides a device for training a classification model, the device comprising a processing circuit and a memory, the processor storing instructions, which, when executed by the processing circuit, enable the device to perform the above-mentioned method for training a classification model.
[0026] One aspect of the present disclosure is a non-transitory computer-readable medium storing a computer program comprising instructions that, when executed by a processing circuit, cause the processing circuit to perform any of the methods disclosed herein.
[0027] According to the embodiments of the present disclosure, the present disclosure can effectively improve the accuracy and efficiency of nanopore sequencing signal classification, and has good adaptability to scenarios with constantly changing features.
[0028] The present disclosure compresses the process of sequencing signal classification so that it can be classified from the original electrical signals without base calling, and uses the most advanced time series classification model, which greatly improves efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Other objects, features and advantages will become apparent from the following detailed description of the embodiments made with reference to the accompanying drawings, in which:
[0030] FIG1 shows a flowchart of a method for classifying nanopore sequencing time-series electrical signals according to an embodiment of the present disclosure.
[0031] Figure 2 shows the average ranking of classification accuracy of different algorithms in 109 UCR datasets.
[0032] Figure 3 shows the total training time of different algorithms to complete the 112 UCR data classification tasks.
[0033] FIG4 shows a schematic diagram of dilated convolution and padding terms according to an embodiment of the present disclosure.
[0034] FIG5 shows a schematic diagram of sequence data according to an embodiment of the present disclosure.
[0035] FIG6 shows signals generated by three different hole states according to an embodiment of the present disclosure.
[0036] FIG7 shows a flowchart of a method 700 for training a classification model according to an embodiment of the present disclosure.
[0037] FIG8 shows a schematic diagram of adding a characteristic sequence to the end of a sequencing library according to an embodiment of the present disclosure.
[0038] FIG9 shows the classification accuracy of sequencing signals obtained from two different motor proteins according to an embodiment of the present disclosure.
[0039] FIG10 shows a distribution histogram of comparison and scoring results according to an embodiment of the present disclosure.
[0040] FIG11 is a schematic diagram showing the results of comparing and scoring the predicted barcode sequence with the sequences of four barcodes according to an embodiment of the present disclosure.
[0041] FIG12 schematically shows a schematic block diagram of a device 1200 for executing methods 100 and 700 according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0042] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely illustrative and are not intended to limit the scope of the present disclosure. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0043] The terms used herein are intended only to describe specific embodiments and are not intended to limit the present disclosure. The terms "a," "an," and "the" used herein should also include "plurality" and "multiples," unless the context clearly indicates otherwise. Furthermore, the terms "comprise," "include," and "includes" as used herein indicate the presence of the described features, steps, operations, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, or components.
[0044] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0045] Nanopore sequencing involves immersing a synthetic polymer membrane in an ionic solution. The membrane is covered with engineered transmembrane channel proteins (nanopores). Different voltages are applied across the membrane, creating a voltage difference. Pulled by motor proteins, the DNA strand unwinds through the nanopore, generating base-specific ionic current changes. Nanopore electrical signals are one-dimensional time-series signals. To meaningfully classify or detect anomalies from nanopore sequencing signals, the nanopore signals must be converted into base sequences using base calling techniques, followed by machine learning or deep learning methods for sequence classification. Existing technologies have significant limitations. Base calling often uses complex deep learning models, which are extremely time-consuming to train and infer. These models require advanced hardware and can also exhibit errors in training results. Building on the recent success of convolutional neural networks in time series classification, the MultiRocket algorithm can be used to extract features using a large number of convolutional kernels, followed by a simple linear classifier to achieve state-of-the-art accuracy. This significantly reduces computational time compared to complex algorithms of similar accuracy.
[0046] To solve some of the above problems, the present disclosure proposes a method 100 for classifying nanopore sequencing time series electrical signals. Figure 1 shows a flowchart of the method 100 for classifying nanopore sequencing time series electrical signals according to an embodiment of the present disclosure.
[0047] As shown in FIG1 , in step S102, feature extraction is performed on the nanopore sequencing time series electrical signal to obtain a feature value. The feature value may include at least one of the feature values described below. In one embodiment, feature extraction is performed on the nanopore sequencing time series electrical signal, for example, using a MultiRocket algorithm.
[0048] In the present disclosure, the MultiRocket algorithm is used directly to obtain characteristic values of the nanopore sequencing time series electrical signals without converting them into base sequences. This compresses the sequencing signal classification process, making it possible to classify the raw electrical signals without base interpretation, greatly improving efficiency.
[0049] MultiRocket shares the overall architecture of MiniRocket, using convolutional kernels to transform time series and computing features based on the convolution output (which can then be fed into a linear classifier for classification). Figure 2 shows the average ranking of classification accuracy of different algorithms on 109 UCR datasets, with the bold horizontal line indicating no significant difference in accuracy between the two algorithms. As shown in Figure 2, MultiRocket is significantly more accurate than MiniRocket, and its accuracy is not much worse than the most accurate TSC method to date (e.g., HIVE-COTE 2.0). Figure 3 shows the total training time for different algorithms to complete the classification task on 112 UCR data. As shown in Figure 3, with the same number of 50,000 features, MultiRocket is only 3 times slower than MiniRocket. However, MultiRocket is still significantly faster than all other algorithmic methods.
[0050] The MultiRocket algorithm uses a large number of convolutions, combined with engineering techniques such as dilated convolution and random bias terms, to extract some potentially important eigenvalues, and then imports these eigenvalues into a linear classification machine learning model to complete the time series classification task.
[0051] Specifically, using the MultiRocket algorithm for sequence classification can have the following advantages:
[0052] 1) Addressing the limitations and robustness of existing technologies:
[0053] a. The MultiRocket-based time series classification method solves the problem through data. In theory, it can identify target signals of any form as long as the training set is prepared, so its application scenarios are much wider than traditional solutions.
[0054] b. The time series classification model trained based on data naturally has good robustness and generalization ability. Even if it is not robust in some cases, this problem can be solved by adding such data back into the training set and retraining the model.
[0055] 2) Addressing the low efficiency of existing technologies
[0056] a. Most traditional time series classification methods require complex feature engineering or data preprocessing. MultiRocket uses random convolution kernels to automatically extract features, greatly improving problem-solving efficiency.
[0057] b. Many advanced time series classification methods achieve high classification accuracy by designing very complex learning models. The time complexity and scalability of these solutions are very poor. The method used in this invention uses a fixed-length random convolution kernel to extract features and only uses a very simple linear model for classification. The time efficiency and scalability of this method are incomparable to complex models.
[0058] In an embodiment of the present disclosure, feature extraction using the MultiRocket algorithm may include the following aspects that may be used individually or in combination:
[0059] ● First-order difference: Use the original sequence and the first-order difference sequence as input sequences to obtain different features. For example, the first-order difference sequence of the sequence {1,2,3,4,5} is {2-1,3-2,4-3,5-4} = {1,1,1,1}.
[0060] Convolution kernel: Use a fixed-length convolution kernel. For example, the kernel length can be set to 9, and the weights can be {-1, -1, -1, -1, -1, -1, 2, 2, 2}. In the case of a fixed length, the kernel weights can remain unchanged, but the permutation can be changed. For example, the kernel weights can be {2, -1, -1, -1, -1, -1, 2, 2, -1} or {2, 2, -1, -1, -1, -1, 2, -1, -1}, etc. There are 84 possible permutations of the weights.
[0061] As an example, assuming the convolution kernel is {-1, -1, -1, -1, -1, -1, 2, 2, 2}, and the sequence data is {1, 1, 1, 1, 1, 1, 1, 1}, then the result is (-1*1)+(-1*1)+(-1*1)+(-1*1)+(-1*1)+(-1*1)+(2*1)+(2*1)+(2*1)=0.
[0062] ●Dilated convolution: Add dilated convolution to obtain receptive fields of different scales.
[0063] ●Padding items: You can use 0 for standard padding. The padding is performed alternately, that is, half of the weight / dilated convolution combinations are not padded, and half of the combinations are padded.
[0064] Figure 4 shows a schematic diagram of dilated convolution and padding according to an embodiment of the present disclosure. In the example shown in Figure 4, the dilation distance is 2 and the receptive field length is increased to 5. In addition, Figure 4 also shows an example of padding with zeros.
[0065] ●Bias term: Randomly select certain samples and use the quartiles of their convolution output as bias. The bias term is the only random processing step in the feature extraction process.
[0066] In some embodiments, a pooling operation is performed on the feature sequences obtained in the feature extraction stage. For example, a total of four pooling operators may be provided, wherein the four pooling operators are calculated in parallel. Each feature sequence is calculated by the pooling operator to generate the following four independent feature values, which can be input into the linear classification model.
[0067] PPV: the proportion of positive values in the series
[0068] MPV: the mean of the positive values in the series
[0069] MIPV: the average value of the positive index in the series
[0070] LSPV: Maximum number of consecutive positive values
[0071] Table 1. Calculation of four eigenvalues using the pooling operator
[0072]
[0073]
[0074] Table 1 shows an example of using the pooling operator to calculate four eigenvalues in MultiRocket. For example, in Table 1, the feature sequence is {0, 0, 0, 0, 0, 0, 1, 1, 1, 1}, where there are four consecutive positive values, then PPV = 0.4, MPV = 1, MIPV = 7.5, and LSPV = 4. It should be understood that the number of eigenvalues to be calculated is not limited to 4, and other eigenvalues can be calculated in addition or alternatively.
[0075] Returning to FIG1 , in step S104 of FIG1 , the obtained feature values are input into a linear classification model to classify the nanopore sequencing time series electrical signals. In some examples, the linear classification model may be a ridge regression model.
[0076] According to an embodiment of the present disclosure, method 100 may further include: sequencing data acquisition. This involves acquiring nanopore sequencing time-series electrical signal data containing target signals, i.e., each sequence contains several target signal fragments. The target signal can theoretically be of any form (e.g., detectable by the naked eye). Figure 5 shows a schematic diagram of sequence data according to an embodiment of the present disclosure. In Figure 5, the sequence data includes two mode signals: 1 is the sequencing signal, and 2 is the pore current signal.
[0077] According to an embodiment of the present disclosure, method 100 may further include preprocessing the nanopore sequencing time series electrical signal to convert it into a sequence of a predetermined length. For example, sequencing data preprocessing may include converting one-dimensional sequencing data containing the target signal into a fixed-length signal. This may typically be done by truncation or zero padding, or by other resampling methods such as nearest neighbor interpolation, bilinear interpolation, and cubic convolution interpolation.
[0078] According to an embodiment of the present disclosure, method 100 may further include training the feature extraction and linear classification model, wherein the training includes:
[0079] - Annotate sequencing signals to generate training data;
[0080] - Iteratively extracting features from the training data (e.g., using the MultiRocket algorithm) and inputting the extracted feature values into a linear classification model for classification, so that the classification results match the annotation results; and
[0081] -Stores parameters used in linear classification models and feature extraction (such as the MultiRocket algorithm).
[0082] The above steps can save a linear classification model and feature extraction parameters. You can then use the saved parameters and model to perform time series classification tasks in combination with specific sequencing work, for example:
[0083] ● Refer to the above data preprocessing logic to preprocess the unknown sequence first;
[0084] ●Input the preprocessed sequence data into the trained model for feature extraction and linear model classification.
[0085] In some embodiments, training data production can be performed. The sequencing electrical signal can be used to draw a two-dimensional line graph, as shown in Figure 6. Figure 6 shows the signals generated by three different pore states according to the embodiment of the present disclosure, wherein each data point of the sequencing signal is a voltage value, the unit is picoampere (pA), the horizontal axis is the index of the data point, the sampling rate is 5000 sample / s, and a total of 1s of data is captured in the figure. Have an experienced sequencing engineer mark the sequencing signal, as shown in the figure, normal is a normal sequencing signal, unstable and inactive represent unstable or inactivated electrical signals; or after the sequencing data is converted into a base sequence by basecall software (base calling technology), the sequencing signal label is produced by analyzing the base sequence.
[0086] DNA barcoding enables rapid, accurate, and standardized species identification using a short, standardized DNA fragment. Using different barcode sequences to connect libraries requires identifying which barcodes correspond to which sequencing signals. Since not every sequencing signal can be attached to a barcode, a specific strategy for barcoding sequencing signals is proposed. Because barcode sequences may be attached to either the head or the tail, alignment is performed at both the head and the tail, with the highest score of the two alignments used as the score.
[0087] According to an embodiment of the present disclosure, labeling the sequencing signal may further include:
[0088] -Convert sequencing signals into base sequences;
[0089] - For each base sequence, the first M bases and the last N bases are truncated, and the scores are calculated based on the predetermined rules and the corresponding barcodes. The highest score of the two alignments is taken as the score, where M and N are appropriate values estimated based on the sequencing speed and adapter size, so that the truncated bases completely cover the connected barcode sequence; and
[0090] - Sequencing signals with barcode matching scores exceeding a predetermined score are labeled with corresponding barcode tags.
[0091] The present disclosure also proposes a method for training a classification model. FIG7 shows a flowchart of a method 700 for training a classification model according to an embodiment of the present disclosure.
[0092] According to an embodiment of the present disclosure, the classification model may include: MultiRocket feature extraction and a linear classification model, wherein the linear classification model may be a ridge regression model.
[0093] As shown in FIG7 , in step S702 , the sequencing signal is converted into a base sequence.
[0094] In step S704 of FIG. 7 , for each base sequence, the first M bases and the last N bases are truncated, and then aligned with the corresponding barcode based on predetermined rules for scoring. The highest score of the two alignments is taken as the score. M and N are appropriate values estimated based on sequencing speed and adapter size so that the truncated bases completely cover the connected barcode sequence.
[0095] According to an embodiment of the present disclosure, the predetermined rules include: adding 2 points for matching a base; subtracting 1 point for mismatching a base; an initial gap penalty of 0.5 points; and an extended gap penalty of 0.1 points.
[0096] In step S706 of FIG. 7 , sequencing signals whose barcode alignment scores exceed a predetermined score are labeled with corresponding barcode tags.
[0097] According to an embodiment of the present disclosure, the predetermined score may be 45 points.
[0098] In step S708 of FIG. 7 , the sequencing signals with barcode labels may be preprocessed to obtain training data.
[0099] According to an embodiment of the present disclosure, preprocessing may include: intercepting data from the first m seconds and the last n seconds of all sequencing signals with barcode tags, and using the data as training data, where m>0 and n>0.
[0100] In step S710 of FIG. 7 , training data is input into the classification model to train the classification model.
[0101] The present invention will be further explained in detail below with reference to practical examples. Those skilled in the art and related developers will be able to fully understand and practice the contents of the present invention through these detailed examples. Logical, implementation, and other changes may be flexibly made to the implementation without departing from the main purpose of the present invention. Therefore, the following detailed examples and descriptions are not intended to be limiting, and the scope of the invention is defined by the claims.
[0102] 1. Build hardware and software development environment
[0103] The present invention can be implemented based on the development environment shown in Table 2. Although the present invention can run stably on machines with general CPU configurations, the algorithm module involves reading, writing, and storing large amounts of data, as well as training and testing models, requiring a large amount of data calculations and operations. Running it in a high-performance C software and hardware environment can significantly improve efficiency and stability.
[0104] Table 2. Development environment
[0105]
[0106] 2. Data pre-collection and extraction
[0107] Data was collected using a single-channel nanopore sequencing device and a PC. The target library data was collected using the nanopore sequencing device. The collected signal data and other information were saved as an h5 file structure and stored in real time on the PC hard drive.
[0108] 3. Algorithm
[0109] The time series classification algorithm used in this paper is the open-source MultiRocket algorithm (the input is a one-dimensional time series. Assuming four classification items are required, the output for each time series is one of {0, 1, 2, 3}, which is the result of model judgment). MultiRocket uses a large number of convolutions, combined with engineering techniques such as dilated convolutions and random bias terms, to extract potentially important eigenvalues. These eigenvalues can be imported into a linear classification machine learning model to complete the time series classification task.
[0110] 4. Example 1: Sequencing integrity analysis
[0111] Figure 8 shows a schematic diagram of adding a signature sequence to the end of a sequencing library according to an embodiment of the present disclosure. As shown in the figure, the signature sequence is added to the end of the sequencing library, which generates a specific signal characteristic during the sequencing process, thereby being used to determine the integrity of the sequencing.
[0112] 1) Data annotation and preprocessing
[0113] Each collected sequencing signal is plotted and then handed over to a professional sequencing engineer for data annotation. Since the characteristic signal is at the end, the last 1 second of data can be captured. All data can be divided into a test set and a training set.
[0114] 2) Feature extraction
[0115] Select the sequencing signals of the training set and extract features using the feature extraction method mentioned in the technical solution. Here, the number of features is set to 50,000. The corresponding feature extraction parameters (for example, the coefficient and bias of the dilated convolution) are retained.
[0116] 3) Model training
[0117] Combine the feature values extracted in the previous step with the known labels and input them into the ridge regression model for iterative training. Save the model after training.
[0118] 4) Testing
[0119] Extract features from the test data using the saved feature extraction parameters, input them into the model from the previous step, and calculate the accuracy of the prediction results. You can use different parameters to improve the classification results based on the test results or application scenario, for example, using a different linear model or setting a different number of features.
[0120] Figure 9 shows the classification accuracy of sequencing signals obtained for two different motor proteins according to an embodiment of the present disclosure. As shown, the same model achieved good accuracy under different usage conditions, demonstrating the robustness of the algorithm.
[0121] 5. Example 2: barcode sequence classification
[0122] When connecting libraries using different barcode sequences, it is necessary to classify which signals in the sequencing signals correspond to which barcode.
[0123] 1) Data annotation
[0124] Since not every sequencing signal can be connected to a barcode, the strategy of the present invention is to perform base interpretation on the sequence of each barcode signal and then use an alignment algorithm to calculate the score. The first 150 bases and the last 100 bases can be intercepted and compared with the corresponding barcode for scoring. Based on the above predetermined rules, 2 points are added for each base match, 1 point is deducted for each base mismatch, 0.5 points are penalized for initial gaps, and 0.1 points are penalized for extended gaps. Figure 10 shows a distribution histogram of the alignment scoring results according to an embodiment of the present disclosure. As shown in the figure, only sequences connected to the corresponding barcode can obtain a high score of 45 points or above. Therefore, signals with a barcode matching score of more than 45 points can be selected and labeled with the corresponding barcode label.
[0125] 2) Data preprocessing
[0126] The data before 1s and after 0.5s of all labeled signals can be intercepted and all the data can be used as training sets.
[0127] 3) Feature extraction and model training
[0128] The MultiRocket algorithm is used for feature extraction and model training. The number of feature extraction can be set to 50,000, and the linear classification model can use the ridge regression model.
[0129] 4) Prediction and effect analysis
[0130] Data preprocessing was performed using signals mixed with four different barcodes before importing them into the model. After sequence classification, these classified sequences were base-called and then aligned with the four barcode sequences for scoring. Figure 11 shows a schematic diagram of the scoring results of the predicted barcode sequences compared with the four barcode sequences according to an embodiment of the present disclosure. As shown in the figure, the predicted barcode sequences scored significantly higher than the other groups, demonstrating that the model's classification was effective.
[0131] FIG12 schematically illustrates a schematic block diagram of a device 1200 for executing methods 100 and 700 according to an example embodiment. The device shown in FIG12 can be any device with processing capabilities. It should be noted that the device shown in FIG12 is merely an example and should not limit the functionality or scope of use of the embodiments of the present disclosure.
[0132] As shown in FIG12 , the device 1200 according to this embodiment includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1202 or a program loaded from a storage unit 1208 into a random access memory (RAM) 1203. Various programs and data required for the operation of the device 1200 are also stored in the RAM 1203. The CPU 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0133] Device 1200 may also include one or more of the following components connected to I / O interface 1205: an input section 1206 including a keyboard or mouse, etc.; an output section 1207 including a cathode ray tube (CRT) or liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card, such as a LAN card or modem. Communication section 1209 performs communication processing via a network, such as the Internet. A drive 1210 is also connected to I / O interface 1205 as needed. Removable media 1211, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, etc., is installed in drive 1210 as needed, so that computer programs read from the removable media can be installed in storage section 1208 as needed.
[0134] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1209, and / or installed from a removable medium 1211. When the computer program is executed by the central processing unit (CPU) 1201, the above-mentioned functions defined in the device of the embodiment of the present application are executed.
[0135] It should be noted that the computer-readable medium described in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code embodied on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical cable, RF, or any suitable combination thereof.
[0136] The method and related devices of the present disclosure have been described above in conjunction with preferred embodiments. Those skilled in the art will appreciate that the methods shown above are merely exemplary. The methods of the present disclosure are not limited to the steps and sequence shown above. The devices shown above may include more modules, for example, modules that can be developed or will be developed in the future and can be used for the devices, etc. The various identifiers shown above are merely exemplary and not restrictive, and the present disclosure is not limited to the specific information elements used as examples of these identifiers. Those skilled in the art may make many changes and modifications based on the teachings of the illustrated embodiments.
[0137] The sampling methods, feature extraction algorithms, classification models, and data conversion into images used in this disclosure can all be improved, modified, or replaced. This disclosure can also be extended to areas such as time series clustering, automatic time series annotation, or anomaly detection.
[0138] The program running on the device according to the present disclosure may be a program that controls a central processing unit (CPU) to enable a computer to implement the functions of the embodiments of the present disclosure. The program or the information processed by the program may be temporarily stored in a volatile memory (such as a random access memory RAM), a hard disk drive (HDD), a non-volatile memory (such as a flash memory), or other memory systems.
[0139] The program for realizing each embodiment function of the present disclosure can be recorded on a computer-readable recording medium. The corresponding function can be realized by making a computer system read the program recorded on the recording medium and executing these programs. The so-called "computer system" herein can be a computer system embedded in the device, and can include an operating system or hardware (such as a peripheral device). "Computer-readable recording medium" can be a semiconductor recording medium, an optical recording medium, a magnetic recording medium, a short-term dynamic storage program recording medium or any other recording medium that is computer-readable.
[0140] The various features or functional modules of the devices used in the above embodiments can be implemented or executed by circuits (e.g., single-chip or multi-chip integrated circuits). The circuits designed to perform the functions described in this specification may include a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of the above devices. The general-purpose processor may be a microprocessor, or any existing processor, controller, microcontroller, or state machine. The above circuits may be digital circuits or analog circuits. In the case where new integrated circuit technologies have emerged to replace existing integrated circuits due to advances in semiconductor technology, one or more embodiments of the present disclosure may also be implemented using these new integrated circuit technologies.
[0141] As described above, the embodiments of the present disclosure have been described in detail with reference to the accompanying drawings. However, the specific structure is not limited to the above-mentioned embodiments, and the present disclosure also includes any design changes that do not deviate from the main purpose of the present disclosure. In addition, various modifications can be made to the present disclosure within the scope of the claims, and embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present disclosure. In addition, components with the same effect described in the above-mentioned embodiments can be replaced with each other.
Claims
1. A method for classifying nanopore sequencing time series electrical signals, the method comprising: Performing feature extraction on the nanopore sequencing time series electrical signal to obtain a feature value, wherein the feature value includes at least one of a proportion of positive values (PPV), a mean value of positive values (MPV), a mean value of positive value indexes (MIPV), and a maximum number of consecutive positive values (LSPV); as well as The obtained feature values are input into a linear classification model to classify the nanopore sequencing time series electrical signals.
2. The method according to claim 1, further comprising: The nanopore sequencing time series electrical signal is preprocessed to convert it into a sequence of predetermined length.
3. The method according to claim 2, wherein: The preprocessing includes: using truncation or zero padding, or using a resampling method.
4. The method according to claim 1, further comprising: The MultiRocket algorithm is used to extract features from the nanopore sequencing time series electrical signals.
5. The method according to claim 4, wherein The MultiRocket algorithm uses multiple convolution kernels of fixed length.
6. The method according to claim 1, wherein The linear classification model is a ridge regression model.
7. The method according to claim 1, further comprising training the feature extraction and the linear classification model, wherein The training includes: Annotate sequencing signals to generate training data; Iteratively extracting features from the training data and inputting the extracted feature values into the linear classification model for classification, so that the classification result matches the labeling result; and The linear classification model and parameters used in feature extraction are stored.
8. The method according to claim 7, wherein the marking comprises: converting the sequencing signal into a base sequence; For each base sequence, the first M bases and the last N bases are truncated, and the scores are calculated based on the predetermined rules and the corresponding barcodes. The highest score of the two comparisons is taken as the score, where M and N are appropriate values estimated based on the sequencing speed and the size of the adapter, so that the truncated bases completely cover the connected barcode sequence; and The sequencing signals whose barcode alignment scores exceed the predetermined score are marked with corresponding barcode tags.
9. The method according to claim 4, wherein: The MultiRocket algorithm is directly used to obtain the characteristic value of the nanopore sequencing time series electrical signal without converting it into a base sequence.
10. A method for training a classification model, the method comprising: Convert sequencing signals into base sequences; For each base sequence, the first M bases and the last N bases are cut off and compared with the corresponding barcode based on the predetermined rules. The highest score of the two comparisons is taken as the score. M and N are appropriate values estimated based on the sequencing speed and the size of the adapter, so that the cut bases completely cover the connected barcode sequence. The sequencing signals whose barcode alignment scores exceed the predetermined score are marked with corresponding barcode labels; Preprocess the sequencing signals with barcode labels to obtain training data; The training data is input into a classification model to train the classification model.
11. The method according to claim 10, wherein: The classification model includes: MultiRocket feature extraction and linear classification model.
12. The method according to claim 10, wherein: The preprocessing includes: intercepting data of the first m seconds and the last n seconds of all sequencing signals with barcode labels, and using the data as the training data, wherein m>0 and n>0.
13. The method according to claim 10, wherein: The predetermined rules include: Matching one base adds 2 points; One point is deducted for each base mismatch; A 0.5 point penalty for the initial open shot; and A penalty of 0.1 points will be imposed for extending the open space.
14. The method according to claim 10, wherein: The predetermined score is 45 points.
15. The method according to claim 11, wherein The linear classification model is a ridge regression model.
16. A device for classifying nanopore sequencing time series electrical signals, the device comprising: processing circuits; as well as A memory storing instructions, which, when executed by the processing circuit, cause the apparatus to perform the method according to any one of claims 1 to 9.
17. A device for training a classification model, the device comprising: processing circuits; as well as A memory storing instructions, which, when executed by the processing circuit, cause the apparatus to perform the method according to any one of claims 10 to 15.
18. A non-transitory computer-readable medium storing a computer program, the computer program comprising instructions which, when executed by a processing circuit, cause the processing circuit to perform the method according to any one of claims 1 to 15.