Nanopore sequencing method, training method, electronic equipment and storage medium
By using a sequencing model trained in combination with multiple loss functions in nanopore sequencing and introducing downsampling and upsampling structures into the model, the problem of insufficient nanopore sequencing accuracy and speed in the prior art is solved, and high-precision and fast base recognition are achieved.
Patent Information
- Application Number
- CN202311824058.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-12-27
AI Technical Summary
In the prior art, nanopore sequencing uses high-precision inference based on via current signals, especially due to the influence of noise, electrode drift and base interaction, which leads to slow inference speed and low accuracy.
A sequencing model trained by multiple loss functions is adopted. The model includes a structure that downsamples before encoding and upsamples after encoding the nanopore sequencing signal to improve inference accuracy and speed.
Through combined loss training and lift sampling technology, the accuracy and speed of nanopore sequencing are significantly improved, and fast and accurate base recognition is achieved.
Smart Images

Figure CN120220809A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of machine learning, and in particular, to a nanopore sequencing method, a training method, an electronic device, and a storage medium. Background Art
[0002] Nanopore sequencing has become an important high-throughput sequencing technology in modern genomics; accurate identification of the base sequence on the polymer chain is crucial for determining further analysis and other work. However, the original ionic current signal through the pore is often affected by various factors, such as noise, electrode drift, and interactions between bases. Therefore, high-precision inference based on the ionic current signal through the pore in nanopore sequencing is a challenging problem and is crucial for its wide application.
[0003] In the prior art, machine learning algorithms are used for nanopore sequencing to infer the base sequence on the polymer chain, but the existing machine learning algorithms have the problem of high algorithm complexity, which severely limits the inference speed, and the inference accuracy still needs to be improved. Summary of the Invention
[0004] In view of this, the present disclosure provides a nanopore sequencing method, a training method for a sequencing model, a nanopore sequencing device, a training device for a sequencing model, an electronic device, a storage medium, and a computer program product.
[0005] According to one aspect of the present disclosure, there is provided a nanopore sequencing method, the method comprising:
[0006] Obtaining a nanopore sequencing signal;
[0007] Inputting the nanopore sequencing signal into a sequencing model for base recognition to obtain a base sequence corresponding to the nanopore sequencing signal; wherein, the sequencing model is a model jointly trained by multiple loss functions, and the sequencing model includes a structure for downsampling before encoding the nanopore sequencing signal and upsampling after encoding.
[0008] In a possible implementation, the sequencing model includes: a shared encoder and a connectionist temporal classification (CTC) decoder, wherein the shared encoder is used to encode the nanopore sequencing signal; the CTC decoder is used to decode the encoding result of the shared encoder based on the beam search of CTC to obtain a base sequence corresponding to the nanopore sequencing signal.
[0009] In a possible implementation, the shared encoder includes one or more of a downsampling layer, an encoding layer, an upsampling layer, or a weighted residual connection layer; wherein, the downsampling layer is used to reduce the sampling rate of the nanopore sequencing signal, the encoding layer is used to encode the signal after the sampling rate is reduced, the upsampling layer is used to increase the sampling rate of the output result of the encoding layer, and the weighted residual connection layer is used to perform weighted summation on the nanopore sequencing signal and the output result after the sampling rate is increased to obtain the encoded result of the shared encoder.
[0010] In a possible implementation, the method further includes: segmenting the nanopore sequencing signal to obtain multiple segments of signals;
[0011] The inputting the nanopore sequencing signal into a sequencing model for base recognition to obtain the base sequence corresponding to the nanopore sequencing signal includes:
[0012] Inputting each segment of the multiple segments of signals into the sequencing model for base recognition to obtain the base sequence corresponding to each segment of signal;
[0013] Splicing the base sequences corresponding to each segment of signal to obtain the base sequence corresponding to the nanopore sequencing signal.
[0014] According to another aspect of the present disclosure, there is provided a method for training a sequencing model, the method including:
[0015] Obtaining training data; wherein, the training data includes: multiple groups of nanopore sequencing training signals and the standard base sequence corresponding to each group of nanopore sequencing training signals;
[0016] Inputting each group of nanopore sequencing training signals into a preset model to obtain the sequencing result corresponding to each group of nanopore sequencing training signals; wherein, the preset model includes a structure for downsampling before encoding and upsampling after encoding each group of nanopore sequencing training signals;
[0017] Based on the sequencing result corresponding to each group of nanopore sequencing training signals and the corresponding standard base sequence, jointly training the preset model using multiple loss functions until a preset termination condition is reached, and determining a sequencing model according to the preset model corresponding to when the preset termination condition is reached, the sequencing model being used to perform base recognition on a nanopore sequencing signal to obtain the base sequence corresponding to the nanopore sequencing signal.
[0018] In a possible implementation, the preset model includes: a shared encoder, a Connectionist Temporal Classification (CTC) decoder, and an attention decoder; wherein, the shared encoder is used to encode each group of nanopore sequencing training signals; the CTC decoder is used to decode the encoding result of the shared encoder based on the beam search of CTC to obtain a first base sequence corresponding to each group of nanopore sequencing training signals; the attention decoder is used to decode the encoding result of the shared encoder based on the attention mechanism to obtain a second base sequence corresponding to each group of nanopore sequencing training signals.
[0019] The multiple loss functions include: the loss function of the CTC decoder determined based on the first base sequence corresponding to each group of nanopore sequencing training signals and the standard base sequence, and the loss function of the attention decoder determined based on the second base sequence corresponding to each group of nanopore sequencing training signals and the standard base sequence.
[0020] In a possible implementation, the shared encoder includes one or more of a downsampling layer, an encoding layer, an upsampling layer, or a weighted residual connection layer; wherein, the downsampling layer is used to reduce the sampling rate of each group of nanopore sequencing training signals, the encoding layer is used to encode the signals after the sampling rate is reduced, the upsampling layer is used to increase the sampling rate of the output result of the encoding layer, and the weighted residual connection layer is used to perform weighted summation on each group of nanopore sequencing training signals and the output result after the sampling rate is increased to obtain the encoding result of the shared encoder.
[0021] In a possible implementation, determining the sequencing model according to the preset model corresponding to when the preset termination condition is reached includes:
[0022] Taking the shared encoder and the CTC decoder corresponding to when the preset termination condition is reached as the shared encoder and the CTC decoder of the sequencing model respectively.
[0023] According to another aspect of the present disclosure, a nanopore sequencing device is provided, and the device includes:
[0024] A first acquisition module, configured to acquire nanopore sequencing signals;
[0025] A first sequencing module, configured to input the nanopore sequencing signals into a sequencing model for base recognition to obtain a base sequence corresponding to the nanopore sequencing signals; wherein, the sequencing model is a model obtained by joint training with multiple loss functions, and the sequencing model includes a structure for downsampling before encoding the nanopore sequencing signals and upsampling after encoding.
[0026] According to another aspect of the present disclosure, there is provided an apparatus for training a sequencing model, the apparatus comprising:
[0027] A second acquisition module, configured to acquire training data; wherein the training data includes: multiple groups of nanopore sequencing training signals and a standard base sequence corresponding to each group of nanopore sequencing training signals;
[0028] A second sequencing module, configured to input each group of nanopore sequencing training signals into a preset model to obtain a sequencing result corresponding to each group of nanopore sequencing training signals; wherein the preset model includes a structure for downsampling before encoding and upsampling after encoding each group of nanopore sequencing training signals;
[0029] A training module, configured to jointly train the preset model based on the sequencing result corresponding to each group of nanopore sequencing training signals and the corresponding standard base sequence by using multiple loss functions until a preset termination condition is reached, and determine a sequencing model according to the preset model corresponding to when the preset termination condition is reached, the sequencing model being used to perform base recognition on nanopore sequencing signals to obtain a base sequence corresponding to the nanopore sequencing signals.
[0030] According to another aspect of the present disclosure, there is provided an electronic device, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0031] According to another aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0032] According to another aspect of the present disclosure, there is provided a computer program product, comprising computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0033] In the embodiments of the present disclosure, the sequencing model is a model obtained by jointly training with multiple loss functions, so as to improve the inference accuracy of the sequencing model; and the sequencing model includes a structure for downsampling before encoding and upsampling after encoding nanopore sequencing signals, so as to effectively improve the inference speed of the sequencing model. In this way, by inputting the acquired nanopore sequencing signals into the sequencing model, base recognition can be performed quickly and accurately to obtain a base sequence corresponding to the nanopore sequencing signals, thereby realizing high-precision and fast nanopore sequencing.
[0034] Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Brief Description of the Drawings
[0035] The drawings included in and constituting a part of the specification illustrate exemplary embodiments, features, and aspects of the present disclosure together with the specification, and are used to explain the principles of the present disclosure.
[0036] Figure 1 A flowchart showing a nanopore sequencing method according to an embodiment of the present disclosure.
[0037] Figure 2 A schematic structural diagram showing a shared encoder according to an embodiment of the present disclosure.
[0038] Figure 3 A flowchart showing the process of nanopore sequencing using a sequencing model according to an embodiment of the present disclosure.
[0039] Figure 4 A flowchart showing a nanopore sequencing method according to an embodiment of the present disclosure.
[0040] Figure 5 A flowchart showing a training method according to an embodiment of the present disclosure.
[0041] Figure 6 A flowchart showing the process of training a preset model according to an embodiment of the present disclosure.
[0042] Figure 7 A structural diagram showing a nanopore sequencing device according to an embodiment of the present disclosure.
[0043] Figure 8 A structural diagram showing a training device for a sequencing model according to an embodiment of the present disclosure.
[0044] Figure 9 A schematic structural diagram showing an electronic device according to an embodiment of the present disclosure. Detailed Description of the Embodiments
[0045] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.
[0046] References to "one embodiment" or "some embodiments" etc. described in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present disclosure. Thus, statements such as "exemplary", "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0047] In the present disclosure, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may mean: including the case where A exists alone, the case where A and B exist simultaneously, and the case where B exists alone, where A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item)" or its similar expression means any combination of these items, including any combination of a single item or plural items. For example, at least one (item) of a, b, or c may mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or multiple.
[0048] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following specific implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without certain specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0049] Nanopore sequencing is a biological nanopore sequencing method. It can embed a nanopore (protein pore or solid-state pore) in an insulating artificial membrane to form an ion channel. The two sides of the artificial membrane are filled with electrolyte solutions, and two electrodes are arranged on both sides of the artificial membrane. The potential difference between the two electrodes forms a transmembrane current in the pore of the nanopore. When a polymer chain (such as single-stranded deoxyribonucleic acid (DNA), ribonucleic acid (RNA), protein, etc.) passes through the nanopore, due to the different impedances of different monomers on the polymer chain (such as bases A, T, G, C on the DNA strand), the transmembrane current is modulated when the polymer chain passes through the nanopore, so that the transmembrane current signal can be detected. This transmembrane current corresponds to the monomer sequence on the polymer chain, and thus the monomer sequence on the polymer chain can be deduced by detecting the change of the transmembrane current; for example, the base sequence on the polymer chain can be detected.
[0050] Nanopore sequencing has become an important high-throughput sequencing technology in modern genomics; the accurate identification of the base sequence on the polymer chain is crucial for determining further analysis and other work. However, the original transmembrane current signal is often affected by various factors, such as noise, electrode drift, and the interaction between bases. Therefore, high-precision inference based on the transmembrane current signal in nanopore sequencing is a challenging problem and is crucial for its wide application; in related technologies, machine learning algorithms are used for nanopore sequencing to deduce the base sequence on the polymer chain; however, the existing machine learning algorithms have the problem of high algorithm complexity, which severely limits the inference speed. For example, the Transformer algorithm has reached a milestone height in inference accuracy, but its high complexity severely limits the operation speed. In addition, the algorithms in related technologies all adopt a single loss function training method, which restricts the further improvement of inference accuracy.
[0051] Therefore, the embodiments of the present disclosure provide a solution for nanopore sequencing (for detailed description, see below). Based on joint loss training and upsampling and downsampling, it can overcome the deficiencies of existing nanopore sequencing algorithms, break through the bottleneck of existing algorithms, and improve the accuracy and speed of nanopore sequencing.
[0052] First, the nanopore sequencing method provided by the embodiments of the present disclosure will be specifically described below.
[0053] Figure 1 The flowchart of a nanopore sequencing method according to an embodiment of the present disclosure is shown. This method can be executed by an electronic device or a processor with data processing capabilities, such as Figure 1 As shown, this method may include the following steps:
[0054] Step 101, obtain a nanopore sequencing signal.
[0055] Among them, the nanopore sequencing signal can be the transmembrane current signal detected when the polymer chain to be sequenced passes through the nanopore. For example, a nanopore sequencer can be used to sequence a DNA strand through a nanopore, and the transmembrane current signal of the DNA strand can be detected. This transmembrane current signal is the nanopore sequencing signal.
[0056] Exemplarily, the transmembrane current signal detected when the polymer chain passes through the nanopore can be preprocessed, and the preprocessed signal can be used as the nanopore sequencing signal. Considering that the detected transmembrane current signal may be a current signal with poor filtering quality and high noise, the original detected transmembrane current signal can be preprocessed; exemplarily, the original current signal can be preprocessed such as cleaning and normalization, so as to filter out the current information with poor quality and high noise, and use the preprocessed high-quality transmembrane current signal as the nanopore sequencing signal.
[0057] Exemplarily, after obtaining the nanopore sequencing signal, the nanopore sequencing signal can also be segmented to obtain multiple segments of signals. Exemplarily, there may be a certain overlapping region between adjacent segments of the multiple segments of signals, where the size of the overlapping region can be set according to requirements and is not limited herein.
[0058] Step 102: Input the nanopore sequencing signal into a sequencing model for base recognition to obtain the base sequence corresponding to the nanopore sequencing signal.
[0059] Among them, the sequencing model is a model obtained by jointly training with multiple loss functions, and the sequencing model includes a structure for downsampling before encoding the nanopore sequencing signal and upsampling after encoding. Thus, the base recognition accuracy of the sequencing model can be effectively improved, and at the same time, the computational complexity of the sequencing model can be greatly reduced.
[0060] In a possible implementation manner, the sequencing model may include: a shared encoder and a Connectionist Temporal Classification (CTC) decoder. Among them, the shared encoder is used to encode the nanopore sequencing signal; the CTC decoder is used to decode the encoding result of the shared encoder based on the beam search of CTC to obtain the base sequence corresponding to the nanopore sequencing signal. It can be understood that the shared encoder and the CTC decoder are obtained by jointly training with multiple loss functions, and the training process is described below. Exemplarily, the number of shared encoders can be multiple, that is, the sequencing model may include multiple shared encoders, and different shared encoders can be connected in sequence front and back. In this way, the nanopore sequencing signal can pass through the multiple shared encoders in sequence to obtain the final encoding result.
[0061] In a possible implementation, the structure for downsampling the nanopore sequencing signal before encoding in the sequencing model is used to reduce the sampling rate of the nanopore sequencing signal to be encoded before the sequencing model encodes the nanopore sequencing signal. For example, the nanopore sequencing signal can be regarded as composed of multiple frames of sub-signals. The structure for downsampling can perform weighted summation on adjacent predetermined frames of sub-signals in the nanopore sequencing signal to combine them into one frame of sub-signal, thereby reducing the sampling rate and decreasing the data volume of the nanopore sequencing signal to be encoded. Thus, considering that the inference speed of the sequencing model is limited by the encoding complexity in the sequencing model, and the encoding complexity is affected by the length of the nanopore sequencing signal, therefore, before encoding the nanopore sequencing signal, the sampling rate of the nanopore sequencing signal to be encoded is reduced, thereby greatly reducing the encoding complexity and improving the inference speed of the sequencing model. The structure for upsampling the nanopore sequencing signal after encoding in the sequencing model is used to increase the sampling rate of the encoding result after the sequencing model encodes the nanopore sequencing signal. For example, when the sequencing model encodes the nanopore sequencing signal, it can generate feature vectors for each frame of sub-signal. The structure for upsampling can perform a replication operation on the generated feature vectors for each frame of sub-signal, thereby increasing the sampling rate and increasing the data volume of the encoding result.
[0062] Exemplarily, the sampling rate reduced by the structure for downsampling the nanopore sequencing signal before encoding can be the same as the sampling rate increased by the structure for upsampling the nanopore sequencing signal after encoding. For example, the structure for downsampling can perform weighted summation on adjacent N frames of sub-signals in the nanopore sequencing signal to combine them into one frame of sub-signal. Exemplarily, the downsampling layer can be configured with N weight scalars. The specific values of the N weight variables can be obtained through learning during the training process. The N weight scalars represent the summation weights corresponding to each sub-signal in the adjacent N frames of sub-signals. Correspondingly, the structure for upsampling can replicate the generated feature vectors for each frame of sub-signal N times, where N is an integer greater than 1, so as to achieve proportional reduction and increase of the data volume before and after encoding.
[0063] As an example, the shared encoder includes one or more of a downsampling layer, an encoding layer, an upsampling layer, or a weighted residual connection layer. For example, the shared encoder includes a downsampling layer, an encoding layer, an upsampling layer, and a weighted residual connection layer. Among them, the downsampling layer is a structure for downsampling the nanopore sequencing signal before encoding, and is used to reduce the sampling rate of the nanopore sequencing signal. The encoding layer is used to encode the signal after the sampling rate is reduced. For example, it can extract features from the nanopore sequencing signal after the sampling rate is reduced to generate corresponding feature vectors. Exemplarily, the encoding layer can be a Recurrent Neural Network (RNN), a Convolutional Neural Network (CNN), a Transformer, a Conformer, etc. For example, the encoding layer can be a Transformer layer. The upsampling layer is a structure for upsampling the nanopore sequencing signal after encoding, and is used to increase the sampling rate of the output result of the encoding layer. The weighted residual connection layer is used to perform weighted summation on the nanopore sequencing signal and the output result after the sampling rate is increased to obtain the encoding result of the shared encoder. Among them, the weight value of the nanopore sequencing signal and the weight value of the output result after the sampling rate is increased can be determined in the model training stage. In this way, by reducing the sampling rate of the nanopore sequencing signal through the downsampling layer, the amount of data input to the encoding layer can be reduced, thereby reducing the amount of data that the encoding layer needs to process and effectively improving the operation speed of the encoding layer. After the encoding layer encodes the signal with the reduced sampling rate, the sampling rate of the output result of the encoding layer is increased through the upsampling layer, so that the final encoding result of the shared encoder can represent more detailed information, effectively improving the decoding accuracy and effect of the subsequent CTC decoder. It should be noted that one or more shared encoders can be configured in the sequencing model as needed, and this is not limited. For example, Figure 2 FIG. shows a schematic structural diagram of a shared encoder according to an embodiment of the present disclosure. As Figure 2 shown, the number of the shared encoders in the sequencing model can be 6. Among them, each shared encoder includes a downsampling layer, an encoding layer, an upsampling layer, and a weighted residual connection layer. Among them, the input data is the above-mentioned obtained nanopore sequencing signal, and the output data is the final encoding result.
[0064] As an example, the CTC decoder can include a linear layer. For example, it can be composed of a single linear layer. The encoding result of the shared encoder can be input into this linear layer, and the linear layer decodes based on the beam search of CTC to generate a base sequence, and this base sequence is the base sequence corresponding to the above-mentioned obtained nanopore sequencing signal inferred.
[0065] For example, Figure 3The flowchart of nanopore sequencing by a sequencing model according to an embodiment of the present disclosure is shown. As Figure 3 shown, the sequencing model includes six shared encoders and a CTC decoder. Each shared encoder includes a downsampling layer, a Transformer encoding layer, an upsampling layer, and a weighted residual connection layer. When performing nanopore sequencing, the nanopore sequencing signal is input into the downsampling layer and the weighted residual connection layer of the first shared encoder. Then, the signal with a reduced sampling rate output by the downsampling layer is input into the Transformer encoding layer of the first shared encoder for encoding. Subsequently, the output result of the Transformer encoding layer is input into the upsampling layer of the first shared encoder. After that, the result with an increased sampling rate output by the upsampling layer is input into the weighted residual connection layer of the first shared encoder. The weighted residual connection layer performs weighted summation on the nanopore sequencing signal and the result with an increased sampling rate output by the upsampling layer to obtain the encoding result of the first shared encoder, and this encoding result is input into the second shared encoder. In the second shared encoder, the downsampling layer reduces the sampling rate of this encoding result, and then the signal with a reduced sampling rate output by the downsampling layer is input into the Transformer encoding layer of the second shared encoder for encoding. Subsequently, the output result of the Transformer encoding layer is input into the upsampling layer of the second shared encoder. After that, the result with an increased sampling rate output by the upsampling layer is input into the weighted residual connection layer of the second shared encoder. The weighted residual connection layer performs weighted summation on the encoding result of the first shared encoder and the result with an increased sampling rate output by the upsampling layer to obtain the encoding result of the second shared encoder, and this encoding result is input into the third shared encoder. Similarly, the third, fourth, fifth, and sixth shared encoders also perform the above processing in sequence. Finally, the encoding result of the sixth shared encoder is used as the final encoding result and input into the CTC decoder. The CTC decoder consists of a linear layer and performs decoding processing on the final encoding result of the shared encoder based on the beam search of CTC, and outputs the base sequence corresponding to the nanopore sequencing signal.
[0066] In a possible implementation, when the nanopore sequencing signal includes multiple segments of signals, each segment of the multiple segments of signals can be separately input into the sequencing model for base recognition to obtain the base sequence corresponding to each segment of the signal; and the base sequences corresponding to each segment of the signal are spliced to obtain the base sequence corresponding to the nanopore sequencing signal. Exemplarily, when splicing, there may be an overlapping part between the base sequences corresponding to adjacent segments of signals. For any segment of signal, half of the overlapping part between the base sequence corresponding to this segment of signal and the base sequence corresponding to the adjacent segment of signal can be retained as a part of the base sequence corresponding to this segment of signal. In this way, by traversing all segments of signals, the spliced base sequence, that is, the base sequence corresponding to the nanopore sequencing signal, can be obtained.
[0067] In the embodiments of the present disclosure, the sequencing model is a model obtained by joint training with multiple loss functions, so as to improve the inference accuracy of the sequencing model; and the sequencing model includes a structure for downsampling before encoding the nanopore sequencing signal and upsampling after encoding, so as to effectively improve the inference speed of the sequencing model. In this way, based on joint loss training and up and downsampling, the deficiencies of existing nanopore sequencing algorithms can be overcome, the bottleneck of existing algorithms can be broken through. By inputting the obtained nanopore sequencing signal into the sequencing model, base recognition can be quickly and accurately performed to obtain the base sequence corresponding to the nanopore sequencing signal, thereby realizing high-precision and fast nanopore sequencing.
[0068] Figure 4 The flowchart showing a nanopore sequencing method according to an embodiment of the present disclosure is as Figure 4 shown. First, sequencing data is obtained, then the obtained sequencing data is preprocessed, and then the preprocessed sequencing data is encoded, decoded, downsampled, upsampled, etc. using the sequencing model. Finally, the processing results are spliced to obtain the final complete sequencing result. Exemplarily, the sequencing data can be the above-mentioned nanopore sequencing signal, and the nanopore sequencing signal includes multiple segments of signals. Then, the multiple segments of signals can be first obtained, and then the multiple segments of signals are preprocessed such as cleaning and normalizing; furthermore, each segment of the multiple segments of signals can be encoded, decoded, downsampled, upsampled, etc. through the sequencing model to obtain the base sequence corresponding to each segment of the signal; finally, the base sequences corresponding to each segment of the signal are spliced to obtain the final complete sequencing result, that is, the base sequence corresponding to the entire nanopore sequencing signal.
[0069] The training method of the sequencing model provided by the embodiments of the present disclosure will be specifically described below. Through this training method, the above-mentioned Figure 1 sequencing model can be generated.
[0070] Figure 5A flow chart of a training method according to an embodiment of the present disclosure is shown. The method can be executed by an electronic device or a processor with processing capabilities, such as Figure 5 As shown, the method may include the following steps:
[0071] Step 501: Obtain training data.
[0072] The training data includes: multiple groups of nanopore sequencing training signals and standard base sequences corresponding to each group of nanopore sequencing training signals; illustratively, the nanopore sequencing training signals can be through-hole current signals when a polymer chain passes through a nanopore, and the standard base sequence of the through-hole current signal is known.
[0073] Exemplarily, nanopore sequencing training signals with known base sequences can be obtained from an existing database as training data; or a nanopore sequencer can be used to collect a through-pore current signal when a polymer chain passes through a nanopore, and the base sequence corresponding to the through-pore current signal can be determined using existing standard methods to obtain training data; or training data can be obtained using other methods, which are not limited to this.
[0074] Exemplarily, each set of nanopore sequencing training signals may be preprocessed and / or segmented, wherein possible implementations of the preprocessing and segmented processing may refer to the related description in the above step 101.
[0075] Step 502: Input each set of nanopore sequencing training signals into a preset model to obtain sequencing results corresponding to each set of nanopore sequencing training signals.
[0076] In this step, each group of nanopore sequencing training signals in the training data can be sequentially input into the preset model, and the preset model performs base recognition on each group of nanopore sequencing training signals to generate a base sequence corresponding to each group of nanopore sequencing training signals, that is, a sequencing result corresponding to each group of nanopore sequencing training signals. Exemplarily, in the case where a group of nanopore sequencing training signals includes multiple segments of signals, each segment of the multiple segments of signals can be respectively input into the preset model for base recognition to obtain the base sequence corresponding to each segment of the signal; the base sequences corresponding to each segment of the signal are spliced to obtain the base sequence corresponding to the group of nanopore sequencing training signals.
[0077] In a possible implementation, the preset model may include: a shared encoder, a CTC decoder, and an attention decoder (Attention based Encoder-Decoder, AED); wherein, the shared encoder is used to encode each group of nanopore sequencing training signals; the CTC decoder is used to decode the encoding result of the shared encoder based on the beam search of CTC to obtain a first base sequence corresponding to each group of nanopore sequencing training signals; the attention decoder is used to decode the encoding result of the shared encoder based on the attention mechanism to obtain a second base sequence corresponding to each group of nanopore sequencing training signals.
[0078] Among them, the preset model includes a structure for downsampling before encoding each group of nanopore sequencing training signals and upsampling after encoding. Exemplarily, the structure for downsampling each group of nanopore sequencing training signals in the preset model can be used to reduce the sampling rate of the nanopore sequencing training signals to be encoded before the preset model encodes the nanopore sequencing training signals, and the structure for upsampling after encoding the nanopore sequencing training signals in the preset model is used to increase the sampling rate of the encoding result after the preset model encodes the nanopore sequencing training signals. For the specific description of the structure for downsampling before encoding and upsampling after encoding, reference can be made to the relevant expressions of the structure for downsampling before encoding and upsampling after encoding in the sequencing model in step 102 above, which will not be elaborated here.
[0079] Exemplarily, the number of shared encoders can be multiple, that is, the preset model may include multiple shared encoders, and different shared encoders can be connected in sequence before and after. In this way, the nanopore sequencing training signals can pass through the multiple shared encoders in sequence, so as to obtain the final encoding result. As an example, the shared encoder includes one or more of a downsampling layer, an encoding layer, an upsampling layer, or a weighted residual connection layer; for example, each shared encoder may include a downsampling layer, an encoding layer, an upsampling layer, and a weighted residual connection layer; wherein, the downsampling layer is used to reduce the sampling rate of each group of nanopore sequencing training signals, the encoding layer is used to encode the signals after reducing the sampling rate, the upsampling layer is used to increase the sampling rate of the output result of the encoding layer, and the weighted residual connection layer is used to perform weighted summation on each group of nanopore sequencing training signals and the output result after increasing the sampling rate to obtain the encoding result of the shared encoder. For the specific description of each layer in the shared encoder, reference can be made to the relevant expressions of the shared encoder in the sequencing model in step 102 above, which will not be elaborated here.
[0080] Exemplarily, the CTC decoder may include a linear layer. For example, it may consist of a single linear layer. The encoded result of the shared encoder can be input into this linear layer, and the linear layer decodes based on the beam search of CTC to generate a base sequence, which is the base sequence corresponding to each group of nanopore sequencing training signals inferred.
[0081] Exemplarily, the attention decoder may be a bidirectional decoder. For example, it may include three layers each of a forward decoder and a backward decoder. Alternatively, the attention decoder may be a unidirectional decoder. For example, it may include three layers of forward decoders. Among them, the bidirectional decoder can model the encoded result of the shared encoder from both the forward and backward directions to enhance robustness; the unidirectional decoder can read the encoded result of the shared encoder from left to right, with a faster decoding speed.
[0082] As an example, in the preset model, the CTC decoder and the attention decoder can be set in parallel; the CTC decoder is used for CTC decoding, and the goal is to directly correspond each group of nanopore sequencing training signals to the corresponding standard base sequence; the attention decoder performs autoregressive decoding, focusing on the data that is more critical for base recognition among numerous input data, reducing the attention to other data, and even filtering out irrelevant data; in this way, through the CTC decoder and the attention decoder, the efficiency and accuracy of the preset model for base recognition can be improved.
[0083] Step 503: Based on the sequencing result corresponding to each group of nanopore sequencing training signals and the corresponding standard base sequence, jointly train the preset model using multiple loss functions until a preset termination condition is reached, and determine a sequencing model according to the preset model corresponding to when the preset termination condition is reached. The sequencing model is used to perform base recognition on nanopore sequencing signals to obtain the base sequence corresponding to the nanopore sequencing signal.
[0084] In this step, the sequencing model determined through training can be used to perform base recognition on nanopore sequencing signals to obtain the base sequence corresponding to the nanopore sequencing signal. For example, it can be used as the sequencing model described in step 102 above. Figure 1 Among them, the preset termination condition may include: a preset number of iterations, a preset training time, or convergence of the loss function, etc., which is not limited herein.
[0085] Exemplarily, in this step, based on the sequencing result corresponding to each group of nanopore sequencing training signals and the corresponding standard base sequence, multiple loss functions can be used for joint training, iteratively updating the parameters of the preset model until a preset termination condition is reached, and determining the sequencing model according to the parameters of the preset model corresponding to when the preset termination condition is reached.
[0086] As an example, the shared encoder and CTC decoder corresponding to when the preset termination condition is reached can be used as the shared encoder and CTC decoder of the sequencing model respectively; that is, the sequencing model for actual nanopore sequencing includes: an encoder and a CTC decoder, but does not include a self-attention decoder. In this way, the jointly trained model consists of an encoder and a CTC decoder, and beam search based on CTC is used for decoding to accurately obtain the base sequence, which can avoid performing self-attention decoder inference, maximize the simplification of the sequencing model inference, and speed up the inference speed.
[0087] As another example, the values of the N weight scalars configured in the downsampling layer corresponding to when the preset termination condition is reached can be used as the values of these N weight scalars in the downsampling layer of the sequencing model. For example, when the sampling coefficient is 5, the downsampling layer of the preset model is configured with 5 learnable weight scalars. During the training process of the preset model, the values of these 5 weight scalars can be continuously optimized through the multiple loss functions. The values of these 5 weight scalars corresponding to when the preset termination condition is reached are used as the values of these 5 weight scalars in the downsampling layer of the sequencing model. In this way, during the actual nanopore sequencing process, the sub-signals of 5 adjacent frames in the nanopore sequencing signal can be weighted and summed into one frame of sub-signal. Correspondingly, the upsampling layer can directly copy the feature vector of each frame of sub-signal 5 times, so as to realize the proportional reduction and increase of the data volume before and after encoding.
[0088] As an example, the weight value of the nanopore sequencing training signal in the weight residual connection layer of the preset model corresponding to when the preset termination condition is reached can be used as the weight value of the nanopore sequencing signal in the weight residual connection layer of the sequencing model, and the weight value of the output result after increasing the sampling rate in the weight residual connection layer of the preset model corresponding to when the preset termination condition is reached can be used as the weight value of the output result after increasing the sampling rate in the weight residual connection layer of the sequencing model.
[0089] For example, the calculation principle of the weight residual connection layer in the preset model is shown in the following formula:
[0090] y′=(1-c)·x + c·y
[0091] where x represents the nanopore sequencing training signal, y represents the output result after increasing the sampling rate, y′ represents the output result of the entire shared encoder, c is a learnable scalar between 0 and 1, representing the weight value of the output result after increasing the sampling rate, and 1 - c represents the weight value of the nanopore sequencing training signal.
[0092] Exemplarily, the multiple loss functions include: the loss function of the CTC decoder determined based on the first base sequence corresponding to each group of nanopore sequencing training signals and the standard base sequence, and the loss function of the attention decoder determined based on the second base sequence corresponding to each group of nanopore sequencing training signals and the standard base sequence. In this way, the multiple loss functions can be jointly used as the target loss function to train the preset model, that is, the loss function of the CTC decoder and the loss function of the attention decoder are jointly used as the target loss function, so as to effectively improve the base recognition accuracy and efficiency of the trained preset model. Among them, the sum of the weight values of the loss function of the CTC decoder and the loss function of the attention decoder is 1.
[0093] For example, the loss function of the CTC decoder and the loss function of the attention decoder can be jointly used as the target loss function through the following formula:
[0094] L joint (x, y) = λL CTC (x, y) + (1 - λ)L AED (x, y)
[0095] where, L joint () represents the value of the target loss function, L CTC () represents the value of the loss function of the CTC decoder, L AED () represents the value of the loss function of the attention decoder, L CTC and L AED can be calculated using existing formulas; x represents the base sequence output by the decoder, y is the label of the training data (i.e., the standard base sequence), and λ is a hyperparameter between 0 and 1, which is used to adjust the weights of the loss function of the CTC decoder and the loss function of the AED decoder. For example, it can be taken as 0.5.
[0096] In the embodiments of the present disclosure, a plurality of loss functions are used to jointly train a preset model, and a sequencing model is determined according to the preset model corresponding to when a preset termination condition is reached. As an example, a target loss function composed of the loss function of a CTC decoder and the loss function of an AED decoder can be used to jointly train the preset model. In this way, through joint loss training, it can help the preset model converge better, thereby improving the inference accuracy and efficiency of the trained preset model. At the same time, the preset model includes a structure for downsampling before encoding each group of nanopore sequencing training signals and upsampling after encoding, which can effectively improve the inference speed of the trained preset model. In this way, based on joint loss training and configuring an upsampling and downsampling structure in the preset model, the training speed and training effect of the preset model are effectively improved. The sequencing model determined by the preset model corresponding to when the preset termination condition is reached has good inference speed and inference accuracy, can overcome the deficiencies of existing nanopore sequencing algorithms, and break through the bottleneck of existing algorithms. Furthermore, by using the sequencing model to process nanopore sequencing signals, base recognition can be performed quickly and accurately, and the base sequence corresponding to the nanopore sequencing signal can be obtained, thereby realizing high-precision and fast nanopore sequencing.
[0097] Figure 6 FIG. shows a schematic flow diagram of training a preset model according to an embodiment of the present disclosure, as Figure 6As shown, the preset model may include a shared encoder, a CTC decoder, and an attention decoder. During the training process, the training data is input into the shared encoder. The training data may be any one of the above groups of nanopore sequencing training signals. The number of shared encoders may be six. Each shared encoder includes a downsampling layer, a Transformer encoding layer, an upsampling layer, and a weighted residual connection layer. The group of nanopore sequencing training signals is input into the downsampling layer and the weighted residual connection layer of the first shared encoder. Then, the signal with the reduced sampling rate output by the downsampling layer is input into the Transformer encoding layer of the first shared encoder for encoding. Subsequently, the output result of the Transformer encoding layer is input into the upsampling layer of the first shared encoder. After that, the result with the increased sampling rate output by the upsampling layer is input into the weighted residual connection layer of the first shared encoder. The weighted residual connection layer performs weighted summation on the nanopore sequencing training signals and the result with the increased sampling rate output by the upsampling layer to obtain the encoding result of the first shared encoder, and inputs this encoding result into the second shared encoder. In the second shared encoder, the downsampling layer reduces the sampling rate of this encoding result. Then, the signal with the reduced sampling rate output by the downsampling layer is input into the Transformer encoding layer of the second shared encoder for encoding. Subsequently, the output result of the Transformer encoding layer is input into the upsampling layer of the second shared encoder. After that, the result with the increased sampling rate output by the upsampling layer is input into the weighted residual connection layer of the second shared encoder. The weighted residual connection layer performs weighted summation on the encoding result of the first shared encoder and the result with the increased sampling rate output by the upsampling layer to obtain the encoding result of the second shared encoder, and inputs this encoding result into the third shared encoder. Similarly, the third, fourth, fifth, and sixth shared encoders also perform the above processing in sequence. Finally, the encoding result of the sixth shared encoder is used as the final encoding result and is respectively input into the CTC decoder and the attention decoder. The CTC decoder decodes the final encoding result of the shared encoder and calculates the CTC loss value based on the decoded base sequence and the training data label. The attention decoder decodes the final encoding result of the shared encoder and calculates the AED loss value based on the decoded base sequence and the training data label. Furthermore, joint loss training is performed based on the loss value jointly determined by the CTC loss value and the AED loss value, and the parameter values in the preset model are adjusted based on the loss value. In this way, the above operations are repeatedly executed for different groups of nanopore sequencing training signals until a preset termination condition is reached, and the sequencing model is determined according to the preset model corresponding to when the preset termination condition is reached.
[0098] Next, the performance of the sequencing model obtained by using the training method in the embodiments of the present disclosure will be described.
[0099] (1) Compare the performance of different joint training models under different decoding methods. For example, respectively for the preset model and the Rescore model trained in the embodiments of the present disclosure, compare the performance of using a unidirectional decoder and a bidirectional decoder in the nanopore sequencing signal decoding task.
[0100] Since the preset model in the embodiments of the present disclosure is trained using a combination of weighted CTC and AED losses, the CTC decoder can also work independently. That is, in practical applications, the sequencing model only needs to be configured with a CTC decoder without configuring an AED decoder. The embodiments of the present disclosure conduct ablation experiments to compare the performance differences between the pure encoder and the encoder-decoder architecture. Table 1 shows the results of the performance comparison of different joint training models under different decoding methods. As can be seen from Table 1, when using the pure encoder method, both the unidirectional decoder and the bidirectional decoder achieve performance comparable to their respective complete models. When trained to the 35th round, the accuracy of the unidirectional decoder model is 94.16%, and the accuracy of the bidirectional decoder model is 94.24%.
[0101] Table 1. Performance comparison of different joint training models under different decoding methods
[0102]
[0103] (2) Evaluate the performance of the sequencing model trained in the embodiments of the present disclosure and the CRF model on the same test dataset. Table 2 shows the comparison results of the decoding performance between the sequencing model and the CRF model in the embodiments of the present disclosure. As can be seen from Table 2, the sequencing model proposed in the embodiments of the present disclosure achieves a decoding accuracy not lower than the inference accuracy of the CRF model, and is significantly ahead of the CRF model in terms of inference speed.
[0104] Table 2. Comparison of the decoding performance between the sequencing model and the CRF model in the embodiments of the present disclosure
[0105]
[0106] Based on the same inventive concept of the above method embodiments, the embodiments of the present disclosure also provide a nanopore sequencing device and a training device for a sequencing model, which can be used to execute the technical solutions described in the above method embodiments.
[0107] Figure 7 Show a structural diagram of a nanopore sequencing device according to an embodiment of the present disclosure, as Figure 7 shown, the device includes:
[0108] A first acquisition module 701, configured to acquire nanopore sequencing signals;
[0109] The first sequencing module 702 is configured to input the nanopore sequencing signal into a sequencing model for base recognition to obtain a base sequence corresponding to the nanopore sequencing signal; wherein, the sequencing model is a model obtained by jointly training multiple loss functions, and the sequencing model includes a structure for downsampling before encoding the nanopore sequencing signal and upsampling after encoding.
[0110] In the embodiments of the present disclosure, the sequencing model is a model obtained by jointly training multiple loss functions, so as to improve the inference accuracy of the sequencing model; and the sequencing model includes a structure for downsampling before encoding the nanopore sequencing signal and upsampling after encoding, so as to effectively improve the inference speed of the sequencing model. In this way, by inputting the obtained nanopore sequencing signal into the sequencing model, base recognition can be performed quickly and accurately to obtain a base sequence corresponding to the nanopore sequencing signal, thereby realizing high-precision and fast nanopore sequencing.
[0111] In a possible implementation manner, the sequencing model includes: a shared encoder and a connectionist temporal classification (CTC) decoder, wherein the shared encoder is configured to encode the nanopore sequencing signal; the CTC decoder is configured to decode the encoding result of the shared encoder based on the beam search of CTC to obtain a base sequence corresponding to the nanopore sequencing signal.
[0112] In a possible implementation manner, the shared encoder includes one or more of a downsampling layer, an encoding layer, an upsampling layer, or a weighted residual connection layer; wherein, the downsampling layer is configured to reduce the sampling rate of the nanopore sequencing signal, the encoding layer is configured to encode the signal with the reduced sampling rate, the upsampling layer is configured to increase the sampling rate of the output result of the encoding layer, and the weighted residual connection layer is configured to perform weighted summation on the nanopore sequencing signal and the output result with the increased sampling rate to obtain the encoding result of the shared encoder.
[0113] In a possible implementation manner, the first acquisition module 701 is further configured to: segment the nanopore sequencing signal to obtain multiple segments of signals; the first sequencing module 702 is further configured to: respectively input each segment of the multiple segments of signals into the sequencing model for base recognition to obtain a base sequence corresponding to each segment of the signal; and splice the base sequences corresponding to each segment of the signal to obtain a base sequence corresponding to the nanopore sequencing signal.
[0114] Figure 8 FIG. shows a structural diagram of a training device for a sequencing model according to an embodiment of the present disclosure, as Figure 8 shown, the device includes:
[0115] A second acquisition module 801, configured to acquire training data; wherein the training data includes: multiple groups of nanopore sequencing training signals and a standard base sequence corresponding to each group of nanopore sequencing training signals;
[0116] A second sequencing module 802, configured to input each group of nanopore sequencing training signals into a preset model to obtain a sequencing result corresponding to each group of nanopore sequencing training signals; wherein the preset model includes a structure that performs downsampling before encoding each group of nanopore sequencing training signals and upsampling after encoding;
[0117] A training module 803, configured to jointly train the preset model based on the sequencing result corresponding to each group of nanopore sequencing training signals and the corresponding standard base sequence by using multiple loss functions until a preset termination condition is reached, and determine a sequencing model according to the preset model corresponding to when the preset termination condition is reached, where the sequencing model is used to perform base recognition on nanopore sequencing signals to obtain a base sequence corresponding to the nanopore sequencing signals.
[0118] In the embodiments of the present disclosure, the preset model is jointly trained by using multiple loss functions, and the sequencing model is determined according to the preset model corresponding to when the preset termination condition is reached. As an example, the preset model can be jointly trained by an objective loss function composed of the loss function of a CTC decoder and the loss function of an AED decoder. In this way, through joint loss training, it can help the preset model converge better, thereby improving the inference accuracy and efficiency of the trained preset model; at the same time, the preset model includes a structure that performs downsampling before encoding each group of nanopore sequencing training signals and upsampling after encoding, thereby effectively improving the inference speed of the trained preset model. In this way, based on joint loss training and configuring an upsampling and downsampling structure in the preset model, the training speed and training effect of the preset model are effectively improved. The sequencing model determined by the preset model corresponding to when the preset termination condition is reached has good inference speed and inference accuracy, can overcome the deficiencies of existing nanopore sequencing algorithms, and break through the bottleneck of existing algorithms; furthermore, by using the sequencing model to process nanopore sequencing signals, base recognition can be performed quickly and accurately to obtain a base sequence corresponding to the nanopore sequencing signals, thereby realizing high-precision and fast nanopore sequencing.
[0119] In a possible implementation, the preset model includes: a shared encoder, a Connectionist Temporal Classification (CTC) decoder, and an attention decoder; wherein, the shared encoder is configured to encode each group of nanopore sequencing training signals; the CTC decoder is configured to decode the encoding result of the shared encoder based on the beam search of CTC to obtain a first base sequence corresponding to each group of nanopore sequencing training signals; the attention decoder is configured to decode the encoding result of the shared encoder based on the attention mechanism to obtain a second base sequence corresponding to each group of nanopore sequencing training signals; the multiple loss functions include: the loss function of the CTC decoder determined based on the first base sequence corresponding to each group of nanopore sequencing training signals and the standard base sequence, and the loss function of the attention decoder determined based on the second base sequence corresponding to each group of nanopore sequencing training signals and the standard base sequence.
[0120] In a possible implementation, the shared encoder includes one or more of a downsampling layer, an encoding layer, an upsampling layer, or a weighted residual connection layer; wherein, the downsampling layer is configured to reduce the sampling rate of each group of nanopore sequencing training signals, the encoding layer is configured to encode the signals with the reduced sampling rate, the upsampling layer is configured to increase the sampling rate of the output result of the encoding layer, and the weighted residual connection layer is configured to perform weighted summation on each group of nanopore sequencing training signals and the output result with the increased sampling rate to obtain the encoding result of the shared encoder.
[0121] In a possible implementation, the training module 803 is further configured to: respectively use the shared encoder and the CTC decoder corresponding to when the preset termination condition is reached as the shared encoder and the CTC decoder of the sequencing model.
[0122] The above Figure 7 shown nanopore sequencing device, Figure 8 For the technical effects and specific descriptions of the shown sequencing model training device and its various possible implementations, reference can be made to the methods in the above embodiments, which will not be elaborated here.
[0123] It should be understood that the division of each module in the above device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. In addition, the modules in the device can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or the functions of each module of the device. The processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory inside or outside the device. Alternatively, the modules in the device can be implemented in the form of hardware circuits, and the functions of some or all of the modules can be implemented through the design of the hardware circuits. The hardware circuits can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC), and the functions of some or all of the above modules are implemented through the design of the logical relationships of the components in the circuit. Again, in another implementation, the hardware circuit can be implemented through a programmable logic device (PLD). Taking a field programmable gate array (FPGA) as an example, it can include a large number of logic gate circuits, and the connection relationships between the logic gate circuits are configured through a configuration file, so as to implement the functions of some or all of the above modules. All the modules of the above device can be fully implemented in the form of a processor calling software, or fully implemented in the form of hardware circuits, or partially implemented in the form of a processor calling software, and the remaining part is implemented in the form of hardware circuits.
[0124] In an embodiment of the present disclosure, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a CPU, a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a neural-network processing unit (NPU), a tensor processing unit (TPU), etc.; in another implementation, the processor can implement certain functions through the logical relationship of hardware circuits, and the logical relationship of the hardware circuits is fixed or can be reconfigured. For example, the processor is a hardware circuit implemented by an ASIC or a PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above modules.
[0125] It can be seen that each module in the above device can be one or more processors (or processing circuits) configured to implement the methods of the above embodiments, such as: CPU, GPU, NPU, TPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms. In addition, each module in the above device can be fully or partially integrated together, or can be independently implemented, and this is not limited.
[0126] The embodiment of the present disclosure also provides an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, when the processor is configured to execute the instructions, the methods of the above embodiments are implemented. Exemplarily, it can execute the Figure 1 、 Figure 4 or Figure 5 steps of the methods shown.
[0127] Figure 9 A schematic structural diagram of an electronic device according to an embodiment of the present disclosure is shown, as Figure 9 shown, the electronic device may include: at least one processor 901, a communication line 902, a memory 903, and at least one communication interface 904.
[0128] Processor 901 may be a general-purpose central processing unit, a microprocessor, a specific application integrated circuit, or one or more integrated circuits for controlling the execution of the program of the disclosed solution; processor 901 may also include a heterogeneous computing architecture of multiple general-purpose processors, for example, it may be a combination of at least two of a CPU, a GPU, a microprocessor, a DSP, an ASIC, and an FPGA; as an example, processor 901 may be a CPU+GPU or a CPU+ASIC or a CPU+FPGA.
[0129] The communication link 902 may include a pathway for transmitting information between the above-mentioned components.
[0130] The communication interface 904 uses any transceiver or other device for communicating with other devices or communication networks, such as Ethernet, RAN, wireless local area networks (WLAN), etc.
[0131] The memory 903 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory can be independent and connected to the processor through a communication line 902. The memory can also be integrated with the processor. The memory provided in the embodiment of the present disclosure can generally have non-volatility. Among them, the memory 903 is used to store the computer execution instructions for executing the scheme of the present disclosure, and is controlled by the processor 901 to execute. The processor 901 is used to execute the computer-executable instructions stored in the memory 903, so as to implement the method provided in the above embodiment of the present disclosure; illustratively, the above Figure 1 , Figure 4 or Figure 5 The steps of the method shown.
[0132] Optionally, the computer-executable instructions in the embodiments of the present disclosure may also be referred to as application program codes, which is not specifically limited in the embodiments of the present disclosure.
[0133] Exemplarily, the processor 901 may include one or more CPUs. For example, Figure 9 CPU0 in [[ ]]; the processor 901 may also include one CPU and any one of GPU, ASIC, and FPGA. For example, Figure 9 CPU0 + GPU0 or CPU 0 + ASIC0 or CPU0 + FPGA0 in [[ ]].
[0134] Exemplarily, the electronic device may include multiple processors. For example Figure 9 the processor 901 and the processor 907 in [[ ]]. Each of these processors may be a single-CPU processor, a multi-CPU processor, or a heterogeneous computing architecture including multiple general-purpose processors. The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0135] In a specific implementation, as an embodiment, the electronic device may further include an output device 905 and an input device 906. The output device 905 communicates with the processor 901 and can display information in various ways. For example, the output device 905 may be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. For example, it may be a display device such as an in-vehicle HUD, an AR-HUD, or a monitor. The input device 906 communicates with the processor 901 and can receive user input in various ways. For example, the input device 906 may be a mouse, a keyboard, a touch screen device, or a sensing device, etc.
[0136] Embodiments of the present disclosure provide a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the methods in the above embodiments are implemented. Exemplarily, the above Figure 1 、 Figure 4 or Figure 5 shown steps of the method can be executed.
[0137] Embodiments of the present disclosure provide a computer program product. For example, it may include computer-readable code or a non-volatile computer-readable storage medium carrying computer-readable code; when the computer program product runs on a computer, the computer is caused to execute the methods in the above embodiments. Exemplarily, the above Figure 1 、 Figure 4 orFigure 5 Steps of the method shown
[0138] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement aspects of the present disclosure.
[0139] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0140] The computer-readable program instructions described herein may be downloaded to various computing / processing devices from a computer-readable storage medium or may be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0141] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0142] Aspects of the present disclosure are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.
[0143] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0144] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0145] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0146] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A nanopore sequencing method, characterized in that, The method includes: Obtaining nanopore sequencing signals; Inputting the nanopore sequencing signals into a sequencing model for base recognition to obtain a base sequence corresponding to the nanopore sequencing signals; wherein, the sequencing model is a model obtained by jointly training with multiple loss functions, and the sequencing model includes a structure for downsampling before encoding the nanopore sequencing signals and upsampling after encoding.
2. The method according to claim 1, characterized in that, The sequencing model includes: a shared encoder and a Connectionist Temporal Classification (CTC) decoder. The shared encoder is used to encode the nanopore sequencing signals; the CTC decoder is used to decode the encoding result of the shared encoder based on beam search of CTC to obtain a base sequence corresponding to the nanopore sequencing signals.
3. The method according to claim 2, wherein The shared encoder includes one or more of a downsampling layer, an encoding layer, an upsampling layer, or a weighted residual connection layer; wherein, the downsampling layer is used to reduce the sampling rate of the nanopore sequencing signals, the encoding layer is used to encode the signals with the reduced sampling rate, the upsampling layer is used to increase the sampling rate of the output result of the encoding layer, and the weighted residual connection layer is used to perform weighted summation on the nanopore sequencing signals and the output result with the increased sampling rate to obtain the encoding result of the shared encoder.
4. The method according to any one of claims 1 to 3, characterized in that The method further includes: segmenting the nanopore sequencing signals to obtain multiple segments of signals; The step of inputting the nanopore sequencing signals into a sequencing model for base recognition to obtain a base sequence corresponding to the nanopore sequencing signals includes: Inputting each segment of the multiple segments of signals into the sequencing model for base recognition to obtain a base sequence corresponding to each segment of the signals; Splicing the base sequences corresponding to each segment of the signals to obtain a base sequence corresponding to the nanopore sequencing signals.
5. A training method for a sequencing model, characterized in that, The method includes: Obtaining training data; wherein, the training data includes: multiple groups of nanopore sequencing training signals and a standard base sequence corresponding to each group of nanopore sequencing training signals; Inputting each group of nanopore sequencing training signals into a preset model to obtain a sequencing result corresponding to each group of nanopore sequencing training signals; wherein, the preset model includes a structure for downsampling before encoding each group of nanopore sequencing training signals and upsampling after encoding; Based on the sequencing result corresponding to each group of nanopore sequencing training signals and the corresponding standard base sequence, jointly training the preset model with multiple loss functions until a preset termination condition is reached, and determining a sequencing model according to the preset model corresponding to when the preset termination condition is reached. The sequencing model is used to perform base recognition on nanopore sequencing signals to obtain a base sequence corresponding to the nanopore sequencing signals.
6. The method according to claim 5, wherein The preset model includes: a shared encoder, a Connectionist Temporal Classification (CTC) decoder, and an attention decoder; wherein, the shared encoder is used to encode each group of nanopore sequencing training signals; the CTC decoder is used to decode the encoding result of the shared encoder based on the beam search of CTC to obtain a first base sequence corresponding to each group of nanopore sequencing training signals; the attention decoder is used to decode the encoding result of the shared encoder based on the attention mechanism to obtain a second base sequence corresponding to each group of nanopore sequencing training signals. The multiple loss functions include: the loss function of the CTC decoder determined based on the first base sequence corresponding to each group of nanopore sequencing training signals and the standard base sequence, and the loss function of the attention decoder determined based on the second base sequence corresponding to each group of nanopore sequencing training signals and the standard base sequence.
7. The method according to claim 6, wherein The shared encoder includes one or more of a downsampling layer, an encoding layer, an upsampling layer, or a weighted residual connection layer; wherein, the downsampling layer is used to reduce the sampling rate of each group of nanopore sequencing training signals, the encoding layer is used to encode the signals with the reduced sampling rate, the upsampling layer is used to increase the sampling rate of the output result of the encoding layer, and the weighted residual connection layer is used to perform weighted summation on each group of nanopore sequencing training signals and the output result with the increased sampling rate to obtain the encoding result of the shared encoder.
8. The method according to claim 6 or 7, characterized in that, The determining of the sequencing model according to the preset model corresponding to when the preset termination condition is reached includes: Taking the shared encoder and the CTC decoder corresponding to when the preset termination condition is reached as the shared encoder and the CTC decoder of the sequencing model, respectively.
9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to implement the method according to any one of claims 1-4 or the method according to any one of claims 5-8 when executing the instructions stored in the memory.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions, when executed by the processor, implement the method according to any one of claims 1-4 or the method according to any one of claims 5-8.
Citation Information
Patent Citations
Space environment sensing method and device
CN110545373A
Method for quickly identifying single-molecule nanopore sequencing bases based on deep network
CN112183486A
Real nanopore sequencing signal filtering method and device based on neural network
CN112735524A
RNA modification site prediction method of two-way representation model based on attention
CN115424663A
Design and application of sequencing joint for nanopore sequencing
CN115747211A