Basic group identification method and system based on full convolutional network, medium and program
Through the base recognition method based on the full convolutional network, the base calling model is constructed using the DenseNet convolutional neural network, which solves the problems of low signal sampling rate, high noise, accuracy fluctuations and current fluctuations feature recognition in direct RNA sequencing technology, and achieves efficient and stable base recognition and significant improvement in accuracy.
Patent Information
- Application Number
- CN202510154589.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-04
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-12
AI Technical Summary
In direct RNA sequencing technology, RNA passes through nanopores slowly, resulting in low signal sampling rate and high signal noise, affecting the accuracy of base calls; the accuracy fluctuates greatly when sequencing RNA in different species; the accuracy of base calls is insufficient, especially in identifying key genetic elements such as short exons and exon connections; species-specific challenges and difficulty in identifying current fluctuations.
A base recognition method based on a full convolutional network is adopted to construct a base calling model through a DenseNet convolutional neural network, including the first convolutional layer, the largest pooling layer, several dense blocks, the transition layer, the fully connected layer and the CTC decoder. The model optimizes the structure of the base call model by adjusting the size of the convolution kernel and the number of dense layers layer by layer to improve signal processing capability and accuracy.
It significantly improves the accuracy of base calling, enhances the ability to identify unique signal patterns of RNA data in different species, improves the ability to identify and resolve current fluctuation characteristics, and improves the overall performance of direct RNA sequencing.
Smart Images

Figure CN120032718A_ABST
Abstract
Description
[0001] Priority application This application claims priority to the Chinese invention patent application [202410895365.9] "[Base recognition method, system, medium and program based on fully convolutional network]" filed on July 4, 2024, which is incorporated by reference in its entirety. Technical Field
[0002] The present invention belongs to the technical field of base recognition, and in particular relates to a base recognition method, system, medium and program based on a fully convolutional network. Background Art
[0003] The development of third-generation sequencing technology has revolutionized the ability to perform long-read sequencing of genomes and transcriptomes. It is able to directly sequence DNA and RNA and their modifications without fragmentation and amplification. Oxford Nanopore Technology (ONT)'s Direct RNA Sequencing (DRS) technology can directly sequence RNA molecules on a synthetic membrane through a series of protein nanopores. When the motor enzyme slowly passes a single RNA molecule through the nanopore, different nucleotides block the nanopore, causing current fluctuations, which can be recorded by the base calling algorithm to determine the sequence of the RNA molecule.
[0004] Direct RNA sequencing technology can generate long read sequences (up to 21 kb) and aims to obtain complete transcripts or nearly complete RNA genomes. Complete sequencing makes it powerful in measuring the length of RNA poly(A) tails and analyzing the regulatory function of poly(A) tails on gene expression regulation. In addition, direct RNA sequencing can directly detect modifications of RNA molecules, such as m6A, 7-methylguanosine (m7G), pseudouracil (Ψ), and 2'-O-methylation (Nm), because these modifications will produce different current fluctuations.
[0005] In addition, amplification-free library preparation methods avoid amplification or RT bias and are particularly suitable for identifying exon inclusions (exitrons), which are likely to originate from artifacts caused by reverse transcription. These advantages give direct RNA sequencing technology a special advantage in determining the RNA structure of specific subtypes, identifying tRNA and rRNA, and sequencing the genomes of organisms such as RNA viruses, prokaryotes and parasites. Recent studies have also used direct RNA sequencing technology to perform transcriptome or epitranscriptome analysis in a wide range of species, including humans, animals (nematodes, insects), plants, zooplankton and pathogens. Another application of direct RNA sequencing technology is to analyze RNA metabolic dynamics by labeling nascent RNA with base analogs, such as 5-ethynyluracil and 4-thiouracil, followed by nanopore direct RNA sequencing.
[0006] Although direct RNA sequencing (DRS) technology has shown great potential in genome and transcriptome sequencing, its base calling process still faces many challenges.
[0007] The applicant found that DRS had at least the following technical problems: 1) RNA passes through the nanopore slowly Since RNA passes through the nanopore at a slow rate of only 70 bases per second, compared to 450 bases per second for DNA, the signal sampling rate is low and the signal noise is high, affecting the accuracy of base calling.
[0008] 2) When applied to RNA sequencing of different species, the accuracy fluctuates greatly (i.e., the stability is poor) In particular, the accuracy of base calls in studies of some specific species is low. The base call accuracy of DRS is currently about 86%, which has not yet reached the level required for high-precision base calls (above 90%), especially in identifying key genetic elements such as short exons and exon junctions. In addition, existing technologies also have significant deficiencies in identifying and analyzing the complex current fluctuation characteristics generated by modifications of RNA molecules (such as m6A, 7-methylguanosine, pseudouracil, and 2'-O-methylation).
[0009] 3) Insufficient base calling accuracy: The current most advanced DRS base caller RODAN has an accuracy of about 86%, which has not yet reached the level required for high-precision base calling (above 90%). Especially for the study of RNA viruses and pathogens with high mutation rates, the existing base calling accuracy is not enough to reliably record genetic elements such as short exons and exon junctions, affecting the accuracy and depth of the research.
[0010] 4) Species-specific challenges: Although convolutional neural networks (CNNs) have improved the accuracy of base calling to a certain extent, no research has yet developed species-specific base calling methods to address the problem of low accuracy in species-specific studies. Existing methods cannot fully consider the unique signal patterns of RNA data from different species, resulting in low base calling accuracy in studies of specific species.
[0011] 5) Difficulty in identifying current fluctuation characteristics: Direct RNA sequencing technology can detect modifications of RNA molecules (such as m6A, 7-methylguanosine, pseudouracil, and 2'-O-methylation), but these modifications will produce different current fluctuations, increasing the complexity of base calling. Existing technologies have significant deficiencies in identifying and analyzing these complex current fluctuation characteristics, limiting their application in RNA modification detection.
[0012] Prior art Chinese patent application CN202010026283.2 discloses a method for identifying a base sequence, comprising the steps of: reading a data file output by an Oxford Nanopore sequencer and extracting a current signal corresponding to the DNA / RNA molecule to be tested; cutting out a number of current signal segments of a preset length according to a preset overlap rate from the current signal; inputting each of the current signal segments into a preset timing convolutional network model for timing modeling, so as to generate a corresponding base probability matrix for each current signal segment; wherein the base probability matrix is the probability distribution of bases appearing in the current signal segment at each sampling time point; decoding the corresponding base sequence fragments according to each of the base probability matrices, and generating the base sequence according to each base sequence fragment.
[0013] However, the above-mentioned existing technologies also have the problems of low recognition accuracy and poor universality. Summary of the invention
[0014] The object of the present invention is to provide a base recognition method and system based on a fully convolutional network, which can partially solve or alleviate the above-mentioned deficiencies in the prior art and can efficiently and stably recognize base sequences.
[0015] In order to solve the above-mentioned technical problems, the present invention specifically adopts the following technical solutions: A base recognition method based on a fully convolutional network, comprising: S101, constructing a base calling model based on a DenseNet convolutional neural network; the base calling model includes: The first convolutional layer is used to reduce noise on the electrical signal; The maximum pooling layer is used to downsample the electrical signal; Several dense blocks are used to extract features of electrical signals; The transition layer interspersed between dense blocks is used to reduce the dimension of the features output by the dense blocks; The fully connected layer is used to integrate the global features of the feature maps processed by the dense block and transition layer, and perform classification prediction; The CTC decoder is used to decode the prediction results output by the fully connected layer to obtain the sequence of bases; wherein the dense block includes multiple dense layers, and the dense layers include: a batch normalization layer, a first Silu nonlinear unit layer, a 1×1 second convolutional layer, a batch normalization layer again, a second Silu nonlinear unit layer, and an n×1 third convolutional layer, wherein the steps of setting n include: The reference value of the third convolutional layer of the i-th layer is calculated according to the incremental model, wherein the incremental model is: k= kernel_size + i * kernel_size_increment; k is the reference value, kernel_size is the initial convolution kernel size, i is the number of current dense layers, and kernel_size_increment is the convolution kernel increment; When k < set convolution size K MAX , and k is an even number, then n=k+1; When k < set convolution size K MAX , and when k is an odd number, then n=k; When k ≥ set the convolution size K MAX , then n=K MAX ; S102, training the base calling model using the data set according to the current training conditions to obtain the base calling model; wherein the training conditions include: the number of dense layers, the size of the third convolution kernel of the third convolution layer; S103, determining whether the prediction accuracy of the currently obtained base calling model is greater than a set accuracy threshold; if so, outputting the corresponding base calling model; if not, updating the training conditions and returning to S102; S104, collecting the electrical signal of the base to be identified; S105, inputting the electrical signal into the base calling model for identification to obtain a sequence of the base to be identified.
[0016] In some embodiments, the step of updating the training conditions includes: The number of dense layers of at least one dense block is modified, and the size of the third convolution kernel of each dense layer is updated according to the modified number of dense layers.
[0017] In some embodiments, the plurality of dense blocks include: a first dense block, a second dense block, a third dense block, and a fourth dense block; correspondingly, the step of updating the training condition further includes: Calculate the first prediction accuracy of the base calling model obtained in the last training process; Calculating the second prediction accuracy of the base calling model obtained in the current training process; Calculate the accuracy growth rate according to the first prediction accuracy and the second prediction accuracy; When the accuracy increment rate is greater than or equal to a first increment threshold, it is recommended to perform a first modification operation on the number of dense layers of the dense blocks of the first set, wherein the first set includes: a first dense block, a second dense block, a third dense block, and a fourth dense block; When the accuracy increment rate is less than the first increment threshold, it is recommended to perform a second modification operation on the number of dense layers of the dense blocks of the second set, wherein the second set includes: a third dense block and a fourth dense block.
[0018] In some embodiments, the steps include: When the accuracy increment rate is still less than the first increment threshold after at least one second modification operation, it is recommended to modify the pooling size of the maximum pooling layer.
[0019] In some embodiments, the first modification operation or the second modification operation satisfies the following conditions: The number of dense layers of the second dense block is less than the first set value; The number of dense layers of the third dense block is greater than the second set value.
[0020] In some embodiments, K MAX is 99.
[0021] As an improvement, the pooling size of the maximum pooling layer is 10.
[0022] As an improvement, the transition layer is a 1×1 convolutional layer.
[0023] As an improvement, the data set includes a training set, a validation set, and a test set; The training set and validation set include base electrical signals of Arabidopsis thaliana, human, Caenorhabditis elegans, and Escherichia coli; The test set includes base electrical signals of human, Arabidopsis, mouse, Saccharomyces cerevisiae, Populus nigra, Plasmodium berghei, Seneca Valley virus, porcine epidemic diarrhea virus, and porcine reproductive and respiratory syndrome virus.
[0024] As an improvement, the collected electrical signals of the bases to be identified are preprocessed, including: The electrical signal is divided into signal blocks, each signal block contains 4096 signal values; Signal values were normalized using the median absolute deviation method.
[0025] As an improvement, a base calling model is trained using a data set of a certain organism to obtain a species-specific base calling model for the organism.
[0026] As an improvement, the species-specific base calling model was trained using the DEMINERS training function; the training parameters were a learning rate of 0.002, a batch size of 32, and training for 30 epochs.
[0027] The present invention also provides a base recognition system based on a fully convolutional network, comprising: a model construction module for constructing a base calling model based on a DenseNet convolutional neural network; the base calling model sequentially comprises: The first convolutional layer is used to reduce noise on the electrical signal; The maximum pooling layer is used to downsample the electrical signal; Several dense blocks are used to extract features of electrical signals; The transition layer interspersed between dense blocks is used to reduce the dimension of the features output by the dense blocks; The fully connected layer is used to integrate the global features of the feature maps processed by the dense block and transition layer, and perform classification prediction; CTC decoder, used to decode the prediction results output by the fully connected layer to obtain the sequence of bases; wherein the dense block includes multiple dense layers, and the dense layer includes: batch normalization layer, first Silu nonlinear unit layer, 1×1 second convolution layer, batch normalization layer again, second Silu nonlinear unit layer, n×1 third convolution layer; The steps of setting n include: The reference value of the third convolutional layer of the i-th layer is calculated according to the incremental model, wherein the incremental model is: k= kernel_size + i * kernel_size_increment; k is the reference value, kernel_size is the initial convolution kernel size, i is the number of current dense layers, and kernel_size_increment is the convolution kernel increment; When k < set convolution size K MAX , and k is an even number, then n=k+1; When k < set convolution size K MAX , and when k is an odd number, then n=k; When k ≥ set the convolution size K MAX , then n=K MAX; A first training module is used to train the base calling model using a data set according to current training conditions to obtain a base calling model; wherein the training conditions include: the number of dense layers and the size of the third convolution kernel of the third convolution layer; The second training module is used to determine whether the prediction accuracy of the currently obtained base calling model is greater than a set accuracy threshold; if so, output the corresponding base calling model; if not, update the training conditions and return to the first training module; A data acquisition module, used to collect the electrical signal of the base to be identified; The recognition module is used to input the electrical signal into the base calling model for recognition to obtain the sequence of the base to be recognized.
[0028] Beneficial technical effects: In the existing RNA sequencing process, due to various problems such as low RNA signal sampling rate, large signal noise, and limited sample size, it is difficult to improve the prediction accuracy of RNA.
[0029] In this regard, in view of the one-dimensional data characteristics of RNA, this application chooses to introduce DenseNet convolutional neural network for model training, and preferably sets 4 dense blocks, and sets the size of the convolution kernel of the third convolution layer in a layer-by-layer increasing mode between adjacent dense blocks and between dense layers within dense blocks, so as to gradually increase the in-depth learning of the front-to-back correlation of one-dimensional data (i.e., gradually increase the field of view of the third convolution layer). In addition, in order to coordinate the operational pressure of the model, the size of the convolution kernel of the dense layer of the back-end part will also be restricted. For example, once the convolution kernel size reaches 99, it will no longer increase layer by layer.
[0030] In other words, for RNA, a one-dimensional data, this application proposes a training mode that restricts the convolution kernels between dense blocks and within dense blocks, thereby achieving a certain degree of coordination between training accuracy and computational pressure. And, it has been verified that this restrictive and progressive convolution kernel setting can significantly improve prediction accuracy.
[0031] It is understandable that the parameters of the model may change for different types of prediction data sets. For example, targeted update training may be required for different types of objects (such as humans, mice, etc.), and in order to improve the training efficiency of the model during repeated debugging, this application also proposes an iterative training mode that adjusts the learning details of sequence data and the gradient transfer of details. That is, this application focuses on adjusting the model's learning details of sequence data and other information by adjusting the number of dense layers of dense blocks, so as to quickly correct the model's prediction performance.
[0032] Specifically, the restrictive incremental setting of the third convolution kernel size proposed in this application can expand the convolution field of view layer by layer between multiple dense layers. This restrictive layer-by-layer incremental scheme can not only improve the field of view, that is, to conduct in-depth learning of the correlation of the sequence of electrical signals, but also help to maintain the effective transmission of gradients to ensure that the network can be stably trained. In addition, this method of increasing the field of view layer by layer and limiting the field of view size at the back end (that is, limiting the convolution kernel size) can also control the computational pressure during training.
[0033] In summary, the present invention performs base recognition by using a base calling model based on a fully convolutional network, and can solve the problems in the prior art in the following ways: 1. Improve the signal processing capability of the sampling rate: The base calling model architecture design in the present invention can effectively process the high-noise signal caused by the low sampling rate and improve the accuracy of signal analysis.
[0034] 2. Improve base calling accuracy: By introducing the DenseNet convolutional neural network architecture and adjusting the convolutional field of view of the convolutional neural network, the base calling model in the present invention achieves the latest performance level on various data sets such as animals, plants and viruses, and the base calling accuracy is significantly improved.
[0035] Species-specific model training: The base calling model in this invention supports species-specific model training, which can be optimized for the unique signal patterns of different species, significantly improving the accuracy of base calling in studies of specific species.
[0036] In addition to solving the problems existing in the prior art, the present invention also has the following advantages: 1. Advantages of architectural design.
[0037] The base calling model provided by the present invention achieves efficient parameter utilization and overfitting control through the innovative design of the fully convolutional network and dense blocks. The densely connected design enhances feature reuse, reduces the number of parameters, and improves the generalization ability of the model. These architectural advantages enable it to perform well in processing complex electrical signals.
[0038] Efficient feature reuse: The dense block design in the base calling model provided by the present invention enables each layer to receive feature maps from all previous layers, thereby enhancing feature reuse and information flow. This design not only reduces the number of parameters, but also significantly improves the performance and stability of the model.
[0039] Parameter optimization and computational efficiency: The base calling model provided by the present invention increases the number of channels in the dense layer and the size of the convolution kernel to capture a wider range of base position information. The number of channels in the last dense block reaches 1024, and the convolution kernel size is 99. Through this optimization, the model ensures high accuracy while improving computational efficiency, making it suitable for large-scale data processing.
[0040] 2. Wide range of applications.
[0041] The base calling model provided by the present invention performs well not only on a single species dataset, but also on a mixed dataset of multiple species. Its versatility and adaptability make it have wide application potential in different research fields.
[0042] Wide range of application scenarios: Applicable to RNA sequencing of a variety of biological samples, including humans, plants, animals, viruses, etc. In various RNA sequencing applications, whether basic research or clinical diagnosis, Densecall has demonstrated excellent performance and reliability.
[0043] Research on RNA viruses with high mutation rates: In the study of RNA viruses with high mutation rates (such as SARS-CoV-2, PRRSV, SVV, and PEDV), the high accuracy and low error rate of the base calling model provided by the present invention can more reliably identify and record the genetic information of the virus, providing strong support for the study of virus evolution and transmission.
[0044] In summary, the base calling model provided by the present invention significantly improves the base calling accuracy of direct RNA sequencing through an innovative fully convolutional network architecture and species-specific model training, overcomes many shortcomings of the prior art, and demonstrates broad application prospects and significant beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, each element or part is not necessarily drawn according to the actual scale. Obviously, the drawings described below are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without paying creative labor.
[0046] Figure 1 A schematic diagram of the base calling model structure in an exemplary embodiment of the present application; Figure 2 It is a schematic diagram of the overall accuracy effect of the base recognition method in an exemplary embodiment of the present application; Figure 3 It is a schematic diagram of the accuracy and error rate results of Densecall (the base calling model in the present invention), Guppy, and RODAN on different species; Figure 4 It shows that the base calling accuracy of Densecall (the base calling model in the present invention) on the mouse transcriptome increased from 87.84% to 90.44%, exceeding 88.16% of and 84.06% of Guppy; Figure 5 The differences in mapping rates and base call lengths of different identification methods are shown; Figure 6 A first schematic diagram of a model training interface in an exemplary embodiment of the present application is shown; Figure 7 A second schematic diagram of a model training interface in an exemplary embodiment of the present application is shown; Figure 8 A third schematic diagram of a model training interface in an exemplary embodiment of the present application is shown; Fig. 9 Schematic diagram of a base recognition method in an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0048] Herein, suffixes such as "module", "component" or "unit" used to represent elements are only used to facilitate the description of the present invention, and have no specific meanings by themselves. Therefore, "module", "component" or "unit" can be used mixedly.
[0049] In this document, the terms "upper", "lower", "inner", "outer", "front", "back", "one end", "the other end" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance.
[0050] In this document, unless otherwise clearly specified and limited, the terms "installed", "provided with", "connected", etc. should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0051] Herein "and / or" includes any and all combinations of one or more of the associated listed items.
[0052] Herein, "plurality" means two or more than two, ie, it includes two, three, four, five, etc.
[0053] Embodiment 1 In order to improve the processing capability of low sampling rate RNA signals and enhance the accuracy of base calling, the present invention provides a base recognition method based on a fully convolutional network. Figure 1 and Fig. 9 As shown, including: S101, builds a base calling model based on the DenseNet convolutional neural network (also referred to as Densecall in this article).
[0054] DenseNet is a deep convolutional neural network structure. Its core idea is to establish direct connections between each layer of the network to ensure maximum information transfer between layers.
[0055] In this embodiment, Figure 1 As shown in the figure, the specific structure of the base calling model (Densecall) built based on DenseNet includes: 1. The first convolutional layer is used to reduce the noise of the electrical signal (which is one-dimensional data).
[0056] The first convolutional layer performs noise reduction on the input electrical signal through three convolutional layers. The convolutional layer can effectively extract the local features of the signal and reduce noise.
[0057] 2. The maximum pooling layer is used to downsample the electrical signal.
[0058] Preferably, the pooling size of the maximum pooling layer in this embodiment is 10. By downsampling the electrical signal, the dimension of the signal can be reduced, thereby reducing the computational complexity and further removing redundant information.
[0059] 3. Several dense blocks are used to extract the features of electrical signals.
[0060] Preferably, a dense block includes multiple dense layers, and the dense layers include: a batch normalization layer, a first Silu nonlinear unit layer, a 1×1 second convolution layer, another batch normalization layer, a second Silu nonlinear unit layer, and an n×1 third convolution layer, where n is the size of the convolution kernel of the third convolution layer (also referred to as: the third convolution kernel size).
[0061] In this embodiment, the batch normalization layer is used to accelerate training and reduce internal covariate shift. The Silu nonlinear unit is a nonlinear activation function. The role of the 1×1 second convolutional layer is to reduce the feature dimension and reduce the computational complexity. The batch normalization and SiLU layers repeat the batch normalization and SiLU operations again. The n×1 third convolutional layer is used to extract local features.
[0062] In this embodiment, the step of setting n includes: The reference value of the third convolutional layer of the i-th layer is calculated according to the incremental model, wherein the incremental model is: k= kernel_size + i * kernel_size_increment; k is the reference value, kernel_size is the initial convolution kernel size, i is the number of current dense layers, and kernel_size_increment is the convolution kernel increment; When k < set convolution size K MAX , and k is an even number, then n=k+1; When k < set convolution size K MAX , and when k is an odd number, then n=k; When k ≥ set the convolution size K MAX , then n=K MAX ; In this embodiment, on the one hand, the incremental model is used to increase the size of the third convolution kernel of the dense layer layer by layer, and the field of view of the third convolution layer is gradually expanded layer by layer in a layer-by-layer manner to ensure that the model can conduct in-depth learning of one-dimensional data; on the other hand, the convolution size K is set MAX The maximum value of the convolution kernel is limited to further balance the computational pressure during model training. In other words, the present invention proposes a training method for restrictively increasing the convolution kernel between dense layers, and this restrictive increasing training method can achieve a good balance between training accuracy and training data volume.
[0063] 4. The three transition layers interspersed between the four dense blocks are used to reduce the dimension of the features output by the dense blocks.
[0064] After each dense block, except the last one, a 1x1 convolutional layer is used as a transition layer for feature compression to speed up processing.
[0065] 5. The fully connected layer is used to integrate the global features of the feature maps processed by the dense block and transition layer, and perform classification prediction.
[0066] The feature map processed by the dense block and transition layer is input into the fully connected layer, which flattens the multi-dimensional feature map into a one-dimensional vector and integrates the features through a fully connected operation.
[0067] Use log-softmax activation for classification. Log-softmax is an activation function used for multi-classification problems, usually used in the output layer of a neural network. Its function is to convert the raw output (logits) of the model into a probability distribution and take the logarithm for ease of calculation.
[0068] 6. CTC decoder, used to decode the prediction results output by the fully connected layer to obtain the sequence of bases.
[0069] CTC (Connectionist Temporal Classification) is an algorithm for sequence recognition, especially suitable for dealing with inconsistent lengths between input and output. The CTC decoder eliminates blank labels in the sequence through a dynamic programming algorithm and merges consecutive repeated labels to obtain the final sequence output.
[0070] The loss function of this base calling model uses CTC loss for gradient descent, and base calling uses beam search of size 5. The input of this base calling model is: electrical signal; the output is: base sequence corresponding to the electrical signal.
[0071] S102, training the base calling model using the data set according to the current training conditions to obtain a base calling model through training and updating; wherein the training conditions include: the number of dense layers, the size of the third convolution kernel of the third convolution layer; S103, determining whether the prediction accuracy of the currently obtained base calling model is greater than a set accuracy threshold; If yes, the corresponding base calling model is output, and the user can be recommended to use the current base calling model to predict the RNA sequence. If no, the training condition is updated, and the process returns to S102.
[0072] Usually, the training conditions need to be modified multiple times (ie, S102 may be executed multiple times) to train a base calling model that meets the requirements.
[0073] For example, in some embodiments, the training condition is updated by modifying the number of dense layers. For example, in some embodiments, the number of dense layers is modified by doubling the number of dense layers of each dense block. Of course, at this time, the size of the third convolution kernel will also be adaptively changed according to the change in the number of dense layers.
[0074] For another example, in some embodiments, the training conditions are updated by modifying the size of the third convolution kernel, for example, increasing the initial convolution kernel size.
[0075] S104, collecting the electrical signal of the base to be identified; S105, inputting the electrical signal into the base calling model for identification to obtain a sequence of the base to be identified.
[0076] In some embodiments, the step of updating the training conditions includes: The number of dense layers of at least one dense block is modified, and the size of the third convolution kernel of each dense layer is updated according to the modified number of dense layers.
[0077] In some embodiments, the plurality of dense blocks include: a first dense block, a second dense block, a third dense block and a fourth dense block in sequence. Preferably, the number of dense layers of the first, second and third dense blocks increases in sequence, and the number of dense layers of the fourth dense block is less than the number of dense layers of the third dense block.
[0078] In this embodiment, the first and second dense blocks have a relatively small number of dense layers, which are mainly used to learn low-level features of data, and the third and fourth dense blocks have a slightly larger number of dense layers, which are mainly used to learn high-level features of data.
[0079] Furthermore, in some embodiments, the step of updating the training conditions further includes: Calculate the first prediction accuracy of the base calling model obtained in the last training process; Calculating the second prediction accuracy of the base calling model obtained in the current training process; Calculate the accuracy growth rate according to the first prediction accuracy and the second prediction accuracy, wherein the accuracy growth rate = (second prediction accuracy - first prediction accuracy) / first prediction accuracy; When the accuracy increment rate is greater than or equal to a first increment threshold, it is recommended to perform a first modification operation on the number of dense layers of the dense blocks of the first set, wherein the first set includes: a first dense block, a second dense block, a third dense block, and a fourth dense block; For example, in some embodiments, when the accuracy increment rate is greater than or equal to the first increment threshold, it is considered that the current ratio of the number of dense layers of each dense block is relatively appropriate, and the number of dense layers can be adjusted based on the current ratio, such as doubling the number of dense layers of each dense block.
[0080] When the accuracy increment rate is less than the first increment threshold, it is recommended to perform a second modification operation on the number of dense layers of the dense blocks of the second set, wherein the second set includes: a third dense block and a fourth dense block.
[0081] For example, in some embodiments, when the accuracy increment rate is less than the first increment threshold, it is recommended to directly adjust the ratio of the number of dense layers of each dense block, and preferably it is recommended to adjust the number of dense layers of the third dense block and the fourth dense block to adjust the model's learning ability for high-level features.
[0082] In this embodiment, during the model training process, emphasis is placed on adjusting the number of dense layers to adjust the learning dimension of the base calling model for one-dimensional data.
[0083] It is worth noting that the sequence of RNA electrical signals is usually long, and there are often dependencies in the sequence. Increasing the size of the convolution kernel also means that the field of view of the convolution is larger, which helps to improve the convolution layer's ability to learn dependencies.
[0084] In this regard, the restrictive incremental setting of the third convolution kernel size proposed in this application can expand the convolution field of view layer by layer between multiple dense layers. This restrictive layer-by-layer incremental scheme can not only improve the field of view, that is, to conduct in-depth learning of the correlation of the sequence of electrical signals, but also help to maintain the effective transmission of gradients to ensure that the network can be stably trained. In addition, this method of increasing the field of view layer by layer and limiting the field of view size at the back end (that is, limiting the convolution kernel size) can also control the computational pressure during training.
[0085] Furthermore, in order to improve the training efficiency of the model during repeated debugging, this application proposes an iterative training mode that adjusts the learning details of the sequence data and the gradient transfer of the details. That is, this application focuses on adjusting the number of dense layers of the dense block to adjust the model's learning details of the sequence data and other information, thereby being able to quickly correct the model's prediction performance.
[0086] In some embodiments, the steps include: When the accuracy increment rate is still less than the first increment threshold after at least one second modification operation, it is recommended to modify the pooling size of the maximum pooling layer.
[0087] In some embodiments, the first modification operation or the second modification operation satisfies the following conditions: The number of dense layers of the second dense block is less than the first set value; The number of dense layers of the third dense block is greater than the second set value.
[0088] In this embodiment, it is preferred to limit the number of dense layers of the second dense block for learning low-level features and the third dense block for learning high-level features respectively, so as to ensure the complete learning of low-level and high-level features, respectively, and further ensure that the overall detailed content formed by the connection of low-level and high-level features can more fully display the characteristics of the sequence.
[0089] In some embodiments, K MAX is 99.
[0090] As a preferred base calling model, there are 4 dense blocks, namely the first dense block, the second dense block, the third dense block and the fourth dense block; the first dense block includes 6 dense layers; the second dense block includes 12 dense layers; the third dense block includes 24 dense layers; the fourth dense block includes 16 dense layers. Each layer is connected in a feedforward manner and receives the cascaded feature maps of all previous layers to enhance feature reuse and reduce the number of parameters.
[0091] In this embodiment, the shallow layer of the network is usually responsible for extracting low-level features, while the deep layer is responsible for extracting higher-level semantic features. By gradually increasing the number of layers of the dense block, the complexity of the features can be gradually increased, so that the network can better capture features from low to high levels. The present invention preferably configures the dense layers of the four dense blocks of Densecall to [6, 12, 24, 16], and after a large number of experimental verifications on the training data set, it can be proved that this special configuration can indeed achieve a better balance between model performance and computational efficiency. In the first few dense blocks, gradually increasing the number of layers can gradually increase the number of feature maps, thereby capturing more details and complex features. For example, the first dense block has 6 layers, the second dense block has 12 layers, and the third dense block has 24 layers. Reducing the number of layers in the last dense block can reduce the amount of calculation while maintaining the performance of the model. For example, the fourth dense block has 16 layers.
[0092] The existing DenseNet convolutional neural network is widely used in image classification tasks, where the size of the convolution kernel is usually a two-dimensional structure of n×n, which is suitable for processing two-dimensional image data. However, in this scheme, the DenseNet convolutional neural network is applied to one-dimensional base electrical signal recognition to realize the calling of base sequences. In order to adapt to this transformation, we adjust the convolution kernel to a one-dimensional convolution kernel of n×1 accordingly.
[0093] This modification is based on the flexibility and powerful feature learning ability of DenseNet. Through its unique dense connection method, DenseNet enables each layer to be directly connected to all previous layers, thereby enhancing the transfer and reuse of features, reducing feature redundancy, and improving the performance of the model while reducing the number of parameters. When processing one-dimensional signals, this design can also play its advantages, effectively extract local features in the sequence, and enhance the expressiveness of features through dense connections.
[0094] Specifically, in this embodiment, the convolution layer inside each Dense Block receives not only the output of the previous layer, but also the output of all previous layers, and concatenates them together as input. This dense connection method enables each layer to obtain the feature information of all previous layers, thereby improving the utilization of features. In one-dimensional signal processing, this helps capture the timing features and local patterns in the signal.
[0095] In addition, the densely connected structure of DenseNet (i.e., the idea of increasing convolution kernels layer by layer) can also help alleviate the gradient vanishing problem, which is especially important for training deep networks. In this scheme, since the sequence of base electrical signals may be long, deep networks help capture dependencies in long sequences, and the special structure used by DenseNet in this paper helps to maintain the effective transmission of gradients and ensure that the network can be stably trained.
[0096] In practical applications, the two-dimensional convolution kernel of DenseNet is adjusted to a one-dimensional convolution kernel to adapt to the data characteristics of base electrical signals. This includes modifying the parameter settings of the convolution layer to make it suitable for processing one-dimensional input data. At the same time, in some embodiments, other parts of the network, such as the transition layer, can also be adjusted to ensure the rationality and effectiveness of the network structure.
[0097] In general, applying the DenseNet convolutional neural network to the recognition of one-dimensional base electrical signals is expected to achieve good performance in base sequence calling tasks by modifying the convolution kernel to an n×1 one-dimensional convolution kernel and taking advantage of its densely connected characteristics. This innovative application demonstrates the flexibility and universality of DenseNet and also provides new ideas and methods for the field of one-dimensional signal processing.
[0098] The base calling model in this implementation increases the number of channels and kernel size used in the dense layers to capture a wider range of base position information. In the last dense block, the number of channels reaches 1024 and the kernel size is 99.
[0099] More specifically, each dense block includes several dense layers, each dense layer includes an n×1 one-dimensional convolution kernel; the size of the one-dimensional convolution kernel increases with the number of layers until n=K MAX , and the one-dimensional convolution kernels of the subsequent dense layers are all n=K MAX , n is an odd number. In this embodiment, the third convolution kernel size of the first dense layer of the first dense block (that is, the initial convolution kernel size) is set to the initialized 7×1, and the reference value of the current layer can be calculated by the formula k = kernel_size + i * kernel_size_increment; where k is the reference value, kernel_size is the initial convolution kernel size, i is the number of the current dense layer, and kernel_size_increment is the convolution kernel increment (kernel_size_increment=3 in this embodiment); for example, for the 10th dense layer of the second dense block (the total number of layers i=6+10=16), k is 7+3*16=55, that is, the convolution kernel size is 55×1.
[0100] In the present invention, it is also necessary to ensure that the convolution kernel size is an odd number, so if the final result k is an even number, then n = k + 1. For example, for the 11th dense layer of the second dense block (total number of layers i = 6 + 11 = 17), k is 7 + 3 * 17 = 58, n = 58 + 1 = 59, that is, the convolution kernel size is 59 × 1, and so on.
[0101] In this embodiment, the maximum size of the convolution kernel is limited to prevent the convolution kernel from being too large, which leads to excessive consumption of computing resources and possible performance problems. Specifically, in this embodiment, the maximum size threshold of the convolution kernel is 99, that is, the size of the one-dimensional convolution kernel increases with the number of layers until n=99, and the one-dimensional convolution kernels of subsequent dense layers are all 99.
[0102] In addition, the performance evaluation of the base calling model described in this embodiment includes: 1. Evaluation criteria: Sequence consistency (accuracy = M / (M+S+I+D)) is used for evaluation, where M is the number of matching bases, S is the number of mismatches, I is the number of insertions, and D is the number of deletions.
[0103] 2. Alignment and evaluation: Minimap2 (v2.15, -cs) was used for sequence alignment and accuracy assessment, including quantification of the number of mismatches, insertions, and deletions compared with the reference genome.
[0104] Figure 6-Figure 8 A schematic diagram of the training interface of a base calling model in an exemplary embodiment of the present application is shown.
[0105] In some embodiments, the data set includes a training set, a validation set, and a test set.
[0106] In order to improve the generalization and accuracy of the base calling model, the training set in this embodiment includes data from Arabidopsis thaliana, humans, Caenorhabditis elegans, and Escherichia coli. A total of one million signal blocks are used for training and 100,000 are used for validation.
[0107] In this example, PyTorch (v2.0.1) is used to train the sequencing dataset, with a batch size of 32 and 30 epochs. The test set consists of the following parts: 1) RODAN research dataset: These include humans (Homo sapiens), Arabidopsis thaliana, mice (Musmusculus), Saccharomyces cerevisiae (S. cerevisiae S288C), and Populus trichocarpa.
[0108] 2) Published SARS-CoV-2 datasets.
[0109] 3) Datasets generated by the present invention: including Homo sapiens, Plasmodium berghei (P. berghei), Seneca Valley virus (SVV), porcine epidemic diarrhea virus (PEDV) and porcine reproductive and respiratory syndrome virus (PRRSV).
[0110] On these data sets, the performance of the base calling model provided by the present invention is significantly better than the existing technology, as shown in Table 1 (DEMINERS in Table 1 is the base calling model of the present invention).
[0111] Table 1 Comparison of base calling performance of RNA-seq data using different base callers It can be seen that the base calling model in the present invention has significantly improved the base calling accuracy on various data sets. For example, on the test set, the accuracy is significantly higher than Guppy (a toolkit provided by ONT), and it is also better than RODAN on most species. By evaluating on a data set containing multiple species, it not only performs well in overall accuracy, but also shows lower error rates in mismatch, insertion and deletion rates, such as Figure 2 shown.
[0112] Compared with Guppy and RODAN, the base calling model in the present invention showed higher accuracy and lower error rate in most species, such as Figure 3 For example, among the 10 species in the test set, Densecall's base calling accuracy on 9 species exceeded that of RODAN, showing a significant advantage (see Table 1). In RNA viruses with high mutation rates such as SARS-CoV-2, PRRSV, SVV, and PEDV, the base calling accuracy of the base calling model in the present invention is significantly better than the existing technology, and can more accurately identify and record the genetic elements of these viruses.
[0113] It is worth noting that in order to improve the accuracy of base recognition in a specific species, the base recognition model in this embodiment supports species-specific optimization. By using a dataset of a certain organism to train a base calling model, a species-specific base calling model of the organism is obtained, which can be optimized for the unique signal patterns of different species, significantly improving the base calling performance in the study of specific species.
[0114] For example, the species-specific model was trained using the mouse dataset (Mus musculus) of RODAN. The training set included 20,000 electrical signals, and the validation and test sets each contained 4,000 electrical signals. HDF5 data was generated using Taiyaki (v5.3.0), and the species-specific model was trained using the DEMINERS training function. The training parameters were set to a learning rate of 0.002, a batch size of 32, and 30 epochs of training.
[0115] After specific training, the model's base calling accuracy on the mouse transcriptome increased from 87.84% to 90.44%, exceeding RODAN's 88.16% and Guppy's 84.06%. Figure 4 In addition, although the mapping rates are comparable, the base call length of the mouse-specific pattern in this model is 11 bp longer than that in RODAN, as shown in Figure 2. Figure 5 This species-specific optimization significantly improved the base calling performance in the study of a specific species, indicating that the model is highly adaptable and accurate when processing RNA sequencing data from different species.
[0116] Exemplarily, the present embodiment can read the data file output by the Oxford Nanopore sequencer and extract the current signal corresponding to the DNA / RNA molecule to be tested.
[0117] Specifically, the Oxford Nanopore sequencing method is a third-generation single-molecule real-time sequencing technology based on electrical signals, which can directly read double-stranded DNA / RNA molecules and capture current signals. During the sequencing process, the DNA / RNA double-stranded is first connected to the motor protease, and combined with the nanopore protein embedded in the biomembrane to unwind. The motor protease controls the movement of the DNA / RNA double-stranded through the nanopore. During the displacement process, the ionic current in the nanopore will fluctuate with the movement of the nucleic acid in the pore, thereby capturing the fluctuating current signal and storing it in the data file. By connecting to the data file storing the current signal, the current signal corresponding to the DNA / RNA molecule to be tested in the data file is obtained to carry out the subsequent base sequence recognition process.
[0118] Furthermore, before feeding the data into the model, the training and test data sets can be preprocessed using Taiyaki (Oxford Nanopore Technologies, v5.3.0). Each signal chunk contains 4096 signal values, which are normalized using the median absolute deviation method.
[0119] Splitting long electrical signal sequences into multiple signal blocks can improve processing efficiency, reduce memory usage, and allow parallel processing. 4096 is a power of 2, which is convenient for efficient processing in a computer. In actual use, GPU memory usage can be reduced from 6628MB to 4700MB with a batch size of 32, highlighting the efficiency and effectiveness of this method.
[0120] It is understandable that the data in the dataset can also be preprocessed using the above method when training the model.
[0121] Embodiment 2: The present invention also provides a base recognition system based on a fully convolutional network, comprising: A model building module is used to build a base calling model based on a DenseNet convolutional neural network; the base calling model includes: The first convolutional layer is used to reduce noise on the electrical signal; The maximum pooling layer is used to downsample the electrical signal; Several dense blocks are used to extract features of electrical signals; The transition layer interspersed between dense blocks is used to reduce the dimension of the features output by the dense blocks; The fully connected layer is used to integrate the global features of the feature maps processed by the dense block and transition layer, and perform classification prediction; CTC decoder, used to decode the prediction results output by the fully connected layer to obtain the sequence of bases; wherein the dense block includes multiple dense layers, and the dense layer includes: batch normalization layer, first Silu nonlinear unit layer, 1×1 second convolution layer, batch normalization layer again, second Silu nonlinear unit layer, n×1 third convolution layer; The steps of setting n include: The reference value of the third convolutional layer of the i-th layer is calculated according to the incremental model, wherein the incremental model is: k= kernel_size + i * kernel_size_increment; k is the reference value, kernel_size is the initial convolution kernel size, i is the number of current dense layers, and kernel_size_increment is the convolution kernel increment; When k < set convolution size K MAX , and k is an even number, then n=k+1; When k < set convolution size K MAX , and when k is an odd number, then n=k; When k ≥ set the convolution size K MAX , then n=K MAX ; A first training module is used to train the base calling model using a data set according to current training conditions to obtain a base calling model; wherein the training conditions include: the number of dense layers and the size of the third convolution kernel of the third convolution layer; The second training module is used to determine whether the prediction accuracy of the currently obtained base calling model is greater than a set accuracy threshold; if so, output the corresponding base calling model; if not, update the training conditions and return to the first training module; A data acquisition module, used to collect the electrical signal of the base to be identified; The recognition module is used to input the electrical signal into the base calling model for recognition to obtain the sequence of the base to be recognized.
[0122] It can be understood that the system in this embodiment can also implement the method or steps in any of the above embodiments, which will not be repeated here.
[0123] The present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method or steps in any embodiment of the present invention are implemented.
[0124] The present invention also provides a computer program product, which is tangibly stored on a non-transitory computer-readable medium and includes machine-executable instructions for executing the method or steps in any embodiment of the present invention.
[0125] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0126] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a computer terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0127] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.
Claims
1. A base recognition method based on a fully convolutional network, characterized in that: include: S101, constructing a base calling model based on a DenseNet convolutional neural network, wherein the base calling model sequentially includes: The first convolutional layer is used to reduce noise on the electrical signal; The maximum pooling layer is used to downsample the electrical signal; Several dense blocks are used to extract features of electrical signals; The transition layer interspersed between dense blocks is used to reduce the dimension of the features output by the dense blocks; The fully connected layer is used to integrate the global features of the feature maps processed by the dense block and transition layer, and perform classification prediction; The CTC decoder is used to decode the prediction results output by the fully connected layer to obtain the sequence of bases; wherein the dense block includes multiple dense layers, and the dense layers include: a batch normalization layer, a first Silu nonlinear unit layer, a 1×1 second convolutional layer, a batch normalization layer again, a second Silu nonlinear unit layer, and an n×1 third convolutional layer, wherein the steps of setting n include: The reference value of the third convolutional layer of the i-th layer is calculated according to the incremental model, wherein the incremental model is: k= kernel_size + i * kernel_size_increment; k is the reference value, kernel_size is the initial convolution kernel size, i is the number of current dense layers, and kernel_size_increment is the convolution kernel increment; When k < set convolution size K MAX , and k is an even number, then n=k+1; When k < set convolution size K MAX , and when k is an odd number, then n=k; When k ≥ set the convolution size K MAX , then n=K MAX ; S102, training the base calling model using the data set according to the current training conditions to obtain a base calling model through training and updating; wherein the training conditions include: the number of dense layers, the size of the third convolution kernel of the third convolution layer; S103, determining whether the prediction accuracy of the currently obtained base calling model is greater than a set accuracy threshold; if so, outputting the corresponding base calling model; if not, updating the training conditions and returning to S102; S104, collecting the electrical signal of the base to be identified; S105, inputting the electrical signal into the base calling model for identification to obtain a sequence of the base to be identified.
2. A base recognition method based on a fully convolutional network according to claim 1, characterized in that: The step of updating the training conditions comprises: The number of dense layers of at least one dense block is modified, and the size of the third convolution kernel of each dense layer is updated according to the modified number of dense layers.
3. A base recognition method based on a fully convolutional network according to claim 2, characterized in that: The plurality of dense blocks include: a first dense block, a second dense block, a third dense block and a fourth dense block; correspondingly, the step of updating the training conditions further includes: Calculate the first prediction accuracy of the base calling model obtained in the last training process; Calculating the second prediction accuracy of the base calling model obtained in the current training process; Calculate the accuracy growth rate according to the first prediction accuracy and the second prediction accuracy; When the accuracy increment rate is greater than or equal to a first increment threshold, it is recommended to perform a first modification operation on the number of dense layers of the dense blocks of the first set, wherein the first set includes: a first dense block, a second dense block, a third dense block, and a fourth dense block; When the accuracy increment rate is less than the first increment threshold, it is recommended to perform a second modification operation on the number of dense layers of the dense blocks of the second set, wherein the second set includes: a third dense block and a fourth dense block.
4. A base recognition method based on a fully convolutional network according to claim 2, characterized in that: Also includes the steps: When the accuracy increment rate is still less than the first increment threshold after at least one second modification operation, it is recommended to modify the pooling size of the maximum pooling layer.
5. A base recognition method based on a fully convolutional network according to claim 4, characterized in that: The first modification operation or the second modification operation meets the following conditions: The number of dense layers of the second dense block is less than the first set value; The number of dense layers of the third dense block is greater than the second set value.
6. A base recognition method based on a fully convolutional network according to claim 1, characterized in that: K MAX is 99; and / or, the transition layer includes a 1×1 convolutional layer.
7. A base recognition method based on a fully convolutional network according to claim 1, characterized in that: The data set includes a training set, a validation set and a test set; The data of the training set and the validation set include one or more of the following: base electrical signals of Arabidopsis thaliana, human, Caenorhabditis elegans, and Escherichia coli; The data of the test set include one or more of the following: base electrical signals of human, Arabidopsis, mouse, Saccharomyces cerevisiae, Populus nigra, Plasmodium berghei, Seneca Valley virus, porcine epidemic diarrhea virus, and porcine reproductive and respiratory syndrome virus; And / or, the method further comprises: preprocessing the collected electrical signals of the bases to be identified, comprising: The electrical signal is divided into signal blocks, each of which contains 4096 signal values; the signal values are normalized using the median absolute deviation method.
8. A base recognition system based on a fully convolutional network, characterized in that: include: A model building module is used to build a base calling model based on a DenseNet convolutional neural network, wherein the base calling model includes: The first convolutional layer is used to reduce noise on the electrical signal; The maximum pooling layer is used to downsample the electrical signal; Several dense blocks are used to extract features of electrical signals; The transition layer interspersed between dense blocks is used to reduce the dimension of the features output by the dense blocks; The fully connected layer is used to integrate the global features of the feature maps processed by the dense block and transition layer, and perform classification prediction; CTC decoder, used to decode the prediction results output by the fully connected layer to obtain the sequence of bases; wherein the dense block includes multiple dense layers, and the dense layer includes: batch normalization layer, first Silu nonlinear unit layer, 1×1 second convolution layer, batch normalization layer again, second Silu nonlinear unit layer, n×1 third convolution layer; The steps of setting n include: The reference value of the third convolutional layer of the i-th layer is calculated according to the incremental model, wherein the incremental model is: k= kernel_size + i * kernel_size_increment; k is the reference value, kernel_size is the initial convolution kernel size, i is the number of current dense layers, and kernel_size_increment is the convolution kernel increment; When k < set convolution size K MAX , and k is an even number, then n=k+1; When k < set convolution size K MAX , and when k is an odd number, then n=k; When k ≥ set the convolution size K MAX , then n=K MAX ; A first training module is used to train the base calling model using a data set according to current training conditions to train and update the base calling model; wherein the training conditions include: the number of dense layers and the size of the third convolution kernel of the third convolution layer; The second training module is used to determine whether the prediction accuracy of the currently obtained base calling model is greater than a set accuracy threshold; if so, output the corresponding base calling model; if not, update the training conditions and return to the first training module; A data acquisition module, used to collect the electrical signal of the base to be identified; The recognition module is used to input the electrical signal into the base calling model for recognition to obtain the sequence of the base to be recognized.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer program product tangibly stored on a non-transitory computer readable medium and comprising machine executable instructions, characterized in that, The machine executable instructions are used to perform the method according to any one of claims 1-7.
Citation Information
Patent Citations
Base sequence identification method and device and storage medium
CN111243674A
Method for quickly identifying single-molecule nanopore sequencing bases based on deep network
CN112183486A
Remote sensing image ground object identification method and system, and computer readable storage medium
CN112560544A
Damage positioning method and device based on DenseNet network and storage medium
CN114331960A
Convolution with kernel extension and tensor accumulation
CN117413280A