Nucleic acid non-canonical base ultra-high precision detection and positioning method based on third-generation sequencing

By designing random context core training sequences and deep learning models, and combining them with signal redistribution technology, high-precision detection and localization of non-classical bases in nanopore sequencing technology has been achieved. This solves the problem of insufficient model generalization ability in existing technologies and provides a general and scalable tool for analyzing nucleic acid chemical modifications and variants.

CN122435993APending Publication Date: 2026-07-21SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUN YAT SEN UNIV
Filing Date
2026-03-19
Publication Date
2026-07-21

Smart Images

  • Figure CN122435993A_ABST
    Figure CN122435993A_ABST
Patent Text Reader

Abstract

The application discloses a nucleic acid non-canonical base ultra-high precision detection and positioning method based on third-generation sequencing, and belongs to the technical field of gene sequencing. The method first rationally designs and synthesizes standard training data in vitro, and adopts the structure of a random base region-target base-random base region in the core region to provide diversified sequence context background. The data is used for third-generation sequencing of DNA or RNA molecules, the original current signal features and local sequence information corresponding to the target base site are accurately extracted, and the deep learning classification model is trained by taking the features and the information as input. Finally, the trained model is used to realize high-precision identification and positioning of various non-canonical bases (including epigenetic modification, damaged base and artificial pseudo base) at a single molecule level in unknown samples. The method breaks through the limitation of base types in existing methods, has the advantages of strong universality, high precision and good scalability, and provides a universal scheme for nucleic acid modification research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of gene sequencing technology, specifically relating to a method for ultra-high precision detection and localization of non-classical bases in nucleic acids based on third-generation sequencing. Background Technology

[0002] Nanopore sequencing is a third-generation sequencing method that uses changes in electrical current signals to analyze nucleic acid sequences. Its basic principle is that as a single nucleic acid molecule passes through a nanopore protein, different base combinations influence the ion current, forming a time-series signal that can be used to identify base types. Compared to second-generation sequencing technologies that rely on chemical amplification and fluorescent labeling, nanopore sequencing can directly read raw molecular information, offering significant advantages such as longer read lengths, real-time performance, and the ability to simultaneously acquire sequence and modification information. Therefore, it shows unique potential in fields such as epigenetic modification detection, structural variation analysis, and single-molecule dynamics studies. Since base modifications typically introduce distinguishable subtle changes in electrical current signals, nanopore sequencing has been considered an ideal technical approach for directly identifying base modifications at the single-molecule level since its inception. Existing research shows that by performing deep learning analysis on the raw electrical current signals generated during sequencing, the presence of specific modified bases can be inferred without additional chemical processing.

[0003] In current research and applications, nanopore sequencing primarily focuses on detecting a few classic epigenetic modification bases in DNA. Among these, 5-methylcytosine (5mC) and N6-methyladenine (6mA) are the earliest and most extensively studied targets due to their widespread presence in eukaryotes and prokaryotes and their crucial roles in gene expression regulation, development, and disease pathogenesis. Various signal analysis methods based on Hidden Markov Models, deep neural networks, or mixed statistical models have been proposed for these two types of modifications, demonstrating strong detection performance. However, with continuous iterations in sequencing chip structures, current sampling frequencies, and chemical systems, existing methods still face numerous challenges in terms of model generalization ability, cross-platform applicability, and robustness to complex background signals. In recent years, the main development direction of related technologies has focused more on adapting to next-generation nanopore chip architectures, improving the sensitivity and specificity of 5mC detection, and reducing the false positive rate under low coverage conditions, while systematic expansion to other types of base modifications has been relatively limited. At the commercial software level, base identification and modification detection in nanopore sequencing primarily rely on officially provided base identification programs. Taking the latest version of Dorado as an example, while its built-in modification detection function has improved in algorithm efficiency and computational performance, the newly added detection targets are still mainly limited to modifications similar to 5-hydroxymethylcytosine (5hmC) and 4-methylcytosine (4mC), which are structurally similar to 5mC. For other types of base chemical variants, especially modifications with significant structural differences or low frequency of occurrence, there is still a lack of mature, stable, and universal identification schemes. In fact, the types of non-classical bases in DNA molecules far exceed the aforementioned classic epigenetic modifications. On the one hand, cells are continuously affected by endogenous oxidative stress, deamination reactions, and replication errors during normal metabolism, resulting in various forms of base damage in the genome, such as oxidized bases, alkylated bases, and exocyclic adducts. On the other hand, exogenous factors such as ultraviolet radiation, ionizing radiation, and environmental chemicals can also introduce structurally diverse damaged bases, which play an important role in genome stability, mutation accumulation, and disease development. Furthermore, in emerging application fields such as synthetic biology, molecular markers, and nucleic acid information storage, artificially designed and introduced pseudo-bases are receiving increasing attention, and their precise identification at the molecular level is crucial for functional validation and quality control.

[0004] However, the detection of these non-classical bases currently relies mainly on chemical derivatization, enzyme digestion enrichment, or mass spectrometry analysis, which generally have limitations such as complex operation procedures, large sample requirements, and difficulty in achieving single-molecule localization. There is still a lack of a technical solution based on nanopore sequencing that is universal and can achieve high-precision single-molecule recognition. Summary of the Invention

[0005] To address the aforementioned technical challenges, this invention proposes a method for ultra-high precision detection and localization of non-classical bases in nucleic acids based on third-generation sequencing, thereby obtaining a universal single-molecule detection technology that can adapt to various types of non-classical bases and possesses high precision and strong generalization capabilities.

[0006] The present invention provides a method for ultra-high precision detection and localization of non-classical bases in nucleic acids based on third-generation sequencing, comprising the following steps: S1. Constructing standard training data: Designing and synthesizing nucleic acid sequences, wherein the sequences contain a core region, and the structure of the core region is: a first random base region—a target base site—a second random base region, wherein the target base site is a non-classical base to be detected or the corresponding unmodified classical base; S2. Obtain sequencing signal: Perform third-generation sequencing on the nucleic acid sequence obtained in S1 to obtain raw current signal data containing the target base sites; S3. Extract feature signals: Process the raw current signal data, accurately locate the target base site, and extract the raw current signal within a preset length window centered on the site, as well as the local sequence context information corresponding to the window. S4. Training the detection model: Using the original current signal and local sequence context information extracted in S3 as input features, and the base type of the target base site as a label, a deep learning classification model is trained to obtain a trained non-classical base detection model. S5. Detection and localization: Using the trained non-classical base detection model, analyze the third-generation sequencing data of unknown samples to predict and localize the presence and type of non-classical bases.

[0007] The core region in S1 is flanked by fixed auxiliary sequences; and the oligonucleotide sequences used to construct the standard training data are selected from SEQ ID NO. 1 to SEQ ID NO. 6, or variants thereof that have at least 90% sequence identity and retain the function of the core region.

[0008] The first random base region and the second random base region in S1 are each composed of no less than 4 random classical bases; the nucleic acid is DNA or RNA.

[0009] The non-classical bases in S1 include at least two of the following: epigenetically modified bases, damaged bases, and artificially synthesized pseudo-bases.

[0010] The third-generation sequencing in S2 is nanopore sequencing.

[0011] The precise localization in S3 includes: identifying the bases in the original current signal to obtain the sequencing sequence, and aligning it to the designed sequence to determine the target base site; the extraction of feature signals includes obtaining the precisely aligned current signal based on signal redistribution technology.

[0012] The preset length window is m signal points upstream and downstream of the target base site for DNA, and n signal points upstream and downstream for RNA, where n is 2-4 times m and m is an integer greater than or equal to 10; the local sequence context information is the k-mer base sequence corresponding to each signal point, where k is an integer greater than or equal to 5.

[0013] The deep learning classification model in S4 includes convolutional neural network layers and / or Transformer layers.

[0014] The present invention also provides a DNA oligonucleotide pair for constructing the above-mentioned standard training data, wherein the oligonucleotide pair is selected from any one of the following groups: (a) A first chain containing the sequence shown in SEQ ID NO. 1 and a second chain containing the sequence shown in SEQ ID NO. 2; (b) A first chain containing the sequence shown in SEQ ID NO. 3 and a second chain containing the sequence shown in SEQ ID NO. 4.

[0015] Compared with the prior art, the beneficial effects of the present invention are: By designing core training sequences incorporating random context, the model learns to focus on the signal characteristics of the target base itself, rather than the fixed sequence background. This allows it to adapt to a variety of structurally diverse non-classical bases in DNA and RNA, including classically modified, damaged, and artificially pseudo-bases. This method utilizes a deep learning model to extract subtle differences from raw current signals, combined with high-precision signal redistribution techniques, to achieve reliable identification and localization of non-classical bases at the single-molecule level. When detecting new types of non-classical bases, the detection capability can be rapidly expanded simply by synthesizing standard training data containing that base and retraining or fine-tuning the model, without needing to redevelop the entire algorithm. Therefore, this invention provides a universal and scalable tool for analyzing nucleic acid chemical modifications and variants in fields such as epigenetics, genome stability research, environmental toxicology, synthetic biology, and nucleic acid information storage. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments are briefly introduced below.

[0017] Figure 1 This is the overall flowchart of the method (OpenBase) of this invention.

[0018] Figure 2 This is an experimental flowchart for the construction and sequencing of the standard training samples of this invention.

[0019] Figure 3 This is a confusion matrix diagram of the OpenBase detection performance for various non-classical DNA bases in Example 1.

[0020] Figure 4 The graph shows the ROC and PR curves of OpenBase on the official standard dataset in Example 1.

[0021] Figure 5 This is a scatter plot showing the consistency of OpenBase and BS-seq in 5mC detection in Example 1.

[0022] Figure 6 This is a predicted distribution map of 6mA modification in the GATC motif of different bacteria using OpenBase in Example 1.

[0023] Figure 7 This is a confusion matrix diagram of the performance of OpenBase in detecting non-classical RNA bases in Example 2.

[0024] Figure 8 The graph shows the ROC and PR curves of OpenBase on the official RNA standard dataset in Example 2. Detailed Implementation

[0025] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0026] The overall flow of the method of the present invention is as follows: Figure 1 As shown.

[0027] Example 1: Application of OpenBase in DNA Non-classical Base Detection according to Figure 2 The experimental procedure shown first involves constructing standard DNA training data. A DNA template with a core region of "16 random N - target base - 16 random N" is designed and synthesized.

[0028] Specifically, an unmodified control dataset was constructed using oligonucleotide pairs as shown in SEQ ID NO. 1 and SEQ ID NO. 2; a modified training dataset was constructed using oligonucleotide pairs as shown in SEQ ID NO. 3 (containing 5mC) and SEQ ID NO. 4 (containing 6mA). The target base types can be expanded to include various non-classical bases such as 5hmC, 8-oxodG, and BrdU, as well as their corresponding unmodified bases. Double strands were formed through PCR extension, libraries were constructed using the ONT ligation sequencing kit, and data were acquired on an ONT third-generation sequencer.

[0029] After Dorado base identification and Minimap2 alignment to the design sequence, Remora was used for signal redistribution, extracting 100 original signal points upstream and downstream of the target base and their corresponding 9-mer sequences as features. These features were then used to train a deep learning model containing convolutional layers and Transformers.

[0030] On a test set containing 10 classes of non-classical DNA bases, the model achieved an overall recognition accuracy of over 95%, with its confusion matrix as shown below. Figure 3 As shown. The ROC and PR curves were evaluated on the standard single-molecule datasets at 5mC and 6mA provided by ONT. Figure 4 The results showed that the AUC was greater than 0.99. Further validation using real human genome data showed that OpenBase's prediction of 5mC was highly consistent with the BS-seq results. Figure 5 ).

[0031] In bacterial 6mA assays, OpenBase accurately distinguishes modification specificity: only in strains containing the GATC methylation system does this site show a near 100% modification rate, as shown in the results. Figure 6 As shown.

[0032] Example 2: Application of OpenBase in RNA Non-classical Base Detection according to Figure 2 The RNA library construction process involves synthesizing a training RNA sequence containing a core region of "12 random rNs — target base — 12 random rNs". For example, sequences such as SEQ ID NO. 5 (containing m6A) or SEQ ID NO. 6 (containing m5C and m6A) can be used. The library is constructed and sequenced using the ONT direct RNA sequencing kit. In signal processing, 200 signal points upstream and downstream of the target base are extracted as features.

[0033] The trained model exhibits high accuracy on the test set including m6A and m5C, and its confusion matrix is ​​as follows: Figure 7As shown. Validation was performed using the official ONT RNA modification standard data, and the ROC and PR curves are as follows. Figure 8 As shown, this method also demonstrates high accuracy and stability in RNA modification detection.

[0034] The above embodiments demonstrate that the OpenBase method provided by the present invention can systematically and reliably achieve single-molecule level detection and localization of various non-classical bases in DNA and RNA.

[0035] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not describe all details exhaustively, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification.

Claims

1. A method for ultra-high precision detection and localization of non-classical bases in nucleic acids based on third-generation sequencing, characterized in that, Includes the following steps: S1. Constructing standard training data: Designing and synthesizing nucleic acid sequences, wherein the sequences contain a core region, and the structure of the core region is: a first random base region—a target base site—a second random base region, wherein the target base site is a non-classical base to be detected or the corresponding unmodified classical base; S2. Obtain sequencing signal: Perform third-generation sequencing on the nucleic acid sequence obtained in S1 to obtain raw current signal data containing the target base sites; S3. Extract feature signals: Process the raw current signal data, accurately locate the target base site, and extract the raw current signal within a preset length window centered on the site, as well as the local sequence context information corresponding to the window. S4. Training the detection model: Using the original current signal and local sequence context information extracted in S3 as input features, and the base type of the target base site as a label, a deep learning classification model is trained to obtain a trained non-classical base detection model. S5. Detection and localization: Using the trained non-classical base detection model, analyze the third-generation sequencing data of unknown samples to predict and localize the presence and type of non-classical bases.

2. The method for ultra-high precision detection and localization of non-classical bases in nucleic acids based on third-generation sequencing according to claim 1, characterized in that, The core region in S1 is flanked by fixed auxiliary sequences; and the oligonucleotide sequences used to construct the standard training data are selected from SEQ ID NO. 1 to SEQ ID NO. 6, or variants thereof that have at least 90% sequence identity and retain the function of the core region.

3. The method for ultra-high precision detection and localization of non-classical bases in nucleic acids based on third-generation sequencing according to claim 2, characterized in that, The first random base region and the second random base region in S1 are each composed of no less than 4 random classical bases; the nucleic acid is DNA or RNA.

4. The method for ultra-high precision detection and localization of non-classical bases in nucleic acids based on third-generation sequencing according to claim 3, characterized in that, The non-classical bases in S1 include at least two of the following: epigenetically modified bases, damaged bases, and artificially synthesized pseudo-bases.

5. The method for ultra-high precision detection and localization of non-classical bases in nucleic acids based on third-generation sequencing according to claim 1, characterized in that, The third-generation sequencing in S2 is nanopore sequencing.

6. The method for ultra-high precision detection and localization of non-classical bases in nucleic acids based on third-generation sequencing according to claim 1, characterized in that, The precise localization in S3 includes: identifying the bases in the original current signal to obtain the sequencing sequence, and aligning it to the designed sequence to determine the target base site; the extraction of feature signals includes obtaining the precisely aligned current signal based on signal redistribution technology.

7. The method for ultra-high precision detection and localization of non-classical bases in nucleic acids based on third-generation sequencing according to claim 6, characterized in that, The preset length window is m signal points upstream and downstream of the target base site for DNA, and n signal points upstream and downstream for RNA, where n is 2-4 times m and m is an integer greater than or equal to 10; the local sequence context information is the k-mer base sequence corresponding to each signal point, where k is an integer greater than or equal to 5.

8. The method for ultra-high precision detection and localization of non-classical bases in nucleic acids based on third-generation sequencing according to claim 1, characterized in that, The deep learning classification model in S4 includes convolutional neural network layers and / or Transformer layers.

9. A DNA oligonucleotide pair for constructing the standard training data of claim 1, characterized in that, The oligonucleotide pairs are selected from any one of the following groups: (a) A first chain containing the sequence shown in SEQ ID NO. 1 and a second chain containing the sequence shown in SEQ ID NO. 2; (b) A first chain containing the sequence shown in SEQ ID NO. 3 and a second chain containing the sequence shown in SEQ ID NO. 4.