Methods for detecting minority gene variants
The method enhances the detection of rare genetic variants in biological samples by using CRISPR-Cas9 and rolling circle amplification to enrich and identify specific sequences, overcoming the limitations of existing techniques in sensitivity and cost.
Patent Information
- Application Number
- JP2025538618
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-29
- Filing Date
- 2023-12-29
- Publication Date
- 2025-12-25
AI Technical Summary
Existing methods for detecting rare genetic variants in biological samples, such as cell-free DNA, require ultra-deep sequencing due to the low concentration and high fragmentation of cfDNA, leading to high costs and long analysis times, and lack specificity and sensitivity, especially when competing sequences are present.
A method involving nucleic acid isolation, circularization, and CRISPR-Cas9-mediated selective amplification using rolling circle amplification (RCA) to enrich and identify specific nucleic acid sequences, allowing for the detection of rare genetic variants without the need for ultra-deep sequencing.
Enables reliable and efficient detection of rare genetic variants with high specificity and reduced analysis time and cost, facilitating non-invasive monitoring of conditions like cancer and prenatal diagnosis.
Smart Images

Figure 2025542506000001 
Figure 2025542506000002 
Figure 2025542506000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for detecting a first nucleic acid sequence in a biological sample, wherein the biological sample contains a second nucleic acid sequence that may compete for detection of such first nucleic acid sequence. [Background technology]
[0002] The presence of cell-free DNA (cfDNA) was first identified in human peripheral blood samples in 1948 (Mandel and Metais, (1948) "Les acids nucleiques du plasma sanguin chez l'homme" Biologie 3-4:241-243). Since its identification, cell-free DNA has rapidly become the focus of intense interest in the search for clinically applicable biomarkers and in a variety of different research fields.
[0003] One of the greatest strengths of the application of cell-free DNA as a biomarker relates to the possibility of analyzing it using non-invasive sampling, e.g., a blood sample.
[0004] Cell-free DNA can arise from a variety of different mechanisms, for example, it can be released by cells as a mechanism for immune system communication and regulation (Korabecna et al. (2020) Cell-free DNA in plasma as an essential immune system regulator. Ski Rep vol. 10, 17478 DOI: https: / / DOI.org / 10.1038 / s41598-020-74288-2); as a result of cellular apoptosis and necrosis (Ranucci R (2019) Cell-Free DNA: Applications in Different Diseases. Methods Mol Biol. 1909, pp. 3–12 DOI: 10.1007 / 978-1-4939-8973-7_1); in the maternal bloodstream, from the fetus during embryonic development (Ranucci (2019); Bronkhorst et al. (2022) New Perspectives on the Importance of Cell-Free DNA Biology. Diagnostics, Vol. 12(9):2147 DOI: 10.3390 / diagnostics12092147); from pathogens during infection (Blauwkamp et al., (2019) Analytical and clinical validation of a microbial cell-free DNA sequencing test for infectious disease. Nat Microbiol, Vol. 4, pp. 663–674 DOI: https: / / DOI.org / 10.1038 / s41564-018-0349-6); from the human microbiota, as well as from assimilated environmental DNA (Bronkhorst et al., 2022).
[0005] Cell-free DNA is often highly fragmented, with fragment sizes of approximately 160-200 bp, and is present not only in a patient's peripheral blood, but also in a variety of different types of biological samples, such as saliva, sputum, urine, feces, cerebrospinal fluid, and in samples commonly referred to as "liquid biopsies."
[0006] The diversity of possible cell-free DNA derivatives allows its wide application through the selection of specific markers, ranging from oncology to prenatal diagnosis, from pathogen identification to transplant organ monitoring, from the study and management of cardiovascular and chronic diseases to the detection of assimilable environmental DNA to control the presence of pollutants (Bronkhorst et al., 2022).
[0007] In addition to the diversity of possible sources of cell-free DNA, another advantage of its application is related to its half-life, which is generally limited over time. This aspect allows for noninvasive and continuous monitoring of the patient's condition over time, especially in clinical settings. For example, the application of circulating tumor DNA (ctDNA) is well known and widespread in the diagnosis, prognosis, and treatment selection / monitoring stages of cancer patients.
[0008] Generally, the amount of cell-free DNA in the bloodstream is very low in healthy individuals. However, the concentration of cfDNA in blood samples tends to increase during tumor development, especially in the advanced stage, in maternal blood during pregnancy, and in the presence of cardiovascular and chronic diseases.
[0009] Along with cell-free DNA, in recent years there has been particular interest in the study and identification of biomarkers in DNA present in extracellular vesicles (EVs), especially in oncology.
[0010] It has been observed that EVs are particularly stable in plasma, and they can be considered promising markers, for example, in the early stage diagnosis of tumor-bearing patients (Fernando et al., WL (2017) New evidence that a large proportion of human blood plasma cell-free DNA is localized in exosomes. PLoS ONE Vol. 12(No. 8): e0183915 DOI: https: / / DOI.org / 10.1371 / journal.pone.0183915).
[0011] Furthermore, EVs have been shown to be important biomarkers for prognosis. Indeed, studies of EV contents can be very useful in the early identification of the formation of premetastatic niches, as tumor cells have shown a higher activity in secreting exosomes with altered contents (Chin and Wang, (2016) Cancer Tills the Premetastatic Field: Mechanistic Basis and Clinical Implications. Clin Cancer Res. 2016 Aug 1;22(15):3725-33 DOI: 10.1158 / 1078-0432.CCR-16-0028).
[0012] In the last few years, various sensitive and specific methods have been developed to detect cfDNA and its specific biomarkers, such as the "BEAMing Safe-Sequencing System" (BEAMing Safe-SeqS), "Tagged-Amplicon deep Sequencing" (TamSeq), "Cancer Individual Profiling by deep Sequencing" (CAPP-Seq), digital PCR to detect single-base mutations in cfDNA, whole-genome sequencing (WGS) to establish copy number alterations, or the application of CRISPR-Cas systems to cleave targets of interest in a highly specific manner.
[0013] However, the above-mentioned techniques require prior knowledge of the sequence of interest, along with the need for ultra-deep sequencing, in order to be able to identify the first sequence of interest with statistical significance and distinguish it from the second sequence. The need for ultra-deep sequencing limits the repetition of analysis on samples from the same source, for example, the same patient, due to the long analysis time and high cost.
[0014] In principle, technologies can be divided into targeted approaches, which aim to detect mutations in a defined set of genes (e.g., mutations in the EGFR gene are important for the response of patients with non-small cell lung cancer (NSCLC) to tyrosine kinase inhibitor (TKI) blocking), or non-targeted approaches, which require screening of the entire genome (e.g., comparative genomic hybridization on arrays, WGS, or exome sequencing). Targeted approaches usually have higher analytical sensitivity than non-targeted approaches, although considerable efforts are underway to improve the detection limits of the latter.
[0015] Such limitations are primarily caused by the very nature of cfDNA (short half-life, low concentration, and high fragmentation level), which requires significant time and highly optimized workflows, plus ultra-deep sequencing, for meaningful analysis of markers of interest.
[0016] Considering the foregoing, it is necessary to improve the reliability of cfDNA analysis techniques in terms of sensitivity and specificity, as well as to identify protocols that could become standard and more accessible in terms of both applicability and time and cost. Summary of the Invention [Problem to be solved by the invention]
[0017] It is a goal of the present invention to provide a method for detecting rare genetic variants in biological samples that overcomes the limitations of known techniques.
[0018] Within this goal, it is an object of the present invention to provide a method that allows for the detection of specific rare genetic variants in the context of non-mutated sequences, thus reducing the need for ultra-deep sequencing relative to known techniques.
[0019] Another object of the present invention is to provide a method that allows for the detection of exogenous DNA sequences in biological samples.
[0020] Another object of the present invention is to provide a method that allows for the detection of DNA sequences of fetal or embryonic origin in a biological sample.
[0021] Another object of the present invention is to provide a method that is highly reliable and relatively easy to implement. [Means for solving the problem]
[0022] This goal and these and other objects that will become more apparent hereinafter are directed to a method for identifying a first nucleic acid sequence in a biological sample, the method comprising: (i) isolating nucleic acid from a biological sample, wherein the nucleic acid (a) a first nucleic acid sequence, and (b) a second nucleic acid sequence including, steps; (ia) optionally, reverse transcribing RNA present in the nucleic acid; (ii) optionally fragmenting the DNA obtained in step (i) or step (ia) to obtain DNA fragments having lengths comprised between 100 base pairs and 10,000 base pairs; (iii) circularizing the DNA fragments present in the DNA isolated in step (i) or the DNA obtained in step (ii) to obtain a mixture of circularized DNAs comprising circular DNAs containing the "first sequence" and circular DNAs containing the "second sequence"; (iv) contacting the mixture of circularized DNA obtained in step (iii) with a CRISPR system characterized by an sgRNA selected from an sgRNA complementary to the "second sequence" of DNA and an sgRNA containing 1 to 3 nucleotides that are not complementary to the "second sequence" of DNA, thereby obtaining a mixture of DNA containing circular DNA containing the "first sequence" and linear DNA containing the "second sequence"; (v) selectively amplifying the circular DNA obtained in step (iv) by rolling circle amplification (RCA), thereby obtaining an amplification product containing multiple copies of the first DNA sequence; (vi) identifying a first DNA sequence in the amplification product obtained in step (v); This is achieved by a method comprising: DETAILED DESCRIPTION OF THE INVENTION
[0023] In the present invention, the term "first sequence" means a gene sequence selected from the group consisting of a gene sequence of tumor origin, a gene sequence of fetal origin, a gene sequence of embryonic origin, a gene sequence of bacterial origin, a gene sequence of viral origin, a gene sequence of plant origin, and a gene sequence of animal origin.
[0024] In the present invention, the term "second sequence" refers to any gene sequence that may compete in the detection of the marker of interest (the "first sequence").
[0025] The methods of the present invention require the isolation of nucleic acids from biological samples by methods known to those skilled in the art, such as phenol-chloroform nucleic acid extraction, nucleic acid extraction using magnetic beads, and nucleic acid extraction using silica membrane columns.
[0026] The biological sample is selected from the group consisting of blood, saliva, sputum, urine, feces, cerebrospinal fluid, and liquid biopsy.
[0027] Some applications require cell-free biological samples. In such cases, if cell samples are used, the method includes a step of removing cellular parts. For example, if the method is applied to detect cfDNA of tumor origin, the presence of cells will result in excessive presence of "second sequence" DNA, thus impairing the sensitivity and specificity of the method.
[0028] In a preferred embodiment, the "first sequence" is a DNA sequence. Isolated nucleic acids often contain a sequence of interest ("first sequence") in the presence of a significant excess of one or more sequences ("second sequence") that may compete for the detection of the "first sequence". For example, for liquid biopsy in oncology, the "first sequence" is the mutant DNA of tumor origin present in the sample together with an excess of non-mutated DNA ("second sequence"). In cases where contamination is detected, the "first sequence" is exogenous DNA present in the sample together with an excess of endogenous DNA ("second sequence"). In the field of prenatal diagnosis, the "first sequence" is fetal or embryonic DNA in a sample that is mainly composed of maternal DNA.
[0029] The circulating DNA extracted in step (i) of the method is generally sufficiently fragmented that it can be circularized as described in step (iii) of the method of the invention.
[0030] In an embodiment, particularly when DNA is extracted from EVs, the method of the present invention further comprises an optional step of (ii) fragmenting the DNA obtained in step (i) or step (ia) to obtain DNA fragments having lengths between 100 and 10,000 base pairs. Such fragmentation can be achieved by chemical or physical methods known to those skilled in the art, such as sonication and tagmentation.
[0031] The method of the present invention then includes (iii) circularizing DNA fragments present in the DNA isolated in step (i) or the DNA obtained in step (ii) to obtain a mixture of circularized DNAs containing circular DNAs containing the "first sequence" and circular DNAs containing the "second sequence." Circularization of fragments is performed by methods known to those skilled in the art ("Application of Circular Ligase to Provide Template for Rolling Circle Amplification of Low Amounts of Fragmented DNA." Ada N. Nunez, MSFS, Mark F. Kavlick, BS, James M. Robertson, PHD, CFSRU, and Bruce Budowle, PHD, FBI Laboratory, FBI Academy, Quantico, VA 22135).
[0032] In an embodiment, the cyclization in step (iii) is as follows: a. repairing the ends of the DNA fragments; b. adding nucleotides to the 3' end of each fragment; c. c1. A nucleotide complementary to the nucleotide added to one end of the adapter in step b.; c2. Restriction enzyme recognition sequence "α"; c3. A sequence "β" between 15 and 35 base pairs, characterized by the absence of secondary structures or homodimer formation; c4. Any sequence "γ" containing 150 to 1000 base pairs adding an adapter comprising: d. Add ligase and ligate accordingly. - a circular DNA comprising an adapter and a "first sequence"; and - Circular DNA containing the adapter and the "second sequence" obtaining a mixture of circularized DNA comprising is obtained using
[0033] The function of the adapter "γ" sequence is to help promote circularization by reducing conformational entropy. It is a synthetic sequence designed not to hybridize with the "first sequence," "second sequence," or sgRNA of the CRISPR system used under the reaction conditions.
[0034] In step (iv), the circularized DNA mixture obtained in step (iii) is contacted with a CRISPR system characterized by an sgRNA that is substantially complementary to the "second sequence" of the DNA. This system specifically cleaves the "second sequence" and linearizes it accordingly, while leaving the circular DNA containing the "first sequence" intact. Therefore, for the specificity of the method, it is very important to minimize the hybridization of the sgRNA to the "first sequence." This can be achieved using a variety of different approaches known to those skilled in the art. For example, specificity can be optimized by performing the reaction under low salt conditions and / or by using a CRISPR system characterized by an sgRNA designed to maximize discrimination during hybridization. In a preferred embodiment, particularly when the first and second sequences differ by only a single base, specificity can be optimized by using a CRISPR system characterized by a "destabilizing" sgRNA, i.e., an sgRNA containing 1-3 bases, preferably selected from CG residues, that are not complementary to the second sequence.
[0035] In a preferred embodiment, the CRISPR system is a CRISPR-Cas9 system.
[0036] The high specificity of the CRISPR system allows for simultaneous degradation of multiple sequences in a multiplexed manner using several CRISPR systems, each characterized by a complementary sgRNA. In this way, the method of the present invention allows for the simultaneous linearization of several sequences ("second sequences") that may compete with one or more sequences of interest ("first sequences").
[0037] In an embodiment, the "second sequence" is selected from the group consisting of SEQ ID NO:1 to SEQ ID NO:21023.
[0038] The method of the present invention advantageously combines CRISPR technology for linearizing the second sequence with (v) selectively amplifying the circular DNA obtained in step (iv) by rolling circle amplification, thus obtaining an amplification product comprising multiple copies of the first DNA sequence; the combination of steps (iv) and (v) results in highly selective amplification of one or more markers of interest ("first sequences"). The method then requires (vi) identifying the first DNA sequence in the amplification product obtained in step (v).
[0039] Advantageously, selective amplification of the non-linearized sequence in step (iv) using RCA can be obtained with a single primer complementary to the sequence "β" introduced into the circular template in step (iii) of the method. Thus, the method allows multiple "first" sequences to be interrogated in a multiplexed manner without the need to use specific primers for the amplification step.
[0040] The identification in step (vi) can be carried out using a variety of methods, for example, real-time PCR, Illumina sequencing, or nanopore sequencing.
[0041] In a preferred embodiment, the identification of the first sequence in step (vi) is obtained by sequencing the amplification product obtained in step (v).
[0042] Advantageously over known techniques, due to the selective amplification of the marker of interest, identification does not require ultra-deep sequencing.
[0043] The DNA extracted in step (i) and optionally fragmented in step (ii) can be subjected to one or more steps of enrichment for one or more gene sequences of interest before proceeding with circularization in step (iii).
[0044] In a preferred embodiment, the method of the present invention further comprises, prior to step (iii), a step of (ii)a capturing one or both of the "first sequence" and the "second sequence" using a probe complementary thereto.
[0045] Alternatively or additionally, in another embodiment, the method of the present invention comprises, prior to step (iii), a step (ii)b of amplifying one or both of the "first sequence" and the "second sequence." The amplification in step (ii)b is preferably obtained by PCR.
[0046] Therefore, the method of the present invention allows to detect and identify rare genetic variants in biological samples.Advantageously, this method does not require knowledge of the sequence of the variant of interest.The combination of the CRISPR system and RCA amplification, which is designed to cut potentially conflicting sequences, allows rare variants to be amplified with exquisite specificity, which simplifies the procedure compared to currently used methods, and reduces the need for ultra-deep sequencing to identify variants.
[0047] The disclosures in Italian Patent Application No. 102022000027141 from which this application claims priority are incorporated herein by reference.
Claims
1. 1. A method for identifying a first nucleic acid sequence in a biological sample, comprising: (i) isolating nucleic acid from said biological sample, wherein said nucleic acid is (a) a first nucleic acid sequence, and (b) a second nucleic acid sequence the steps of: (ia) optionally, reverse transcribing RNA present in said nucleic acid; (ii) optionally fragmenting the DNA obtained in step (i) or step (ia) to obtain DNA fragments having lengths comprised between 100 base pairs and 10,000 base pairs; (iii) circularizing the DNA fragments present in the DNA isolated in step (i) or the DNA obtained in step (ii) to obtain a mixture of circularized DNAs comprising circular DNAs containing the "first sequence" and circular DNAs containing the "second sequence"; (iv) contacting the mixture of circularized DNA obtained in step (iii) with a CRISPR system characterized by an sgRNA selected from an sgRNA that is complementary to the "second sequence" of DNA and an sgRNA that comprises 1 to 3 nucleotides that are not complementary to the "second sequence" of DNA, thereby obtaining a mixture of DNA comprising circular DNA comprising the "first sequence" and linear DNA comprising the "second sequence"; (v) selectively amplifying the circular DNA obtained in step (iv) by rolling circle amplification, thereby obtaining an amplification product comprising multiple copies of the first DNA sequence; (vi) identifying the first DNA sequence in the amplification product obtained in step (v); A method comprising:
2. The method of claim 1 , wherein the biological sample is selected from a cell-free biological sample and a biological sample from which cellular portions have been removed.
3. 3. The method of claim 1 or 2, wherein the identification of the first sequence in step (vi) is obtained by sequencing the amplification product obtained in step (v).
4. 10. The method of any one of the preceding claims, wherein said "first sequence" is a DNA sequence.
5. 10. The method of any one of the preceding claims, further comprising, prior to step (iii), (ii) capturing one or both of the "first sequence" and the "second sequence" using a probe complementary thereto.
6. 10. The method of any one of the preceding claims, further comprising, prior to step (iii), (ii) amplifying one or both of said "first sequence" and said "second sequence".
7. The method of claim 6, wherein the amplification is obtained by PCR.
8. Step (iii) is: a. repairing the ends of the DNA fragments; b. adding nucleotides to the 3' end of each fragment; c. c1. A nucleotide complementary to the nucleotide added to one end of the adapter in step b.; c2. Restriction enzyme recognition sequence "α"; c3. A sequence "β" comprised between 15 and 35 base pairs, characterized in that it does not form secondary structures or homodimers; c4. Any sequence "γ" containing 150 to 1000 base pairs adding said adapter comprising: d. Add ligase and follow - a circular DNA comprising said adapter and said "first sequence"; and - a circular DNA comprising said adapter and said "second sequence" obtaining the formation of a mixture of circularized DNA comprising 10. The method of any one of the preceding claims, comprising:
9. 10. The method of any one of the preceding claims, wherein the CRISPR system is a CRISPR-Cas9 system.
10. 10. The method of any one of the preceding claims, wherein said "second sequence" is selected from the group consisting of SEQ ID NO: 1 to SEQ ID NO: 21023.
11. 10. The method of any one of the preceding claims, wherein the biological sample is selected from the group consisting of blood, saliva, sputum, urine, feces, cerebrospinal fluid, and liquid biopsy.
12. 10. The method of any one of the preceding claims, wherein the "first sequence" is selected from the group consisting of a gene sequence of tumor origin, a gene sequence of fetal origin, a gene sequence of embryonic origin, a gene sequence of bacterial origin, a gene sequence of viral origin, a gene sequence of plant origin, and a gene sequence of animal origin.