Differential reaction path probe composition for single base mutation specific detection and corresponding detection method
By combining probe composition with lambda phage exonuclease, combined with machine learning algorithms to analyze fluorescent signals, the problem of difficult to distinguish single base mutations in the prior art is solved, and a fast and accurate detection effect is achieved.
Patent Information
- Application Number
- CN202311451350.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to distinguish single-base mutations with high sensitivity and high specificity, especially when the mutation abundance is low and the sequence similarity is high.
Using a probe composition containing an S probe and a hairpin probe T4 or a double-stranded probe A, combined with a lambda phage exonuclease (λexo), different fluorescence signals are generated through different series of strand replacement reactions, which are used to distinguish between mutant and wild-type targets. At the same time, machine learning algorithms are used to analyze fluorescent signal results to improve the accuracy of detection.
High sensitivity and specificity detection of single-base mutations are achieved, accurate results can be obtained within 20 minutes, significantly improving the accuracy and efficiency of the detection.
Smart Images

Figure CN119932168A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a probe composition for specific detection of single-base mutations and a corresponding detection method in the field of biotechnology. Background Art
[0002] Single base mutation (SNV) refers to a change in a single base in the genome, which is related to cancer and genetic diseases. Disease-related SNVs have been recognized as biomarkers for clinical diagnosis and treatment of human cancer. The ability to detect disease-related SNVs in early case analysis, tissue sections or blood is very helpful for monitoring the development of the disease and treatment. However, many disease-related SNVs exist at trace levels, with mutation abundance as low as 0.1%, and due to the high sequence similarity, it is difficult to achieve high sensitivity and specificity to distinguish SNVs.
[0003] At present, the technology used to detect single-base mutations is mainly based on sequencing, PCR and dynamic DNA nanotechnology. However, the equipment cost required for sequencing is high, the data processing is complex, and the sample quality is also too high. Detecting single-base mutations by PCR often results in false positives, and the mutation site is detected at a single site. Dynamic DNA nanotechnology can achieve high specificity, sensitivity, programmability, and robustness, providing a unique solution for distinguishing single-base mutations. Specifically, the single-base mutation distinction method based on Watson-Crick hybridization is mainly achieved through DNA molecular chain displacement reaction. In addition, dynamic DNA nanotechnology can also be combined with some auxiliary enzymes, such as polymerase, CRISPR-Cas12a, endonuclease IV, exonuclease III and other auxiliary enzymes. These studies improve the discrimination of single-base mutations by expanding the thermodynamic or kinetic differences of nucleic acid hybridization or by utilizing the binding and catalytic specificity of enzymes to different substrates. These methods usually identify mutants and wild types by artificially setting signal thresholds, and the net incremental ratio of the signals generated by the two, that is, the DF discrimination factor, is used as a parameter of the evaluation method. Although these dynamic DNA-based technologies have made progress in distinguishing single-base mutations, mutant and wild-type sequences usually have the same reaction pathway. Different signal intensities (intensity or rate) are generated only due to their different reaction extents. This causes different concentrations of wild-type and mutant types to produce similar signals, leading to misidentification and making it difficult to detect low-abundance mutations.
[0004] Dynamic DNA nanotechnology not only has the ability to manipulate DNA molecular reactions, but also can manipulate DNA molecular reactions to generate unique kinetic signals, such as single-molecule kinetic fingerprint signals, dissipative reaction kinetics, DNA timer, etc. This characteristic kinetic signal can be used for molecular recognition and identification, and the characteristic kinetic signal provides high-resolution and high-sensitivity information, which helps to accurately identify and identify in complex biological molecular mixtures. By combining these signals with computer algorithms and pattern recognition technology, highly accurate molecular analysis and biomarker detection can be achieved.
[0005] The power of machine learning algorithms lies in their ability to model complex network structures and extract valuable information from observed data to achieve accurate estimation and prediction of actual samples. This technology has been successfully applied in many fields, including image recognition, especially in biological and medical image analysis. In addition, the flexibility of machine learning algorithms also makes them suitable for analyzing two-dimensional data, such as spectral data, kinetic fingerprint signals, HRM curves, etc. In these fields, machine learning algorithms have shown excellent performance, especially when dealing with complex kinetic data. In the single-base mutation-specific detection system, the kinetic data obtained provides important information about the sample, but these data can be very complex and multidimensional. By introducing machine learning algorithms into this process, we can better understand and analyze these data, and conduct experiments and predictions accurately. Summary of the invention
[0006] The technical problem to be solved by the present invention is how to distinguish single-base mutations with high sensitivity and high specificity.
[0007] In order to solve the above technical problems, the present invention first provides a probe composition for specific detection of single-base mutations.
[0008] The probe composition for specific detection of single-base mutations provided by the present invention has a target object for detection being a mutant target object and / or a wild-type target object, wherein the mutant target object is a single-stranded DNA containing a mutation region, wherein the mutation region contains a mutant nucleotide, and the mutation region is composed of the mutant nucleotide, the 5' flank of the mutant nucleotide, and the 3' flank of the mutant nucleotide; the wild-type target object contains a single-stranded DNA named a wild region, wherein the wild region contains a wild-type nucleotide corresponding to the mutant nucleotide, and the difference in nucleotide sequence between the mutation region and the wild region is only the mutant nucleotide and the wild-type nucleotide;
[0009] The probe composition is X1 or X2:
[0010] X1: contains S probe and hairpin probe T4;
[0011] X2: contains S probe and double-stranded probe A;
[0012] The S probe is a double-stranded DNA formed by connecting a single-stranded DNA named messenger strand T2, a single-stranded DNA named protective strand T3, and a single-stranded DNA named complementary strand T1, wherein the 5' end of the complementary strand T1 is reverse complementary to the mutation region, the 5' end of the messenger strand T2 is modified with a phosphate group, and at least 21 nucleotides from the 3' end of the messenger strand T2 are identical to the 3' end nucleotide sequence of the 5' flank; the single-stranded DNA formed by connecting the 5' end of the messenger strand T2 and the 3' end of the protective strand T3 in the S probe is reverse complementary to the complementary strand T1;
[0013] The hairpin probe T4 is a single-stranded DNA consisting of a messenger chain T2 region, a stem-loop region and a messenger chain T2 binding region from the 5' end to the 3' end, the nucleotides of the messenger chain T2 region and the messenger chain T2 binding region are reversely complementary so that the single-stranded DNA forms a hairpin structure; the nucleotides of the messenger chain T2 binding region and the messenger chain T2 are reversely complementary, the 5' end of the hairpin probe T4 is modified with a fluorescent group, and the messenger chain T2 binding region is modified with a quenching group;
[0014] The double-stranded probe A is a double-stranded DNA formed by connecting a single-stranded DNA named fluorescent chain T5, a single-stranded DNA named quenching chain T6, and a single-stranded DNA named complementary chain T7; the single-stranded DNA formed by connecting the 5' end of the fluorescent chain T5 and the 3' end of the quenching chain T6 is reversely complementary to the complementary chain T7; at least 19 nucleotides from the 3' end of the fluorescent chain T5 are the same as the 3' end sequence of the 5' flank and the 3' end sequence of the messenger chain T2; the 3' end of the fluorescent chain T5 is modified with a fluorescent group, and the 5' end of the quenching chain T6 is modified with a quenching group;
[0015] The quencher group can quench the fluorescence signal of the fluorescent group.
[0016] In the above-mentioned probe composition, the complementary chain T1 can complementarily bind to the messenger chain T2 and the protection chain T3 in sequence from 5' to 3' direction, and can also complementarily bind to the target and the protection chain T3 in sequence from 5' to 3' direction, and when the target is complementarily bound to the protection chain T3, all nucleotides of the mutant target can complementarily bind to the protection chain T3, and other nucleotides of the wild-type target except the nucleotides at the mutation site can complementarily bind to the protection chain T3.
[0017] In the above probe composition, the quenching group may be BHQ1 (BHQ-1 carboxylic acid, BHQ-1CarboxylicAcid); the fluorescent group may be FAM (5-carboxyfluorescein, 5-carboxyfluorescein).
[0018] In the above probe composition, the molar ratio of the S probe to the hairpin probe T4 may be 1:1, and the molar ratio of the S probe to the double-stranded probe A may be 1:1.
[0019] In the above-mentioned single-base mutation-specific detection probe composition, the nucleotide sequence of the mutation region is SEQ ID No.1, the nucleotide sequence of the wild region is SEQ ID No.2, the S probe is formed by connecting a single-stranded DNA with a nucleotide sequence of SEQ ID No.3, a single-stranded DNA with a nucleotide sequence of SEQ ID No.4, and a single-stranded DNA with a nucleotide sequence of SEQ ID No.5, the nucleotide sequence of the hairpin probe T4 is SEQ ID No.6, and the double-stranded probe A is formed by connecting a single-stranded DNA with a nucleotide sequence of SEQ ID No.7, a single-stranded DNA with a nucleotide sequence of SEQ ID No.8, and a single-stranded DNA with a nucleotide sequence of SEQ ID No.9.
[0020] The present invention also provides a nucleic acid detection method, which comprises reacting the above-mentioned probe composition with a sample to be tested and lambda phage nuclease (λexo), detecting a fluorescent signal, and identifying whether the sample to be tested contains a mutant target or a wild-type target according to the fluorescent signal.
[0021] In the above method, whether the sample to be tested contains mutant targets and / or wild-type targets is identified based on the fluorescence signal. If the fluorescence value of the fluorescence signal first increases and then decreases, the sample to be tested contains mutant targets but no wild-type targets; if the fluorescence value of the fluorescence signal first increases and then remains stable, the sample to be tested contains wild-type targets but no mutant targets; if the fluorescence value of the fluorescence signal remains unchanged, the sample to be tested contains neither wild-type targets nor mutant targets.
[0022] The detection method using the above-mentioned single-base mutation-specific detection probe composition includes mixing the single-base mutation-specific detection probe composition with the sample to be tested and λ phage nuclease (λexo), detecting the fluorescence signal, and judging whether the sample to be tested contains the target, and the judgment criteria are as follows: if there is a mutant target in the sample to be tested, the fluorescence value first increases and then decreases; if there is a wild-type target in the sample to be tested, the fluorescence value increases and then remains stable; if there is no target in the sample to be tested, the fluorescence value does not change significantly.
[0023] The above method may also include a step of analyzing the detected fluorescence signal results using a machine learning algorithm and then outputting the results.
[0024] The fluorescence signal results are output as a fluorescence intensity-time trajectory CSV (Comma-Separated Values, CSV) file.
[0025] The machine learning algorithm can be selected from any one of support vector machines (SVM), k-nearest neighbors (KNN), random forests, neural networks, etc.
[0026] The output results use four performance parameters, namely accuracy, receiver operating characteristic curve, Precision-Recall (PR) curve, and F1-score, as result analysis criteria.
[0027] The present invention also provides a device for specific detection of single-base mutations, the device comprising:
[0028] M1, a fluorescence signal acquisition module, used to obtain the fluorescence signal of the sample to be tested using the above-mentioned probe combination;
[0029] M2, a result output module, is used to output whether the sample to be tested contains a mutant target or a wild-type target after analyzing the fluorescence signal through a machine learning algorithm.
[0030] The present invention also provides a reagent or a kit for specific detection of single-base mutations, wherein the reagent or the kit comprises the above-mentioned probe composition and lambda phage nuclease.
[0031] The present invention also provides a product, which consists of X1 and X2; X1 is the single-base mutation-specific detection probe composition and lambda phage nuclease, and X2 is the reagent and / or instrument required for detection.
[0032] In order to solve the above technical problems, the present invention also provides the use of the probe combination in detecting single base mutations.
[0033] In the present invention, the inventors utilize the special properties of the λ bacteriophage exonuclease (λexo) to recognize and digest the 5' phosphate end of the nucleic acid, triggering the digestion reaction of the enzyme, and construct a single-base mutation-specific detection probe composition, as well as a detection method using the above probe. The detection method of the present invention utilizes λexo to drive the wild-type target and the mutant target to undergo different series of chain displacement reactions with the designed probe, forming different optimal substrates of the λexonuclease (λExo) enzyme and then digesting the probe to produce different fluorescent signals, which are strictly corresponding to the wild-type target and the mutant target nucleic acid sequence, and highly reliable to distinguish the wild-type target, the mutant target and the non-target sequence. The present invention also combines the designed single-base mutation-specific detection probe composition with a machine learning algorithm program, and obtains the kinetic data of the target molecule by reasonably designing a single-base mutation detection specific probe, and inputs the obtained kinetic data into the machine learning algorithm to obtain the detection result. The present invention can directly identify unknown samples and obtain the detection results within 20 minutes, and confirm whether the target molecule is contained by the machine learning algorithm program, which greatly improves the accuracy and efficiency of the detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Schematic diagram of the single-base mutation-specific probe detection mode in Example 1, Figure 1 In (A), the detection target is a mutant sample of the EGFR-L858R gene, and the probe composition consists of an S probe and a hairpin probe T4; Figure 1 In (B), the detection target is a wild-type sample of the EGFR-L858R gene, and the probe composition consists of an S probe and a hairpin probe T4; Figure 1 In (C), the detection target is a mutant sample of the EGFR-L858R gene, and the probe composition consists of an S probe and a double-stranded probe A; Figure 1 In (D), the detection target is a wild-type sample of the EGFR-L858R gene, and the probe composition consists of an S probe and a double-stranded probe A; Figure 1 (E) is the fluorescence intensity-time trajectory of the mutant sample of the EGFR-L858R gene with different concentrations as the detection target. The probe composition consists of S probe and hairpin probe T4. Figure 1 (F) is the fluorescence intensity-time trajectory of the wild-type sample with different concentrations of EGFR-L858R gene as the detection target, and the probe composition consists of S probe and hairpin probe T4. Figure 1 (G) is the fluorescence intensity-time trajectory of the mutant sample of EGFR-L858R gene with a target concentration of 100 nM. The probe composition consists of S probe and double-stranded probe A. Figure 1(H) is the fluorescence intensity-time trajectory of the wild-type sample of the EGFR-L858R gene with a target concentration of 100 nM. The probe composition consists of S probe and double-stranded probe A.
[0035] Figure 2 The process and performance parameters of the fluorescence intensity-time trajectory CSV file in Example 2 input into the machine learning algorithm program. Among them, Figure 2 (A) is the flow chart of the input of the fluorescence intensity-time trajectory CSV file into the machine learning algorithm program. Figure 2 (B) Performance parameter diagram of the CSV file of fluorescence intensity-time trajectories for mutant and wild-type detection input into the machine learning algorithm program.
[0036] Figure 3 This is the fluorescence intensity-time trajectory of the asymmetric PCR system of the EGFR-L858R gene plasmid in Example 3, where the probe composition consists of the S probe and the double-stranded probe A. Figure 3 (A) is the fluorescence intensity-time trajectory when the target is the mutant plasmid asymmetric PCR system of EGFR-L858R gene, Figure 3 (B) is the fluorescence intensity-time trajectory when the target is the wild-type plasmid asymmetric PCR system of EGFR-L858R gene. Figure 3 (C) is a parameter diagram of the performance of the machine learning algorithm program inputted from the CSV file of the fluorescence intensity-time trajectory of the asymmetric PCR system for the EGFR-L858R gene plasmid. DETAILED DESCRIPTION
[0037] The present invention is further described in detail below in conjunction with specific embodiments, and the examples provided are only for illustrating the present invention, rather than for limiting the scope of the present invention. The examples provided below can be used as a guide for further improvements by those of ordinary skill in the art, and do not constitute a limitation of the present invention in any way.
[0038] The quantitative tests in the following examples were all repeated three times, and the results were averaged.
[0039] The experimental methods in the following examples are conventional methods unless otherwise specified. The materials, reagents, etc. used in the following examples are commercially available unless otherwise specified.
[0040] The reaction buffer environment 1×λExo buffer in the following examples is a product of New England Biolabs (NEB) with a product number of M0262S. The reagent contains nuclease λexonuclease (λExo).
[0041] The nucleic acid sequences in the following examples are synthesized products of Shanghai Sangon Biotechnology Co., Ltd.
[0042] Example 1
[0043] This example detects EGFR-L858R target molecules, including wild-type and mutant types, based on a single-base mutation-specific detection probe composition. The probe composition for specific detection of single-base mutations in this embodiment has a target object for detection as a mutant target object and / or a wild-type target object, wherein the mutant target object is a single-stranded DNA containing a mutation region, wherein the mutation region (nucleotide sequence is SEQ ID No.1) contains a mutant nucleotide, and the mutation region is composed of the mutant nucleotide (position 22 of SEQ ID No.1), the 5' flank of the mutant nucleotide (positions 1-21 of SEQ ID No.1, which are the same as positions 1-21 of SEQ ID No.2), and the 3' flank of the mutant nucleotide (positions 22-28 of SEQ ID No.1, which are the same as positions 22-28 of SEQ ID No.2); and the wild-type target object contains a single-stranded DNA containing a wild-type region, wherein the wild-type region (WT in Table 1, nucleotide sequence is SEQ ID No.2) contains a wild-type nucleotide (position 22 of SEQ ID No.2) corresponding to the mutant nucleotide, and the difference in nucleotide sequence between the mutation region and the wild-type region is only that the mutant nucleotide (position 22 of SEQ ID No.1) and the wild-type nucleotide (position 22 of SEQ ID No.2) are the same as the 5' flank of the mutant nucleotide (position 22 of SEQ ID No.1). No.22nd);
[0044] That is, the mutant type of the specific EGFR-L858R target molecule is the single-stranded DNA molecule named SNV in Table 1 (i.e., the mutant target, including SEQ ID No. 1), and the wild type of the EGFR-L858R target molecule is the single-stranded DNA molecule named WT in Table 1 (i.e., the wild-type target, including SEQ ID No. 2), and the 22nd nucleotide of SEQ ID No. 1 is mutated from T to G compared with SEQ ID No. 2. The test samples used in this embodiment are the artificially synthesized single-stranded DNA molecules SNV and WT, and a blank control is set.
[0045] This embodiment uses two probe compositions for detection, namely, probe composition 1 comprising S probe and hairpin probe T4, and probe composition 2 comprising S probe and double-stranded probe A. The detection is specifically as follows:
[0046] The probe composition 1 comprises an S probe and a hairpin probe T4:
[0047] The S probe is a double-stranded DNA formed by connecting a single-stranded DNA named messenger strand T2 (SEQ ID No.4), a single-stranded DNA named protective strand T3 (SEQ ID No.5) and a single-stranded DNA named complementary strand T1 (SEQ ID No.3), the 5' end of the complementary strand T1 is reverse complementary to the mutation region (the 1st to 28th positions of SEQ ID No.3 are reverse complementary to SEQ ID No.1), the 5' end of the messenger strand T2 is modified with a phosphate group (the 1st position of SEQ ID No.4 is modified with a phosphate group), and at least 21 nucleotides (21 in this embodiment) from the 3' end of the messenger strand T2 are identical to the 3' end nucleotide sequence of the 5' flank (the 6th to 26th positions of SEQ ID No.4 are identical to the 1st to 21st positions of SEQ ID No.1); the single-stranded DNA formed by connecting the 5' end of the messenger strand T2 and the 3' end of the protective strand T3 in the S probe is reversely complementary to the complementary strand T1 (the 8th to 33rd positions of SEQ ID No.3 are identical to the 5' end nucleotide sequence of SEQ ID No.1). No.4 is reverse complementary, and positions 34-54 of SEQ ID No.3 are reverse complementary to SEQ ID No.5);
[0048] Therefore, the complementary strand T1 can not only complementarily bind to the messenger strand T2 and the protection strand T3 in sequence from the 5' to 3' direction (positions 8-33 of SEQ ID No.3 are reverse complementary to SEQ ID No.4, and positions 34-54 of SEQ ID No.3 are reverse complementary to SEQ IDNo.5), but also complementarily bind to the target and the protection strand T3 in sequence from the 5' to 3' direction, and when the target is complementarily bound to the protection strand T3, all nucleotides of the mutant target can complementarily bind to the protection strand T3 (positions 1-28 of SEQ ID No.3 are reverse complementary to SEQ ID No.1), and other nucleotides of the wild-type target except the nucleotide at the mutation site can complementarily bind to the protection strand T3 (positions 1-28 are reverse complementary to other nucleotides of SEQ ID No.2 except the nucleotide at position 22).
[0049] The hairpin probe T4 (SEQ ID No. 6) is a single-stranded DNA consisting of the messenger chain T2 region (SEQ ID No. 6 positions 1-19, which are the same as SEQ ID No. 4 positions 8-26), the stem-loop region (SEQ ID No. 6 positions 20-35) and the messenger chain T2 binding region (SEQ ID No. 6 positions 36-61) from the 5' end to the 3' end, the nucleotides of the messenger chain T2 region and the messenger chain T2 binding region are reversely complementary so that the single-stranded DNA forms a hairpin structure (SEQ ID No. 6 positions 1-19 are reversely complementary to positions 36-54); the messenger chain T2 binding region and the nucleotides of the messenger chain T2 are reversely complementary (SEQ ID No. 6 positions 36-61 are reversely complementary to SEQ ID No. 4), the 5' end of the hairpin probe T4 is modified with a fluorescent group (SEQ ID No. 6 position 1 is modified with a fluorescent group FAM), and the messenger chain T2 binding region is modified with a quenching group (SEQ ID No.6 is modified with a quenching group BHQ1) at position 54, and the quenching group can quench the fluorescence signal of the fluorescent group.
[0050] The probe composition 2 comprises an S probe and a double-stranded probe A:
[0051] The S probe is the same as the S probe in the probe composition 1 described above.
[0052] The double-stranded probe A is a double-stranded DNA formed by connecting a single-stranded DNA named fluorescent chain T5 (SEQ ID No.7), a single-stranded DNA named quenching chain T6 (SEQ ID No.8), and a single-stranded DNA named complementary chain T7 (SEQ ID No.9); the single-stranded DNA formed by connecting the 5' end of the fluorescent chain T5 and the 3' end of the quenching chain T6 is reversely complementary to the complementary chain T7 (the 1st to 18th positions of SEQ ID No.9 are reversely complementary to SEQ ID No.8, and the 19th to 37th positions of SEQ ID No.9 are reversely complementary to SEQ ID No.7); at least 19 nucleotides (19 in this embodiment) from the 3' end of the fluorescent chain T5 are identical to the 3' end sequence of the 5' flank and the 3' end sequence of the messenger chain T2 (SEQ ID No.7 is identical to the 3rd to 21st positions of SEQ ID No.1, i.e., the 8th to 26th positions of SEQ ID No.4); the 3' end of the fluorescent chain T5 is modified with a fluorescent group (SEQ ID No.7 is modified with a fluorescent group FAM at position 19), and the 5' end of the quencher chain T6 is modified with a quencher group (SEQ ID No.8 is modified with a fluorescent group BHQ1 at position 1), and the quencher group can quench the fluorescent signal of the fluorescent group.
[0053] According to Table 1, the required single-base mutation-specific detection probe composition 1 (comprising an S probe and a hairpin probe T4 or an S probe, wherein the S probe is assembled from SEQ ID No. 3, SEQ ID No. 4 and SEQ ID No. 5 in the sequence list, and the hairpin probe T4 is assembled from a single strand of SEQ ID No. 6 in the sequence list) is synthesized.
[0054] According to Table 1, the required single-base mutation-specific detection probe composition 2 (comprising an S probe and a double-stranded probe A, wherein the S probe is assembled from SEQ ID No. 3, SEQ ID No. 4 and SEQ ID No. 5 in the sequence list, and the double-stranded probe A is assembled from SEQ ID No. 7, SEQ ID No. 8 and SEQ ID No. 9 in the sequence list) is synthesized.
[0055] Table 1 Nucleotide sequences
[0056]
[0057]
[0058] In SEQ ID No.4, P represents a phosphate group, and the A at position 1 is phosphorylated. In SEQ ID No.6, FAM represents 5-carboxyfluorescein, BHQ1 represents BHQ-1carboxylic acid, A at position 1 is modified with FAM, and T at position 54 is modified with BHQ1. In SEQ ID No.7, FAM represents 5-carboxyfluorescein, and C at position 19 is modified with FAM. In SEQ ID No.8, BHQ1 represents BHQ-1carboxylic acid, and G at position 1 is modified with BHQ1.
[0059] The detection method of the present invention is an identification method for detecting a target object using the above-mentioned probe composition (probe composition 1 or probe composition 2) and lambda phage exonuclease (λexo).
[0060] 1. Detection using a probe composition 1 comprising an S probe and a hairpin probe T4
[0061] The detection method using the probe composition 1 of the present invention comprises the following steps:
[0062] S1. Add the sample to be tested and exonuclease (λExo) to the reaction system containing the above-mentioned probe composition to form a detection system, and record the fluorescence intensity.
[0063] The samples to be tested are synthetic EGFR-L858R mutant and wild type.
[0064] S1-1 Establishing the probe system
[0065] The probe composition 1 is composed of S probes (T1, T2, T3) and hairpin probe T4.
[0066] Establish a probe system (20 μL): The reaction buffer environment is 1×λExo buffer, and the single-base mutation specific detection probe composition 1 (S probe and hairpin probe T4 are assembled separately, and the molar ratio of S probe and hairpin probe T4 is 1:1) is added to make the concentration of the single-base mutation specific detection probe composition 1 in the system 1 μM. Make the total volume of the system equal to 20 μL, and the remaining volume is supplemented with ddH2O.
[0067] S1-2 detection
[0068] The samples to be tested are wild-type samples (single-stranded DNA molecule WT, the concentrations in the detection system are set to 80nM, 50nM, and 20nM, respectively), mutant samples (single-stranded DNA molecule SNV, the concentrations in the detection system are set to 80nM, 50nM, and 20nM, respectively), and a blank control is set.
[0069] 20μL detection system: Take 2μL of 1μM single-base mutation-specific detection probe composition 1 (composed of S probe and hairpin probe T4) to make its concentration in the detection system 100nM, add 10μL of the sample to be tested, and add λExo to make its concentration in the system 50U / mL. If the total volume of the system is less than 20μL, supplement it with ddH2O. Place it in a real-time fluorescence detector for detection at 37°C, set the detection time interval to 5 seconds per cycle, 240 cycles, and a detection time of 20 minutes to obtain the fluorescence intensity-time trajectory.
[0070] The fluorescence value-time trajectory of the reaction process is captured and recorded by a real-time fluorescence instrument and output as a fluorescence intensity-time trajectory CSV (Comma-Separated Values, CSV) file.
[0071] S2. Identify whether the sample to be tested contains a mutant target or a wild-type target based on the fluorescence signal.
[0072] If the sample to be tested contains a mutant target of the EGFR-L858R gene, see Figure 1(A), the mutant target binds to the S probe to generate a TMSD reaction to release a messenger chain T2 with a phosphate group modified at the 5' end, and the messenger chain T2 binds to the hairpin probe T4 to generate a TMSD reaction, and the hairpin probe T4 modified with a fluorescent group and a quenching group changes from an assembled form to an open form, and the fluorescent group changes from being quenched to being unquenched, and the fluorescence value increases significantly; then, the messenger chain T2 bound to the opened hairpin probe T4 is digested by a nuclease (λExo) because the 5' end is modified with a phosphate group, and the opened hairpin probe T4 is restored to an assembled form, and the fluorescent group changes from being unquenched to being quenched, and the fluorescence value decreases significantly; that is, the fluorescence value increases significantly first and then decreases significantly, and the mutant target of the EGFR-L858R gene is present in the sample to be tested;
[0073] If the wild-type target of EGFR-L858R gene exists in the sample to be tested, refer to Figure 1 (B), part of the wild-type target combines with the S probe to generate a TMSD reaction to release the messenger chain T2 with a phosphate group modified at the 5' end, and the messenger chain T2 combines with the hairpin probe T4 to generate a TMSD reaction, and the hairpin probe T4 modified with a fluorescent group and a quenching group changes from an assembled form to an open form, and the fluorescent group changes from being quenched to being unquenched, and the fluorescence value rises; then, the messenger chain T2 bound to the opened hairpin probe T4 is digested by the nuclease (λExo) because the 5' end is modified with a phosphate group, and part of the wild-type target competes with the opened hairpin probe T4 for binding, resulting in the opened hairpin probe T4 being unable to return to an assembled form, the fluorescent group is not quenched, and the fluorescence value is stable; that is, if the fluorescence value rises and then remains stable, then the wild-type target of the EGFR-L858R gene is present in the sample to be tested;
[0074] If the target is not present in the sample to be tested (neither mutant target nor wild-type target), the fluorescence value will not change significantly.
[0075] Therefore, the identification of whether the sample to be tested contains mutant targets and / or wild-type targets based on the fluorescence signal can be: if the fluorescence value of the fluorescence signal first increases and then decreases, the sample to be tested contains mutant targets but no wild-type targets; if the fluorescence value of the fluorescence signal first increases and then remains stable, the sample to be tested contains wild-type targets but no mutant targets; if the fluorescence value of the fluorescence signal remains unchanged, the sample to be tested contains neither wild-type targets nor mutant targets.
[0076] Figure 1 (E) is the fluorescence intensity-time trajectory of the S probe and hairpin probe T4 detecting the target as a single base mutation. Figure 1(F) is the fluorescence intensity-time trajectory of the S probe and hairpin probe T4 detecting the single-base wild-type target. The results show that the single-base mutation detection specific probe can identify and distinguish the EGFR-L858R mutant and wild-type.
[0077] 2. Detection using probe composition 2 comprising S probe and double-stranded probe A
[0078] The detection method using the probe composition 2 of the present invention comprises the following steps:
[0079] Z1. Add the sample to be tested and exonuclease (λExo) to the reaction system containing the above-mentioned probe composition 2 to form a detection system, and record the fluorescence intensity.
[0080] The samples to be tested are synthetic EGFR-L858R mutant and wild type.
[0081] Z1-1 Establishing the probe system
[0082] Probe composition 2 is composed of S probes (T1, T2, T3) and double-stranded probes A (T5, T6, T7).
[0083] Establish probe system (20 μL): The reaction buffer environment is 1×λExo buffer, and the single-base mutation specific detection probe composition 2 (S probe and double-stranded probe A are assembled separately, and the molar ratio of S probe and double-stranded probe A is 1:1) is added to make the concentration of single-base mutation specific detection probe composition 2 in the system 1 μM. Make the total volume of the system equal to 20 μL, and the remaining volume is supplemented with ddH2O.
[0084] Z1-2 Testing
[0085] The samples to be tested are wild-type samples (single-stranded DNA molecule WT, the concentration in the detection system is set to 100 nM) and mutant samples (single-stranded DNA molecule SNV, the concentration in the detection system is set to 100 nM), and a blank control is set.
[0086] 20μL detection system: Take 2μL of single-base mutation-specific detection probe composition 2 (composed of S probe and double-stranded probe A) with a concentration of 1μM, make its concentration in the detection system 100nM, add 10μL of the sample to be tested, and add λExo to make its concentration in the system 50U / mL. If the total volume of the system is less than 20μL, supplement it with ddH2O. Place it in a real-time fluorescence detector for detection at 37°C, set the detection time interval to 5 seconds per cycle, 240 cycles, and the detection time to 20 minutes to obtain the fluorescence intensity-time trajectory.
[0087] The fluorescence value-time trajectory of the reaction process is captured and recorded by a real-time fluorescence instrument and output as a fluorescence intensity-time trajectory CSV (Comma-Separated Values, CSV) file.
[0088] Z2. Identify whether the sample to be tested contains a mutant target or a wild-type target based on the fluorescence signal.
[0089] If the sample to be tested contains a mutant target of the EGFR-L858R gene, see Figure 1 (C), the mutant target combines with the S probe to generate a TMSD reaction to release a messenger chain T2 with a phosphate group modified at the 5' end, and the messenger chain T2 combines with the double-stranded probe A to generate a TMSD reaction to release a fluorescent chain T5 modified with a fluorescent group, and the fluorescent group changes from being quenched to being unquenched, and the fluorescence value increases significantly; the double-stranded probe A that undergoes the TMSD reaction contains the messenger chain T2 with a phosphate group modified at the 5' end, which is regarded as the optimal substrate for the nuclease (λExo), and the nuclease (λExo) digests the messenger chain T2, and the released fluorescent chain can be recombined with the complementary chain T7 to form the double-stranded probe A, and the fluorescent group changes from being unquenched to being quenched by the quenching group on the quenching chain T6, and the fluorescence value decreases significantly; that is, the fluorescence value increases significantly first and then decreases significantly, and the mutant target of the EGFR-L858R gene is present in the sample to be tested;
[0090] If the wild-type target of EGFR-L858R gene exists in the sample to be tested, refer to Figure 1 (D), the wild-type target combines with the S probe to generate a TMSD reaction to release a messenger chain T2 modified with a phosphate group at the 5' end, and the messenger chain T2 combines with the double-stranded probe A to generate a TMSD reaction to release a fluorescent chain T5 modified with a fluorescent group, and the fluorescent group changes from being quenched to being unquenched, and the fluorescence value increases significantly; the double-stranded probe A that undergoes the TMSD reaction contains the messenger chain T2 modified with a phosphate group at the 5' end, which is regarded as the optimal substrate for the nuclease (λExo), and the nuclease (λExo) digests the messenger chain T2, and part of the wild-type target competes with the released fluorescent chain T5 for binding to the double-stranded probe A, resulting in the release of the fluorescent chain cannot return to the double-stranded probe A, the fluorescent group is not quenched, and the fluorescence value is stable; that is, the fluorescence value remains stable after rising, and the wild-type target of the EGFR-L858R gene is present in the sample to be tested;
[0091] If the target is not present in the sample to be tested (neither mutant target nor wild-type target), the fluorescence value will not change significantly.
[0092] Therefore, the identification of whether the sample to be tested contains mutant targets and / or wild-type targets based on the fluorescence signal can be: if the fluorescence value of the fluorescence signal first increases and then decreases, the sample to be tested contains mutant targets but no wild-type targets; if the fluorescence value of the fluorescence signal first increases and then remains stable, the sample to be tested contains wild-type targets but no mutant targets; if the fluorescence value of the fluorescence signal remains unchanged, the sample to be tested contains neither wild-type targets nor mutant targets.
[0093] Figure 1 (G) is the fluorescence intensity-time trajectory of the S probe and double-stranded probe A detecting the target as a single-base mutant. Figure 1 (H) is the fluorescence intensity-time trajectory of the S probe and the double-stranded probe A detecting the single-base wild-type target. The results show that the single-base mutation detection specific probe can identify and distinguish the EGFR-L858R mutant and wild-type.
[0094] Example 2
[0095] In this embodiment, based on the first embodiment, the detected fluorescence signal result is analyzed by a machine learning algorithm and the result is output.
[0096] Figure 2 (A) is a process of analyzing the detected fluorescence signal results with a machine learning algorithm and then outputting the results: according to the experimental results of Example 1, the fluorescence value-time trajectory of the wild type and the single base mutant type is loaded into the machine learning algorithm program as a data set in CSV format. The loaded data set is first normalized, and then the data set is divided into a training set and a test set according to a ratio of 7:3.
[0097] The machine learning analysis program includes a data loading module, an algorithm analysis module, and a performance parameter module.
[0098] The algorithm analysis module used is analyzed by four different machine learning algorithms, specifically the following four:
[0099] Support vector machines (SVM);
[0100] k-nearest neighbors (KNN);
[0101] Random Forest;
[0102] Neural Networks.
[0103] The above training set data is loaded into the algorithm analysis modules (four types) through the loading data module to build a classification model. The performance parameter module is the result obtained by the four machine learning algorithm analysis modules, and the accuracy, receiver operating characteristic curve, Precision-Recall (PR) curve, and F1-score are used as the result analysis criteria.
[0104] The test set is used as a validation of the classification model, and the performance parameters of the test set are obtained after algorithm analysis, thereby determining whether the sample to be tested is wild type or single base mutation type.
[0105] The machine learning analysis program is run on Python 3.10.9.
[0106] The machine learning analysis program code is taken from sklearn and PyTorch.
[0107] The present invention can directly identify unknown samples and obtain test results within 20 minutes. Figure 2 (B) is a performance parameter diagram of the mutant and wild-type fluorescence intensity-time trajectory CSV file input into the machine learning algorithm program. The results show that the single-base mutation detection specific probe combined with the machine learning algorithm program analysis can achieve 100% accuracy in distinguishing between wild-type and mutant types.
[0108] Example 3
[0109] According to the experimental results of Example 1 and Example 2, the inventors constructed an analytical method for detecting single-base mutations. In order to further verify the performance of this method in detecting real samples, the inventors constructed a mutant plasmid of the EGFR-L858R gene (containing the nucleotide sequence shown in SEQ ID No. 1) and a wild-type plasmid of the EGFR-L858R gene (containing the nucleotide sequence shown in SEQ ID No. 2). The constructed plasmids were synthesized by Shanghai Bioengineering Co., Ltd.
[0110] In order to verify the performance of the detection of real samples, the specific steps are as follows:
[0111] First, it is necessary to obtain the mutant and wild-type target fragments of the EGFR-L858R gene, and it is necessary to amplify the specific fragments of the mutant plasmid of the EGFR-L858R gene and the wild-type plasmid of the EGFR-L858R gene to obtain single-stranded DNA, which specifically includes the following steps: amplify the target gene using asymmetric PCR technology, amplify the target gene, and use Q5 High-Fidelity2X Master Mix (Cat. No. M0494). It is necessary to add the components of the asymmetric PCR system (50 μL) in Table 2 to the EP tube respectively, mix well, and place the EP tube in the PCR instrument to amplify the EGFR-L858R mutant and wild-type genes respectively.
[0112] Table 2 Asymmetric PCR system (50 μL)
[0113] Components Dosage Final concentration Enzyme-free water 19μL Q5 High-Fidelity 2X Master 25μL Upstream primer 2.5μL 1000nM Downstream primer 2.5μL 100nM Plasmid (wild type / mutant) 1μL
[0114] The upstream primer sequence for amplifying the target fragment of the EGFR-L858R gene is FP, and the downstream primer sequence is RP. The specific primer sequences are shown in Table 3.
[0115] Table 3 Primer sequences
[0116] name Sequence (5'-3') FP ACCGCAGCATGTCAAGATCA RP TTTGCCTCCTTCTGCATGGT
[0117] The asymmetric PCR program was: 98°C for 30 s, cycling phase: 98°C for 10 s, 60°C for 30 s, 72°C for 30 s, 10 cycles, cycling phase: 98°C for 10 s, 50°C for 30 s, 72°C for 30 s, 30 cycles, 72°C for 3 min, 4°C hold.
[0118] Three parallel controls were performed for each gene.
[0119] The samples to be tested are components of the asymmetric PCR system in the complementary chain T1, which are the following two types:
[0120] (1) Asymmetric PCR system for the mutant plasmid of EGFR-L858R gene;
[0121] (2) Asymmetric PCR system using wild-type plasmid of EGFR-L858R gene.
[0122] This example experiments the effect of a device for specific detection of single-base mutations. The device comprises:
[0123] M1, a fluorescence signal acquisition module, is used to acquire the fluorescence signal of the sample to be tested using the S probe and the hairpin probe T4 in the probe composition 1 in Example 1. The details are as follows:
[0124] M1-1 Establishing the probe system
[0125] Establish a probe system (20 μL): The reaction buffer environment is 1×λExo buffer, and the single-base mutation specific detection probe composition 1 is added to make the concentration of the single-base mutation specific detection probe composition 1 in the system 2 μM. Make the total volume of the system equal to 20 μL, and the remaining volume is supplemented with ddH2O.
[0126] M1-2 Testing System
[0127] 20μL detection system: Take 2μL of the probe system established by M1-1 above with a concentration of 2μM, make its concentration in the system 200nM, add 10μL of the components of the asymmetric PCR system, and add λExo to make its concentration in the system 50U / mL. If the total volume of the system is less than 20μL, supplement it with ddH2O. Place it in a real-time fluorescence detector for detection at 37°C, set the detection time interval to 5 seconds per cycle, the number of cycles to 240, the detection time to 20 minutes, obtain the fluorescence intensity-time trajectory, and output it as a CSV file.
[0128] M2, a result output module, is used to output whether the sample to be tested contains a mutant target or a wild-type target after analyzing the fluorescence signal through a machine learning algorithm.
[0129] The details are as follows:
[0130] The CSV files in Example 1 and Example 2 are used as training sets for the machine learning algorithm program, and the CSV file in Example 3 is used as a test set for the machine learning algorithm program. First, the loaded data set is normalized, and the training set data is loaded into the four machine learning algorithms to build a classification model. The test set is used as a verification of the classification model, and the performance parameters of the test set are obtained after algorithm analysis. Thus, it is determined whether the sample to be tested is wild type or single base mutant.
[0131] The results are as follows Figure 3 As shown, it is shown that the single-base mutation specific detection probe combination combined with the machine learning algorithm program analysis can realize the detection of real samples, and the detection accuracy rate reaches 100%. The present invention can directly identify unknown samples and obtain the test results within 20 minutes. It confirms whether the target molecule is contained through the machine learning algorithm program, which greatly improves the detection accuracy and efficiency.
[0132] The present invention has been described in detail above. It will be apparent to those skilled in the art that the present invention may be implemented in a wide range under equivalent parameters, concentrations and conditions without departing from the spirit and scope of the present invention and without the need for unnecessary experimentation. Although the present invention provides specific embodiments, it should be understood that further improvements may be made to the present invention. In short, according to the principles of the present invention, this application intends to include any changes, uses or improvements to the present invention, including changes made by conventional techniques known in the art that depart from the scope disclosed in this application. Applications of some of the basic features may be made within the scope of the following appended claims.
Claims
1. A probe composition for specific detection of single-base mutations, characterized in that: The target detected by the probe composition is a mutant target and / or a wild-type target, wherein the mutant target is a single-stranded DNA containing a mutation region, wherein the mutation region contains a mutant nucleotide, and the mutation region consists of the mutant nucleotide, the 5' flank of the mutant nucleotide, and the 3' flank of the mutant nucleotide; the wild-type target contains a single-stranded DNA named a wild region, wherein the wild region contains a wild-type nucleotide corresponding to the mutant nucleotide, and the difference in nucleotide sequence between the mutation region and the wild region is only the mutant nucleotide and the wild-type nucleotide; The probe composition is X1 or X2: X1: contains S probe and hairpin probe T4; X2: contains S probe and double-stranded probe A; The S probe is a double-stranded DNA formed by connecting a single-stranded DNA named messenger strand T2, a single-stranded DNA named protective strand T3, and a single-stranded DNA named complementary strand T1, wherein the 5' end of the complementary strand T1 is reverse complementary to the mutation region, the 5' end of the messenger strand T2 is modified with a phosphate group, and at least 21 nucleotides from the 3' end of the messenger strand T2 are identical to the 3' end nucleotide sequence of the 5' flank; the single-stranded DNA formed by connecting the 5' end of the messenger strand T2 and the 3' end of the protective strand T3 in the S probe is reverse complementary to the complementary strand T1; The hairpin probe T4 is a single-stranded DNA consisting of a messenger chain T2 region, a stem-loop region and a messenger chain T2 binding region from the 5' end to the 3' end, the nucleotides of the messenger chain T2 region and the messenger chain T2 binding region are reversely complementary so that the single-stranded DNA forms a hairpin structure; the nucleotides of the messenger chain T2 binding region and the messenger chain T2 are reversely complementary, the 5' end of the hairpin probe T4 is modified with a fluorescent group, and the messenger chain T2 binding region is modified with a quenching group; The double-stranded probe A is a double-stranded DNA formed by connecting a single-stranded DNA named fluorescent chain T5, a single-stranded DNA named quenching chain T6, and a single-stranded DNA named complementary chain T7; the single-stranded DNA formed by connecting the 5' end of the fluorescent chain T5 and the 3' end of the quenching chain T6 is reversely complementary to the complementary chain T7; at least 19 nucleotides from the 3' end of the fluorescent chain T5 are the same as the 3' end sequence of the 5' flank and the 3' end sequence of the messenger chain T2; the 3' end of the fluorescent chain T5 is modified with a fluorescent group, and the 5' end of the quenching chain T6 is modified with a quenching group; The quencher group can quench the fluorescence signal of the fluorescent group.
2. The single-base mutation-specific detection probe composition according to claim 1, characterized in that: The fluorescent group is FAM, and the quenching group is BHQ1.
3. The single-base mutation-specific detection probe composition according to claim 1 or 2, characterized in that: The molar ratio of the S probe to the hairpin probe T4 is 1:1, and the molar ratio of the S probe to the double-stranded probe A is 1:
1.
4. The single-base mutation specific detection probe composition according to any one of claims 1 to 3, characterized in that: The nucleotide sequence of the mutation region is SEQ ID No.1, the nucleotide sequence of the wild region is SEQ ID No.2, the S probe is formed by connecting a single-stranded DNA with a nucleotide sequence of SEQ ID No.3, a single-stranded DNA with a nucleotide sequence of SEQ ID No.4, and a single-stranded DNA with a nucleotide sequence of SEQ ID No.5, the nucleotide sequence of the hairpin probe T4 is SEQ ID No.6, and the double-stranded probe A is formed by connecting a single-stranded DNA with a nucleotide sequence of SEQ ID No.7, a single-stranded DNA with a nucleotide sequence of SEQ ID No.8, and a single-stranded DNA with a nucleotide sequence of SEQ ID No.
9.
5. A nucleic acid detection method, characterized in that: The method comprises reacting the probe composition of any one of claims 1 to 4 with a sample to be tested and lambda phage nuclease, detecting a fluorescent signal, and identifying whether the sample to be tested contains a mutant target or a wild-type target according to the fluorescent signal.
6. The method according to claim 5, characterized in that The method also includes the step of analyzing the detected fluorescence signal results using a machine learning algorithm and then outputting the results.
7. The method according to claim 6, characterized in that The machine learning algorithm is selected from any one of support vector machine, nearest neighbor method, random forest and neural network.
8. A device for specific detection of single-base mutations, characterized in that: The device comprises: M1, a fluorescence signal acquisition module, used to obtain the fluorescence signal of the sample to be tested using the probe composition according to any one of claims 1 to 4; M2, a result output module, is used to output whether the sample to be tested contains a mutant target or a wild-type target after analyzing the fluorescence signal through a machine learning algorithm.
9. A reagent or kit for specific detection of single-base mutations, characterized in that: The reagent or kit comprises the probe composition according to any one of claims 1 to 4 and lambda phage exonuclease.
10. Use of the probe composition according to any one of claims 1 to 4 in detecting single base mutations.
Citation Information
Cited By
Standard probe, method, system and medium for reducing spatial effect of gene chip detection signal
CN121438943A