Kit for detecting breast cancer through plasma miRNA isomer next generation sequencing
By constructing a second-generation sequencing library for plasma miRNA isomers, the problem of insufficient sensitivity and depth of miRNA detection in the prior art is solved, high sensitivity and deep detection of breast cancer-related miRNAs are achieved, the detection efficiency of breast cancer is improved, and high accuracy and sensitivity are achieved through machine learning models.
Patent Information
- Application Number
- CN202510018455.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The existing miRNA detection technology has problems such as low sensitivity, low relative sequencing depth and low specificity, which limits its application in breast cancer diagnosis.
By constructing a second-generation sequencing library of plasma miRNA isomers, selecting the miRNAs that need to be amplified are used to construct the library, achieving high sensitivity, depth and specific detection of low-expressed miRNAs.
High sensitivity and deep detection of breast cancer-related miRNAs were achieved, which improved the detection efficiency of breast cancer, and built through machine learning models to achieve 93.48% accuracy, 95.65% sensitivity and 91.30% specificity.
Smart Images

Figure CN119932185A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of biomedical technology, and in particular relates to a kit for detecting breast cancer by second-generation sequencing of plasma miRNA isoforms. Background Art
[0002] MicroRNA (miRNA) is a type of non-coding small RNA with a length of 18 to 25 nucleotides. MiRNA directly and indirectly regulates the expression of most genes and participates in a series of life activities. MiRNA is closely related to the occurrence and development of tumors. More and more studies have shown that miRNA plays an important regulatory role in the occurrence and development of tumors. Malignant tumors are the result of the interaction between genetic and environmental factors, and environmental factors play a greater role. Genetic diagnosis of cancer has great limitations. It can only discover susceptibility genes and cannot be used as biomarkers for the diagnosis of malignant tumors. At the same time, environmental factors cannot be monitored and can only be used as risk factors for malignant tumors. MiRNA is a large class of regulatory factors between the changing environment and the unchanging genetic material. It is the main bridge connecting environmental factors and genetic factors. It has a relatively stable structure, abnormal expression appears earlier, and it is easier to accurately distinguish tumor types. These characteristics make miRNA a biomarker for tumor etiology, diagnosis, progression, recurrence and treatment outcomes.
[0003] However, due to the disadvantages of miRNA, such as only twenty or so bases and low levels in the blood, it is difficult to detect it, and there is still a lack of technology to accurately detect microRNA. qPCR, microarrays, and small RNA sequencing (RNA-seq) are commonly used to study the expression of miRNA in tissues. But they all have defects to varying degrees. The main problems of microarrays are low sensitivity and relatively long turnaround time, while qPCR is not easy to detect a large number of miRNAs. In many studies, the results of circulating miRNA studies have extremely low reproducibility. The test results of different laboratories are not comparable, and may even produce opposite results. A review of 11 studies showed that of the 31 miRNAs associated with heart failure identified in one study, only 5 miRNAs could be repeated in another study, but no miRNA could be repeated in more than two studies. This fully demonstrates that the existing miRNA qPCR detection technology has serious defects. This greatly limits its application in clinical cancer diagnosis. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a kit for detecting breast cancer by second-generation sequencing of plasma miRNA isoforms, which constructs a second-generation sequencing library by selecting miRNA that needs to be amplified, thereby realizing the detection of low-expressed miRNA, so that the sequencing results have the advantages of high sensitivity, high relative sequencing depth and high specificity.
[0005] The present invention provides a kit for detecting breast cancer by second-generation sequencing of plasma miRNA isoforms, comprising primers for amplifying artificial miRNA1 to artificial miRNA6 and primers for amplifying at least one of the following miRNA isoforms: hsa-miR-21-5p, hsa-miR-223-3p, hsa-miR-223-5p, hsa-miR-186-5p, hsa-miR-18a-5p, hsa-miR-146b-5p, hsa-miR-624-5p, hsa-miR-106b-5p, hsa-miR-340-5p, hsa-miR-20a-5p, hsa-miR-451a, hsa-mi R-7976, hsa-miR-2355-3p, hsa-miR-301a-3p, hsa-miR-144-5p, hsa-miR-151a-3p, hsa-miR-3200-5p, hsa-miR-1537-3p, hsa-miR-500a-5p, hsa-mi R-127-3p, hsa-miR-570-3p, hsa-miR-130b-5p, hsa-miR-503-5p, hsa-miR-551a, hsa-miR-409-3p, hsa-miR-330-3p, hsa-miR-889-3p, hsa-miR-625 -5p, hsa-miR-542-3p, hsa-miR-582-3p, hsa-miR-381-3p, hsa-miR-495-3p, hsa-miR-103a-1-5p, hsa-miR-450b-5p, hsa-miR-429, hsa-miR-576-5p , hsa-miR-148b-3p, hsa-miR-320c, hsa-miR-4286, hsa-miR-126-3p, hsa-miR-152-3p, hsa-miR-144-3p, hsa-miR-195-5p, hsa-let-7a-5p, hsa-miR -378f, hsa-miR-126-5p, hsa-miR-26a-5p, hsa-miR-29a-3p, hsa-miR-181a-5p, hsa-miR-32-5p, hsa-miR-142-3p, hsa-miR-29c-3p, hsa-miR-424-5 p, hsa-miR-192-5p, hsa-miR-143-3p, hsa-miR-30c-5p, hsa-miR-146a-5p, hsa-miR-101-3p, hsa-miR-19b-3p, hsa-miR-33b-5p, hsa-miR-378a-3p,hsa-miR-22-3p, hsa-miR-107, hsa-miR-497-5p, hsa-miR-15a-3p, hsa-miR-188-5p, hsa-let-7d-3p, hsa-miR-132-3p, hsa-miR-151a-5p, hsa-miR-194-5p, h sa-miR-99a-5p, hsa-miR-125b-5p, hsa-miR-25-3p, hsa-miR-103a-3p, hsa-miR-1285-3p, hsa-miR-7977, hsa-miR-30b-5p, hsa-miR-363-3p, hsa-miR-93-5p , hsa-miR-375-3p, hsa-miR-99b-5p, hsa-miR-193b-3p, hsa-miR-324-3p, hsa-miR-193a-3p, hsa-miR-342-3p, hsa-miR-484, hsa-miR-532-3p, hsa-miR-210- 3p, hsa-miR-2110, hsa-miR-296-5p, hsa-miR-1307-5p, hsa-miR-19a-3p, hsa-miR-139-5p, hsa-miR-3665, hsa-miR-RG-84, hsa-miR-4454, hsa-let-7b-5p;,
[0006] The primers for amplifying miRNA isoforms include nucleotide sequences as shown in SEQ ID NO: 1 to SEQ ID NO: 97;
[0007] The nucleotide sequences of the primers for amplifying artificial miRNA1 to artificial miRNA6 are shown in SEQ ID NO: 123 to SEQ ID NO: 128, respectively.
[0008] Preferably, the kit further comprises a second pre-amplification PCR primer pair and / or a third pre-amplification PCR primer pair;
[0009] The second pre-amplification PCR primer pair includes a transition primer and a reverse primer for amplifying miRNA isoforms;
[0010] The nucleotide sequence of the transition primer is shown in SEQ ID NO: 99;
[0011] The nucleotide sequence of the reverse primer for amplifying the miRNA isoform is shown in SEQ ID NO: 100;
[0012] The third pre-amplification PCR primer pair includes a 5' universal primer and a 3' universal primer;
[0013] The nucleotide sequence of the 5' universal primer is shown in SEQ ID NO: 101;
[0014] The nucleotide sequence of the 3' universal primer is shown in SEQ ID NO:102.
[0015] Preferably, the kit further comprises primers for adding sequencing adapters and primers for adding barcode tags;
[0016] The forward primer of the primer adding the sequencing adapter is obtained by sequentially connecting the DNA fragment shown in SEQ ID NO: 103, the I5 Index sequence and the DNA fragment sequence shown in SEQ ID NO: 104;
[0017] The reverse primer of the primer for adding the sequencing adapter is obtained by sequentially connecting the DNA fragment shown in SEQ ID NO: 105, the 7Index sequence and the DNA fragment sequence shown in SEQ ID NO: 106;
[0018] The forward primer of the primer with barcode tag is obtained by sequentially connecting the DNA fragment shown in SEQ ID NO: 107, the I5Index sequence and the DNA fragment sequence shown in SEQ ID NO: 108;
[0019] The reverse primer of the primer with barcode tag is obtained by sequentially connecting the DNA fragment shown in SEQ ID NO: 109, the I5Index sequence and the DNA fragment sequence shown in SEQ ID NO: 110.
[0020] Preferably, the kit further comprises a reverse transcription primer;
[0021] The nucleotide sequence of the reverse transcription primer is shown in SEQ ID NO:111.
[0022] The present invention provides application of the kit in constructing a breast cancer second-generation sequencing library.
[0023] The present invention provides a method for constructing a breast cancer second-generation sequencing library, comprising the following steps:
[0024] RNA from breast cancer samples was reverse transcribed to obtain cDNA;
[0025] Using the cDNA as a template, performing a first PCR pre-amplification using the primers described above, to obtain a first pre-amplification product;
[0026] Using the first pre-amplification product as a template, and using the second pre-amplification PCR primer, perform a second PCR pre-amplification to obtain a second pre-amplification product;
[0027] Using the second pre-amplification product as a template, and using the third pre-amplification PCR primers, a third PCR pre-amplification is performed to obtain a third pre-amplification product;
[0028] Using the third pre-amplification product as a template, and using the primers with sequencing adapters to perform a first PCR amplification to obtain a PCR product containing an internal unique double tag;
[0029] The PCR product with the internal unique double tag is used as a template, and the primers with the barcode tag are used for the second PCR amplification to obtain double unique double tag PCR products, and the samples are mixed to obtain a sequencing library.
[0030] The present invention provides a miRNA isoform composition related to breast cancer diagnosis, comprising nucleotide sequences shown as SEQ ID NO: 129 to SEQ ID NO: 210.
[0031] The present invention provides application of the miRNA isoform composition in constructing a breast cancer prediction model.
[0032] Preferably, the machine learning classifier of the breast cancer prediction model includes a support vector classifier.
[0033] The present invention provides an application of primers for amplifying the miRNA isoform composition in preparing a kit for diagnosing breast cancer.
[0034] The present invention provides a kit for detecting breast cancer by second-generation sequencing of plasma miRNA isoforms, including amplification primers of miRNA isoforms closely related to breast cancer diagnosis. Therefore, the kit has the advantages of high sensitivity, high relative sequencing depth and high specificity, and can detect miRNA that cannot be detected by conventional methods. Moreover, due to significant amplification, high-expression miRNA is avoided. For the same sequencing depth, for the same low-expression miRNA, the relative sequencing depth of the library constructed by the kit of the present invention can be much higher than that of conventional technology. The kit of the present invention greatly improves the detection efficiency of breast cancer.
[0035] The present invention also provides a miRNA isoform composition related to breast cancer diagnosis. The present invention uses an SVM classifier based on the second-generation sequencing data of the miRNA isoform composition to construct a breast cancer prediction model, which shows good characteristics in distinguishing cancer and normal samples, with an accuracy of 93.48%, a sensitivity of 95.65%, and a specificity of 91.30%. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a flowchart for analyzing the results of next-generation sequencing of breast cancer miRNA isoforms;
[0037] Figure 2 is the relationship between NGS reads and input concentration of artificial miRNA;
[0038] Figure 3 It is a volcano plot of differential expression of miRNA isoforms;
[0039] Figure 4 is the heat map of differential expression of miRNA isoforms;
[0040] Figure 5 Confusion matrix and classification statistics for SVM classifier;
[0041] Figure 6 ROC curve for the SVM machine learning classifier. DETAILED DESCRIPTION
[0042] The present invention provides a kit for detecting breast cancer by second-generation sequencing of plasma miRNA isoforms, comprising primers for amplifying artificial miRNA1 to artificial miRNA6 and primers for amplifying at least one of the following miRNA isoforms: hsa-miR-21-5p, hsa-miR-223-3p, hsa-miR-223-5p, hsa-miR-186-5p, hsa-miR-18a-5p, hsa-miR-146b-5p, hsa-miR-624-5p, hsa-miR-106b-5p, hsa-miR-340-5p, hsa-miR-20a-5p, hsa-miR-451a, hsa-mi R-7976, hsa-miR-2355-3p, hsa-miR-301a-3p, hsa-miR-144-5p, hsa-miR-151a-3p, hsa-miR-3200-5p, hsa-miR-1537-3p, hsa-miR-500a-5p, hsa-mi R-127-3p, hsa-miR-570-3p, hsa-miR-130b-5p, hsa-miR-503-5p, hsa-miR-551a, hsa-miR-409-3p, hsa-miR-330-3p, hsa-miR-889-3p, hsa-miR-625 -5p, hsa-miR-542-3p, hsa-miR-582-3p, hsa-miR-381-3p, hsa-miR-495-3p, hsa-miR-103a-1-5p, hsa-miR-450b-5p, hsa-miR-429, hsa-miR-576-5p , hsa-miR-148b-3p, hsa-miR-320c, hsa-miR-4286, hsa-miR-126-3p, hsa-miR-152-3p, hsa-miR-144-3p, hsa-miR-195-5p, hsa-let-7a-5p, hsa-miR -378f, hsa-miR-126-5p, hsa-miR-26a-5p, hsa-miR-29a-3p, hsa-miR-181a-5p, hsa-miR-32-5p, hsa-miR-142-3p, hsa-miR-29c-3p, hsa-miR-424-5 p, hsa-miR-192-5p, hsa-miR-143-3p, hsa-miR-30c-5p, hsa-miR-146a-5p, hsa-miR-101-3p, hsa-miR-19b-3p, hsa-miR-33b-5p, hsa-miR-378a-3p,hsa-miR-22-3p, hsa-miR-107, hsa-miR-497-5p, hsa-miR-15a-3p, hsa-miR-188-5p, hsa-let-7d-3p, hsa-miR-132-3p, hsa-miR-151a-5p, hsa-miR-194-5p, h sa-miR-99a-5p, hsa-miR-125b-5p, hsa-miR-25-3p, hsa-miR-103a-3p, hsa-miR-1285-3p, hsa-miR-7977, hsa-miR-30b-5p, hsa-miR-363-3p, hsa-miR-93-5p , hsa-miR-375-3p, hsa-miR-99b-5p, hsa-miR-193b-3p, hsa-miR-324-3p, hsa-miR-193a-3p, hsa-miR-342-3p, hsa-miR-484, hsa-miR-532-3p, hsa-miR-210- 3p, hsa-miR-2110, hsa-miR-296-5p, hsa-miR-1307-5p, hsa-miR-19a-3p, hsa-miR-139-5p, hsa-miR-3665, hsa-miR-RG-84, hsa-miR-4454, hsa-let-7b-5p;,
[0043] The primers for amplifying miRNA isoforms include nucleotide sequences as shown in SEQ ID NO: 1 to SEQ ID NO: 97;
[0044] The nucleotide sequences of the primers for amplifying artificial miRNA1 to artificial miRNA6 are shown in SEQ ID NO: 123 to SEQ ID NO: 128, respectively.
[0045] In the present invention, the primer for amplifying miRNA isomers is obtained by sequentially connecting the universal sequence shown in nucleotide sequence such as SEQ ID NO:98 (ATAGACTCCTCGCATAGCCTCATGAGTC) and the 5' end partial sequence of any one of the miRNA isomers. The 5' end partial sequence of the miRNA isomer refers to the sequence of 10 to 14 nt in length before the 5' end of the miRNA isomer, which can be 11nt, 12nt and 14nt. The primer ensures the specific amplification of microRNAs while also amplifying different types of isomers of all specific microRNAs, and also has the flexibility of the measurement object.
[0046] In the present invention, the primers for amplifying artificial miRNA1 to artificial miRNA6 are obtained by sequentially connecting the universal sequence shown in SEQ ID NO:98 and the 5' end partial sequence of any one of artificial miRNA1 to artificial miRNA6. The 5' end partial sequence of any one of artificial miRNA1 to artificial miRNA6 refers to the sequence of 10 to 14 nt in length before the 5' end of any one of artificial miRNA1 to artificial miRNA6, which can be 11 nt, 12 nt and 14 nt. The nucleotide sequences of artificial miRNA1 to artificial miRNA6 are shown in SEQ ID NO:117 to SEQ ID NO:122. The artificial miRNA1 to artificial miRNA6 are microRNA reference substances, which are added to breast cancer plasma samples and normal human plasma samples as internal labels to correct errors. At the same time, the added concentration is known, and it can also be used for absolute quantification of other natural miRNA isomers and calibration of experimental samples.
[0047] In the present invention, the kit preferably further comprises a second pre-amplification PCR primer pair and / or a third pre-amplification PCR primer pair. The second pre-amplification PCR primer pair comprises a transition primer and a reverse primer for amplifying miRNA isoforms. The nucleotide sequence of the transition primer is shown in SEQ ID NO: 99; the nucleotide sequence of the reverse primer for amplifying miRNA isoforms is shown in SEQ ID NO: 100. The third pre-amplification PCR primer pair comprises a 5' universal primer and a 3' universal primer. The nucleotide sequence of the 5' universal primer is shown in SEQ ID NO: 101; the nucleotide sequence of the 3' universal primer is shown in SEQ ID NO: 102.
[0048] In the present invention, the kit preferably further comprises a primer for adding a sequencing adapter and a primer for adding a barcode tag. The forward primer of the primer for adding a sequencing adapter is obtained by sequentially connecting the DNA fragment shown in SEQ ID NO: 103, the I5 Index sequence and the DNA fragment sequence shown in SEQ ID NO: 104; the reverse primer of the primer for adding a sequencing adapter is obtained by sequentially connecting the DNA fragment shown in SEQ ID NO: 105, the 7 Index sequence and the DNA fragment sequence shown in SEQ ID NO: 106. The forward primer of the primer for adding a barcode tag is obtained by sequentially connecting the DNA fragment shown in SEQ ID NO: 107, the I5 Index sequence and the DNA fragment sequence shown in SEQ ID NO: 108. The reverse primer of the primer for adding a barcode tag is obtained by sequentially connecting the DNA fragment shown in SEQ ID NO: 109, the I5 Index sequence and the DNA fragment sequence shown in SEQ ID NO: 110. The amplification of the primer for adding a sequencing adapter and the primer for adding a barcode tag is conducive to increasing the sequencing tags of each sample, making it easy to identify the sequencing results of each sample and not easy to confuse.
[0049] In the present invention, the I5 Index sequence and the I7 Index sequence refer to the I5 Index and I7 Index recorded in Table 2 of the patent publication number CN118421787A. In the present invention, the I5 Index sequence and the I7 Index sequence are combined in a set. The screening method of the I5 Index sequence and the I7 Index sequence is preferably to randomly generate a short sequence of 10 unique bases, remove the complementary sequence, and then screen out the double unique double tags according to the following standards. The screening criteria preferably include that the same base is not repeated more than three times; each sequence is not severely complementary to other sequences; each sequence is at least 3 bases different from other sequences; the two sequences of the same double unique double tag do not affect the specific amplification of the primers after combining with the surrounding sequences, that is, do not increase the possibility of primer dimers; by calculating the score of the pairing of the two sequences, select the pair with the smallest score (i.e., the best specificity) as a set of combined labeling forward and reverse primers. A total of 976 pairs of double unique double tags for labeling high-throughput samples of second-generation sequencing were obtained through the above screening. Because different dual unique double-label sequences differ by at least 3 bases, even if there are certain sequencing errors in sequencing, under the condition of allowing one base mismatch, it still maintains its uniqueness and will not become other unique double-index tags, so the accuracy of the tag is very high. This also shows higher superiority than conventional methods, because conventional methods have much more tag sequences, such as 2000 sequences for 1000 samples, and the probability of wrong tagging is at least 5 times that of our method. The present invention has experimentally analyzed a large number of samples, with a data volume of more than 400G, and no mismatched second-generation sequencing reads were found.
[0050] In the present invention, the kit preferably further comprises a reverse transcription primer. The nucleotide sequence of the reverse transcription primer is shown in SEQ ID NO: 109. The reverse transcription primer is used for reverse transcription of total RNA.
[0051] The present invention has no particular limitation on the source of the primers, and any gene synthesis method known in the art may be used.
[0052] The present invention provides application of the kit in constructing a breast cancer second-generation sequencing library.
[0053] The present invention provides a method for constructing a breast cancer second-generation sequencing library, comprising the following steps:
[0054] RNA from breast cancer samples was reverse transcribed to obtain cDNA;
[0055] Using the cDNA as a template and using the primers to perform a first PCR pre-amplification to obtain a first pre-amplification product;
[0056] Using the first pre-amplification product as a template, and using the second pre-amplification PCR primer to perform a second PCR pre-amplification to obtain a second pre-amplification product;
[0057] Using the second pre-amplification product as a template and the third pre-amplification PCR primers to perform a third PCR pre-amplification to obtain a third pre-amplification product;
[0058] Using the third pre-amplification product as a template, and using the primers with sequencing adapters to perform a first PCR amplification to obtain a PCR product containing an internal unique double tag;
[0059] The PCR product with the internal unique double tag is used as a template, and the primers with the barcode tag are used for the second PCR amplification to obtain double unique double tag PCR products, and the samples are mixed to obtain a sequencing library.
[0060] The present invention extracts total RNA from breast cancer samples respectively, and obtains cDNA through reverse transcription.
[0061] The present invention has no particular limitation on the method for extracting total RNA, and any method for extracting total RNA known in the art may be used, such as a commercial kit method.
[0062] In the present invention, the reverse transcription includes PolyA reaction, denaturation reaction and reverse transcription reaction. The system of the PolyA reaction is preferably 20 μl, including the following reagents: 5× reverse transcription buffer 4 μl, 10 mM ATP 2 μl, 5000 U / μl PolyA enzyme 1 μl, 40000 U / μl RNase inhibitor 0.5 μl, RNA sample 12.5 μl. The conditions of the PolyA reaction are preferably 37°C for 30 min and 65°C for 20 min. The system of the denaturation reaction is preferably 20 μl, including the following reagents: 10 mM dNTPs 1.5 μl, 10 μM reverse transcription primer (USRTPn) 1.5 μl, Poly A reaction product 17 μl. The primer of the reverse transcription is preferably USRTPn, and the corresponding nucleotide sequence is such as CCTCCATCCGAGACACACGATTGATGGTTTTTTTTTTTTTTTTTTVN, SEQ ID NO: 111). The conditions of the denaturation reaction are preferably 65°C for 5 minutes, then taken out and immediately placed in an ice bath 1 second before the end of the reaction, and then placed in an ice bath for 1 minute. The system of the reverse transcription reaction is preferably 30 μl, including the following reagents: 5× reverse transcription buffer 2 μl, 1.6M trehalose 4.5 μl, 1mg / μl Actinomycin D 1.2 μl, T4gp32 / RecA / ATP mixed solution 1.5 μl, 40000U / μl RNA Inhibitor 0.3 μl, 50U / μl Maxima H reverse transcriptase 1.5 μl, and denaturation reaction product 19 μl. The preparation method of the T4gp32 / RecA / ATP mixed solution is preferably prepared according to the number of samples of 2, 10μg / μl T4gp32 0.6 μl, 2μg / μl Tth RecA 0.2 μl, 100mM ATP 0.24 μl, and 1× reverse transcription buffer 1.96 μl. The conditions of the reverse transcription reaction are preferably 42°C for 15 min, 50°C for 30 min, 55°C for 30 min, 60°C for 30 min, 65°C for 30 min, and 85°C for 5 min.
[0063] After obtaining the cDNA, the present invention uses the cDNA as a template and utilizes the amplification primer set to perform a first PCR pre-amplification to obtain a first pre-amplification product.
[0064] In the present invention, the reaction system for the first PCR pre-amplification is preferably a 20 μl system, including the following reagents: 2×Boost mix 10 μl, 0.2 μg / μl Tth RecA 1 μl, 1 μM amplification primer set 1.5 μl, cDNA 7.5 μl. The components and preparation method of 2×Boostmix can be found in Patent No.: ZL 20191021982
[0065] 7.4, the specific quantitative PCR reaction mixture described in Example 1 in the patent titled "A specific quantitative PCR reaction mixture, a miRNA quantitative detection kit, and a detection method", but the difference is that the 2×boostmix (containing UDG) is prepared with a dNTPs mixture that does not contain dUTP. The reaction program for the first PCR pre-amplification is preferably ① 25°C 10min, ② 95°C 10min, ③ (95°C 10s, 55°C 10min) 3 cycles, ④ (95°C 10s, 50°C 10min) 3cycles, ⑤ (95°C 10s, 45°C 10min) 2 cycles, ⑥ (95°C 10S, 40°C 10min) 2 cycles, ⑦ (95°C 10S, 37°C 10min) 2 cycles, ⑧ (95°C 10S, 60°C 2min 72°C 10min) 1 cycle, ⑨ the program runs to 72°C 5min, then take out the PCR tube and immediately put it in an ice bath. The first PCR pre-amplification is conducive to amplifying a large number of miRNA isomers from the reverse transcription product. The first pre-amplification product is purified and then treated with EXO I enzyme to remove the PCR primers in the system.
[0066] In the present invention, the breast cancer sample preferably also includes any three of the artificial miRNA1 to artificial miRNA6, and the first pre-amplification PCR is performed using the corresponding primers for amplifying any three of the artificial miRNA1 to artificial miRNA6. At the same time, the normal healthy human plasma sample is added with the amplification primers of the remaining three artificial miRNAs as internal standards for PCR amplification. The present invention has no special restrictions on the PCR amplification method, and the reaction system and reaction procedure of the first PCR pre-amplification can be used.
[0067] After obtaining the first pre-amplification product, the present invention uses the first pre-amplification product as a template, and uses the transition primer in the amplification primer set and the reverse primer for amplifying microRNA isomers to perform a second PCR pre-amplification to obtain a second pre-amplification product.
[0068] In the present invention, the reaction system of the second PCR pre-amplification is preferably 20 μl, including the following reagents: 10 μl of the 2×Boost mix, 1 μl of 10 μm transition primer (USEXPnb), 1 μl of 10 μm IsomiR primer, 1 μl of 0.2 μg / μl Tth RecA, and 7 μl of the first pre-amplification product. The reaction program of the second PCR pre-amplification is preferably ①25°C 10min, ②95°C 10min, ③(95°C 10s, 65°C 1min) 3cycles, ④(95°C 10s, 62°C 1min) 3cycles, ⑤(95°C 10s, 58°C 2min) 2 cycles, ⑥(95°C 10S, 60°C 2min) 2 cycles, ⑦(95°C 10S, 60°C 2min 72°C 10min) 1 cycle, ⑧The program is run to 72°C 5min and then the PCR tube is taken out for ice bath. The second pre-amplification product is preferably purified by magnetic beads. The second PCR pre-amplification is performed by a transition primer and a reverse primer of the microRNA isomer, with the purpose of introducing a binding site for the 3' universal primer and the 5' universal primer.
[0069] The second pre-amplification product is used as a template, and the 5' universal primer and the 3' universal primer in the amplification primer set are used to perform a third PCR pre-amplification, and the obtained third pre-amplification product is a microRNA isomer.
[0070] In the present invention, the reaction system of the third PCR pre-amplification is preferably 20 μl, including the following reagents: 10 μl of the 2×Boost mix, 1 μl of 10 μm URP, 1 μl of 10 μm UFP, 1 μl of 0.2 μg / μl Tth RecA, and 7 μl of the second pre-amplification product. The reaction procedure of the third PCR pre-amplification is preferably ① 95°C for 10 min, ② (95°C for 10 s, 65°C for 1 min), 12 cycles, ④ 72°C for 10 min, and ⑤ 72°C for 5 min followed by an ice bath. The third pre-amplification product is purified and treated with EXO I enzyme to remove the PCR primers.
[0071] In the present invention, after obtaining the third pre-amplification product, it is preferred to use the third pre-amplification product as a template to perform qPCR amplification so as to obtain the expression level of miRNA isomers as part of quality control. The forward primer of the qPCR amplification is preferably the 5' universal primer. The reverse primer of the qPCR amplification is preferably the 3' universal primer. The probe of the qPCR amplification is preferably the LNAFAM probe, and the corresponding nucleotide sequence is shown in SEQ ID NO:112 (ACC+AT+CA+AT+CG+TG+TG, + represents locked nucleic acid). The reaction system of the qPCR amplification is preferably 10 μl, preferably including the following reagents: 0.08 μl of the third pre-amplification product diluted by times, 5 μl of 2× DNA polymerase mixture, a final concentration of 0.2 μM forward primer and a final concentration of 0.2 μM reverse primer and a final concentration of 0.2 μM probe, and ddH2O is filled with 10 μl. The reaction program of the qPCR amplification is preferably 95°C for 10 min; 95°C for 30 s, 65°C for 1 min, and 40 cycles.
[0072] After obtaining the third pre-amplification product, the present invention uses the third pre-amplification product as a template and uses primers with sequencing adapters to perform a first PCR amplification to obtain a PCR product containing an internal unique double tag.
[0073] In the present invention, the reaction system of the first PCR amplification is preferably 30 μl, comprising the following steps: 15 μl of 2×PCR enzyme (containing UDG and UTP), 0.5 μl of each forward and reverse primer for 10 μM plus internal double unique double label, 2 μl of the third pre-amplification product, and the remainder of water. The reaction program of the PCR amplification is preferably ① 95°C 10min, ② (95°C 15s, 62°C 30s, 72°C 1min) 3cycles, ③ (95°C 15s, 64°C 30s, 72°C 1min) 2cycles, ④ (95°C 15s, 68°C 30s, 72°C 1min) 11cycles, ⑤ (95°C 15s, 72°C 20min) 1cycle, ⑥ The program runs to 72°C 18min and then ice baths are performed.
[0074] Using the PCR product with the unique double tag as a template, the present invention uses the primer with the barcode tag to perform a second PCR amplification to obtain a double unique double tag PCR product, mix the samples, and obtain a sequencing library.
[0075] In the present invention, the reaction system of the second PCR amplification is preferably 30 μl, including the following steps: 15 μl of 2×PCR enzyme (containing UDG and UTP), 0.5 μl of each forward and reverse primer for 10 μM plus internal double unique double label, 2 μl of the third pre-amplification product, and the rest of water. The reaction program of the PCR amplification is preferably ① 95°C 10min, ② (95°C 15s, 62°C 30s, 72°C 1min) 3cycles, ③ (95°C 15s, 64°C 30s, 72°C 1min) 2cycles, ④ (95°C 15s, 68°C 30s, 72°C 1min) 11cycles, ⑤ (95°C 15s, 72°C 20min) 1cycle, ⑥ The program runs to 72°C 18min and then ice baths are performed.
[0076] In the present invention, the PCR product containing the dual unique dual tags is preferably mixed and then precipitated, and the product precipitation is used to remove the PCR primers with ExoI enzyme to obtain a sequencing library. Removing the PCR primers refers to removing the forward and reverse primers of the dual unique dual tag amplification primers that are unreacted in the above-mentioned PCR process. The purpose of removing the PCR primers is to prevent the downstream sequencing reaction of the PCR primers.
[0077] In the present invention, the PCR product containing dual unique dual tags is obtained based on the dual unique dual index tagging technology for second generation sequencing multiplex developed by the present invention. The dual unique dual index tagging technology for second generation sequencing multiplex uses cDNA as a template, and uses PCR to add internal unique dual tags (IUDI) to both ends of the cDNA fragment, and then uses PCR amplification means to add external unique dual tags and sequencing adapters to both ends of the PCR product obtained in the above steps to obtain PCR products carrying dual unique dual tags, mix, remove PCR primers, and obtain an amplified library for second generation sequencing analysis. After the second generation sequencing is completed, the second generation sequencing raw data is split into the corresponding number of samples when the mixed sample is split according to the dual unique dual tag sequence, and after splitting, it is processed by removing irrelevant sequences to identify and quantify micro RNA isomers.
[0078] In the present invention, the PCR product containing dual unique dual tags is obtained based on the dual unique dual index tagging technology for second generation sequencing multiplex developed by the present invention. The dual unique dual index tagging technology for second generation sequencing multiplex uses cDNA as a template, and uses PCR to add internal unique dual tags (IUDI) to both ends of the cDNA fragment, and then uses PCR amplification means to add external unique dual tags and sequencing adapters to both ends of the PCR product obtained in the above steps to obtain PCR products carrying dual unique dual tags, mix, remove PCR primers, and obtain an amplified library for second generation sequencing analysis. After the second generation sequencing is completed, the second generation sequencing raw data is split into the corresponding number of samples when the mixed sample is split according to the dual unique dual tag sequence, and after splitting, it is processed by removing irrelevant sequences to identify and quantify micro RNA isomers.
[0079] The present invention provides a miRNA isoform composition related to breast cancer diagnosis, comprising nucleotide sequences shown as SEQ ID NO: 129 to SEQ ID NO: 210.
[0080] In the present invention, the miRNA isoform composition is the miRNA isoforms remaining after removing the highly correlated (Pearson's correlation coefficient, Pearson's correlation coefficient, greater than or equal to 0.75) miRNA isoforms among the miRNA isoforms with significant differential expression (calibrated p value less than 0.05) between breast cancer and normal people, and is used for machine learning model construction.
[0081] The present invention provides application of the miRNA isoform composition in constructing a breast cancer prediction model.
[0082] In the present invention, the machine learning classifier of the breast cancer prediction model preferably includes a support vector classifier. Preferably, the parameters of the SVM algorithm are optimized by grid search, and the numerical range of the parameters used is gamma=2 (-8 to 1), cost=2 (0 to 4). The prediction model is validated by 10-fold cross validation. After obtaining the prediction model, the model is preferably evaluated. The evaluation criteria preferably include accuracy and kappa. Accuracy, kappa and other evaluation indicators are preferably described by confusion matrix and ROC curve.
[0083] In the embodiment of the present invention, the sensitivity of the breast cancer prediction model constructed by the second generation sequencing library constructed by the kit of the present invention according to the second generation sequencing results is 0.913, which can correctly identify 91.3% of the actual breast cancer cases without any false positives. The AUC value is 0.9924 and the threshold is 0.6461. The analysis of the ROC curve shows that the SVM classifier performs very well in distinguishing cancer and non-cancer cases, and achieves a strong balance between sensitivity and specificity at the selected threshold.
[0084] The present invention provides an application of primers for amplifying the miRNA isoform composition in preparing a kit for diagnosing breast cancer.
[0085] In the present invention, the primer for amplifying the miRNA isoform composition is obtained by sequentially connecting the universal sequence shown in nucleotide sequence SEQ ID NO: 98 (ATAGACTCCTCGCATAGCCTCATGAGTC) and the 5' end partial sequence of any one of the miRNA isoforms. The 5' end partial sequence of the miRNA isoform refers to the sequence of 10 to 14 nt in length before the 5' end of the miRNA isoform, which can be 11 nt, 12 nt and 14 nt.
[0086] In the present invention, the method for diagnosing breast cancer is based on sequencing the sample to be tested based on the primers, the sequencing data is brought into a breast cancer prediction model based on a support vector classifier, and whether the sample to be tested has a risk of breast cancer is judged according to the prediction result of the breast cancer prediction model: when the prediction result is higher than the threshold value, it indicates that the sample to be tested has a high risk of breast cancer and further medical diagnosis can be performed; when the prediction result is lower than the threshold value, it indicates that the sample to be tested has a low risk of breast cancer.
[0087] The kit for detecting breast cancer by next-generation sequencing of plasma miRNA isoforms provided by the present invention is described in detail below in conjunction with the examples, but they should not be construed as limiting the scope of protection of the present invention.
[0088] In this technical solution, there are 100 breast cancer samples and 100 normal human clinical samples, and the breast cancer samples (100) and the normal human clinical samples (100) are from the Cancer Hospital of the Chinese Academy of Medical Sciences.
[0089] RNA extraction kit was purchased from Thermo Fisher Scientific.
[0090] Example 1
[0091] RNA extraction kit was used to extract breast cancer samples and normal human clinical samples. After the concentration and quality of total RNA were tested by nucleic acid quantitative detector, they were stored at -20℃ for future use.
[0092] Example 2
[0093] Reverse transcription and pre-amplification of IsomiRs from breast cancer samples, including the following steps:
[0094] 1. PolyA reaction:
[0095] The reaction system was prepared according to 20 μl of one sample. Each 20 μl reaction system included the following reagents: 4 μl of 5× reverse transcription buffer, 2 μl of ATP (10 mM), 1 μl of Poly A enzyme (5000 U / μl), 0.5 μl of RNase inhibitor (40000 U / μl), and 12.5 μl of total RNA. The prepared reaction system was subjected to Poly A reaction under the following conditions: 37°C for 30 min, 65°C for 20 min; the reaction was sealed and stored at -5°C (for heat inactivation).
[0096] 2. Reverse Transcription Reaction
[0097] A reaction system was prepared according to 20 μl of one sample, and each 20 μl reaction system specifically included the following reagents: 1.5 μl of 10 mM dNTPs, 1.5 μl of 10 μM reverse transcription primer (USRTPn, CCTCCATCCGAGACACACGATTGATGGTTTTTTTTTTTTTTTTTTVN, SEQ ID NO: 111), and 17 μl of PolyA template; the prepared reaction system was placed at 65° C. for 5 min to perform a denaturation reaction, and was taken out 1 second before the end of the treatment, immediately placed in an ice bath, placed in an ice bath for 1 min, and centrifuged;
[0098] Note: The dNTPs do not contain dUTP, otherwise the reverse transcription product cDNA will be degraded; USEXPnb primers are purified by magnetic beads.
[0099] The denatured reaction product was used to prepare a reverse transcription reaction system, and a 30 μl reaction system was prepared for one sample, specifically including the following reagents: 2 μl 5× reverse transcription buffer, 4.5 μl 1.6 M trehalose, 1.2 μl Actinomycin D (1 mg / μl), 1.5 μl T4gp32 / RecA / ATP mixed solution, 0.3 μl RNase inhibitor (40000 U / μl), 1.5 μl Maxima H reverse transcriptase (50 U / μl), and 19 μl of the denatured reaction product; wherein the T4gp32 / RecA / ATP mixed solution was prepared according to the number of two samples by adding the following reagents (μl): 0.6 μl T4gp32 (10 μg / μl), 0.2 μl Tth RecA (2 μg / μl), 0.24 μl ATP (100 mM), 1× reverse transcription buffer 1.96 μl. The prepared reaction system was placed under the following conditions for reverse transcription reaction: 42°C for 15 min, 50°C for 30 min, 55°C for 30 min, 60°C for 30 min, 65°C for 30 min, 85°C for 5 min.
[0100] 3. Pre-amplification
[0101] (I) First pre-amplification PCR reaction:
[0102] The pre-amplification PCR reaction solution was prepared according to the reaction system configuration of 20 μl per sample. The specific reaction system of 20 μl was as follows: 2×Boost mix*10 μl, Tth RecA (0.2 μg / μl)1 μl, Pre-IsomiR mix* (1 μM)1.5 μl, reverse transcription product 7.5 μl;
[0103] Among them, 2×Boost mix* (containing UDG) is prepared with a dNTPs mixture without dUTP; 2×Boost mix is specifically referred to in patent number: ZL 201910219827.4, and the patent name is a specific quantitative PCR reaction mixture, miRNA quantitative detection kit and detection method, which is a specific quantitative PCR reaction mixture described in Example 1; Pre-IsomiRmix* (1μM): 10μl of each of 97 primers (specific sequences are shown in Table 1) with a mother solution concentration of 100μM, and then 20μl H2O (Nuclease-Free) is added to prepare a primer mix with a final concentration of 1μm (1000μl);
[0104] Table 1 Artificial miRNA sequences and their pre-amplification primer sequences
[0105]
[0106]
[0107]
[0108] Set the PCR instrument's reaction program according to the following conditions: ①25℃10min, ②95℃10min, ③(95℃10s, 55℃10min)3cycles, ④(95℃10s, 50℃10min)3cycles, ⑤(95℃10s, 45℃10min)2cycles, ⑥(95℃10s, 40℃10min)2cycles, ⑦(95℃10s, 37℃10min)2cycles, ⑧(95℃10s, 60℃2min, 72℃10min)1cycle, ⑨After the program runs to 72℃5min, take out the PCR tube and immediately immerse it in an ice box to terminate the Taq Activity of DNA polymerase; after the liquid freezes (about 3 minutes), place the PCR tube on a 96-well insulation module (pre-frozen to -40°C); add an equal volume of 20μl chloroform and immediately vortex until the ice melts (about 1 minute); the centrifuge program is 12000rpm, 4°C, 15min centrifugation, use a pipette to aspirate the supernatant (generally 18μl) into a labeled new PCR tube; centrifuge on a handheld centrifuge until all samples reach the bottom of the tube, open the lid of the PCR tube after chloroform extraction, place it in a PCR instrument at 50°C for 10 minutes to completely evaporate the chloroform; add 2.5μl EXO I enzyme to each tube, mix it upside down, centrifuge it, place it in the PCR instrument, set the program to 37°C for 4min, and pause when there are 5 seconds left in the program; set the PCR program and run: 37°C for 4min, 80°C for 1min; centrifuge until all samples reach the bottom of the tube.
[0109] (ii) Second pre-amplification PCR reaction:
[0110] The reaction system was prepared according to 20 μl per sample. The 20 μl reaction system specifically included the following reagents: 10 μl of 2×Boostmix, 1 μl of transition primer (USEXPnb, TCTACAGATCCTGGCCTCTGACTCCAGGATCTGTAGACCTCCATC CGAGACACACGAT, SEQ ID NO: 99) purified by 10 μm magnetic beads, 1 μl of IsomiR primer (IsomiRupb, GTTTGTTGCTACGCTCAGAATCCTAAGCGTAGCAACAAACATAGACTCCTCGCATAGCCTCATGAG TC, SEQ ID NO: 100) purified by 10 μm magnetic beads, 1 μl of Tth RecA (0.2 μg / μl), and 7 μl of the first pre-amplification PCR product;
[0111] 2×Boost mix* (containing UDG) is prepared with a dNTPs mixture without dUTP, and the formula is the same as above;
[0112] Set the PCR instrument to program Touch Down PCR: ①25℃10min, ②95℃10min, ③(95℃10s, 65℃1min)3cycles, ④(95℃10s, 62℃1min)3cycles, ⑤(95℃10s, 58℃2min)2cycles, ⑥(95℃10s, 60℃2min)2cycles, ⑦(95℃10s, 60℃2min, 72℃10min)1cycle, ⑧After the program runs to 72℃5min, take out the PCR tube and immediately immerse it in the ice box, stop Taq activity; after the liquid freezes (about 3 minutes), place the PCR tube on a 96-well insulation module (pre-frozen to -40°C); add an equal volume of 20 μl chloroform and immediately vortex until the ice melts (about 1 minute); centrifuge at 12000 rpm and 4°C for 15 minutes, use a pipette to draw the supernatant (usually 18 μl) into a labeled new PCR tube; centrifuge until all samples reach the bottom of the tube, open the lid of the PCR tube after chloroform extraction, and place it in P Place the tube on a PCR machine at 50°C for 10 min to completely evaporate the chloroform; add 4 μl of washed streptavidin magnetic beads to every 20 μl of the reaction solution (mix the beads with a vortexer before use and use immediately); set the oscillator to 500 rpm at room temperature for 30 min; fully suspend the beads with a vortexer and incubate at 50°C for 3 min in a PCR machine; place the tube on a magnetic stand for about 1 min after the incubation to adsorb the beads, and use a pipette to aspirate the liquid (try not to aspirate the beads) into another newly labeled PCR tube.
[0113] (III) The third pre-amplification PCR reaction:
[0114] The reaction system was prepared according to 20 μl of one sample. The 20 μl reaction system specifically included the following reagents: 2×Boost mix* 10 μl, 10 μm URP (CAGAATCCTAAGCGTAGCAACAAAC, SEQ ID NO: 101) 1 μl, 10 μm UFP (GCCTCTGACTCCAGGATCTGTAGAC, SEQ ID NO: 102) 1 μl, Tth RecA (0.2 μg / μl) 1 μl, and 7 μl of the second pre-amplification PCR product; wherein, 2×Boost mix* (containing UDG) was prepared with a dNTPs mixed solution without dUTP.
[0115] The PCR instrument was set up with the following program: ① 95℃ for 10 min, ② (95℃ for 10 s, 65℃ for 1 min), 12 cycles, ④ 72℃ for 10 min, ⑤ 72℃ for 5 min. After that, the PCR tube was taken out and immediately immersed in a program cooling box containing isopropanol stored at -80℃ to terminate the activity of TaqDNA enzyme (to avoid non-specific amplification caused by temperature drop). After the liquid was frozen (about 3 min), the PCR tube was placed on a 96-well insulation module (pre-frozen to -40℃). Add an equal volume of 20 μl chloroform and immediately vortex until the ice melts (about 1 min); centrifuge program is 12000 rpm, 4°C, 15 min centrifugation, use a pipette to aspirate the supernatant (generally 18 μl) into a labeled new PCR tube; centrifuge until all samples reach the bottom of the tube, open the lid of the PCR tube after chloroform extraction, place it in the PCR instrument at 50°C for 10 min to completely evaporate the chloroform; add 2.5 μl EXO I (Thermolabile) mixture to each reaction (20 μl); set the PCR program and run: 37°C 4 min, 80°C 1 min; take 5 μl and dilute it 10 times with 0.1×TE as a PCR template for subsequent detection.
[0116] Example 3
[0117] qPCR detection of pre-amplification products
[0118] qPCR amplification detection: USQ-miR DNA polymerase mixture was prepared, containing 0.2 μM (final concentration) of forward primer (UFP: GCCTCTGACTCCAGGATCTGTAGAC, SEQ ID NO: 102), universal reverse primer (URP: GCCTCTGACTCCAGGATCTGTAGAC, SEQ ID NO: 101) and 0.2 μM (final concentration) of LNAFAM probe (ACC+AT+CA+AT+CG+TG+TG, SEQ ID NO: 112, + refers to locked nucleic acid, referred to as LNA), and the amount of PCR template used in 10 μl PCR reaction system was 0.08 μl of 10-fold dilution of the third pre-amplification PCR product;
[0119] The USQ-miR DNA polymerase mixture also contains 2×qPCR premix, the formula of which is: Tris-HCl pH 8.8, 75mM; (NH4)2SO4, 20mM; Triton-100, 0.1%; MgCl2, 2.5mM; dNTPs, 200μM; Trehalose, 200mM; Taq DNA polymerase, 50U / ml.
[0120] The PCR cycle parameters were: 95°C for 10 min, followed by 95°C for 30 s and 65°C for 1 min for 40 cycles.
[0121] Example 4
[0122] PCR amplification of pre-amplification product plus barcode and adapter (sequencing connector):
[0123] (1) Design of primers for adding barcode and sequencing adapter:
[0124] The forward primer is designed as follows:
[0125] Add the sequence overlapping with the forward primer of the sequencing adapter (SEQ ID NO: 103) + IUDI (I5 Index) + the sequence partially overlapping with the universal sequence at the 5' end of the cDNA reverse transcription product (SEQ ID NO: 104);
[0126] The reverse primer is designed as follows:
[0127] Add the sequence overlapping with the reverse primer of the sequencing adapter (SEQ ID NO: 105) + IUDI (I7 Index) + the sequence partially overlapping with the 3' end sequence of the IsomiR primer (SEQ ID NO: 106).
[0128] AATGATACGGCGACCACCGAGATCTACACtacgaatcttACACTCTTTCCCTACACGACGCCTTCCGATCT (SEQ ID NO: 113);
[0129] CTGTCTCTTATACACATCTCCGAGCCCACGAGACaccaagttacCTCGGAGATGTGTATAAGAGAC AG (SEQ ID NO: 114).
[0130] (2) PCR template: The third pre-amplification PCR product was undiluted and one tube was prepared for each sample;
[0131] (3) Preparation of PCR reaction solution: 15 μl of 2× PCR enzyme (containing UDG and UTP), 12 μl of water, 0.5 μl of forward primer (10 μM) with barcode and sequencing adapter, 0.5 μl of reverse primer (10 μM) with barcode and sequencing adapter, 2 μl of the third pre-amplification PCR product, total volume 30 μl.
[0132] (4) Quantitative PCR with barcode and sequencing adapter:
[0133] Each third pre-amplification PCR product was subjected to quantitative PCR using a specific forward primer with barcode and sequencing adapter and a reverse primer with barcode and sequencing adapter. The reaction procedure was as follows: 37°C for 10 min, ① 95°C for 10 min, ② (95°C for 15 s, 62°C for 30 s, 72°C for 1 min) for 3 cycles, ③ (95°C for 15 s, 64°C for 30 s, 72°C for 1 min) for 2 cycles, ④ (95°C for 15 s, 68°C for 30 s, 72°C for 1 min) for 45 cycles, and the fluorescence signal was collected at 72°C for 1 min.
[0134] (5) Repeat the above PCR reaction, but use the dilution factor and logarithmic cycle number of the third pre-amplification PCR product of each sample determined by the above qPCR, and directly amplify using a common PCR instrument. Try to use the same parameters as the above qPCR program, including the temperature rise rate and fall rate. Determine the logarithmic cycle number of the barcode PCR reaction.
[0135] (6) Ordinary PCR amplification program: ① 95℃10min, ② (95℃15s, 62℃30s, 72℃1min) 3 cycles, ③ (95℃15s, 64℃30s, 72℃1min) 2 cycles, ④ (95℃15s, 68℃30s, 72℃1min) 11 cycles, ⑤ (95℃15s, 72℃20min) 1 cycle, ⑥ After the program runs to 72℃18min, take out the PCR tube and immediately put it on ice to terminate the enzyme activity. *Set the fluorescence signal to be collected at 72℃1min.
[0136] (7) Add an equal volume of 30 μl of chloroform on ice (chloroform was placed on ice for 30 min) and vortex (about 1 min).
[0137] (8) Centrifuge at 12,000 rpm and 4°C for 15 min. Transfer 25 μl of the supernatant to a new tube (do not allow the tip to come into contact with chloroform; keep a portion of the supernatant).
[0138] (9) Centrifuge the sample until all samples are at the bottom of the tube. Place the supernatant in a PCR instrument and heat at 50°C for 10 min to completely evaporate the chloroform, otherwise chloroform will inhibit downstream enzyme reactions.
[0139] (10) Add 2.5 μl of diluted EXOI (Thermolabile) solution to each PCR reaction.
[0140] (11) Mix by inverting, centrifuge for 20 min at 37°C and 10 min at 42°C.
[0141] (12) Treat at 60°C for 15 min to inactivate ExoI.
[0142] (13) Run 3% agarose gel for 45–60 min with a 50 bp marker to see whether the primer band disappears.
[0143] Example 5
[0144] Quantitative qPCR and amplification of adapter PCR products (2×PCR enzyme contains UTP, not UDG):
[0145] Quantitative qPCR of barcode and adapter PCR products:
[0146] The PCR product sample obtained in the above experiment was diluted 50 times as a template, and the primers used were designed as follows: forward primer: sequence containing the I5 sequencing adapter (SEQ ID NO: 107) + external unique double tag (I5 Index) + sequence overlapping with the 5' end of IUDI (SEQ ID NO: 108);
[0147] Reverse primer: contains the sequence of Nextera I7 sequencing adapter (SEQ ID NO: 109) + external unique double tag (I7 Index sequence) + overlap with the 5' end of IUDI in the above reverse primer (SEQ ID NO: 110);
[0148] Wherein, the I5 Index and I7 Index sequences are selected from the set of I5 Index and I7 Index sequences in Table 2, but are different from the I5 Index and I7 Index sequences involved in the first PCR amplification;
[0149] Examples of primers for the second PCR with external unique dual tag (OUDI) are as follows:
[0150] Forward primer: AATGATACGGCGACCACCGAGATCTACACtacgaatcttACACTCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID NO: 115);
[0151] Reverse primer: CTGTCTCTTATACACATCTCCGAGCCCACGAGACaccaagttacCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO: 116);
[0152] Table 2 I5 sequencing adapter and Nextera I7 sequencing adapter
[0153] Connector Name sequence Sequence number I5 Sequencing Adapter AATGATACGGCGACCACCGAGATCTACA SEQ ID NO: 107 NexteraI7 Sequencing Adapter CTGTCTTCTTATACACATCTCCGAGCCCACGAGA SEQ ID NO: 109
[0154] The qPCR reaction program is as follows: ① 95℃10min, ② (95℃15s, 62℃30s, 72℃1min) 3 cycles, ③ (95℃15s, 64℃30s, 72℃1min) 2cycles, ④ (95℃15s, 68℃30s, 72℃1min) 45cycles, set at 72℃1min to collect fluorescence signals. qPCR was repeated 3 times for each sample;
[0155] According to the quantitative qPCR results above, the dilution factor was calculated, and conventional PCR amplification (low cycle number to prevent the introduction of human errors) was performed with 6 wells per sample;
[0156] Common PCR amplification conditions: ①95℃10min, ②(95℃15s, 62℃30s, 72℃1min) 3 cycles, ③(95℃15s, 64℃30s, 72℃1min) 2cycles, ④(95℃15s, 68℃30s, 72℃1min) 11cycles to be determined, ⑤(95℃15s, 72℃20min) 1cycle, ⑥After the program runs to 72℃18min, take out the PCR tube and put it on ice immediately to terminate Taq activity. *The number of cycles is determined by the above qPCR (if it cannot be continued, store at 4℃, or if it is to be stored for a long time at -20℃);
[0157] For each PCR reaction, 2.5 μl of diluted EXOI (Thermolabile) was added;
[0158] Invert and mix, centrifuge for 20 min at 37°C, then 10 min at 42°C;
[0159] Treat at 60°C for 15 min to inactivate ExoI;
[0160] All samples were placed on ice, and the PCR samples of 6 wells of each sample were mixed, and then quantified by qPCR (each sample was diluted 100,000 times. Three replicates were performed starting from the dilution, and 45 cycles were performed).
[0161] Example 6
[0162] Precipitation and gel recovery of mixed samples of PCR products with barcode and adapter:
[0163] According to the qPCR quantitative results, mix all samples in equal amounts and vortex to mix. Make two tubes according to the following experimental steps;
[0164] Take 700μl of the mixed sample with the adapter into a 1.5ml EP tube; add 77μl of 3M pH 5.2 sodium acetate solution; add 500μl of isopropanol and mix well (the above samples need to be placed on ice); -20℃ or -80℃ for 1h, put the centrifuge tube into the centrifuge (with the lid handle facing outward), centrifuge at 15000g for 30min at 4℃; the white DNA precipitate is at the bottom of the tube facing outward, and the supernatant is carefully poured out. The remaining supernatant is carefully removed with a gun, and the DNA precipitate should not be touched during the process to avoid DNA removal; add 500μl of 70% room temperature ethanol, leave at room temperature for 5min; centrifuge at 15000g for 30min 4℃; the white DNA precipitate is facing outward at the bottom of the tube, and the supernatant is carefully poured out. The remaining supernatant is carefully removed with a P200 gun; it is placed flat in an ultra-clean workbench (the lid of the centrifuge tube should not be covered), and the air is turned on for about 10 minutes; 60μl TE is added to dissolve; 1.5% agarose gel is prepared (the gel thickness is about 1cm, which can hold 15μl of sample, and the length is twice as long as usual, i.e. about 15cm); 3-4 wells of gel are run for recovery. Note: the dye band should run to the bottom of the gel, otherwise the size of the DNA cannot be fully separated; when cutting the gel, the smaller the gel strip, the better, but the main band should be strictly included; the recovered DNA is dissolved in 60μl TE, the DNA concentration is determined, the gel is run for identification, and it is placed at -20℃ for use; the mixed solution of all samples is used for precipitation and gel recovery. If the PCR product is of high purity, gel recovery is not required. After the product is precipitated, the PCR primer is removed with the above-mentioned ExoI to obtain a micro RNA isomer library. The established microRNA isoform library was subjected to NGS high-throughput sequencing to obtain sequencing data results.
[0165] Example 7
[0166] Building a machine learning model
[0167] Based on the sequencing data results in Example 6, the t-Test P value of the expression difference of each IsomiR in 100 breast cancer samples and 100 normal samples was first calculated, and the IsomiRs were arranged from small to large according to the P value. The first 239 IsomiRs of all isomers arranged from small to large P values were selected, and the IsomiRs that were highly correlated with other IsomiRs were removed after further calculating the correlation between different IsomiRs. The remaining data were machine-learned and classified using different classifiers such as SVM, KNN, RF, CART and IDA to find the best classifier. A variety of classifiers were used, and the data were first divided into two parts: 80% of the part was used to train the model, and 20% of the part was used to verify the model. According to the accuracy and Kappa value, the best classifier result was the Support Vector Machine (SVM) algorithm.
[0168] Use the SVM algorithm to establish a machine learning model for breast cancer auxiliary diagnosis: First, divide the above second-generation sequencing data into two parts, 80% of which is used to train the model (training set) and 20% of which is used to verify the model (test set). The samples are divided into training set and test set. Repeated samples only exist in the training set or the test set, that is, different repeats of the same sample cannot exist in both the training set and the test set, otherwise it will lead to information leakage and the model evaluation will be too high.
[0169] Optimization of SVM algorithm model: The parameters of SVM algorithm need to be debugged to find the best parameters. Use grid search to optimize SVM algorithm parameters. The parameter value range is gamma = 2 (-8:1) , cost = 2 (0:4) . In this way, gamma has 10 values and cost has 5 values, that is, 50 combinations. Each combination undergoes 10-fold cross validation, that is, the training set is divided into ten parts, 9 of which are used as training data and 1 as test data in turn for testing. Each test will result in a corresponding error rate. Ten tests give the average error rate of each combination. The gamma / cost combination with the lowest average error rate is the optimal parameter of the SVM algorithm. Since the final diagnostic model is obtained after 500 (50×10) experiments, overfitting can be avoided. Overfitting is a phenomenon in which the trained model performs well on the training set but performs poorly on the test set.
[0170] Model evaluation: There are many different metrics to evaluate machine learning algorithms. The default evaluation criteria for classification problems are accuracy and kappa. Kappa is similar to accuracy, but it is calibrated by a random baseline of the data set. The kappa value represents the consistency and the accuracy of the classification. A positive value close to 1 represents better consistency. Usually, 0.75 or more represents a satisfactory consistency result, and 0.8-1 is almost completely consistent. Accuracy, kappa and other evaluation indicators can be described by the confusion matrix and ROC curve.
[0171] like Figure 1 As shown: 200 samples, including 100 breast cancer samples and 100 normal samples, were labeled with 200 pairs of internal unique dual indexes (IUDI), and then a sequencing adapter containing external unique dual indexes (OUDI, ACCAAGTTAC and AAGATTCGTA) was added and mixed together and sequenced using NovaSeq TMPaired-end read sequencing was performed on the XPlus platform. The result read count was: 69.64M; the base count was: 10.45G; the quality control score Q30 (%, the ratio of the number of bases with a quality value greater than 30 (error rate less than 0.1%) in the original sequence to the total number of bases) was: 71.5. The unique double-tag sequence was split into 200 files, corresponding to 200 samples. After trimming and removing sequences unrelated to mature miRNAs such as tags and adapter sequences, mature miRNA sequence data was obtained. The IsoMiRmap software was run to process each file to obtain the miRNA isoform sequence and its read count. Then the DESeq2 software was used to perform differential expression analysis of miRNA isoforms. DESeq2 is used to process the statistical analysis of RNA-seq data, identify genes that are differentially expressed under different conditions, take into account biological variability, and provide corrected p-values. Data with a corrected p-value greater than 0.05 were filtered, and data with a calibrated p-value less than 0.05 were retained. Calculate the correlation coefficient between miRNA isoforms, and filter out other miRNA isoforms with correlation coefficients greater than 0.5. Use different classifiers for machine learning classification of the remaining data to find the best classifier. A variety of classifiers were used, and the data was first divided into two parts: 80% of the part was used to train the model, and 20% of the part was used to verify the model. The best result was the Support Vector Machine (SVM) algorithm.
[0172] like Figure 2 As shown: The points in the figure represent the number of NGS reads and the corresponding input concentration values. The Z score is used to detect outliers, that is, points with a Z score greater than 3 are considered outliers (red) and excluded from the linear regression analysis to avoid distorting the relationship between the number of NGS reads and the input concentration. Data points without outliers (light blue) are used to fit the linear model. The equation in the figure represents the linear relationship between the number of NGS reads (x-axis) and the input concentration (y-axis). It shows that as the number of NGS reads increases, the input concentration also increases proportionally (the slope is 1.17e-07). Linear regression analysis shows that the NGS reads of artificial miRNAs have a relatively strong correlation with the input concentration (correlation coefficient 0.86, R 2 =0.74). The extremely small P value (6.26e-82) indicates that the observed relationship is highly statistically significant, meaning that the probability of this correlation occurring by chance is very low. This is of great significance for quality control or optimization in sequencing experiments, that is, the miRNA isoform NGS reads are positively correlated with the content of miRNA isoforms in plasma, and miRNA isoforms in plasma can be quantified.
[0173] like Figure 3Figure 4: The miRNA isoforms with significant differential expression in the DESeq2 analysis are shown. The X-axis represents the log2-transformed fold change of miRNA isoform expression under the two conditions. Negative values on the left and positive values on the right indicate that the miRNA isoforms are up-regulated and down-regulated in breast cancer, respectively. The Y-axis represents the statistical significance of the differential expression of miRNA isoforms, with larger values indicating more significant differences. Genes with low p-values are displayed higher on the graph. Each black dot represents a gene in the dataset. The dots show the relationship between the magnitude of the fold change and the statistical significance of each gene. MiRNA isoforms with differential expression p-values (padj) < 0.05 and fold changes greater than 1.5-fold (|log2FoldChange|>0.584) are indicated by red dots and are potential candidates for further biological studies, such as potential biomarkers or targets for further investigation. The blue numbers next to the red dots correspond to significant miRNA isoforms. Table 4 shows the miRNA isoforms with differential expression of 1.5-fold or more.
[0174] like Figure 4 Shown: The expression heatmap shows the differential expression levels of the top 30 isomiRs (i.e., isoforms of microRNA) with the highest variance. The data was processed by the DESeq2 package and analyzed using variance stabilizing transformation (VST). The data were further clustered by row and column clustering, respectively, based on the expression pattern of isomiRs and different samples. Row clustering can reveal isomiRs that behave similarly under different conditions. By calibrating the expression level of each isomiR with the average value in all samples, the color ratio of each row in the figure is adjusted independently, and it can be quickly seen which isomiRs are upregulated or downregulated. Column clustering shows the similarities or differences in expression responses between breast cancer groups and normal groups. Blue indicates lower expression levels, and yellow to red indicate higher expression levels. At the top of each column, 'C' represents breast cancer (red) and 'N' represents normal (blue).
[0175] like Figure 5Shown: This confusion matrix and its associated classification statistics were generated by a support vector machine (SVM) classifier that was trained using 10-fold cross validation. The model's hyperparameters (gamma and cost) were tuned and predictions were made on the test dataset. The performance metrics summarize the model's ability to classify samples into two categories: cancer and normal. The SVM classifier achieved high accuracy (93.48%), strong sensitivity (95.65%), and specificity (91.30%), indicating that it performs well in distinguishing cancer from normal samples. The Kappa coefficient of 0.8696 further supports the reliability of the model, while the McNemar test indicates that there is no significant imbalance between the two error rates. Overall, the classifier performed very well on the given dataset. The confusion matrix provides a summary of the prediction results: True Positives (TP) (cancer predicted as cancer): 22; False Positives (FP) (normal predicted as cancer): 2; True Negatives (TN) (normal predicted as normal): 21; False Negatives (FN) (cancer predicted as normal): 1; This matrix is used to calculate performance indicators such as accuracy, sensitivity, and specificity. Accuracy: 0.9348 (93.48%). This means that the model correctly classified 93.48% of the test samples. 95% Confidence Interval: (0.821, 0.9863). This means that with 95% confidence, the true accuracy will fall within this interval. No Information Rate (NIR): 0.5. This is the accuracy when the most common category is always predicted (baseline accuracy). P value (Acc>NIR): 2.311e-10. This is a significance test used to compare whether the model accuracy is better than NIR. The result shows that the model is significantly better than random guessing. Kappa coefficient: 0.8696. The Kappa coefficient measures the agreement between the predicted class and the actual class and excludes the influence of chance agreement. A Kappa of 0.8696 indicates excellent agreement. McNemar's test P value: 1. This test evaluates the significant difference between the two error rates (cancer vs. normal). A P value of 1 indicates that there is no significant difference between the two error rates. Sensitivity (Recall, Cancer): 0.9565 (95.65%). The proportion of actual cancer cases correctly identified by the model. Specificity: 0.9130 (91.30%). The proportion of actual normal cases correctly identified by the model. Positive Predictive Value (PPV): 0.9167 (91.67%). The probability that a sample predicted as cancer is actually cancer. Negative Predictive Value (NPV): 0.9545 (95.45%). The probability that a sample predicted as normal is actually normal. Prevalence: 0.5. The proportion of actual positive (cancer) cases in the test data. Detection rate: 0.4783. The proportion of actual positive cases that were correctly identified (22 / 46 samples). Detection prevalence: 0.5217. The proportion of positive (cancer) predicted by the model, that is, the number of times the model predicted cancer. Balanced accuracy: 0.9348.The average of sensitivity and specificity provides a balanced measure for imbalanced datasets.
[0176] like Figure 6 Figure 1: Shows a ROC curve (Receiver Operating Characteristic Curve) for a SVM (Support Vector Machine) classifier for predicting cancer. The ROC curve represents the performance of a classifier by plotting the True Positive Rate (Sensitivity) and the False Positive Rate (1-Specificity) at different thresholds. True Positive Rate (Y-axis): This represents the sensitivity of the classifier, which is the proportion of actual positive (cancer) cases that the model correctly identifies. False Positive Rate (X-axis): This represents 1-Specificity, which is the proportion of negative (non-cancer) cases that the model incorrectly classifies as positive. ROC Curve: This colorful curve shows the trade-off between sensitivity and specificity of the classifier at different thresholds. A perfect classifier's curve would fit snugly in the upper left corner of the graph (True Positive Rate = 1, False Positive Rate = 0) Diagonal Line: The dotted line represents the performance of a random classifier (AUC = 0.5), which is a baseline for comparison. The further the ROC curve is from this line, the better the classifier performs. AUC (Area Under the Curve): The AUC value (here 0.9924) is a single scalar value that summarizes the overall performance of the classifier. The closer the AUC is to 1, the better the classifier performs, while an AUC of 0.5 indicates random performance. Cutoff (threshold): The threshold (here 0.6461) is the optimal threshold at which the classifier achieves a balance between sensitivity and specificity, determined by the custom `opt.cut` function. Sensitivity: At the optimal threshold, the classifier has a sensitivity of 0.913, which means that it is able to correctly identify 91.3% of actual cancer cases. Specificity: A specificity of 1 indicates that the classifier is able to correctly identify all non-cancer cases without any false positives at the optimal threshold. The ROC curve evaluates the performance of the SVM classifier in distinguishing cancer from non-cancer cases, visualizing the trade-off between sensitivity and specificity, and thus helping to choose the best classification threshold (Cutoff). A higher AUC value indicates that the classifier performs very well and achieves a good balance between true positives and false positives. AUC: represents the area under the curve, which is used to summarize the overall performance of the classifier. By analyzing this ROC curve, it can be concluded that the SVM classifier performs very well in distinguishing cancer from non-cancer cases, achieving a strong balance between sensitivity and specificity at the selected threshold.
[0177] Cutoff (threshold): represents the optimal threshold when the classifier strikes a balance between sensitivity and specificity.
[0178] Sensitivity and Specificity: These values show the performance of the classifier at the optimal threshold, helping the user understand the classifier's ability to discriminate at this threshold.
[0179] Table 3 Artificial miRNA sequences and their pre-amplification primer sequences
[0180]
[0181]
[0182] As shown in Table 3, as an experimental control, artificial miRNAs do not exist in the human body, but their sequences and pre-amplification primers are as similar as possible to natural miRNAs. We added artificial miRNA-1, artificial miRNA-2 and artificial miRNA-3 to all breast cancer plasma samples, and added another three artificial miRNAs to all normal human plasma samples. The concentrations of the three artificial miRNA markers are high concentration (1.0E-04μM), medium concentration (5.0E-05μM) and low concentration (1.0E-05μM). That is, different artificial miRNAs are added as internal tags in different samples. In this way, even if the samples are mislabeled during the experiment, these unique miRNA tags can be used to correct this common human error. At the same time, because the concentrations of artificial miRNAs are known, they can be used to absolutely quantify other natural miRNAs and calibrate experimental samples. The results of the second-generation sequencing (Table) show that, in the comparison between manual records and artificial miRNA markers, among 55 breast cancers, 5 results could not determine whether the two were consistent, which may be due to the low concentration of artificial miRNA; 1 result did not have an artificial miRNA marker, which may have been omitted; and one result was inconsistent, indicating human error during the experiment. Apart from this, the other markers were correct. This labeling method is feasible and can ensure accurate tracking of samples, especially when dealing with a large number of samples or conducting high-throughput experiments, which can effectively reduce the risk of sample misplacement or cross-contamination.
[0183] Table 4 Six artificial miRNA markers tracking the next generation sequencing results of breast cancer and normal human samples
[0184]
[0185]
[0186]
[0187]
[0188]
[0189] Table 5. miRNA isoforms with differential expression of 1.5 times or more
[0190]
[0191] Table 6 shows the miRNA isoforms that are significantly differentially expressed (calibrated p value less than 0.05) between breast cancer and normal subjects, after removing the highly correlated (Pearson's correlation coefficient, greater than or equal to 0.75) miRNA isoforms, and the remaining miRNA isoforms are used for machine learning model construction.
[0192] Table 6. miRNA isoforms used for machine learning
[0193]
[0194]
[0195] This embodiment refers to the prior art (see the literature: Pliatsika V.et al. (2018) MINTbase v2.0: a comprehensive database for tRNA-derived fragments that includes nuclear and mitochondrial fragments from all The Cancer Genome Atlas projects. Nucleic Acids Res., 46, D152–D159) to name the unique isomer name of each isomer. The unique isomer name in Table 6 can also be called the isomer "license plate". Each isomiR has a unique "license plate", and each "license plate" corresponds to a specific isomiR. This "license plate" serves as an identifier of the isomiR and is used to accurately track and distinguish each miRNA isomer during sequencing or analysis. The number separated in the middle is the number of bases of this miRNA isomer.
[0196] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A kit for detecting breast cancer by next-generation sequencing of plasma miRNA isoforms, characterized in that: Primers for amplifying artificial miRNA1 to artificial miRNA6 and primers for amplifying at least one of the following miRNA isomers: hsa-miR-21-5p, hsa-miR-223-3p, hsa-miR-223-5p, hsa-miR-186-5p, hsa-miR-18a-5p, hsa-miR-146b-5p, hsa-miR-624-5p, hsa-miR-106b-5p, hsa-miR-340-5p, hsa-miR-20a-5p, hsa-miR-451a, hsa-miR-7976, hsa-miR-2355-3p, hsa-miR-301a-3p, hsa-miR-144-5p, hsa-miR-151a-3p, hsa-miR-3200-5p, hsa-miR-1537-3p, hsa-miR-500a-5p, hsa-miR-127-3p, hsa-miR-570-3p, hsa-miR-130b-5p, hsa-miR-503-5p, hsa-miR-551a, hsa-miR-409-3p, hsa-miR-330-3p, hsa-miR-889-3p, hsa-miR-625-5p, hsa-miR-542-3p, hsa-miR-582-3p, hsa-miR-381-3p, hsa-miR-495-3p, hsa-miR-103a-1-5p, hsa-miR-450b-5p, hsa-miR-429, hsa-miR-576-5p, hsa-miR-148b-3p, hsa-miR-320c, hsa-miR-4286, hsa-miR-126-3p, hsa-miR-152-3p, hsa-miR-144-3p, hsa-miR-195-5p, hsa-let-7a-5p, hsa-miR-378f, hsa-miR-126-5p, hsa-miR-26a-5p, hsa-miR-29a-3p, hsa-miR-181a-5p, hsa-miR-32-5p, hsa-miR-142-3p, hsa-miR-29c-3p, hsa-miR-424-5p, hsa-miR-192-5p, hsa-miR-143-3p, hsa-miR-30c-5p, hsa-miR-146a-5p, hsa-miR-101-3p, hsa-miR-19b-3p, hsa-miR-33b-5p, hsa-miR-378a-3p, hsa-miR-22-3p, hsa-miR-107, hsa-miR-497-5p,hsa-miR-15a-3p, hsa-miR-188-5p, hsa-let-7d-3p, hsa-miR-132-3p, hsa-miR-151a-5p, hsa-miR-194-5p, hsa-miR-99a-5p, hsa-miR-125b-5p, hsa-miR-25-3p, hsa-miR-103a-3p, hsa-miR-1285-3p, hsa-miR-7977, hsa-miR-30b-5p, hsa-miR-363-3p, hsa-miR-93-5p, hsa-miR-375-3p, and hsa- miR-99b-5p, hsa-miR-193b-3p, hsa-miR-324-3p, hsa-miR-193a-3p, hsa-miR-342-3p, hsa-miR-484, hsa-miR-532-3p, hsa-miR-210-3p, hsa-miR-2110, hsa-miR-296-5p, hsa-miR-1307-5p, hsa-miR-19a-3p, hsa-miR-139-5p, hsa-miR-3665, hsa-miR-RG-84, hsa-miR-4454, and hsa-let-7b-5p; The primers for amplifying miRNA isoforms include nucleotide sequences as shown in SEQ ID NO: 1 to SEQ ID NO: 97; The nucleotide sequences of the primers for amplifying artificial miRNA1 to artificial miRNA6 are shown in SEQ ID NO: 123 to SEQ ID NO: 128, respectively.
2. The kit according to claim 1, characterized in that The kit further comprises a second pre-amplification PCR primer pair and / or a third pre-amplification PCR primer pair; The second pre-amplification PCR primer pair includes a transition primer and a reverse primer for amplifying miRNA isoforms; The nucleotide sequence of the transition primer is shown in SEQ ID NO: 99; The nucleotide sequence of the reverse primer for amplifying the miRNA isoform is shown in SEQ ID NO: 100; The third pre-amplification PCR primer pair includes a 5' universal primer and a 3' universal primer; The nucleotide sequence of the 5' universal primer is shown in SEQ ID NO: 101; The nucleotide sequence of the 3' universal primer is shown in SEQ ID NO:
102.
3. The kit according to claim 1, characterized in that The kit also includes a primer for adding a sequencing adapter and a primer for adding a barcode tag; The forward primer of the primer adding the sequencing adapter is obtained by sequentially connecting the DNA fragment shown in SEQ ID NO: 103, the I5 Index sequence and the DNA fragment sequence shown in SEQ ID NO: 104; The reverse primer of the primer for adding the sequencing adapter is obtained by sequentially connecting the DNA fragment shown in SEQ ID NO: 105, the 7Index sequence and the DNA fragment sequence shown in SEQ ID NO: 106; The forward primer of the barcode-tagged primer is obtained by sequentially connecting the DNA fragment shown in SEQ ID NO: 107, the I5 Index sequence, and the DNA fragment sequence shown in SEQ ID NO: 108; The reverse primer of the primer with barcode tag is obtained by sequentially connecting the DNA fragment shown in SEQ ID NO: 109, the I5 Index sequence and the DNA fragment sequence shown in SEQ ID NO:
110.
4. The kit according to any one of claims 1 to 3, characterized in that The kit also includes a reverse transcription primer; The nucleotide sequence of the reverse transcription primer is shown in SEQ ID NO:
111.
5. Use of the kit according to any one of claims 1 to 5 in constructing a breast cancer next-generation sequencing library.
6. A method for constructing a breast cancer second-generation sequencing library, characterized in that: The following steps are involved: RNA from breast cancer samples was reverse transcribed to obtain cDNA; Using the cDNA as a template, and using the primers in claim 1 to perform a first PCR pre-amplification to obtain a first pre-amplification product; Using the first pre-amplification product as a template, and using the second pre-amplification PCR primer in claim 3 to perform a second PCR pre-amplification to obtain a second pre-amplification product; Using the second pre-amplification product as a template, and using the third pre-amplification PCR primers described in claim 3 to perform a third PCR pre-amplification to obtain a third pre-amplification product; Using the third pre-amplification product as a template, and using the primers with sequencing adapters as described in claim 4 to perform a first PCR amplification, to obtain a PCR product containing an internal unique double tag; Using the PCR product with the internal unique double tag as a template, a second PCR amplification is performed using the primers with the barcode tag as described in claim 4 to obtain double unique double tag PCR products, and the samples are mixed to obtain a sequencing library.
7. A miRNA isoform composition associated with breast cancer diagnosis, characterized in that: The miRNA isoform composition includes the nucleotide sequences shown in SEQ ID NO: 129 to SEQ ID NO:
210.
8. Use of the miRNA isoform composition according to claim 7 in constructing a breast cancer prediction model.
9. The use according to claim 8, characterized in that: The machine learning classifier of the breast cancer prediction model includes a support vector classifier.
10. Use of primers for amplifying the miRNA isoform composition according to claim 7 in preparing a kit for diagnosing breast cancer.
Citation Information
Patent Citations
Specific quantitative PCR reaction mixed liquor, miRNA quantitative detection kit, and detection method
CN109957611A
Internal reference substance for detecting bladder cancer serum miRNA and its detection primers and use
CN103602747A
Expression of miRNAs in placental tissue
CN104080911A
Marker miR-126-3P for HER-2 positive breast cancer tissue, and application and diagnosis kit of marker miR-126-3P
CN106636375A
Primer pair and kit for detecting hsacirc0006411
CN116555424A
Cited By
System and kit for early auxiliary diagnosis of breast cancer and detection of minimal residual lesions in various stages
CN121087183A
Systems and kits for early adjuvant diagnosis of breast cancer and detection of minimal residual disease at various stages
CN121087183B
Breast cancer plasma transporter miRNA marker and application thereof
CN121428099A