A kit for detecting breast cancer by plasma miRNA isomer second-generation sequencing
By constructing a next-generation sequencing detection method for plasma miRNA isoforms, and combining specific primers and a machine learning classifier, the challenge of detecting low-expression miRNAs in plasma was solved, achieving high-sensitivity and high-specificity breast cancer diagnosis and improving detection accuracy.
Patent Information
- Application Number
- CN202510018455.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-01-06
AI Technical Summary
Existing technologies are insufficient for efficiently detecting low-expression miRNAs in plasma, resulting in inadequate sensitivity and specificity in breast cancer diagnosis. Conventional methods suffer from poor reproducibility and low sensitivity.
The plasma miRNA isoform next-generation sequencing detection method was adopted. By selecting specific miRNA primers and amplification steps to construct sequencing libraries, and combining them with a machine learning classifier, high sensitivity and high specificity of low-expression miRNAs were achieved.
It improved the sensitivity and specificity of breast cancer detection, significantly increased detection efficiency, and achieved an accuracy of 93.48%, a sensitivity of 95.65%, and a specificity of 91.30%.
Smart Images

Figure CN119932185B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of biomedicine, and particularly relates to a kit for detecting breast cancer by plasma miRNA isomer second-generation sequencing. BACKGROUND
[0002] microRNA (miRNA) is a kind of non-coding small molecule RNA with a length of 18-25 nucleotides. miRNA directly and indirectly regulates the expression of most genes and participates in a series of life activities. miRNA is closely related to the occurrence and development of tumors. More and more studies have shown that miRNA plays an important regulatory role in the occurrence and development of tumors. Malignant tumor is the result of the interaction of genetic factors and environmental factors, and the environmental factors play a greater role. Genetic diagnosis of cancer has great limitations, which can only find susceptible genes and cannot be used as a biomarker for diagnosis of malignant tumors. At the same time, environmental factors are also unmonitorable and can only be used as a risk factor for malignant tumors. miRNA is a major bridge connecting environmental factors and genetic factors, and is a major regulatory factor between changing environment and unchanging genetic material. miRNA is relatively stable in structure, appears earlier in abnormal expression, and is easier to accurately distinguish tumor types. These characteristics make miRNA a biomarker for tumor etiology, diagnosis, progression, recurrence and treatment outcome.
[0003] However, due to the disadvantages of miRNA, such as only twenty bases and low level in blood, it is difficult to detect miRNA, and there is still a lack of accurate detection technology for microRNA. qPCR, microarray and small RNA sequencing (RNA-seq) are commonly used to study the expression of miRNA in tissues. However, they all have different degrees of defects. The main problem of microarray is low sensitivity and relatively long turnaround time, while qPCR is not easy to detect a large number of miRNAs. In many studies, the repeatability of circulating miRNA research is extremely low. The detection results of different laboratories are not comparable, and even may be opposite. A review of 11 studies showed that only 5 miRNAs out of 31 miRNAs identified in one study as related to heart failure could be repeated in another study, but no miRNA could be repeated in more than two studies. This fully illustrates that the existing miRNA qPCR detection technology has serious defects. This greatly limits its application in clinical cancer diagnosis. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a kit for detecting breast cancer by plasma miRNA isomer second-generation sequencing. By selecting the miRNA to be amplified, a second-generation sequencing library is constructed, the detection of low-expression miRNA is realized, and the sequencing results have the advantages of high sensitivity, high relative sequencing depth and high specificity.
[0005] The application provides a kit for detecting breast cancer by plasma miRNA isomer second-generation sequencing, which comprises primers for amplifying artificial miRNA1-artificial miRNA6 and primers for amplifying at least one miRNA isomer of hsa-miR-21-5p, hsa-miR-223-3p, hsa-miR-223-5p, hsa-miR-186-5p, hsa-miR-18a-5p, hsa-miR-146b-5p, hsa-miR-624-5p, hsa-miR-106b-5p, hsa-miR-340-5p, hsa-miR-20a-5p, hsa-miR-451a, hsa-miR-7976, hsa-miR-2355-3p, hsa-miR-301a-3p, hsa-miR-144-5p, hsa-miR-151a-3p, hsa-miR-3200-5p, hsa-miR-1537-3p, hsa-miR-500a-5p, hsa-miR-127-3p, hsa-miR-570-3p, hsa-miR-130b-5p, hsa-miR-503-5p, hsa-miR-551a, hsa-miR-409-3p, hsa-miR-330-3p, hsa-miR-889-3p, hsa-miR-625-5p, hsa-miR-542-3p, hsa-miR-582-3p, hsa-miR-381-3p, hsa-miR-495-3p, hsa-miR-103a-1-5p, hsa-miR-450b-5p, hsa-miR-429, hsa-miR-576-5p, hsa-miR-148b-3p, hsa-miR-320c, hsa-miR-4286, hsa-miR-126-3p, hsa-miR-152-3p, hsa-miR-144-3p, hsa-miR-195-5p, hsa-let-7a-5p, hsa-miR-378f, hsa-miR-126-5p, hsa-miR-26a-5p, hsa-miR-29a-3p, hsa-miR-181a-5p, hsa-miR-32-5p, hsa-miR-142-3p, hsa-miR-29c-3p, hsa-miR-424-5p, hsa-miR-192-5p, hsa-miR-143-3p, hsa-miR-30c-5p, hsa-miR-146a-5p, hsa-miR-101-3p, hsa-miR-19b-3p, hsa-miR-33b-5p, hsa-miR-378a-3p,hsa-miR-22-3p, hsa-miR-107, hsa-miR-497-5p, hsa-miR-15a-3p, hsa-miR-188-5p, hsa-let-7d-3p, hsa-miR-132-3p, hsa-miR-151a-5p, hsa-miR-194-5p, hsa-miR-99a-5p, hsa-miR-125b-5p, hsa-miR-25-3p, hsa-miR-103a-3p, hsa-miR-1285-3p, hsa-miR-7977, hsa-miR-30b-5p, hsa-miR-363-3p, hsa-miR-93-5p, hsa-miR-375-3p, hsa-miR-99b-5p, hsa-miR-193b-3p, hsa-miR-324-3p, hsa-miR-193a-3p, hsa-miR-342-3p, hsa-miR-484, hsa-miR-532-3p, hsa-miR-210-3p, hsa-miR-2110, hsa-miR-296-5p, hsa-miR-1307-5p, hsa-miR-19a-3p, hsa-miR-139-5p, hsa-miR-3665, hsa-miR-RG-84, hsa-miR-4454, hsa-let-7b-5p;
[0006] The primer for amplifying the miRNA isomers includes a nucleotide sequence as shown in SEQ ID NO: 1 to SEQ ID NO: 97.
[0007] The nucleotide sequences of the primers for amplifying the artificial miRNAs 1 to 6 are shown in SEQ ID NO: 123 to SEQ ID NO: 128, respectively.
[0008] Preferably, the kit further comprises a second pair of pre-amplification PCR primers and / or a third pair of pre-amplification PCR primers.
[0009] The second pair of pre-amplification PCR primers comprises a transition primer and a reverse primer for amplifying the miRNA isomers.
[0010] The nucleotide sequence of the transition primer is shown in SEQ ID NO: 99.
[0011] The nucleotide sequence of the reverse primer for amplifying the miRNA isomers is shown in SEQ ID NO: 100.
[0012] The third pair of pre-amplification PCR primers comprises a 5' universal primer and a 3' universal primer.
[0013] The nucleotide sequence of the 5' universal primer is shown as SEQ ID NO: 101;
[0014] The nucleotide sequence of the 3' universal primer is shown as SEQ ID NO: 102.
[0015] Preferably, the kit further comprises a sequencing adapter primer and a barcode-labeled primer;
[0016] The forward primer of the sequencing adapter primer is obtained by sequentially connecting a DNA fragment shown as SEQ ID NO: 103, an I5 Index sequence, and a DNA fragment sequence shown as SEQ ID NO: 104;
[0017] The reverse primer of the sequencing adapter primer is obtained by sequentially connecting a DNA fragment shown as SEQ ID NO: 105, an I7 Index sequence, and a DNA fragment sequence shown as SEQ ID NO: 106;
[0018] The forward primer of the barcode-labeled primer is obtained by sequentially connecting a DNA fragment shown as SEQ ID NO: 107, an I5 Index sequence, and a DNA fragment sequence shown as SEQ ID NO: 108;
[0019] The reverse primer of the barcode-labeled primer is obtained by sequentially connecting a DNA fragment shown as SEQ ID NO: 109, an I5 Index sequence, and a DNA fragment sequence shown as SEQ ID NO: 110.
[0020] Preferably, the kit further comprises a reverse transcription primer;
[0021] The nucleotide sequence of the reverse transcription primer is shown as SEQ ID NO: 111.
[0022] The application provides an application of the kit in constructing a breast cancer next-generation sequencing library.
[0023] The application provides a method for constructing a breast cancer next-generation sequencing library, comprising the following steps:
[0024] RNA of a breast cancer sample is subjected to reverse transcription to obtain cDNA;
[0025] The cDNA is used as a template to perform first PCR pre-amplification using the primers to obtain a first pre-amplification product;
[0026] The first pre-amplification product is used as a template to perform second PCR pre-amplification using the second pre-amplification PCR primers to obtain a second pre-amplification product;
[0027] Third pre-amplification PCR is performed using the second pre-amplification product as a template and the third pre-amplification PCR primer, and a third pre-amplification product is obtained;
[0028] First PCR amplification is performed using the third pre-amplification product as a template and the primer with a sequencing adapter, and a PCR product containing an inner unique dual tag is obtained;
[0029] Second PCR amplification is performed using the PCR product with the inner unique dual tag as a template and the primer with a barcode tag, and a dual unique dual tag PCR product is obtained, and a sequencing library is obtained by mixing samples.
[0030] The application provides a miRNA isomer composition related to breast cancer diagnosis, which comprises a nucleotide sequence shown in SEQ ID NO: 129-SEQ ID NO: 210.
[0031] The application provides application of the miRNA isomer composition in construction of a breast cancer prediction model.
[0032] Preferably, the machine learning classifier of the breast cancer prediction model comprises a support vector classifier.
[0033] The application provides application of a primer for amplifying the miRNA isomer composition in preparation of a breast cancer diagnosis kit.
[0034] The application provides a kit for detecting breast cancer by plasma miRNA isomer second-generation sequencing, which comprises an amplification primer of a miRNA isomer closely related to breast cancer diagnosis, and therefore, the kit has the advantages of high sensitivity, high relative sequencing depth and high specificity, can detect miRNAs that cannot be detected by conventional methods, and because of significant amplification, high-expression miRNAs are avoided, and at the same time, the same sequencing depth can be used for low-expression miRNAs, and the relative sequencing depth of the library constructed by using the kit can be much higher than that of conventional technology. The kit greatly improves the detection efficiency of breast cancer.
[0035] The application also provides a miRNA isomer composition related to breast cancer diagnosis, and the application constructs a breast cancer prediction model based on SVM classification of second-generation sequencing data of the miRNA isomer composition, and has good performance in distinguishing cancer and normal samples, with an accuracy of 93.48%, a sensitivity of 95.65% and a specificity of 91.30%. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 It is a flowchart for result analysis of breast cancer miRNA isomer second-generation sequencing.
[0037] Figure 2 For the relationship between NGS reads and input concentration for artificial miRNAs;
[0038] Figure 3 For volcano plot of miRNA isoform differential expression;
[0039] Figure 4 For heatmap of miRNA isoform differential expression;
[0040] Figure 5 For confusion matrix and classification statistics for SVM classifier;
[0041] Figure 6 For ROC curve for SVM machine learning classifier. DETAILED DESCRIPTION
[0042] The application provides a kit for detecting breast cancer by plasma miRNA isomer second-generation sequencing, which comprises primers for amplifying artificial miRNA1-artificial miRNA6 and primers for amplifying at least one miRNA isomer of hsa-miR-21-5p, hsa-miR-223-3p, hsa-miR-223-5p, hsa-miR-186-5p, hsa-miR-18a-5p, hsa-miR-146b-5p, hsa-miR-624-5p, hsa-miR-106b-5p, hsa-miR-340-5p, hsa-miR-20a-5p, hsa-miR-451a, hsa-miR-7976, hsa-miR-2355-3p, hsa-miR-301a-3p, hsa-miR-144-5p, hsa-miR-151a-3p, hsa-miR-3200-5p, hsa-miR-1537-3p, hsa-miR-500a-5p, hsa-miR-127-3p, hsa-miR-570-3p, hsa-miR-130b-5p, hsa-miR-503-5p, hsa-miR-551a, hsa-miR-409-3p, hsa-miR-330-3p, hsa-miR-889-3p, hsa-miR-625-5p, hsa-miR-542-3p, hsa-miR-582-3p, hsa-miR-381-3p, hsa-miR-495-3p, hsa-miR-103a-1-5p, hsa-miR-450b-5p, hsa-miR-429, hsa-miR-576-5p, hsa-miR-148b-3p, hsa-miR-320c, hsa-miR-4286, hsa-miR-126-3p, hsa-miR-152-3p, hsa-miR-144-3p, hsa-miR-195-5p, hsa-let-7a-5p, hsa-miR-378f, hsa-miR-126-5p, hsa-miR-26a-5p, hsa-miR-29a-3p, hsa-miR-181a-5p, hsa-miR-32-5p, hsa-miR-142-3p, hsa-miR-29c-3p, hsa-miR-424-5p, hsa-miR-192-5p, hsa-miR-143-3p, hsa-miR-30c-5p, hsa-miR-146a-5p, hsa-miR-101-3p, hsa-miR-19b-3p, hsa-miR-33b-5p, hsa-miR-378a-3p,hsa-miR-22-3p, hsa-miR-107, hsa-miR-497-5p, hsa-miR-15a-3p, hsa-miR-188-5p, hsa-let-7d-3p, hsa-miR-132-3p, hsa-miR-151a-5p, hsa-miR-194-5p, hsa-miR-99a-5p, hsa-miR-125b-5p, hsa-miR-25-3p, hsa-miR-103a-3p, hsa-miR-1285-3p, hsa-miR-7977, hsa-miR-30b-5p, hsa-miR-363-3p, hsa-miR-93-5p, hsa-miR-375-3p, hsa-miR-99b-5p, hsa-miR-193b-3p, hsa-miR-324-3p, hsa-miR-193a-3p, hsa-miR-342-3p, hsa-miR-484, hsa-miR-532-3p, hsa-miR-210-3p, hsa-miR-2110, hsa-miR-296-5p, hsa-miR-1307-5p, hsa-miR-19a-3p, hsa-miR-139-5p, hsa-miR-3665, hsa-miR-RG-84, hsa-miR-4454, hsa-let-7b-5p;
[0043] The primer for amplifying the miRNA isomers includes the nucleotide sequences shown in SEQ ID NO: 1 to SEQ ID NO: 97.
[0044] The nucleotide sequences of the primers for amplifying the artificial miRNAs 1 to 6 are shown in SEQ ID NO: 123 to SEQ ID NO: 128, respectively.
[0045] In the present application, the primer for amplifying the miRNA isomers is obtained by sequentially connecting a universal sequence shown in the nucleotide sequence of SEQ ID NO: 98 (ATAGACTCCTCGCATAGCCTCATGAGTC) and a 5' end partial sequence of any one of the miRNA isomers. The 5' end partial sequence of the miRNA isomers refers to a sequence of 10 to 14 nt in length at the 5' end of the miRNA isomers, and can be 11 nt, 12 nt, and 14 nt. The primer ensures specific amplification of microRNAs while also amplifying different types of isomers of all specific microRNAs, and has flexibility in the determination object.
[0046] In the present application, the primer for amplifying the artificial miRNA1-artificial miRNA6 is obtained by sequentially connecting a universal sequence as shown in the nucleotide sequence of SEQ ID NO: 98 and a 5' end partial sequence of any one of the artificial miRNA1-artificial miRNA6. The 5' end partial sequence of any one of the artificial miRNA1-artificial miRNA6 refers to a sequence of 10-14 nt in length at the 5' end of any one of the artificial miRNA1-artificial miRNA6, which can be 11 nt, 12 nt and 14 nt. The nucleotide sequence of the artificial miRNA1-artificial miRNA6 is as shown in SEQ ID NO: 117-SEQ ID NO: 122. The artificial miRNA1-artificial miRNA6 is a microRNA control product, which is added in breast cancer plasma samples and normal human plasma samples as an internal label, used to correct errors, and its added concentration is known, and can also be used to absolutely quantify other natural miRNA isomers and calibrate experimental samples.
[0047] In the present application, the kit preferably further comprises a second pre-amplification PCR primer pair and / or a third pre-amplification PCR primer pair. The second pre-amplification PCR primer pair comprises a transition primer and a reverse primer for amplifying miRNA isomers. The nucleotide sequence of the transition primer is as shown in SEQ ID NO: 99; the nucleotide sequence of the reverse primer for amplifying miRNA isomers is as shown in SEQ ID NO: 100. The third pre-amplification PCR primer pair comprises a 5' universal primer and a 3' universal primer. The nucleotide sequence of the 5' universal primer is as shown in SEQ ID NO: 101; the nucleotide sequence of the 3' universal primer is as shown in SEQ ID NO: 102.
[0048] In the present application, the kit preferably further comprises primers for adding sequencing adapters and primers for adding barcode tags. The forward primer of the primer for adding sequencing adapters is obtained in sequence by connecting a DNA fragment shown in SEQ ID NO: 103, an I5 Index sequence, and a DNA fragment sequence shown in SEQ ID NO: 104; the reverse primer of the primer for adding sequencing adapters is obtained in sequence by connecting a DNA fragment shown in SEQ ID NO: 105, a 7 Index sequence, and a DNA fragment sequence shown in SEQ ID NO: 106. The forward primer of the primer for adding barcode tags is obtained in sequence by connecting a DNA fragment shown in SEQ ID NO: 107, an I5 Index sequence, and a DNA fragment sequence shown in SEQ ID NO: 108. The reverse primer of the primer for adding barcode tags is obtained in sequence by connecting a DNA fragment shown in SEQ ID NO: 109, an I5 Index sequence, and a DNA fragment sequence shown in SEQ ID NO: 110. The amplification of the primer for adding sequencing adapters and the primer for adding barcode tags is conducive to increasing the sequencing tags of each sample, easy to identify the sequencing results of each sample, and not easy to be confused.
[0049] In the present application, the I5 Index sequence and the I7 Index sequence refer to the I5 Index and I7 Index recorded in Table 2 of the patent with publication number CN118421787A. In the present application, the I5 Index sequence and the I7 Index sequence are combined in a set. The screening method of the I5 Index sequence and the I7 Index sequence preferably generates 10-base unique short sequences randomly, removes complementary sequences, and then screens double unique dual labels according to the following criteria. The screening criteria preferably include that the same base is not repeated more than three times; each sequence has no serious complementarity with other sequences; each sequence is different from other sequences by at least 3 bases; the two sequences of the same double unique dual label do not affect the specific amplification of the primer after being combined with the surrounding sequences, that is, they do not increase the possibility of primer dimers; by calculating the score of the two sequence pairs, the pair with the smallest score (i.e., the best specificity) is selected as the set of positive and negative primer labels. A total of 976 pairs of double unique dual labels for labeling high-throughput samples of next-generation sequencing are obtained through the above screening. Based on different double unique dual label sequences, there are at least 3 bases different, so even if there is a certain sequencing error in sequencing, it still maintains its uniqueness under the condition of allowing one base mismatch, and will not become other unique dual index labels, so the labeling accuracy is very high. This also shows more superiority than the conventional method, because the labeling sequence of the conventional method is much more, such as 2000 sequences for 1000 samples, and the probability of error labeling is at least 5 times that of our method. The present application analyzes a large number of samples through experiments, and has a data volume of more than 400G, and no mismatched next-generation sequencing reads are found.
[0050] In the present application, the kit preferably further comprises a reverse transcription primer. The nucleotide sequence of the reverse transcription primer is shown as SEQ ID NO: 109. The reverse transcription primer is used for reverse transcription of total RNA.
[0051] The present application does not have special restrictions on the source of the above-mentioned primer, and the gene synthesis method known in the art can be used.
[0052] The present application provides an application of the kit in constructing a breast cancer next-generation sequencing library.
[0053] The present application provides a method for constructing a breast cancer next-generation sequencing library, comprising the following steps:
[0054] The RNA of the breast cancer sample is reverse transcribed to obtain cDNA;
[0055] The cDNA is used as a template, and the primer is used for first PCR pre-amplification to obtain a first pre-amplification product;
[0056] The first pre-amplification product is used as a template to perform second PCR pre-amplification by using the second pre-amplification PCR primer, thereby obtaining a second pre-amplification product;
[0057] The second pre-amplification product is used as a template to perform third PCR pre-amplification by using the third pre-amplification PCR primer, thereby obtaining a third pre-amplification product;
[0058] The third pre-amplification product is used as a template to perform first PCR amplification by using the primer for adding a sequencing adapter, thereby obtaining a PCR product containing an inner unique dual tag;
[0059] The PCR product containing the inner unique dual tag is used as a template to perform second PCR amplification by using the primer for adding a barcode tag, thereby obtaining a dual unique dual tag PCR product, which is mixed to obtain a sequencing library.
[0060] The present application extracts total RNA from breast cancer samples, and obtains cDNA through reverse transcription.
[0061] The method for extracting total RNA is not particularly limited, and a method well known in the art can be used, for example, a commercial kit method.
[0062] In the present application, the reverse transcription includes a PolyA reaction, a denaturation reaction and a reverse transcription reaction. The system of the PolyA reaction is preferably 20 μl, including the following reagents: 5x reverse transcription buffer 4 μl, 10 mM ATP 2 μl, 5000 U / μl PolyA enzyme 1 μl, 40000 U / μl RNAase inhibitor 0.5 μl, RNA sample 12.5 μl. The conditions of the PolyA reaction are preferably 37°C for 30 min, 65°C for 20 min. The system of the denaturation reaction is preferably 20 μl, including the following reagents: 10 mM dNTPs 1.5 μl, 10 μM reverse transcription primer (USRTPn) 1.5 μl, Poly A reaction product 17 μl. The primer of the reverse transcription is preferably USRTPn, and the corresponding nucleotide sequence is as follows: CCTCCATCCGAGACACACGATTGATGGTTTTTTTTTTTTTTTTTTVN, SEQ ID NO: 111). The conditions of the denaturation reaction are preferably 65°C for 5 min, and the sample is taken out 1 s before the end of the reaction and immediately placed in an ice bath for 1 min. The system of the reverse transcription reaction is preferably 30 μl, including the following reagents: 5x reverse transcription buffer 2 μl, 1.6 M trehalose 4.5 μl, 1 mg / μl Actinomycin D 1.2 μl, T4 gp32 / RecA / ATP mixture 1.5 μl, 40000 U / μl RNAase inhibitor 0.3 μl, 50 U / μl Maxima H reverse transcriptase 1.5 μl, denaturation reaction product 19 μl. The preparation method of the T4 gp32 / RecA / ATP mixture is preferably prepared according to the number of samples, 10 μg / μl T4 gp32 0.6 μl, 2 μg / μl Tth RecA 0.2 μl, 100 mM ATP 0.24 μl, 1x reverse transcription buffer 1.96 μl. The conditions of the reverse transcription reaction are preferably 42°C for 15 min, 50°C for 30 min, 55°C for 30 min, 60°C for 30 min, 65°C for 30 min, and 85°C for 5 min.
[0063] After obtaining the cDNA, the present application uses the cDNA as a template and uses the amplification primer set to perform first PCR pre-amplification to obtain a first pre-amplification product.
[0064] In the present application, the reaction system of the first PCR pre-amplification is preferably a 20 μl system, including the following reagents: 2x Boost mix 10 μl, 0.2 μg / μl Tth RecA 1 μl, 1 μM amplification primer set 1.5 μl, cDNA 7.5 μl. The components and preparation method of the 2x Boost mix can be found in the patent number: ZL 20191021982
[0065] 7.4, the specific quantitative PCR reaction mixture in the embodiment 1 of the patent with the patent name of a specific quantitative PCR reaction mixture, miRNA quantitative detection kit and detection method, but the difference is that the 2x boostmix (containing UDG) is prepared with dNTPs mixture without dUTP. The reaction procedure of the first PCR pre-amplification is preferably ① 25℃ 10min, ② 95℃ 10min, ③ (95℃ 10s, 55℃ 10min) 3 cycles, ④ (95℃ 10s, 50℃ 10min) 3 cycles, ⑤ (95℃ 10s, 45℃ 10min) 2 cycles, ⑥ (95℃ 10s, 40℃ 10min) 2 cycles, ⑦ (95℃ 10s, 37℃ 10min) 2 cycles, ⑧ (95℃ 10s, 60℃ 2min 72℃ 10min) 1 cycle, ⑨ the PCR tube is taken out after the procedure is run to 72℃ 5min and is immediately ice-bathed. The first PCR pre-amplification is beneficial to amplify a large amount of miRNA isomer from the reverse transcription product. The first pre-amplification product is treated with EXO I enzyme after purification, and the purpose is to remove the PCR primer in the system.
[0066] In the present application, the breast cancer sample preferably further comprises any three of the artificial miRNA1-artificial miRNA6, and the first pre-amplification PCR is carried out by using the corresponding primers for amplifying any three of the artificial miRNA1-artificial miRNA6. At the same time, the amplification primers of the remaining three artificial miRNAs are added to the normal healthy human plasma sample as internal standards for PCR amplification. The method of the present application for the PCR amplification is not particularly limited, and the reaction system and reaction procedure of the first PCR pre-amplification can be used.
[0067] After obtaining the first pre-amplification product, the present application uses the transition primer in the amplification primer set and the reverse primer for amplifying the micro ribonucleic acid isomer as a template to carry out the second PCR pre-amplification, and obtains the second pre-amplification product.
[0068] In the present application, the reaction system of the second PCR pre-amplification is preferably 20 μl, including the following reagents: 10 μl of the 2x Boost mix, 1 μl of 10 μm transition primer (USEXPnb), 1 μl of 10 μm IsomiR primer, 1 μl of 0.2 μg / μl Tth RecA, 7 μl of the first pre-amplification product. The reaction procedure of the second PCR pre-amplification is preferably ① 25 ℃ for 10 min, ② 95 ℃ for 10 min, ③ (95 ℃ for 10 s, 65 ℃ for 1 min) for 3 cycles, ④ (95 ℃ for 10 s, 62 ℃ for 1 min) for 3 cycles, ⑤ (95 ℃ for 10 s, 58 ℃ for 2 min) for 2 cycles, ⑥ (95 ℃ for 10 s, 60 ℃ for 2 min) for 2 cycles, ⑦ (95 ℃ for 10 s, 60 ℃ for 2 min 72 ℃ for 10 min) for 1 cycle, ⑧ the program is run to 72 ℃ for 5 min, and then the PCR tube is taken out for ice bath. The second pre-amplification product is preferably purified by magnetic beads. The second PCR pre-amplification is amplified by the transition primer and the reverse primer of the micro ribonucleic acid isomer, and the purpose is to introduce the binding site of the 3' universal primer and the 5' universal primer.
[0069] The second pre-amplification product is used as a template, and the 5' universal primer and the 3' universal primer in the amplification primer group are used for third PCR pre-amplification, and the third pre-amplification product obtained is a micro ribonucleic acid isomer.
[0070] In the present application, the reaction system of the third PCR pre-amplification is preferably 20 μl, including the following reagents: 10 μl of the 2x Boost mix, 1 μl of 10 μm URP, 1 μl of 10 μm UFP, 1 μl of 0.2 μg / μl Tth RecA, 7 μl of the second pre-amplification product. The reaction procedure of the third PCR pre-amplification is preferably ① 95 ℃ for 10 min, ② (95 ℃ for 10 s, 65 ℃ for 1 min) for 12 cycles, ④ 72 ℃ for 10 min, ⑤ 72 ℃ for 5 min, and then ice bath is performed. The third pre-amplification product is treated by EXO I enzyme after purification, and the purpose is to remove the PCR primer.
[0071] In the present application, after obtaining the third pre-amplification product, preferably, qPCR amplification is performed with the third pre-amplification product as a template, so as to obtain the expression level of miRNA isomer as part of quality control. The forward primer of the qPCR amplification is preferably a 5' universal primer. The reverse primer of the qPCR amplification is preferably the 3' universal primer. The probe of the qPCR amplification is preferably an LNA FAM probe, and the corresponding nucleotide sequence is shown in SEQ ID NO: 112 (ACC+AT+CA+AT+CG+TG+TG, + represents a locked nucleic acid). The reaction system of the qPCR amplification is preferably 10 μl, and preferably includes the following reagents: 0.08 μl of 2-fold diluted third pre-amplification product, 5 μl of 2x DNA polymerase mixture, 0.2 μM of final concentration of forward primer and 0.2 μM of final concentration of reverse primer and 0.2 μM of final concentration of probe, and 10 μl of ddH2O. The reaction program of the qPCR amplification is preferably 95°C for 10 min; 95°C for 30 s, 65°C for 1 min, 40 cycles.
[0072] After obtaining the third pre-amplification product, the present application uses primers with sequencing adapters for first PCR amplification with the third pre-amplification product as a template to obtain a PCR product containing an internal unique double tag.
[0073] In the present application, the reaction system of the first PCR amplification is preferably 30 μl, including the following steps: 2x PCR enzyme (containing UDG and UTP) 15 μl, 10 μM of each of the forward and reverse primers with internal double unique double tags 0.5 μl, third pre-amplification product 2 μl, and the balance of water. The reaction program of the PCR amplification is preferably ① 95°C for 10 min, ② (95°C for 15 s, 62°C for 30 s, 72°C for 1 min) 3 cycles, ③ (95°C for 15 s, 64°C for 30 s, 72°C for 1 min) 2 cycles, ④ (95°C for 15 s, 68°C for 30 s, 72°C for 1 min) 11 cycles, ⑤ (95°C for 15 s, 72°C for 20 min) 1 cycle, and ⑥ the program is run to 72°C for 18 min, followed by ice bath.
[0074] In the present application, the internal unique double tag PCR product is used as a template, and the second PCR amplification is performed with the barcode-labeled primer to obtain a double unique double tag PCR product, and the sample is mixed to obtain a sequencing library
[0075] In the present application, the reaction system of the second PCR amplification is preferably 30 μl, including the following steps: 2x PCR enzyme (containing UDG and UTP) 15 μl, 10 μM plus internal double unique double label forward and reverse primers each 0.5 μl, third pre-amplification product 2 μl, and the rest water. The reaction program of the PCR amplification is preferably ① 95℃ 10 min, ② (95℃ 15 s, 62℃ 30 s, 72℃ 1 min) 3 cycles, ③ (95℃ 15 s, 64℃ 30 s, 72℃ 1 min) 2 cycles, ④ (95℃ 15 s, 68℃ 30 s, 72℃ 1 min) 11 cycles, ⑤ (95℃ 15 s, 72℃ 20 min) 1 cycle, ⑥ program runs to 72℃ 18 min, and then ice bath is carried out.
[0076] In the present application, the PCR product containing double unique double labels is preferably precipitated after mixing. The product precipitate is removed by ExoI enzyme to obtain a sequencing library. Removing the PCR primer refers to removing the unreacted double unique double label amplification primer forward and reverse primers in the above PCR process. The purpose of removing the PCR primer is to prevent the downstream sequencing reaction of the PCR primer.
[0077] In the present application, the PCR product containing double unique double labels is obtained based on the double unique double index labeling technology for multiplexing of second-generation sequencing developed by the present application. The double unique double index labeling technology for multiplexing of second-generation sequencing is to add internal unique double labels (IUDI) to both ends of cDNA fragments by PCR, and then add external unique double labels and sequencing adapters to both ends of the PCR product obtained in the above step by PCR amplification means, to obtain a PCR product carrying double unique double labels, mix, and remove the PCR primer to obtain an amplification library for second-generation sequencing analysis. After the second-generation sequencing is completed, the raw data of the second-generation sequencing is split into the corresponding sample number when the sample is mixed according to the double unique double label sequence, and after splitting, irrelevant sequences are removed, and the micro ribonucleic acid isomers are identified and quantified.
[0078] In the present application, the PCR product containing double unique double labels is obtained based on the double unique double index labeling technology for multiplexing of second-generation sequencing developed by the present application. The double unique double index labeling technology for multiplexing of second-generation sequencing is to add internal unique double labels (IUDI) to both ends of cDNA fragments by PCR, and then add external unique double labels and sequencing adapters to both ends of the PCR product obtained in the above step by PCR amplification means, to obtain a PCR product carrying double unique double labels, mix, and remove the PCR primer to obtain an amplification library for second-generation sequencing analysis. After the second-generation sequencing is completed, the raw data of the second-generation sequencing is split into the corresponding sample number when the sample is mixed according to the double unique double label sequence, and after splitting, irrelevant sequences are removed, and the micro ribonucleic acid isomers are identified and quantified.
[0079] The present application provides a miRNA isomer composition related to breast cancer diagnosis, comprising nucleotide sequences as shown in SEQ ID NO: 129-SEQ ID NO: 210.
[0080] In the present application, the miRNA isomer composition is the miRNA isomer with significant differential expression (calibration p value less than 0.05) in breast cancer and normal people, and the remaining miRNA isomer after removing the miRNA isomer with high correlation (Pearson's correlation coefficient greater than or equal to 0.75) is used for machine learning model construction.
[0081] The present application provides the application of the miRNA isomer composition in the construction of breast cancer prediction model.
[0082] In the present application, the machine learning classifier of the breast cancer prediction model preferably comprises a support vector classifier. Preferably, the parameters of the SVM algorithm are optimized by grid search, and the numerical ranges of the parameters are gamma = 2 (-8-1) and cost = 2 (0-4). The prediction model is predicted by 10-fold cross-validation. After obtaining the prediction model, the model is preferably evaluated. The evaluation criteria preferably include accuracy (Accuracy) and kappa (Kappa). The accuracy, kappa and other evaluation indicators are preferably described by confusion matrix (Confusion Matrix) and ROC curve.
[0083] In the embodiment of the present application, the second-generation sequencing library constructed by the kit of the present application, according to the second-generation sequencing results, the sensitivity of the breast cancer prediction model constructed is 0.913, which can correctly identify 91.3% of the actual breast cancer cases without any false positive. The AUC value is 0.9924, and the threshold is 0.6461. By analyzing the ROC curve, it is shown that the SVM classifier performs very well in distinguishing cancer and non-cancer cases, and a strong balance between sensitivity and specificity is achieved at the selected threshold.
[0084] The present application provides the application of a primer for amplifying the miRNA isomer composition in the preparation of a diagnostic kit for breast cancer.
[0085] In the present application, the primer for amplifying the miRNA isomer composition is obtained by sequentially connecting a universal sequence as shown in SEQ ID NO: 98 (ATAGACTCCTCGCATAGCCTCATGAGTC) and a 5' end partial sequence of any one of the miRNA isomers. The 5' end partial sequence of the miRNA isomer refers to a sequence of 10-14 nt in length at the 5' end of the miRNA isomer, which can be 11 nt, 12 nt or 14 nt.
[0086] In the present application, the method for diagnosing breast cancer is based on the primer for sequencing the sample to be tested, bringing the sequencing data into a breast cancer prediction model constructed based on a support vector classifier, and determining whether the sample to be tested has a risk of breast cancer according to the prediction result of the breast cancer prediction model: when the prediction result is higher than the threshold value, it means that the sample to be tested has a higher risk of breast cancer, which can be further diagnosed by medicine; when the prediction result is lower than the threshold value, it means that the sample to be tested has a lower risk of breast cancer.
[0087] The kit for detecting breast cancer by plasma miRNA isomer next-generation sequencing provided by the present application will be described in detail below in combination with examples, but they should not be understood as limiting the scope of protection of the present application.
[0088] In the present technical solution, 100 breast cancer samples and 100 normal clinical samples are used, and the breast cancer samples (100) and the normal clinical samples (100) are derived from the Cancer Hospital of Chinese Academy of Medical Sciences.
[0089] The RNA extraction kit is purchased from Thermo Fisher.
[0090] Example 1
[0091] The breast cancer samples and the normal clinical samples are extracted by the RNA extraction kit, and after the total RNA concentration and quality are detected by the nucleic acid quantification detector, they are stored at -20℃ for standby.
[0092] Example 2
[0093] The reverse transcription and pre-amplification of the breast cancer sample IsomiRs include the following steps:
[0094] I. PolyA reaction:
[0095] Prepare the reaction system according to 20 μl of sample number, and each 20 μl of reaction system contains the following reagents: 5× reverse transcription buffer 4 μl, ATP (10 mM) 2 μl, PolyA enzyme (5000 U / μl) 1 μl, RNase inhibitor (40000 U / μl) 0.5 μl, total RNA 12.5 μl; the prepared reaction system is subjected to PolyA reaction under the following conditions: 37°C for 30 min, 65°C for 20 min; film-sealed-5°C storage (heat inactivation).
[0096] II. Reverse transcription reaction
[0097] Prepare the reaction system according to 20 μl of sample number, and each 20 μl of reaction system contains the following reagents: 10 mM dNTPs 1.5 μl, 10 μM reverse transcription primer (USRTPn, CCTCCATCCGAGACACACGATTGATGGTTTTTTTTTTTTTTTTTTVN, SEQ ID NO: 111) 1.5 μl, PolyA template 17 μl; the prepared reaction system is subjected to denaturation reaction by being treated at 65°C for 5 min, taken out 1 s before the end of the treatment, immediately ice-bath, ice-bath for 1 min, centrifugation;
[0098] Note: The dNTPs do not contain dUTP, otherwise the reverse transcription product cDNA will be degraded; the USEXPnb primer is purified by magnetic beads.
[0099] Prepare the reverse transcription reaction system from the above denaturation reaction product, and prepare the reaction system according to 30 μl of sample number, which contains the following reagents: 5× reverse transcription buffer 2 μl, 1.6 M trehalose 4.5 μl, Actinomycin D (1 mg / μl) 1.2 μl, T4 gp32 / RecA / ATP mixture 1.5 μl, RNase inhibitor (40000 U / μl) 0.3 μl, Maxima H reverse transcriptase (50 U / μl) 1.5 μl, and the above denaturation reaction product 19 μl; wherein the T4 gp32 / RecA / ATP mixture is prepared according to 2 sample numbers by adding the following reagents (μl): T4 gp32 (10 μg / μl) 0.6 μl, Tth RecA (2 μg / μl) 0.2 μl, ATP (100 mM) 0.24 μl, 1× reverse transcription buffer 1.96 μl. The prepared reaction system is subjected to reverse transcription reaction under the following conditions: 42°C for 15 min, 50°C for 30 min, 55°C for 30 min, 60°C for 30 min, 65°C for 30 min, and 85°C for 5 min.
[0100] III. Pre-amplification
[0101] (I) The first pre-amplification PCR reaction:
[0102] The pre-amplification PCR reaction liquid prepared according to 20 μl of sample number 20 μl reaction system configuration, 20 μl reaction system specific: 2*10 μl Boost mix, Tth RecA (0.2 μg / μl) 1 μl, Pre-IsomiR mix* (1 μM) 1.5 μl, reverse transcription product 7.5 μl;
[0103] Among them, 2*Boost mix (containing UDG) is prepared with dNTPs mixture without dUTP; 2*Boost mix is specifically described in the specific quantitative PCR reaction mixture of patent number: ZL 201910219827.4, the patent name is a kind of miRNA quantitative detection reagent and detection method; Pre-IsomiR mix* (1 μM) : take 10 μl of 97 primers (specific sequence see table 1) of mother liquor concentration 100 μM, then add 20 μl H2O (Nuclease-Free), that is, the primer mix of final concentration 1 μm (1000 μl) is prepared;
[0104] Table 1 Artificial miRNA sequence and pre-amplification primer sequence
[0105]
[0106]
[0107]
[0108] PCR instrument reaction program was set according to the following conditions: 1) 25℃ for 10 min, 2) 95℃ for 10 min, 3) (95℃ for 10 s, 55℃ for 10 min) for 3 cycles, 4) (95℃ for 10 s, 50℃ for 10 min) for 3 cycles, 5) (95℃ for 10 s, 45℃ for 10 min) for 2 cycles, 6) (95℃ for 10 s, 40℃ for 10 min) for 2 cycles, 7) (95℃ for 10 s, 37℃ for 10 min) for 2 cycles, 8) (95℃ for 10 s, 60℃ for 2 min, 72℃ for 10 min) for 1 cycle, 9) after running to 72℃ for 5 min, the PCR tube was taken out and immediately immersed in an ice box to terminate the activity of Taq DNA polymerase; after the liquid was frozen (about 3 min), the PCR tube was placed on a 96-well incubation module (previously frozen to -40℃); an equal volume of 20 μl chloroform was added, and vortexed immediately to melt the ice block (about 1 min); the centrifuge program was 12000 rpm, 4℃, 15 min centrifugation, and the supernatant (usually 18 μl) was taken out using a pipette and placed in a new labeled PCR tube; the samples were centrifuged to the bottom of the tube using a portable centrifuge, the lid of the chloroform-extracted PCR tube was opened, and the tube was placed in a PCR instrument for 50℃ treatment for 10 min to completely volatilize the chloroform; 2.5 μl EXO I enzyme was added to each tube, mixed well, centrifuged, and placed in a PCR instrument, and the program was set as 37℃ for 4 min, and paused when the program was left for 5 seconds; the PCR program was set and run: 37℃ for 4 min, 80℃ for 1 min; centrifuged to the bottom of the tube.
[0109] (II) Second pre-amplification PCR reaction:
[0110] The reaction system was prepared according to 20 μl for 1 sample, and the 20 μl reaction system specifically included the following reagents: 2×Boostmix 10 μl, 10 μm magnetic bead-purified transition primer (USEXPnb, TCTACAGATCCTGGCCTCTGACTCCAGGATCTGTAGACCTCCATCCGAGACACACGAT, SEQ ID NO: 99) 1 μl, 10 μm magnetic bead-purified IsomiR primer (IsomiRupb, GTTTGTTGCTACGCTCAGAATCCTAAGCGTAGCAACAAACATAGACTCCTCGCATAGCCTCATGAGTC, SEQ ID NO: 100) 1 μl, Tth RecA (0.2 μg / μl) 1 μl, first pre-amplification PCR product 7 μl;
[0111] The 2×Boost mix* (containing UDG) was prepared using a dNTPs mixture without dUTP, and the formula was the same as above;
[0112] PCR instrument setting program Touch Down PCR: ① 25℃ 10 min, ② 95℃ 10 min, ③ (95℃ 10 s, 65℃ 1 min) 3 cycles, ④ (95℃ 10 s, 62℃ 1 min) 3 cycles, ⑤ (95℃ 10 s, 58℃ 2 min) 2 cycles, ⑥ (95℃ 10 s, 60℃ 2 min) 2 cycles, ⑦ (95℃ 10 s, 60℃ 2 min, 72℃ 10 min) 1 cycle, ⑧ program runs to 72℃ 5 min, then take out the PCR tube, immediately immerse in an ice box, stop Taq activity; after the liquid is frozen (about 3 min), place the PCR tube on a 96-well incubation module (previously frozen to -40℃); add an equal volume of 20 μl chloroform, immediately vortex to ice block melt (about 1 min); centrifuge at 12000 rpm, 4℃ for 15 min, use a pipette to take the supernatant part (usually take 18 μl) to a new labeled PCR tube; use a portable centrifuge to separate all samples to the bottom of the tube, open the lid of the chloroform-extracted PCR tube, and place it in a PCR instrument at 50℃ for 10 min to completely volatilize the chloroform; add 4 μl of streptavidin magnetic beads (mix well before use and use immediately) to each 20 μl of the reaction solution; on a shaker, set the speed to 500 rpm, room temperature, shake for 30 min; vortex the magnetic beads well, incubate at 50℃ in a PCR instrument for 3 min; after incubation, place it on a magnetic stand for about 1 min, adsorb the magnetic beads, and use a pipette to take the liquid (try not to suck the magnetic beads) to another new labeled PCR tube.
[0113] (III) Third pre-amplification PCR reaction:
[0114] Prepare a reaction system of 20 μl for 1 sample, and the 20 μl reaction system specifically includes the following reagents: 2×Boostmix*10 μl, 10 μm URP (CAGAATCCTAAGCGTAGCAACAAAC, SEQ ID NO: 101) 1 μl, 10 μm UFP (GCCTCTGACTCCAGGATCTGTAGAC, SEQ ID NO: 102) 1 μl, Tth RecA (0.2 μg / μl) 1 μl, second pre-amplification PCR product 7 μl; wherein, 2×Boost mix* (containing UDG) is prepared with a dNTPs mixture without dUTP.
[0115] PCR instrument setting program: ① 95℃ 10 min, ② (95℃ 10 s, 65℃ 1 min), 12 cycles, ④ 72℃ 10 min, ⑤ 72℃ 5 min, then take out the PCR tube and immediately immerse it in a programmed cooling box containing isopropanol at -80℃ to terminate the activity of Taq DNA enzyme (to avoid non-specific amplification caused by temperature reduction); after the liquid is frozen (about 3 min), place the PCR tube on a 96-well incubation module (previously frozen to -40℃); add an equal volume of 20 μl chloroform and immediately vortex to ice block melting (about 1 min); centrifuge at 12000 rpm, 4℃ for 15 min, and use a pipette to take the supernatant part (usually 18 μl) to a labeled new PCR tube; use a portable centrifuge to separate all samples to the bottom of the tube, open the lid of the PCR tube after chloroform extraction, and place it in a PCR instrument at 50℃ for 10 min to completely volatilize the chloroform; add 2.5 μl EXO I (Thermolabile) mixture to each reaction (20 μl); set the PCR program and run: 37℃ for 4 min, 80℃ for 1 min; take 5 μl and dilute 10 times with 0.1×TE as a PCR template for subsequent detection.
[0116] Example 3
[0117] Pre-amplification product qPCR detection
[0118] qPCR amplification detection: configure USQ-miR DNA polymerase mixture containing 0.2 μM (final concentration) of forward primer (UFP: GCCTCTGACTCCAGGATCTGTAGAC, SEQ ID NO: 102), universal reverse primer (URP: GCCTCTGACTCCAGGATCTGTAGAC, SEQ ID NO: 101) and 0.2 μM (final concentration) of LNA FAM probe (ACC+AT+CA+AT+CG+TG+TG, SEQ ID NO: 112, + means locked nucleic acid, abbreviated as LNA), the amount of PCR template in a 10 μl PCR reaction system is 0.08 μl of 10-fold dilution of the third pre-amplification PCR product;
[0119] The USQ-miR DNA polymerase mixture also contains 2×qPCR premix, the formula is: Tris-HCl pH 8.8, 75 mM; (NH4)2SO4, 20 mM; Triton-100, 0.1%; MgCl2, 2.5 mM; dNTPs, 200 μM; Trehalose, 200 mM; Taq DNA polymerase, 50 U / ml.
[0120] The PCR cycle parameters are: 95℃ for 10 min, then 95℃ for 30 s, 65℃ for 1 min, 40 cycles.
[0121] Example 4
[0122] PCR amplification of pre-amplification product plus barcode and adapter (sequencing adaptor):
[0123] (1) Design of barcode and sequencing adaptor (adapter) primers:
[0124] The design method of the forward primer is as follows:
[0125] The sequence of the forward primer overlapping with the sequencing adaptor (SEQ ID NO: 103) + IUDI (I5 Index) + the sequence partially overlapping with the 5' end universal sequence part of the cDNA reverse transcription product (SEQ ID NO: 104);
[0126] The design method of the reverse primer is as follows:
[0127] The sequence of the reverse primer overlapping with the sequencing adaptor (SEQ ID NO: 105) + IUDI (I7 Index) + the sequence partially overlapping with the 3' end sequence of the IsomiR primer (SEQ ID NO: 106).
[0128] AATGATACGGCGACCACCGAGATCTACACtacgaatcttACACTCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID NO: 113);
[0129] CTGTCTCTTATACACATCTCCGAGCCCACGAGACaccaagttacCTCGGAGATGTGTATAAGAGAC AG (SEQ ID NO: 114).
[0130] (2) PCR template: the third pre-amplification PCR product without dilution, one tube for each sample;
[0131] (3) Preparation of PCR reaction solution: 2x PCR enzyme (containing UDG and UTP) 15 μl, water 12 μl, barcode and sequencing adaptor forward primer (10 μM) 0.5 μl, barcode and sequencing adaptor reverse primer (10 μM) 0.5 μl, third pre-amplification PCR product 2 μl, total volume 30 μl.
[0132] (4) Quantitative PCR with barcode and sequencing adaptor:
[0133] Each third pre-amplification PCR product was subjected to quantitative PCR with specific forward primer with barcode and sequencing adaptor and reverse primer with barcode and sequencing adaptor, and the reaction program was as follows: 37°C for 10 min, ① 95°C for 10 min, ② (95°C for 15 s, 62°C for 30 s, 72°C for 1 min) for 3 cycles, ③ (95°C for 15 s, 64°C for 30 s, 72°C for 1 min) for 2 cycles, ④ (95°C for 15 s, 68°C for 30 s, 72°C for 1 min) for 45 cycles, and the fluorescence signal was collected at 72°C for 1 min.
[0134] (5) The above PCR reaction was repeated, but the third pre-amplification PCR product of each sample determined by the above qPCR was directly amplified by a general PCR instrument with the dilution fold and the number of logarithmic phase cycles determined by the above qPCR. Various parameters including temperature rising rate and falling rate were as far as possible the same as those in the above qPCR program. The number of logarithmic phase cycles of the barcode addition PCR reaction was determined.
[0135] (6) General PCR amplification program: ① 95°C for 10 min, ② (95°C for 15 s, 62°C for 30 s, 72°C for 1 min) for 3 cycles, ③ (95°C for 15 s, 64°C for 30 s, 72°C for 1 min) for 2 cycles, ④ (95°C for 15 s, 68°C for 30 s, 72°C for 1 min) for 11 cycles, ⑤ (95°C for 15 s, 72°C for 20 min) for 1 cycle, ⑥ the PCR tube was taken out after the program ran to 72°C for 18 min, and was immediately placed on ice to terminate the enzyme activity. The fluorescence signal was collected at 72°C for 1 min.
[0136] (7) An equal volume of 30 μl chloroform (chloroform was previously placed on ice for 30 min) was added on ice, and vortexed (about 1 min).
[0137] (8) Centrifuged at 12000 rpm and 4°C for 15 min, and 25 μl supernatant was taken to a new tube (Tip did not touch chloroform, and a part of the supernatant was reserved).
[0138] (9) The sample was completely separated to the bottom of the tube by a small centrifuge, and the supernatant was placed in a PCR instrument at 50°C for 10 min to completely volatilize chloroform, otherwise chloroform would inhibit the downstream enzyme reaction.
[0139] (10) 2.5 μl diluted EXOI (Thermolabile) solution was added to each PCR reaction.
[0140] (11) The mixture was inverted and mixed, and small centrifuged, and incubated at 37°C for 20 min and 42°C for 10 min.
[0141] (12) Inactivated ExoI by treating at 60°C for 15 min.
[0142] (13) Run 3% agarose gel for 45-60 min, 50 bp Marker, see if primer band is gone.
[0143] Example 5
[0144] Quantitative qPCR and amplification of adapter added PCR products (2x PCR enzyme with UTP, without UDG):
[0145] Quantitative qPCR of barcode and adapter added PCR products:
[0146] The PCR product sample obtained in the above experiment was diluted 50-fold as a template, and the design method of the primers used was as follows:
[0147] Reverse primer: sequence containing Nextera I7 sequencing adapter (SEQ ID NO: 109) + external unique dual tag (I7 Index sequence) + partial overlap with the 5' end of the IUDI in the above reverse primer (SEQ ID NO: 110);
[0148] wherein the I5 Index and I7 Index sequences are selected from the set of I5 Index and I7 Index sequences in Table 2, but are different from the I5 Index and I7 Index sequences involved in the first PCR amplification;
[0149] Examples of primers for the second PCR adding external unique dual tag (OUDI) are as follows:
[0150] Forward primer: AATGATACGGCGACCACCGAGATCTACACtacgaatctTACACTCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID NO: 115);
[0151] Reverse primer: CTGTCTCTTATACACATCTCCGAGCCCACGAGACaccaagttacTTCGGAGATGTGTATAAGAGACAG (SEQ ID NO: 116);
[0152] Table 2 I5 sequencing adapter and Nextera I7 sequencing adapter
[0153] Adapter name Sequence Sequence number I5 sequencing adapter AATGATACGGCGACCACCGAGATCTACA SEQ ID NO: 107 Nextera I7 sequencing adapter CTGTCTCTTATACACATCTCCGAGCCCACGAGA SEQ ID NO: 109
[0154] qPCR reaction program as follows: ① 95℃ 10min, ② (95℃ 15s, 62℃ 30s, 72℃ 1min) 3 cycles, ③ (95℃ 15s, 64℃ 30s, 72℃ 1min) 2 cycles, ④ (95℃ 15s, 68℃ 30s, 72℃ 1min) 45 cycles, set 72℃ 1min to collect fluorescence signal. qPCR set 3 repeats for each sample;
[0155] According to the above quantitative qPCR results, calculate dilution factor, ordinary PCR amplification (low cycle number, to prevent the introduction of human error), 6 holes for each sample;
[0156] Ordinary PCR amplification conditions: ① 95℃ 10min, ② (95℃ 15s, 62℃ 30s, 72℃ 1min) 3 cycles, ③ (95℃ 15s, 64℃ 30s, 72℃ 1min) 2 cycles, ④ (95℃ 15s, 68℃ 30s, 72℃ 1min) 11 cycles to be determined, ⑤ (95℃ 15s, 72℃ 20min) 1 cycle, ⑥ run the program to 72℃ 18min, then take out the PCR tube and immediately place it on ice to terminate Taq activity. The cycle number is determined by the above qPCR (if it cannot continue to do, save at 4℃, or if you want to store for a long time -20℃);
[0157] Add 2.5 μl of diluted EXOI (Thermolabile) to each PCR reaction;
[0158] Mix well, 37℃ 20min, 42℃ 10min;
[0159] 60℃ treatment for 15min, inactivate ExoI;
[0160] All samples are placed on ice, and the PCR samples of 6 holes for each sample are mixed. Then qPCR quantification (dilute 100000 times for each sample. Do 3 repeats from dilution, 45 cycles).
[0161] Example 6
[0162] Precipitation and gel recovery of barcode and adapter added PCR product mixed samples:
[0163] According to the qPCR quantification results, all samples are mixed equally and vortexed. Follow the following experimental steps to do two tubes;
[0164] Take 700 μl of mixed sample with adapter to 1.5 ml EP tube; add 77 μl of 3M pH 5.2 sodium acetate solution; add 500 μl of isopropyl alcohol, mix well (the above samples need to be placed on ice); -20°C or -80°C for 1 h, put the centrifuge tube into the centrifuge (with the lid handle outward), 15000 g centrifugation for 30 min at 4°C; the white DNA precipitate is at the bottom of the tube outward, carefully pour off the supernatant. The remaining supernatant is carefully removed with a gun, do not touch the DNA precipitate during the process to avoid the removal of DNA; add 500 μl of 70% room temperature ethanol, place at room temperature for 5 min; 15000 g centrifugation for 30 min at 4°C; the white DNA precipitate is at the bottom of the tube outward, carefully pour off the supernatant, carefully remove the remaining supernatant with a P200 gun; place on the clean bench (do not cover the centrifuge tube lid), turn on the hair dryer for about 10 min; add 60 μl of TE to dissolve; prepare 1.5% agarose gel (gel thickness is about 1 cm, can hold 15 μl of sample, length is twice normal, about 15 cm); run 3-4 hole gel recovery, note: the dye band should run to the bottom of the gel, otherwise the DNA size cannot be fully separated; when cutting the gel, the smaller the gel strip, the better, but strictly include the main band; the recovered DNA is dissolved with 60 μl of TE, the DNA concentration is determined, the gel is run for identification, and is placed at -20°C for standby; use a mixture of all samples for precipitation and gel recovery, if the PCR product is high in purity, do not do gel recovery, remove the PCR primer with the above ExoI after the product is precipitated, and obtain the tiny ribonucleic acid isomer library. The established tiny ribonucleic acid isomer library is subjected to NGS high-throughput sequencing to obtain sequencing data results.
[0165] Example 7
[0166] Establishment of machine learning model
[0167] Based on the sequencing data results in Example 6, first, the t-Test P value of the expression difference of each IsomiR in 100 breast cancer samples and 100 normal samples is calculated, and the IsomiRs are arranged in ascending order of P value. The first 239 IsomiRs are selected from all the isomers arranged in ascending order of P value, and the IsomiRs highly correlated with other IsomiRs are removed after calculating the correlation between different IsomiRs. The remaining data is subjected to machine learning classification by using different classifiers such as SVM, KNN, RF, CART and IDA, and the best classifier is found. A variety of classifiers are used, and the data is first divided into two parts: 80% of the part is used to train the model, and 20% of the part is used to verify the model. According to the accuracy and Kappa value, the best classifier result is the support vector machine (Support Vector Machine, SVM) algorithm.
[0168] SVM algorithm to establish a machine learning model for breast cancer diagnosis: The above-mentioned second-generation sequencing data is divided into two parts, 80% of which is used to train the model (training set), and 20% of which is used to verify the model (test set). The samples are divided into training set and test set. The repeated samples only exist in the training set or the test set, that is, to ensure that different repeats of the same sample cannot exist in the training set and the test set at the same time, otherwise it will lead to information leakage and the evaluation of the model will be too high.
[0169] Optimization of SVM algorithm model: The parameters of SVM algorithm need to be adjusted to find the best parameters. Grid search is used to optimize the parameters of SVM algorithm. The numerical range of parameters used is gamma = 2 (-8:1) , cost = 2 (0:4) . Thus, there are 10 values of gamma and 5 values of cost, i.e. 50 combinations. Each combination is subjected to 10-fold cross validation, i.e. the training set is divided into ten parts, and each time 9 parts are used as training data and 1 part is used as test data. Each test will give the corresponding error rate. The average error rate of each combination is obtained by ten tests. The gamma / cost combination with the lowest average error rate is the optimal parameter of the SVM algorithm. Since the final diagnostic model is obtained by 500 experiments (50x10), overfitting can be avoided. Overfitting is a phenomenon that the trained model performs well on the training set but poorly on the test set.
[0170] Model evaluation: There are many different indicators to evaluate machine learning algorithms. The default evaluation criteria for classification problems are accuracy and kappa. Kappa is similar to accuracy, but it is calibrated by the random baseline of the data set. Kappa value represents consistency and also the accuracy of classification, positive value close to 1 represents better consistency, generally above 0.75 represents that the consistency result is satisfactory, 0.8-1 is almost completely consistent. Accuracy, kappa and other evaluation indicators can be described by confusion matrix and ROC curve.
[0171] As shown in Figure 1 : 100 breast cancer and 100 normal samples, a total of 200 samples, are labeled with 200 pairs of internal unique dual indexes (IUDI), and then added with sequencing adapters containing external unique dual indexes (OUDI, ACCAAGTTAC and AAGATTCGTA). After mixing, use NovaSeq TMXPlus platform with paired-end reads. The results were: 69.64M reads; 10.45G bases; the quality control score Q30 (% of bases with quality value greater than 30 (error rate less than 0.1%) in the original sequence) was: 71.5. The data were split into 200 files corresponding to 200 samples by unique dual tag sequences. After trimming the sequences irrelevant to mature miRNA such as tag and adapter sequences, the mature miRNA sequence data were obtained. Running the IsoMiRmap software to process each file, the miRNA isomer sequences and their reads were obtained. Then the DESeq2 software was used for miRNA isomer differential expression analysis. DESeq2 is used for statistical analysis of RNA-seq data, identifying differentially expressed genes under different conditions, taking into account biological variability, and providing corrected p values. Data with corrected p values greater than 0.05 were filtered out, and data with calibrated P values less than 0.05 were retained. The correlation coefficient between miRNA isomers was calculated, and other miRNA isomers with a correlation coefficient greater than 0.5 were filtered out. The remaining data were classified by machine learning using different classifiers to find the best classifier. A variety of classifiers were used, and the data were first divided into two parts: 80% of the data were used to train the model, and 20% of the data were used to verify the model. The best result was the Support Vector Machine (SVM) algorithm.
[0172] As shown in Figure 2 : The points in the figure represent NGS reads and the corresponding input concentration values. Z-score is used to detect outliers, i.e. points with Z-score greater than 3 are considered outliers (red) and are excluded from linear regression analysis to avoid distorting the relationship between NGS reads and input concentration. The data points without outliers (light blue) are used to fit the linear model. The equation in the figure represents the linear relationship between NGS reads (x-axis) and input concentration (y-axis). It shows that as the NGS reads increase, the input concentration also increases proportionally (slope 1.17e-07). Linear regression analysis shows that the NGS reads of artificial miRNA have a strong correlation with the input concentration (correlation coefficient 0.86, R 2 = 0.74). The very small P value (6.26e-82) indicates that the observed relationship is statistically highly significant, meaning that the probability of this correlation occurring by chance is very low. This is of great significance for quality control or optimization in sequencing experiments, i.e. the NGS reads of miRNA isomers are positively correlated with the content of miRNA isomers in plasma, which can quantify the miRNA isomers in plasma.
[0173] As shown in Figure 3Figure 2: Heatmap of the top 30 isomiRs (i.e. microRNA isoforms) with the highest variance. The data was processed with the DESeq2 package and analyzed using Variance Stabilizing Transformation (VST). Further analysis was performed on the data by row clustering and column clustering according to the expression pattern of the isomiRs and the different samples. Row clustering reveals isomiRs that behave similarly under different conditions. The color scale of each row in the figure was adjusted independently by calibrating the expression level of each isomiR to the mean value across all samples, allowing a quick view of which isomiRs are up- or down-regulated. Column clustering shows the similarity or difference in expression response between the breast cancer group and the normal group. Blue indicates lower expression levels, and yellow to red indicates higher expression levels. At the top of each column, 'C' stands for breast cancer (red), and 'N' stands for normal (blue).
[0174] As Figure 4 shown: Heatmap of the top 30 isomiRs (i.e. microRNA isoforms) with the highest variance. The data was processed with the DESeq2 package and analyzed using Variance Stabilizing Transformation (VST). Further analysis was performed on the data by row clustering and column clustering according to the expression pattern of the isomiRs and the different samples. Row clustering reveals isomiRs that behave similarly under different conditions. The color scale of each row in the figure was adjusted independently by calibrating the expression level of each isomiR to the mean value across all samples, allowing a quick view of which isomiRs are up- or down-regulated. Column clustering shows the similarity or difference in expression response between the breast cancer group and the normal group. Blue indicates lower expression levels, and yellow to red indicates higher expression levels. At the top of each column, 'C' stands for breast cancer (red), and 'N' stands for normal (blue).
[0175] As Figure 5Shown: The confusion matrix and its related classification statistics were generated by a Support Vector Machine (SVM) classifier trained with 10-fold cross-validation. The model's hyperparameters (gamma and cost) were tuned and predictions were made on the test dataset. The performance metrics summarize the model's ability to classify samples into two classes: cancer and normal. The SVM classifier achieved a high accuracy (93.48%), strong sensitivity (95.65%), and specificity (91.30%), indicating its good performance in distinguishing between cancer and normal samples. The Kappa coefficient of 0.8696 further supports the model's reliability, while the McNemar's test shows no significant imbalance between the two types of errors. Overall, the classifier performs very well on the given dataset. The confusion matrix provides a summary of the predicted results: True Positives (TP) (cancer predicted as cancer): 22; False Positives (FP) (normal predicted as cancer): 2; True Negatives (TN) (normal predicted as normal): 21; False Negatives (FN) (cancer predicted as normal): 1; this matrix is used to calculate performance metrics such as accuracy, sensitivity, specificity, etc. Accuracy: 0.9348 (93.48%). This indicates that the model correctly classified 93.48% of the test samples. 95% Confidence Interval: (0.821, 0.9863). This indicates that with 95% confidence, the true accuracy will fall within this interval. No Information Rate (NIR): 0.5. This is the accuracy of always predicting the most common class (baseline accuracy). P-value (Acc > NIR): 2.311e-10. This is a significance test to compare the model's accuracy against the NIR, and the result indicates that the model significantly outperforms random guessing. Kappa Coefficient: 0.8696. The Kappa coefficient measures the agreement between the predicted and actual classes, excluding the influence of chance agreement. A Kappa of 0.8696 indicates excellent agreement. McNemar's Test P-value: 1. This test assesses the significance of the difference between the two types of errors (cancer vs. normal). A P-value of 1 indicates that there is no significant difference between the two types of errors. Sensitivity (Recall, Cancer): 0.9565 (95.65%). The proportion of actual cancer cases that the model correctly identified. Specificity: 0.9130 (91.30%). The proportion of actual normal cases that the model correctly identified. Positive Predictive Value (PPV): 0.9167 (91.67%). The probability that a sample predicted as cancer is actually cancer. Negative Predictive Value (NPV): 0.9545 (95.45%). The probability that a sample predicted as normal is actually normal. Prevalence: 0.5. The proportion of actual positive (cancer) cases in the test data. Detection Rate: 0.4783. The proportion of actual positive cases that were correctly identified (22 / 46 samples). Detection Prevalence: 0.5217. The proportion of samples that the model predicted as positive (cancer), i.e., the number of times the model predicted cancer. Balanced Accuracy: 0.9348.The average of sensitivity and specificity provides a balanced measure for imbalanced datasets.
[0176] As Figure 6 shown: An ROC curve (Receiver Operating Characteristic curve) is displayed for an SVM (Support Vector Machine) classifier to predict cancer. The ROC curve represents the performance of the classifier by plotting the true positive rate (sensitivity) and false positive rate (1-specificity) at different thresholds. True positive rate (Y-axis): This represents the sensitivity of the classifier, i.e., the proportion of actual positive (cancer) cases that the model correctly identified. False positive rate (X-axis): This represents 1-specificity, i.e., the proportion of negative (non-cancer) cases that the model incorrectly classified as positive. ROC curve: This colored curve demonstrates the trade-off between sensitivity and specificity of the classifier at different thresholds. A perfect classifier's curve would hug the top-left corner of the plot (true positive rate = 1, false positive rate = 0). Diagonal line: The dotted line represents the performance of a random classifier (AUC = 0.5), which is a baseline for comparison. The further the ROC curve is from this line, the better the performance of the classifier. AUC (Area Under the Curve): The AUC value (0.9924 here) is a single scalar value that summarizes the overall performance of the classifier. The closer the AUC is to 1, the better the performance of the classifier, while an AUC of 0.5 represents random performance. Cutoff (Threshold): The threshold value (0.6461 here) is the optimal threshold at which the classifier achieves a balance between sensitivity and specificity, determined by the custom `opt.cut` function. Sensitivity: At the optimal threshold, the sensitivity of the classifier is 0.913, meaning it can correctly identify 91.3% of actual cancer cases. Specificity: The specificity is 1, indicating that the classifier can correctly identify all non-cancer cases without any false positives at the optimal threshold. The ROC curve evaluates the performance of the SVM classifier in distinguishing between cancer and non-cancer cases, visualizing the trade-off between sensitivity and specificity, and thus helping to select the optimal classification threshold (Cutoff). A higher AUC value indicates that the classifier performs very well, achieving a good balance between true positives and false positives. AUC: Represents the area under the curve, used to summarize the overall performance of the classifier. By analyzing this ROC curve, it can be concluded that the SVM classifier performs very well in distinguishing between cancer and non-cancer cases, achieving a strong balance between sensitivity and specificity at the chosen threshold.
[0177] Cutoff (Threshold): Represents the optimal threshold at which the classifier balances between sensitivity and specificity.
[0178] Sensitivity and Specificity: These values show the performance of the classifier at the optimal threshold, helping users understand the classifier's discrimination ability at this threshold.
[0179] Table 3 Artificial miRNA sequences and their pre-amplification primer sequences
[0180]
[0181]
[0182] As shown in Table 3, as experimental controls, here artificial miRNAs are not exist in human, but their sequences and pre-amplification primers are as similar as possible to natural miRNAs. We added artificial miRNA-1, artificial miRNA-2 and artificial miRNA-3 into all breast cancer plasma samples, and added another 3 artificial miRNAs into all normal human plasma samples. The concentrations of the three artificial miRNA markers are high concentration (1.0E-04 μM), medium concentration (5.0E-05 μM) and low concentration (1.0E-05 μM). That is, different artificial miRNAs are added into different samples as internal markers. In this way, even if the samples are mislabeled during the experiment, these unique miRNA markers can be used to correct this common human error. At the same time, because the concentrations of artificial miRNAs are known, they can be used to absolutely quantify other natural miRNAs and calibrate experimental samples. The second-generation sequencing results (Table) show that, compared with manual recording and artificial miRNA markers, in 55 breast cancers, 5 results cannot determine whether they are consistent, which may be due to the low concentration of artificial miRNAs; 1 result has no artificial miRNA marker, which may be missed; one result is inconsistent, suggesting human error during the experiment. In addition, other markers are correct. This labeling method is feasible and can ensure accurate tracking of samples, especially when handling a large number of samples or conducting high-throughput experiments, which can effectively reduce the risk of sample misplacement or cross-contamination.
[0183] Table 4 Second-generation sequencing results of six artificial miRNA markers tracking breast cancer and normal human samples
[0184]
[0185]
[0186]
[0187]
[0188]
[0189] Table 5 miRNA isomers with differential expression of 1.5 times or more
[0190]
[0191] Table 6 shows the miRNA isomiRs that are significantly differentially expressed (calibrated p-value less than 0.05) in breast cancer and normal human, after removing highly correlated (Pearson's correlation coefficient greater than or equal to 0.75) miRNA isomiRs, the remaining miRNA isomiRs are used for machine learning model construction.
[0192] Table 6 miRNA isomiRs for machine learning
[0193]
[0194]
[0195] The isomiR unique name of each isomiR is named according to the method of the prior art (see literature: Pliatsika V. et al. (2018) MINTbase v2.0: a comprehensive database for tRNA-derived fragments that includes nuclear and mitochondrial fragments from all The Cancer Genome Atlas projects. Nucleic Acids Res., 46, D152-D159). The isomiR unique name in Table 6 can also be referred to as the isomiR "license plate", each isomiR has a unique "license plate", and each "license plate" corresponds to a specific isomiR. This "license plate" serves as an identifier for the isomiR, which is used to accurately track and distinguish each miRNA isomiR during sequencing or analysis, and the middle separated number is the base number of this miRNA isomiR.
[0196] The above only describes the preferred embodiments of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can also be made, which should be considered as the protection scope of the present application.
Claims
1. A miRNA isoform composition associated with breast cancer diagnosis, characterized in that, The nucleotide sequences of the miRNA isoform composition are shown in SEQ ID NO:129~SEQ ID NO:
210.
2. The use of the miRNA isoform composition of claim 1 in constructing a breast cancer prediction model.
3. The application according to claim 2, characterized in that, The machine learning classifier for the breast cancer prediction model includes a support vector classifier.
4. The use of a primer for amplifying the miRNA isoform composition of claim 1 in the preparation of a diagnostic kit for breast cancer.
Citation Information
Patent Citations
Specific quantitative PCR reaction mixed liquor, miRNA quantitative detection kit, and detection method
CN109957611A
Internal reference substance for detecting bladder cancer serum miRNA and its detection primers and use
CN103602747A
Marker miR-126-3P for HER-2 positive breast cancer tissue, and application and diagnosis kit of marker miR-126-3P
CN106636375A
Construction method and application of repeatable dual unique double-tag library for next-generation sequencing of microRNA isomers
CN118421787A