SNP-STR composite genetic marker detection kit and application thereof

By combining ASA-PCR technology, an SNP-STR composite genetic marker detection kit that can synchronously identify STR and SNP sites was developed, solving the problems of limited sensitivity and single typing of human cell line identity identification and cross-contamination detection in the prior art, achieving higher detection accuracy and cumulative individual recognition rate.

CN119979690APending Publication Date: 2025-05-13GUANGXI MEDICAL UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510335086.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has problems such as limited sensitivity, single typing and lack of a comparison basis in the identification and cross-contamination of human cell lines, which leads to instability in cell line identification and difficult to accurately detect cross-contamination of cell lines.

Method used

Combined with ASA-PCR technology, a complex amplification system that can synchronously identify STR and SNP loci was established, and a SNP-STR complex genetic marker detection kit was developed, covering 21 STR loci, 12 SNP loci, 1 Y-STR loci and 1 Y-InDel loci. Through multicolor fluorescent marker composite amplification and capillary electrophoresis technology, the typing ability and cumulative individual recognition rate were improved.

Benefits of technology

The accuracy of identity identification of human cell lines and the reliability of cross-contamination detection are achieved, the cumulative individual recognition rate is improved and the identity misjudgment rate is reduced, and the requirements for cross-contamination identification and screening of cell lines in the biomedical field are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119979690A_ABST
    Figure CN119979690A_ABST
Patent Text Reader

Abstract

The invention discloses an SNP-STR (single nucleotide polymorphism-short tandem repeat) composite genetic marker detection kit and application thereof. The kit comprises specific primers for amplifying 21 STR gene loci, 12 SNP gene loci linked with the STR gene loci, a Y-STR gene locus, Y-Indel and Amel. The SNP-STR composite genetic marker has the advantages as follows: (1) the SNP-STR composite genetic marker meets the ANSI Standard standard for cell line identification, can be parallelly compared with data in an existing conventional mainstream database, and has good universality; (2) an SNP-STR composite amplification system is established and optimized, synchronous amplification of the two genetic markers is realized, the typing ability and cumulative individual recognition rate of a detection system are improved, and the method has important significance on identity identification and cross contamination identification of human cell lines; (3) a Y chromosome signal loss phenomenon caused by AMEL-Y gene deletion or mutation is overcome, and the accuracy of a human cell line sex determination result is improved; and (4) the kit is convenient and fast to operate and low in cost in the practical application process, and the detection speed and efficiency are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of biotechnology, and relates to a method for identifying cross-contamination of human cell lines established by combining ASA-PCR technology, in particular to a SNP-STR composite genetic marker detection kit and application thereof. Background Art

[0002] Human cell lines are important biological resources widely used in basic and clinical medical research, drug screening and production, gene therapy, vaccine production and other fields. However, cross-contamination of human cell lines and cell line identity errors, such as the contamination caused by mixing or replacing human cell lines with other human or non-human cell lines, have caused huge time and economic losses to biomedical research and development and production. Since the first report of cell cross-contamination in 1968, many famous journals such as Science, Nature, and The EMBO Journal have reported that cell lines used in biological resource preservation institutions, research institutions or biomedical fields are contaminated by other cell lines. According to the German Collection of Microorganisms and Cell Cultures, one of the three largest cell culture collections in the world, the cross-contamination rate of cell lines in the international biomedical field is as high as 18%-36%. Ye et al., Huang et al., and Bian et al., respectively reported in the internationally renowned journals FASEB Journal, PLOS ONE, and SCIENTIFIC REPORTS that the cross-contamination rate of cell lines in my country is as high as 25%-46%. Taking only two widely used misidentified cell lines, INT Taking 407 and HEp-2 as examples, direct economic losses of US$713 million and indirect economic losses of US$3.5 billion have been caused. Therefore, cell line identity identification and cross-contamination identification are the top priorities in the research and production quality control of the biopharmaceutical field. At present, both China's "Chinese Pharmacopoeia" (CHP) and the international "United States Pharmacopoeia" (USP) and "European Pharmacopoeia" (EP) clearly stipulate that in the production management and quality control of biological products, cell banks must undergo a comprehensive identification and verification process. In addition, cell identification / verification reports are essential materials for the application of clinical trials (IND) and marketing approval (NDA) in the research and development process of biological products; these reports are also mandatory items for regulatory agencies such as the Center for Drug Evaluation (CDE) of the State Food and Drug Administration and the U.S. Food and Drug Administration (FDA).

[0003] At present, the identification and cross-contamination detection of human cell lines completely rely on the STR (short tandem repeat) typing detection system used in the field of judicial / public security evidence identification. STR, also known as microsatellite sequence, is a short tandem repeat sequence that exists in large quantities in human genomic DNA. The repeating unit is 2-6 nucleotides. Due to its high polymorphism, small fragments, easy amplification, and similar amplification conditions of each locus, it can realize the automated detection of composite amplification. It has the advantages of sensitivity, accuracy, speed, and large amount of information, and is therefore very suitable for establishing a DNA database. In the 1990s, the US FBI selected 13 autosomal STR loci for the establishment of a DNA data index system - CODIS CSF1PO, vWA, TH01, TPOX, D3S1358, D5S818, D7S820, D8S1179, D13S317, D16S539, D18S51, D21S11 and FGA. The system has been used by DNA databases in many countries and applied to the field of cell identity identification. The current cell line authentication standard ANSI Standard (ASN-0002) stipulates that when using STR typing to identify the identity and cross-contamination of cell lines, the above 13 CODIS loci must be tested.

[0004] However, the mutation rate of STR loci is about 5×10 -3 , cell lines often show high-frequency microsatellite instability (microsatellite instability high, MSI-H phenomenon, which refers to the presence of DNA mismatch repair system (mismatch repair, MMR) defects in cell lines, resulting in unstable STR sequences), loss of heterozygosity (i.e., the phenomenon that it is difficult to distinguish from homozygotes due to the loss of one allele), loss of Y caused by AMEL-Y gene deletion or mutation, and the subculture conditions of cell lines (such as excessive passage or excessive dilution) and the differences in cell line culture conditions among institutions will accumulate genetic drift. When the above phenomena lead to incomplete consistency in STR typing between the sample to be tested and the reference sample, it is often difficult to distinguish whether it is caused by genetic drift of the cell line itself or cross-contamination. For example, the STR matching rate of the high-frequency microsatellite unstable cell line LS174T and its own subtype cell line LS74T-HM7 is only 66%, while the SNP matching rate of the two is as high as 99%. All of the above factors are likely to lead to the defect of unstable detection effect that relies solely on STR loci. In addition, the loss of Y chromosome signal caused by AMEL-Y gene deletion or mutation in human cell lines is very common. Yu et al. reported in the top international journal Nature that by comparing the STR gender determination results with the gender annotated on the cell lines, they found that 34% of the male cell lines had to be identified as female due to the loss of AMEL signals.

[0005] In the prior art, STR and SNP are tested separately for cell line identity and cross-contamination detection. TM 21 ID System (Chinese patent application number CN201410076618.6) detects 20 autosomal STR loci + 1 AMEL gender identification site; Basepoint Cognition uses Goldeneye DNA ID System 20A to detect 19 autosomal STR loci + 1 AMEL gender identification site; DSMZ uses 24 SNP sites to detect cell line identity; Chinese patent application number CN201410145409.2 "A mitochondrial SNP fluorescent labeling composite amplification kit and its application" involves the simultaneous detection of 61 mitochondrial SNP sites. The defects of the above detection methods are limited detection sensitivity and single typing. In addition, the databases of global professional cell line collections and identification institutions such as ATCC, DSMZ, JCRB, Cellosaurus, etc. only provide STR loci data comparison services for human cell lines. Even if SNP loci are used for cell line identity identification, the authenticity of the cell line identity cannot be determined due to the lack of a public database of cell line SNP genetic markers, that is, there is no basis for comparison. Summary of the invention

[0006] Technical problem to be solved: In order to overcome the shortcomings of the prior art, a composite amplification system capable of simultaneously identifying STR and SNP sites is established based on multicolor fluorescent labeling composite amplification and capillary electrophoresis, combined with ASA-PCR technology, and cross-contamination detection of human cell lines is performed to improve the cumulative individual recognition rate of human cell lines and reduce their identity misjudgment rate, thereby meeting the requirements for cross-contamination identification and screening of cell lines in the biomedical field. In view of this, the present invention provides a SNP-STR composite genetic marker detection kit and its application.

[0007] Technical solution: A SNP-STR composite genetic marker detection kit, the kit comprising specific primers for amplifying 21 STR loci, 12 SNP loci linked to the above STR loci, 1 Y-STR locus, Y-Indel and Amel; wherein the 21 STR loci are CSF1PO, D3S1358, D5S818, D7S820, D8S1179, D13S317, D16S539, D18S51, D21S11, FGA, TH01, TPOX, vWA, D10S1248, D2S441, D12S391, D 22S1045, D19S433, D1S1656, D2S1338 and D6S1043; 12 SNP loci are rs17651965, rs17077990, rs57346531, rs9531308, rs11642858, rs13413321, rs58390469, rs16887642, rs4847015, rs25768, rs2246512 and rs6736691; 1 Y-STR locus is DYS391, and 1 Y-InDel locus is rs199815934. The above loci information is shown in Table 1, except for DYS391, Y-Indel and Amel:

[0008] Table 1 SNP-STR composite genetic marker locus information

[0009]

[0010]

[0011] AFR: African, African population; AMR: American, American population; EAS: East Asian, East Asian population; EUR: European, European population; SAS: South Asian, South Asian population.

[0012] Preferably, the SNP-STR composite genetic markers are: rs17651965-CSF1PO, rs17077990-D3S1358, rs25768-D5S818, rs16887642-D7S820, rs57346531-D8S1179, rs9531308-D13S317, rs11642858-D16S539, rs4847015-D1S1656, rs6736691-D2S1338, rs13413321-TPOX, rs2246512-D10S1248 and rs58390469-D2S441.

[0013] Preferably, the specific primer sequences of the kit are: rs11642858-D16S539, SEQ ID NOs: 1-3; rs2246512-D10S1248, SEQ ID NOs: 4-6; rs9531308-D13S317, SEQ ID NOs: 7-9; rs25768-D5S818, SEQ ID NOs: 10-12; rs58390469-D2S441, SEQ ID NOs: 13-15; rs57346531-D8S1179, SEQ ID NOs: 16-18; Y-Indel, SEQ ID NOs: 19-20; Amel, SEQ ID NOs: 21-22; D6S1043, SEQ ID NOs: 23-24; D12S391, SEQ ID NOs: 25-26; D18S51, SEQ ID NOs: NO:27-28; D22S1045, SEQ ID NO:29-30; D19S433, SEQ ID NO:31-32; TH01, SEQ ID NO:33-34; rs4847015-D1S1656, SEQ ID NO:35-37; rs16887642-D7S820, SEQ ID NO:38-40; rs13413321-TPOX, SEQ ID NO:41-43; rs17077990-D3S1358, SEQ ID NO:44-46; rs6736691-D2S1338, SEQ ID NO:47-49; rs17651965-CSF1PO, SEQ ID NO:50-52; FGA, SEQ ID NO:53-54; DYS391, SEQ ID NO:55-56; vWA, SEQ ID NO:57-58; D21S11, SEQ ID NO:59-60.

[0014] Preferably, the specific primers are divided into two groups, wherein rs11642858-D16S539, rs2246512-D10S1248, rs9531308-D13S317, rs25768-D5S818, rs58390469-D2S441, rs57346531-D8S1179, D6S1043, D12S391, D18S51, D22S1045, D19S433 and TH0 1 is detection panel A; rs4847015-D1S1656, rs16887642-D7S820, rs13413321-TPOX, rs17077990-D3S1358, rs6736691-D2S1338, rs17651965-CSF1PO, FGA, DYS391, vWA, and D21S11 are detection panel B; both detection panel A and detection panel B contain Y-Indel and Amel.

[0015] Preferably, five fluorescent dyes are used to label the specific primers of detection panel A and detection panel B respectively, and the 5' end of at least one primer in each pair of primers is labeled; wherein the fluorescent dyes are FAM, HEX, SUM, LYN or PUR, and the internal standard is orange fluorescent SIZ.

[0016] Preferably, the concentrations of the specific primers are as follows: SEQ ID NO:1, 0.28 μM; SEQ ID NO:2, 0.13 μM; SEQ ID NO:3, 0.15 μM; SEQ ID NO:4, 0.25 μM; SEQ ID NO:5, 0.15 μM; SEQ ID NO:6, 0.1 μM; SEQ ID NO:7, 0.3 μM; SEQ ID NO:8, 0.15 μM; SEQ ID NO:9, 0.15 μM; SEQ ID NO:10, 0.6 μM; SEQ ID NO:11, 0.3 μM; SEQ ID NO:12, 0.3 μM; SEQ ID NO:13, 0.4 μM; SEQ ID NO:15, 0.15 μM; SEQ ID NO:16, 0.4 μM; SEQ ID NO:17, 0.2 μM; SEQ ID NO:18, 0.2 μM; SEQ ID NO:19, 0.06 μM; SEQ ID NO:20, 0.06 μM; SEQ ID NO:21, 0.06 μM; SEQ ID NO:22, 0.06 μM; SEQ ID NO:23, 0.84 μM; SEQ ID NO:24, 0.84 μM; SEQ ID NO:25, 0.18 μM; SEQ ID NO:26, 0.18 μM; SEQ ID NO:27, 0.132 μM; SEQ ID NO:28, 0.132 μM; SEQ ID NO:29, 0.18 μM; SEQ ID NO:30, 0.18 μM; SEQ ID NO:31, 0.08 μM; SEQ ID NO:32, 0.08 μM; SEQ ID NO:33, 0.08 μM; SEQ ID NO:34, 0.08 μM; SEQ ID NO:35, 0.26 μM; SEQ ID NO:36, 0.16 μM; SEQ ID NO:37, 0.1 μM; SEQ ID NO:38, 0.3 μM; SEQ ID NO:39, 0.15 μM; SEQ ID NO:40, 0.15 μM; SEQ ID NO:41, 0.25 μM; SEQ ID NO:42, 0.11 μM; SEQ ID NO:43, 0.14 μM; SEQ ID NO:44, 0.2 μM; SEQ ID NO:45, 0.1 μM; SEQ ID NO:46, 0.1 μM; SEQ ID NO:47, 0.35 μM; SEQ ID NO:48, 0.25 μM; SEQ ID NO:49, 0.1 μM; SEQ ID NO:50, 0.4 μM; SEQ ID NO:51, 0.2μM; SEQ ID NO:52, 0.2μM; SEQ ID NO:53, 0.072μM; SEQ ID NO:54, 0.072μM; SEQ ID NO:55, 0.09μM; SEQ ID NO:56, 0.09μM; SEQ ID NO:57, 0.15μM; SEQ ID NO:58, 0.15μM; SEQ ID NO:58, 0.15μM NO: 59, 0.16 μM; SEQ ID NO: 60, 0.16 μM. .

[0017] The sequences and information of the specific primers of the present invention are shown in Table 2:

[0018] Table 2 Specific primer sequence information of SNP-STR composite genetic marker loci

[0019]

[0020]

[0021] Note: In the primer sequences in Table 2, the underlined bases are the bases in the primers that match the SNP, and the lowercase bases are the mismatched bases introduced according to the ASA principle.

[0022] Preferably, the components other than the primers of the kit are reaction buffer, hot start Taq enzyme, ultrapure water, Allelic Ladder and fluorescent molecular internal standard. The reaction buffer is Tris-HCl 10mM, KCl 50mM, MgCl22.0 mM, dNTPs0.2mM, BSA 0.9mg / mL. The fluorescent molecular internal standard is AGCU Marker SIZ-500.

[0023] Preferably, the conditions of the amplification reaction of the kit are: pre-denaturation at 95°C for 5 minutes; denaturation at 94°C for 25 seconds, annealing at 60°C for 40 seconds, extension at 65°C for 60 seconds, 30 cycles; final extension at 60°C for 10 minutes. The specific conditions are shown in Table 3:

[0024] Table 3 Amplification procedures of the kit of the present invention

[0025]

[0026] Application of any of the above-mentioned SNP-STR composite genetic marker detection kits in the identification of human cell line identity and cell cross-contamination in the biomedical field.

[0027] Application of any of the above-mentioned SNP-STR composite genetic marker detection kits in improving the cumulative individual recognition rate of the detection system.

[0028] Application of any of the above-mentioned SNP-STR composite genetic marker detection kits in constructing a cell line genetic locus database.

[0029] The design idea of ​​the kit of the present invention is to use the STR loci of the DNA joint index system specified by the cell line identification standard ANSI Standard (ASN-0002), as well as the STR loci covered by the cell line database of identification institutions such as ATCC, DSMZ, and JCRB. In order to improve the detection efficiency, the STR loci such as D10S1248 and D22S1045 covered by Expanded UScore loci, European recommended loci, etc. are included. The SNPs linked to the above STR loci are screened using the information of the 1000Genomes database to obtain 12 SNPs linked to the above STR loci. According to the ASA technical principle, a composite amplification system is constructed, and the analysis software is designed. The accuracy, sensitivity, mixed sample, and allele fragment mobility accuracy of the detection system are tested. The system is used to investigate and calculate genetic data from non-kinship cell lines, and the cell line identity identification ability and individual recognition ability of the identification system are evaluated.

[0030] The principle of ASA technology is to design two primers at the mutation site of the DNA sequence, whose 3' bases are complementary to the mutation site template and the normal wild-type allele template base, respectively, and design a common primer at the other end of the DNA, and artificially introduce mismatched bases at the 2nd to 5th bases of the 3' end of the primer to improve the specificity of primer extension. This method has been used for point mutation detection of various diseases, and has the advantages of simple operation, cost-effectiveness, high sensitivity, multiple detection sites, and short time consumption. When searching the existing technical documents, it was found that there have been no reports on the use of such methods to identify the authenticity of human cell lines and cross-contamination of cell lines.

[0031] Specifically, in order to ensure that the STR loci of the kit of the present invention are compatible with the current ANSI Standard (ASN-0002) and the cell line DNA database of professional cell line collection and authentication institutions, 13 CODIS core loci (CSF1PO, D3S1358, D5S818, D7S820, D8S1179, D13S317, D16S539, D18S51, D21S11, FGA, TH01, TPOX, and vWA) that meet the requirements of ANSI Standard (ASN-0002) and are compatible with the databases of global professional cell line collection and authentication institutions such as ATCC, DSMZ, and JCRB were screened using bioinformatics technology. In order to improve the detection efficiency of the system, the Expanded US core loci, European recommended loci, and European Standard The STR loci covered by the PowerPlex CS7 test kit include D10S1248, D2S441, D12S391, D22S1045, D19S433, D1S1656, D2S1338, and D6S1043 covered by the PowerPlex CS7 test kit. The 1000Genomes database information was used to screen SNPs linked to STRs. The screening principle was that the minor allele mutation rate of the SNP loci in the global population was higher than 0.1 (MAF>0.10); the target amplicon length was within 500bp. To this end, 12 SNP loci closely linked to the above STR loci were screened and obtained: rs17651965-CSF1PO, rs17077990-D3S1358, rs25768-D5S818, rs16887642-D7S820, rs57346531-D8S1179, rs9531308-D13S317, rs11642858-D16S539, rs4847015-D1S1656, rs6736691-D2S1338, rs13413321-TPOX, rs2246512-D10S1248 and rs58390469-D2S441. The above SNP loci can be detected together with the STR loci. Since human cell lines have a high rate of Y signal loss due to AMEL-Y deletion or mutation, a Y-STR locus DYS391 and a Y-Indel locus rs199815934 were introduced in addition to the AMEL-Y locus to achieve accurate gender identification.

[0032] Beneficial effects: (1) The SNP-STR composite genetic marker detection kit of the present invention covers the DNA joint index system STR loci specified in the human cell line identity identification standard ANSI Standard (ASN-0002) and the STR loci covered by the cell line databases of global professional cell line collection and identification institutions such as ATCC, DSMZ, JCRB, Cellosaurus, etc., so it can be parallel compared with the data in the above databases and has good universality; (2) Combined with ASA-PCR technology, the SNP-STR composite amplification system is established and optimized to achieve synchronous amplification of two genetic markers, improve the typing ability and cumulative individual recognition rate of the detection system, which is of great significance for the identity identification and cross-contamination identification of human cell lines; (3) The kit introduces two sex auxiliary sites, one Y-STR locus DYS391 and one Y-InDel locus, to overcome the Y chromosome signal loss phenomenon caused by AMEL-Y gene deletion or mutation, and improve the accuracy of the sex identification results of human cell lines; (4) The kit is convenient and fast to operate in actual application, low cost, and greatly improves the detection speed and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a gene locus information map of the kit of the present invention;

[0034] Figure 2 is a locus arrangement diagram of the kit of the present invention;

[0035] Figure 3 is the electrophoresis pattern of the loci of the kit of the present invention, wherein A is the result of detection panel A, and B is the result of detection panel B;

[0036] Figure 4 is the accuracy verification result of the locus fragment mobility of the kit of the present invention, wherein A is the result of the detection panel A, and B is the result of the detection panel B;

[0037] Figure 5 is the electrophoresis diagram of the kit of the present invention on the cell line A549, wherein A is the result of the detection panel A, and B is the result of the detection panel B;

[0038] Figure 6 is the sensitivity verification result of the kit of the present invention;

[0039] Figure 7is a heat map of sensitivity verification of the kit of the present invention, wherein the left picture of A is 9948 detection panel A, and the right picture is 9948 detection panel B, the left picture of B is cell line A 549 detection panel A, and the right picture is cell line A 549 detection panel B, the left picture of C is cell line MCF7 detection panel A, and the right picture is the detection results of alleles of each locus of cell line MCF7 detection panel B when the template concentration is 1ng-0.0625ng;

[0040] Figure 8 It is a result diagram of the test kit of the present invention for the mixedness of cell lines A549 and MCF7, wherein A is the result of the test panel A, and B is the result of the test panel B;

[0041] Fig. 9 is the electrophoresis diagram of the kit of the present invention on the cell line HeLa, wherein A is the result of detection panel A, and B is the result of detection panel B;

[0042] Fig.10 It is the electrophoresis diagram of the kit of the present invention on the cell line HL-7702 known to be cross-contaminated by HeLa, wherein A is the result of detection panel A, and B is the result of detection panel B. DETAILED DESCRIPTION

[0043] The following examples further illustrate the content of the present invention, but should not be construed as limiting the present invention. Without departing from the spirit and essence of the present invention, modifications and substitutions made to the method, steps or conditions of the present invention all belong to the scope of the present invention. Unless otherwise specified, the technical means used in the examples are conventional means well known to those skilled in the art.

[0044] Example 1

[0045] like Figure 1As shown, a SNP-STR composite genetic marker detection kit comprises specific primers for amplifying 21 STR loci, 12 SNP loci linked to the above STR loci, 1 Y-STR locus, Y-Indel and Amel; wherein the 21 STR loci are CSF1PO, D3S1358, D5S818, D7S820, D8S1179, D13S317, D16S539, D18S51, D21S11, FGA, TH01, TPOX, vWA, D10S1248, D2S441, D12S391, D2 2S1045, D19S433, D1S1656, D2S1338 and D6S1043; 12 SNP loci are rs17651965, rs17077990, rs57346531, rs9531308, rs11642858, rs13413321, rs58390469, rs16887642, rs4847015, rs25768, rs2246512 and rs6736691; 1 Y-STR locus is DYS391; 1 Y-InDel locus is rs199815934. The above loci information is shown in Table 1, except for DYS391, Y-Indel and Amel:

[0046] Table 1 SNP-STR composite genetic marker locus information

[0047]

[0048] AFR: African, African population; AMR: American, American population; EAS: East Asian, East Asian population; EUR: European, European population; SAS: South Asian, South Asian population.

[0049] The SNP-STR composite genetic markers were: rs17651965-CSF1PO, rs17077990-D3S1358, rs25768-D5S818, rs16887642-D7S820, rs57346531-D8S1179, rs9531308-D13S317, rs11642858-D16S539, rs4847015-D1S1656, rs6736691-D2S1338, rs13413321-TPOX, rs2246512-D10S1248, and rs58390469-D2S441.

[0050] The specific primer sequences of the kit are: rs11642858-D16S539, SEQ ID NOs: 1-3; rs2246512-D10S1248, SEQ ID NOs: 4-6; rs9531308-D13S317, SEQ ID NOs: 7-9; rs25768-D5S818, SEQ ID NOs: 10-12; rs58390469-D2S441, SEQ ID NOs: 13-15; rs57346531-D8S1179, SEQ ID NOs: 16-18; Y-Indel, SEQ ID NOs: 19-20; Amel, SEQ ID NOs: 21-22; D6S1043, SEQ ID NOs: 23-24; D12S391, SEQ ID NOs: 25-26; D18S51, SEQ ID NOs: 2 NO:27-28; D22S1045, SEQ ID NO:29-30; D19S433, SEQ ID NO:31-32; TH01, SEQ ID NO:33-34; rs4847015-D1S1656, SEQ ID NO:35-37; rs16887642-D7S820, SEQ ID NO:38-40; rs13413321-TPOX, SEQ ID NO:41-43; rs17077990-D3S1358, SEQ ID NO:44-46; rs6736691-D2S1338, SEQ ID NO:47-49; rs17651965-CSF1PO, SEQ ID NO:50-52; FGA, SEQ ID NO:53-54; DYS391, SEQ ID NO:55-56; vWA, SEQ ID NO:57-58; D21S11, SEQ ID NO:59-60.

[0051] The specific primers are divided into two groups, among which rs11642858-D16S539, rs2246512-D10S1248, rs9531308-D13S317, rs25768-D5S818, rs58390469-D2S441, rs57346531-D8S1179, D6S1043, D12S391, D18S51, D22S1045, D19S433 and TH01 are Detection panel A; rs4847015-D1S1656, rs16887642-D7S820, rs13413321-TPOX, rs17077990-D3S1358, rs6736691-D2S1338, rs17651965-CSF1PO, FGA, DYS391, vWA, and D21S11 are detection panel B; both detection panel A and detection panel B contain Y-Indel and Amel.

[0052] like Figure 2 As shown, five fluorescent dyes are used to label the specific primers of detection panel A and detection panel B respectively, and the 5' end of at least one primer in each pair of primers is labeled; wherein, the fluorescent dyes are FAM, HEX, SUM, LYN or PUR, and the internal standard is orange fluorescent SIZ.

[0053] The concentrations of the specific primers are as follows: SEQ ID NO:1, 0.28 μM; SEQ ID NO:2, 0.13 μM; SEQ ID NO:3, 0.15 μM; SEQ ID NO:4, 0.25 μM; SEQ ID NO:5, 0.15 μM; SEQ ID NO:6, 0.1 μM; SEQ ID NO:7, 0.3 μM; SEQ ID NO:8, 0.15 μM; SEQ ID NO:9, 0.15 μM; SEQ ID NO:10, 0.6 μM; SEQ ID NO:11, 0.3 μM; SEQ ID NO:12, 0.3 μM; SEQ ID NO:13, 0.4 μM; SEQ ID NO:15, 0.15 μM; SEQ ID NO:16, 0.4 μM; SEQ ID NO:17, 0.2 μM; SEQ ID NO:18, 0.2 μM; SEQ ID NO:19, 0.06 μM; SEQ ID NO:20, 0.06 μM; SEQ ID NO:21, 0.06 μM; SEQ ID NO:22, 0.06 μM; SEQ ID NO:23, 0.84 μM; SEQ ID NO:24, 0.84 μM; SEQ ID NO:25, 0.18 μM; SEQ ID NO:26, 0.18 μM; SEQ ID NO:27, 0.132 μM; SEQ ID NO:28, 0.132 μM; SEQ ID NO:29, 0.18 μM; SEQ ID NO:30, 0.18 μM; SEQ ID NO:31, 0.08 μM; SEQ ID NO:32, 0.08 μM; SEQ ID NO:33, 0.08 μM; SEQ ID NO:34, 0.08 μM; SEQ ID NO:35, 0.26 μM; SEQ ID NO:36, 0.16 μM; SEQ ID NO:37, 0.1 μM; SEQ ID NO:38, 0.3 μM; SEQ ID NO:39, 0.15 μM; SEQ ID NO:40, 0.15 μM; SEQ ID NO:41, 0.25 μM; SEQ ID NO:42, 0.11 μM; SEQ ID NO:43, 0.14 μM; SEQ ID NO:44, 0.2 μM; SEQ ID NO:45, 0.1 μM; SEQ ID NO:46, 0.1 μM; SEQ ID NO:47, 0.35 μM; SEQ ID NO:48, 0.25 μM; SEQ ID NO:49, 0.1 μM; SEQ ID NO:50, 0.4 μM; SEQ ID NO:51, 0.2 μM; SEQ ID NO:52, 0.2μM; SEQ ID NO:53, 0.072μM; SEQ ID NO:54, 0.072μM; SEQ ID NO:55, 0.09μM; SEQ ID NO:56, 0.09μM; SEQ ID NO:57, 0.15μM; SEQ ID NO:58, 0.15μM; SEQ ID NO:59, 0.16μM; SEQ ID NO: 60, 0.16μM. .

[0054] The sequences and information of the specific primers of the present invention are shown in Table 2:

[0055] Table 2 Specific primer sequence information of SNP-STR composite genetic marker loci

[0056]

[0057]

[0058] Note: In the primer sequences in Table 2, the underlined bases are the bases in the primers that match the SNP, and the lowercase bases are the mismatched bases introduced according to the ASA principle.

[0059] The components other than the primers in the kit are reaction buffer, hot start Taq enzyme, ultrapure water, AllelicLadder and fluorescent molecular internal standard. The reaction buffer is Tris-HCl 10mM, KCl 50mM, MgCl22.0 mM, dNTPs 0.2mM, BSA 0.9mg / mL. The fluorescent molecular internal standard is AGCU Marker SIZ-500.

[0060] The conditions of the amplification reaction of the kit are: pre-denaturation at 95°C for 5 minutes; denaturation at 94°C for 25 seconds, annealing at 60°C for 40 seconds, extension at 65°C for 60 seconds, 30 cycles; final extension at 60°C for 10 minutes. The specific conditions are shown in Table 3:

[0061] Table 3 Amplification procedures of the kit of the present invention

[0062]

[0063] Example 2 Accuracy Verification

[0064] Taking cell lines A549, MCF-7 and control DNA 9948 as examples, the accuracy of the kit described in Example 1 was verified, and the amplification procedure is shown in Table 3. The typing test results of cell lines A549, MCF-7, and 9948 are shown in Table 4. The accuracy of the SNP typing results of A549 and MCF-7 was verified by sequencing, and the SNP genotyping of the control DNA sample 9948 was obtained by bioinformatics analysis. Other commercially available kits were used to verify the accuracy of the STR typing results.

[0065] Table 4 SNP-STR typing results accuracy verification

[0066]

[0067] SNP-STR genotyping was obtained by allele-specific amplification polymerase chain reaction (ASA-PCR)-based technology, while individual STR genotyping was performed using commercial test kits. Single nucleotide polymorphisms (SNPs) of A549 and MCF-7 cell lines were identified by sequencing, and SNP genotyping of control DNA sample 9948 was obtained by bioinformatics analysis.

[0068] Example 3 Study on the Accuracy of Allele Fragment Mobility

[0069] The electrophoresis pattern of alleles of each locus of the kit described in Example 1 is as follows Figure 3 As shown, further research on the accuracy of allele gene fragment mobility was carried out, such as Figure 4 As shown, except for D21S11 (0.1560 bp), the standard deviation of the mobility of allele fragments at each locus ranged from 0.0553 to 0.1276 bp (<0.15 bp).

[0070] Example 4 Sensitivity study

[0071] The sensitivity of the composite amplification system of the kit described in Example 1 was studied. The results showed that when the amount of the standard 9948 and the cell line A549 template was 0.0625 and above, each detection site could obtain 100% accurate typing. When the amount of the cell line MCF-7 template was as low as 0.125ng and 0.0625, 98.89% and 94.44% of the sites could obtain accurate typing (see Figure 5 and Figure 6 ).

[0072] Example 5 Application in mixed samples

[0073] The composite amplification system of the kit described in Example 1 was used for mixed sample detection, and two mixing systems were adopted: 1) Two control DNA samples, 9948 and 9947A, were selected and gradient mixed according to the mass ratio (ng), with the ratios of 1:1 (0.5ng+0.5ng), 1:3 (0.25ng+0.75ng) and 1:9 (0.1ng+0.9ng), respectively, and the total template amount was fixed at 1ng; 2) Genomic DNA extracted from cell lines A549 and MCF-7 were mixed at 1:1, 1:3 and 1:9, and the total template amount was 1ng.

[0074] like Figure 8 The results showed that all alleles of low-proportion components in all mixed systems could be accurately detected (RFU>50) and were not affected by the main components.

[0075] Example 6

[0076] The DNA samples of cell lines A549 and MCF-7 and the control DNA sample 9948 were subjected to typing detection using the composite amplification system of the kit described in Example 1. The electrophoresis patterns are shown in Figure 7 And Table 4. The SNP-STR typing results are accurate. We used a commercial detection kit to detect STR genotyping alone, and the results were consistent with the results of the present invention. We used sequencing to verify the single nucleotide polymorphism (SNP) typing of A549 and MCF-7, and the results were consistent with the results of the present invention. The SNP genotyping of DNA sample 9948 was obtained by bioinformatics analysis, and the typing results were consistent with the present invention.

[0077] Example 7

[0078] The kit described in Example 1 was used to conduct genetic testing on 96 cell line samples from 75 unrelated patients, and the polymorphism data such as the haplotype distribution and frequency of the SNP-STR composite detection system of the present invention in the cell line were obtained to identify the true identity of the cell line and detect the cell line with the wrong identity. The identity identification ability and cumulative individual recognition ability of the established SNP-STR detection system for the cell line were evaluated (see Tables 5-10 for details).

[0079] The SNP-STR detection system based on ASA-PCR technology of the present invention was used to simultaneously perform typing detection of two genetic markers, SNP and STR, on 96 cell line samples derived from 75 unrelated patient individuals. In parallel, a commercial STR kit was used to compare the accuracy of STR typing. The results showed that the SNP and STR genetic loci detected by the SNP-STR typing detection system were accurate, the mutation rate of the detected SNP typing loci was low, and the genetic stability was high. For tumor cell lines with differences in STR typing, they can be further distinguished based on the SNP typing detected simultaneously by the ASA-PCR detection system, thereby improving the typing ability. rs11642858-D16S539,rs2246512-D10S1248,rs9531308-D13S317,rs25768-D5S8 18,rs58390469-D2S441,rs57346531-D8S1179,rs4847015-D1S1656,rs16887642- The number of SNP-STR alleles detected at D7S820, rs13413321-TPOX, rs17077990-D3S1358, rs6736691-D2S1338 and rs17651965-CSF1PO sites were 10, 12, 13, 12, 16, 14, 16, 19, 9, 12, 15, 14. The number of corresponding STR loci alleles was only 7, 8, 8, 7, 11, 10, 15, 15, 7, 8, 12, 9. Tables 5 and 6 show the above SNP-STR allele frequencies and statistical calculation results, respectively, and Tables 7 and 8 show the corresponding STR allele frequencies and statistical calculation results. The numbers of alleles detected in the other nine STR loci vWA, TH01, FGA, D22S1045, D21S11, D19S433, D18S51, D12S391, and D6S1043 in the system are 8, 5, 10, 7, 13, 14, 13, 14, and 14, respectively (see Table 9 for details); the allele frequencies and statistical calculation results of the SNP loci in the system are shown in Table 10 for details. According to the calculation of the above 75 unrelated cell lines, the cumulative individual identification rate of the SNP-STR typing system based on ASA-PCR technology is 0.99999999999999999999999999999887308399190763, while the corresponding cumulative individual identification rate relying on STR typing is only 0.9999999999999999999999999995594166439684952. The above results show that the SNP-STR composite genetic marker detection based on ASA-PCR has a higher identity identification ability for human cell lines compared with STR detection.

[0080] Table 5 rs11642858-D16S539, rs2246512-D10S1248, rs9531308-D13S317, rs25768-D5S818, rs58390469-D2S441,

[0081] Haplotype frequencies and statistical parameters of the rs57346531-D8S1179 locus

[0082]

[0083]

[0084] PD: Power of Discrimination, individual recognition ability; PIC: polymorphism information content,

[0085] Polymorphism information content; He: expected heterozygosity, polymorphism information content; PM: matching probability,

[0086] Random matching probability.

[0087] Table 6 rs4847015-D1S1656,

[0088] Haplotype frequencies and statistical parameters of the rs16887642-D7S820, rs13413321-TPOX, rs17077990-D3S1358, rs6736691-D2S1338, and rs17651965-CSF1PO loci

[0089]

[0090]

[0091] PD: Power of Discrimination, individual recognition ability; PIC: polymorphism information content,

[0092] Polymorphism information content; He: expected heterozygosity, polymorphism information content; PM: matching probability,

[0093] Random matching probability.

[0094] Table 7 Haplotype frequencies and statistical parameters of STR loci such as D16S539, D10S1248, D13S317, D5S818, D2S441, and D8S1179 corresponding to the SNP-STR typing detection system in cell lines derived from 75 unrelated patients

[0095]

[0096] PD: Power of Discrimination, individual recognition ability; PIC: polymorphism information

[0097] content, polymorphism information content; He: expected heterozygosity, polymorphism information content; PM:

[0098] matching Probability, random matching probability.

[0099] Table 8 Haplotype frequencies and statistical parameters of STR loci such as D1S1656, D7S820, TPOX, D3S1358, D2S1338, CSF1PO corresponding to the SNP-STR typing detection system in cell lines derived from 75 unrelated patients

[0100]

[0101]

[0102] PD: Power of Discrimination, individual recognition ability; PIC: polymorphism information

[0103] content, polymorphism information content; He: expected heterozygosity, polymorphism information content; PM:

[0104] matching Probability, random matching probability.

[0105] Table 9 D6S1043, ...

[0106] Single STR loci such as D12S391, D18S51, D22S1045, D19S433, TH01, FGA, vWA, D21S11

[0107] Ploid frequencies and statistical parameters

[0108]

[0109] PD: Power of Discrimination, individual recognition ability; PIC: polymorphism information content, polymorphism information content; He: expected heterozygosity, polymorphism information content; PM: matching Probability, random matching probability.

[0110] Table 10 Haplotype frequencies and statistical parameters of SNP loci in the SNP-STR typing detection system in cell lines derived from 75 unrelated patients

[0111]

[0112] PD: Power of Discrimination, individual recognition ability; PIC: polymorphism information content, polymorphism information content; He: expected heterozygosity, polymorphism information content; PM: matching Probability, random matching probability.

[0113] Example 8

[0114] The kit described in Example 1 was used to perform genotyping on the high microsatellite instability (MSI-H) cell line HeLa and the HL-7702 cell line known to be contaminated by the HeLa cell line. The results are as follows: Fig. 9 , Fig.10As shown. Since the HeLa cell line has the characteristics of MSI-H, its short tandem repeat sequence (STR) is prone to accumulation of somatic mutations during continuous passage, leading to genetic drift. Although the HL-7702 cell line has been reported many times in the literature as a HeLa cell contamination strain, the STR genetic markers and commercial STR detection kits described in the present invention both show that the STR typing of the two still shows differences at 5 STR loci, including D10S1248, D6S1043, D22S1045, CSF1PO and FGA (the matching rate is only 94.87%). The SNP typing data obtained simultaneously by the kit described in Example 1 of the present invention show that the two cell lines maintain completely consistent genotypes at all detection sites (the matching rate reaches 100%). Therefore, it is confirmed that when the STR typing of the MSI-H cell line differs due to genetic instability, its SNP typing can still be used as a reliable basis for cell line identity / cross-contamination identification; that is, the use of the kit described in the present invention achieves simultaneous amplification of two genetic markers, improves the typing ability and cumulative individual recognition rate of the detection system, and improves the accuracy of the identity identification and cross-contamination identification results of the human cell line.

Claims

1. A SNP-STR composite genetic marker detection kit, characterized in that: The kit comprises specific primers for amplifying 21 STR loci, 12 SNP loci linked to the above STR loci, 1 Y-STR locus, Y-Indel and Amel; wherein the 21 STR loci are CSF1PO, D3S1358, D5S818, D7S820, D8S1179, D13S317, D16S539, D18S51, D21S11, FGA, TH01, TPOX, vWA, D10S1248, D2S441, D12S391, D22S1045, D19S433, D1S1656, D2S1338 and D6S1043; 12 SNP loci were rs17651965, rs17077990, rs57346531, rs9531308, rs11642858, rs13413321, rs58390469, rs16887642, rs4847015, rs25768, rs2246512 and rs6736691; 1 Y-STR locus was DYS391; One Y-InDel locus is rs199815934.

2. The SNP-STR composite genetic marker detection kit according to claim 1, characterized in that: The SNP-STR composite genetic markers are: rs17651965-CSF1PO, rs17077990-D3S1358, rs25768-D5S818, rs16887642-D7S820, rs57346531-D8S1179, rs9531308-D13S317, rs11642858-D16S539, rs4847015-D1S1656, rs6736691-D2S1338, rs13413321-TPOX, rs2246512-D10S1248, and rs58390469-D2S441, and both were detected simultaneously.

3. The SNP-STR composite genetic marker detection kit according to claim 1, characterized in that: The specific primer sequences of the kit are: rs11642858-D16S539, SEQ ID NO: 1-3; rs2246512-D10S1248, SEQ ID NO: 4-6; rs9531308-D13S317, SEQ ID NO: 7-9; rs25768-D5S818, SEQ ID NO: 10-12; rs58390469-D2S441, SEQ ID NO:13-15; rs57346531-D8S1179, SEQ ID NO:16-18; Y-Indel, SEQ ID NO:19-20; Amel, SEQ ID NO:21-22; D6S1043, SEQ ID NO:23-24; D12S391, SEQ ID NO:25-26; D18S51, SEQ ID NO:27-28; D22S1045, SEQ ID NO:29-30; D19S433, SEQ ID NO:31-32; TH01, SEQ ID NO:33-34; rs4847015-D1S1656, SEQ ID NO:35-37; rs16887642-D7S820, SEQ ID NO:38-40; rs13413321-TPOX, SEQ ID NO:41-43; rs17077990-D3S1358, SEQ ID NO:44-46; rs6736691-D2S1338, SEQ ID NO:47-49; rs17651965-CSF1PO, SEQ ID NO:50-52; FGA, SEQ ID NO:53-54; DYS391, SEQ ID NO:55-56; vWA, SEQ ID NO:57-58; D21S11, SEQ ID NO:59-60.

4. The SNP-STR composite genetic marker detection kit according to claim 3, characterized in that: The specific primers are divided into two groups, among which rs11642858-D16S539, rs2246512-D10S1248, rs9531308-D13S317, rs25768-D5S818, rs58390469-D2S441, rs57346531-D8S1179, D6S1043, D12S391, D18S51, D22S1045, D19S433 and TH01 are Detection panel A; rs4847015-D1S1656, rs16887642-D7S820, rs13413321-TPOX, rs17077990-D3S1358, rs6736691-D2S1338, rs17651965-CSF1PO, FGA, DYS391, vWA, and D21S11 are detection panel B; both detection panel A and detection panel B contain Y-Indel and Amel.

5. The SNP-STR composite genetic marker detection kit according to claim 4, characterized in that: Five fluorescent dyes are used to label the specific primers of detection panel A and detection panel B respectively, and the 5' end of at least one primer in each pair of primers is labeled; wherein the fluorescent dyes are FAM, HEX, SUM, LYN or PUR, and the orange fluorescent SIZ is selected as the internal standard.

6. The SNP-STR composite genetic marker detection kit according to claim 3, characterized in that: The concentrations of the specific primers are as follows: SEQ ID NO:1, 0.28 μM; SEQ ID NO:2, 0.13 μM; SEQ ID NO:3, 0.15 μM; SEQ ID NO:4, 0.25 μM; SEQ ID NO:5, 0.15 μM; SEQ ID NO:6, 0.1 μM; SEQ ID NO:7, 0.3 μM; SEQ ID NO:8, 0.15 μM; SEQ ID NO:9, 0.15 μM; SEQ ID NO:10, 0.6 μM; SEQ ID NO:11, 0.3 μM; SEQ ID NO:12, 0.3 μM; SEQ ID NO:13, 0.4 μM; SEQ ID NO:15, 0.15 μM; SEQ ID NO:16, 0.4 μM; SEQ ID NO:17, 0.2 μM; SEQ ID NO:18, 0.2 μM; SEQ ID NO:19, 0.06 μM; SEQ ID NO:20, 0.06 μM; SEQ ID NO:21, 0.06 μM; SEQ ID NO:22, 0.06 μM; SEQ ID NO:23, 0.84 μM; SEQ ID NO:24, 0.84 μM; SEQ ID NO:25, 0.18 μM; SEQ ID NO:26, 0.18 μM; SEQ ID NO:27, 0.132 μM; SEQ ID NO:28, 0.132 μM; SEQ ID NO:29, 0.18 μM; SEQ ID NO:30, 0.18 μM; SEQ ID NO:31, 0.08 μM; SEQ ID NO:32, 0.08 μM; SEQ ID NO:33, 0.08 μM; SEQ ID NO:34, 0.08 μM; SEQ ID NO:35, 0.26 μM; SEQ ID NO:36, 0.16 μM; SEQ ID NO:37, 0.1 μM; SEQ ID NO:38, 0.3 μM; SEQ ID NO:39, 0.15 μM; SEQ ID NO:40, 0.15 μM; SEQ ID NO:41, 0.25 μM; SEQ ID NO:42, 0.11 μM; SEQ ID NO:43, 0.14 μM; SEQ ID NO: 44, 0.2 μM; SEQ ID NO: 45, 0.1 μM; SEQ ID NO: 46, 0.1 μM; SEQ ID NO: 47, 0.35 μM; SEQ ID NO: 48, 0.25 μM; SEQ ID NO: 49, 0.1 μM; SEQ ID NO: 50, 0.4 μM; SEQ ID NO: 51, 0.2 μM; NO: 52, 0.2 μM; SEQ ID NO: 53, 0.072 μM; SEQ ID NO: 54, 0.072 μM; SEQ ID NO: 55, 0.09 μM; SEQ ID NO: 56, 0.09 μM; SEQ ID NO: 57, 0.15 μM; SEQ ID NO: 58, 0.15 μM; SEQ ID NO: 59, 0.16 μM; SEQ ID NO: 60, 0.16 μM.

7. The SNP-STR composite genetic marker detection kit according to claim 1, characterized in that: The conditions of the amplification reaction of the kit are: pre-denaturation at 95°C for 5 minutes; denaturation at 94°C for 25 seconds, annealing at 60°C for 40 seconds, extension at 65°C for 60 seconds, 30 cycles; final extension at 60°C for 10 minutes.

8. Use of the SNP-STR composite genetic marker detection kit according to any one of claims 1 to 7 in the identification of human cell line identity and cell cross-contamination in the field of biomedicine.

9. Use of the SNP-STR composite genetic marker detection kit according to any one of claims 1 to 7 in improving the cumulative individual recognition rate of a detection system.

10. Use of the SNP-STR composite genetic marker detection kit according to any one of claims 1 to 7 in constructing a cell line genetic locus database.

Citation Information

Patent Citations

  • Combined amplification system with 21 short tandem repeats and kit

    CN103820559A

  • Mitochondria SNP (single nucleotide polymorphism) fluorescence-labeling multiple amplification kit and application thereof

    CN103898226A