A method for simultaneously identifying species and genetic relationship of experimental monkeys based on NGS technology
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF ZOOLOGY GUANGDONG ACAD OF SCI
- Filing Date
- 2026-07-03
- Publication Date
- 2026-08-04
AI Technical Summary
[0010]针对现有毛细管电泳STR分型技术亲缘关系鉴定引物分辨能力有限、不兼容物种同步鉴定、STR亲缘关系鉴定分析质量控制难、自动化程度低、检测通量偏低等缺陷,本发明提供一种基于NGS技术同时实现实验猴物种与亲缘关系鉴定的方法,该方法将SNP物种鉴定与STR基因分型整合,具有一测多用,检测通量高的优势
基于自主筛选的物种特异性SNP鉴定位点,建立标准化物种鉴别方法。通过检测该位点基因型,可精准判定样本物种:该位点基因型检测出C/C判定为纯合食蟹猴,该位点检测出T/T判定为纯合猕猴,基因型C/T判定为杂交个体。该技术解决了毛细管电泳STR分型技术难以准确甄别杂交个体的短板,实现物种鉴定的标准化、精准化,为后续亲缘关系鉴定提供了准确的物种信息来源。
Smart Images

Figure SMS_2 
Figure SMS_3 
Figure SMS_4
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomolecular detection, genotyping, and non-human primate genetic identification technology, specifically involving a method for simultaneously identifying the species and kinship of experimental monkeys based on NGS technology. Background Technology
[0002] Crab-eating macaques and rhesus macaques are non-human primates (NHPs) used in biomedical research. They are widely used in core areas such as new drug development, human disease model construction, drug toxicology and safety evaluation, and basic life science research. They are important experimental animals that bridge basic research and clinical applications.
[0003] In the large-scale breeding system of experimental monkeys, kinship identification is the core foundation. By clarifying the kinship of individuals, problems such as population degradation and mixed genetic backgrounds caused by inbreeding can be effectively avoided, ensuring the genetic stability and germplasm purity of the experimental monkey population. At the same time, it can complete the precise genetic tracing and pedigree establishment of individuals, and realize precise breeding management.
[0004] At the same time, accurate species identification is equally indispensable. Both cynomolgus monkeys and macaques belong to the genus *Macaques*, and interspecific hybrids are easily produced during breeding. These hybrids have mixed genetic backgrounds and cannot be used for preclinical research. Using hybrid monkeys in experiments would directly lead to distorted animal experimental results and invalid experimental data, seriously affecting the reliability of new drug development and scientific research achievements. Therefore, rapidly and accurately distinguishing between cynomolgus monkeys, macaques, and their hybrids is a necessary prerequisite for germplasm screening, population purification, and quality assurance of laboratory animals.
[0005] Currently, short tandem repeat (STR) kinship identification technology is well-developed and widely used in human populations, with a complete testing system and judgment criteria. However, reliable genetic identification technology systems for cynomolgus monkeys and rhesus monkeys are extremely lacking, and there is no unified standardized identification technology in the industry. Existing detection methods mainly rely on traditional capillary electrophoresis STR typing technology, and the accuracy and stability of genetic identification results need to be improved. The capillary electrophoresis STR typing technology has the following drawbacks: 1. Limited STR loci in capillary electrophoresis STR typing technology leads to limited resolution: Capillary electrophoresis STR typing technology has a limited number of STR loci and limited locus polymorphism. Since CPI is obtained by multiplying the paternity index of each locus, the fewer the STR loci, the weaker the ability to distinguish kinship, resulting in poor resolution. 2. STR typing based on capillary electrophoresis cannot be used for species identification: This technique can only identify STR relationships and cannot simultaneously detect SNPs for species identification. To perform SNP detection for species identification, another technique (Sanger sequencing, qPCR, or NGS) must be used.
[0006] 3. Quality control of kinship identification using capillary electrophoresis STR typing technology is difficult: 1) Capillary electrophoresis STR typing technology relies solely on fragment length differences for typing, which cannot distinguish sequence-level polymorphisms. The typing resolution is insufficient, the effective genetic information is limited, and typing confusion and misjudgment are prone to occur; 2) The peak diagram of capillary electrophoresis STR typing technology is easily affected by artificial artifacts such as stutter peaks (PCR slippage), pull-up peaks (spectral crosstalk), and non-specific peaks; 3) Capillary electrophoresis STR typing technology cannot achieve quantitative evaluation of STR typing confidence Q value, and cannot automatically screen out low-quality sequencing STR typing. As a result, STR defect typing directly participates in CPI calculation and STR typing site matching statistics, thus the rigor of kinship determination is insufficient and the risk of misjudgment is high; 4) The peak typing of capillary electrophoresis STR typing technology depends on ladder calibration, which is easily affected by batch differences.
[0007] 4. The capillary electrophoresis STR typing technology lacks an automated analysis process, has a high degree of reliance on manual labor, and is prone to human error: The current capillary electrophoresis STR typing technology lacks a fully automated reporting and analysis system. It requires manual interpretation of peak diagrams, manual sorting of genotypes, manual counting of matching loci, and manual calculation of the parentage index CPI. The operation is cumbersome, highly subjective, and prone to human error, and it is difficult to achieve large-scale batch testing.
[0008] 5. Low throughput of capillary electrophoresis STR typing technology: Each STR locus in each sample of capillary electrophoresis STR typing technology requires a sequencing reaction, resulting in a limited number of samples that can be detected in parallel. The throughput is a significant bottleneck, making it difficult to meet the sample volume requirements for large-scale animal population tracing and precise breeding management.
[0009] In summary, existing capillary electrophoresis STR genotyping technology cannot simultaneously achieve high-precision species identification, high-reliability kinship typing, and automated analysis. The industry urgently needs a method with higher resolution, higher intelligence, and scalability for genetic identification of experimental monkey species and kinship. Summary of the Invention
[0010] To address the shortcomings of existing capillary electrophoresis STR genotyping techniques, such as limited primer resolution for kinship identification, incompatible species simultaneous identification, difficulty in quality control of STR kinship identification analysis, low automation, and low throughput, this invention provides a method based on NGS technology to simultaneously identify the species and kinship of experimental monkeys. This method integrates SNP species identification with STR genotyping, offering advantages such as multiple uses in one test and high throughput.
[0011] The present invention aims to: 1. Develop a primer panel based on NGS technology to simultaneously identify SNP species and STR kinship in experimental monkeys, integrating SNP species identification with STR genotyping for multiple uses in one test; 2. Establish a species identification method for experimental monkeys based on NGS technology using SNPs, which can distinguish between cynomolgus monkeys, macaques, and their hybrids; 3. Establish NGS technology for STR kinship identification based on experimental monkey species, and introduce Q-value filtering into kinship analysis based on STR genotyping to improve the accuracy and reliability of kinship determination; 4. We have independently developed a fully intelligent bioinformatics analysis and report generation system to replace manual interpretation of results, reduce human error, and improve work efficiency.
[0012] 5. Leveraging the high-throughput advantage of NGS technology, it supports rapid detection of large-scale samples.
[0013] The first objective of this invention is to provide a set of primer panels for SNP species identification and STR phylogenetic identification in experimental monkeys based on NGS technology, comprising the following primer pairs: F: CAAGCATGGAGATGGTCTGGTT (SEQ ID NO.1), R: GTACATGCCTCTTTGTTGCAGTG (SEQ ID NO.2); F: GCACTGCTAAGGCTTCTATCACA (SEQ ID NO.3), R: GGTAGTGACATGTGCTCACTGT (SEQ ID NO.4); F: TGAGCCTCAGAATTACCCCAGT (SEQ ID NO.5), R: TCACTTGAACCTGTGAGGCG (SEQ ID NO.6); F: CAGTTACGAGGAGGGTTGACATC (SEQ ID NO.7), R: TGAGACAGTGGCATAAACCAGG (SEQ ID NO.8); F:GTGATGGAAAAGAATCGGGACAG(SEQ ID NO.9),R:CATCAACATCACCCCAACACCT(SEQID NO.10); F:CCCTGGTTCTGAGGTTTTTGGA(SEQ ID NO.11),R:TTCGGGTTCTCCAAAGAGACAG(SEQID NO.12); F:CACCTGTCTCAATCCAAGACAAATC(SEQ ID NO.13),R:ACAGGCTATCTATCTATCTATTTATTTATCATCT(SEQ ID NO.14); F:CCATCACTTACTGGCAATGTAACC(SEQ ID NO.15),R:TGCTGGAAACTGATAAGGGCTTTA(SEQ ID NO.16); F:GGATCATGAAAGGGCATGAGGA(SEQ ID NO.17),R:ACTCCCTTCTTCCCTCTCACAG(SEQID NO.18); F:ACTTGGAAAGTATGCTGCCTCTG(SEQ ID NO.19),R:GGATCACTTGAACCTGGGAGATG(SEQID NO.20); F:AACAAAGGAGGCAGTGAGCATC(SEQ ID NO.21),R:AGTGAGCTGAGATCACGTCACT(SEQID NO.22); F:CCCAGGAGTTCCAATTTCTCCA(SEQ ID NO.23),R:TCTGATAAGGGCTTGATATCCAGG(SEQID NO.24); F:GCCAAGGATGGTGAGTTACTCA(SEQ ID NO.25),R:TGGTAGTGATGTGGCCCTAAGT(SEQID NO.26); F:CAAGTTCTAACATCACGTCCCTCT(SEQ ID NO.27),R:GTACCTGAGGTCATCAGGACATTC(SEQ ID NO.28); F: CTCTTCCACTGATTCTGCCCAT (SEQ ID NO.29), R: GTGACACAGAAACAGTCTGGGA (SEQ ID NO.30); F: GTCCATAGTGGTGCTTCTCCAT (SEQ ID NO. 31), R: GCATCTGTGTGGATTTGGGGTA (SEQ ID NO. 32); F: CCCAAGTGGGTCCAAGTGGCT (SEQ ID NO. 33), R: GGATAGGGCAACAGAGAAGAT (SEQ ID NO. 34); F: AGCCCAGATATCCCCAAGATCTC (SEQ ID NO. 35), R: CAAGTGATGGCCCAAATTTGGC (SEQ ID NO. 36); F: CCCCTATTCTTAGGAAATAAACCCTGA (SEQ ID NO. 37), R: CACAGTTGAAATCCTCTACCCAGAT (SEQ ID NO. 38); F: AGGATTCTCCAGGCAAATAGAACC (SEQ ID NO. 39), R: GACATACACCATTGGCTCCCAT (SEQ ID NO. 40).
[0014] The second objective of this invention is to provide a kit for identifying experimental monkey species using SNPs (species identification between cynomolgus monkeys and macaques) and STRs (phylogenetic relationships between cynomolgus monkeys and macaques), which contains the aforementioned primer Panel.
[0015] The third objective of this invention is to provide a method for simultaneously identifying the species and phylogenetic relationships of experimental monkeys based on NGS technology, comprising: (1) Collect samples from experimental monkeys and extract genomic DNA. The experimental monkeys are cynomolgus monkeys, macaques, or their hybrids. (2) Using the aforementioned primer Panel as the amplification primer and genomic DNA as the template, the first round of multiplex PCR targeted amplification was performed. The amplification product was purified by magnetic beads to remove non-specific amplification fragments and primer dimers. Then, the sequencing adapter and barcode tag were connected by the second round of multiplex PCR to construct a sequencing library. Finally, the original FASTQ sequencing data were obtained by high-throughput sequencing. (3) Perform quality control trimming and short fragment filtering on the original FASTQ sequencing data, compare it with the monkey reference genome using BWA-MEM, and process it with samtools to obtain a high-quality BAM file; (4) Then use SNP to perform species-specific identification analysis to determine the base at position 56049049 of chr11. The cynomolgus monkey is C and the macaque is T. If it is a hybrid species, the position is CT. (5) The kinship is then determined by comparing the STR genotype with the high-confidence filtering site; (6) Calculate the unit point paternity index PI using Python script and multiply it to obtain the cumulative paternity index CPI. Convert the relative paternity probability RCP using the Essen-Möller formula and classify the kinship determination criteria according to the CPI and RCP thresholds. (7) Customize Python scripts to integrate all test data and automatically generate standardized kinship test reports in batches.
[0016] Preferably, the two-round multiplex PCR reaction system in step (2) is: The first round of multiplex PCR reaction system consisted of: 13 μL of Nuclease-Free Water, 5 μL of primer panel with a final concentration of 100 nM, 40 ng of genomic DNA per 2 μL of reaction, and 10 μL of Amplicon Mix. First-round multiplex PCR amplification program: Hot-lid temperature: 105℃; Pre-denaturation: 95℃, 3 min 30 s, 1 cycle; Denaturation-extension cycle: 98℃, 20 s; 60℃, 1 min; 65℃, 1 min; Number of cycles: 24 cycles; Final extension: 72℃, 5 min, 1 cycle; Incubation: 4℃, constant temperature storage; The second round of multiplex PCR reaction system consisted of: 10 μL of PCR product purified in the first round, 2 μL of 10 μM Indexed Primer (stock solution concentration 10 μM, containing sequencing adapters and barcodes), 15 μL of PCR Master Mix, and 3 μL of Nuclease-Free Water. Second-round multiplex PCR amplification program: Heat cap temperature: 105℃; Pre-denaturation: 95℃, 3 min 30 s, 1 cycle; Amplification cycles: 98℃, 20 s; 58℃, 1 min; 72℃, 30 s; Number of cycles: 9 cycles; Final extension: 72℃, 5 min, 1 cycle; Incubation: 4℃, constant temperature storage.
[0017] Preferably, step (4) is: using bcftools to determine the base variation at the chr11:56049049 site, the cynomolgus monkey is C, the macaque is T, and if it is a hybrid species, the site is CT coexisting.
[0018] Preferably, step (5) is as follows: the target STR region is genotyped using GangSTR software to obtain the allele fragment length, genotype, and genotyping confidence Q value for each locus. Before analysis, only high-confidence loci with Q=1 are retained, and low-quality loci with Q<1, ambiguous genotyping, and high background noise are automatically removed to construct a reliable STR genotype matrix. According to Mendel's laws of inheritance, the STR loci of offspring and parents are compared: if any allele of the offspring overlaps with the homologous locus fragment of the parent, it is determined to be a genetically matched locus; if both alleles of the offspring do not overlap with the homologous locus fragments of the parents, it is determined to be a genetically mismatched locus; if more than 3 valid mismatched loci are detected, no kinship is determined; if less than or equal to 3 mismatched loci are considered as natural gene mutations, kinship is determined, including parent-child and kinship relationships.
[0019] Preferably, step (6) is: using a customized Python script to calculate the comprehensive paternal index (CPI) for all paired samples. For each STR locus, the formula for calculating the paternal index PI is: when homozygous-homozygous matches, PI = 1 / p; when heterozygous matches, PI = 1 / (2p), where p is the population allele frequency obtained from the pre-calculated frequency database. When there is no shared allele, PI is set to the mutation rate (μ = 0.001). The cumulative paternity index (CPI) is the product of the PI values when Q = 1 for all high-confidence loci. Then, the relative paternal probability (RCP) is calculated based on the Essen-Möller formula: RCP = CPI / (CPI + 1). The kinship is determined by combining the CPI and RCP values. The larger the CPI value, the closer the RCP is to 1, and the stronger the kinship support.
[0020] This invention discloses an integrated technical solution for species identification and kinship determination in experimental monkeys based on NGS technology. Addressing the shortcomings of existing capillary electrophoresis STR typing methods for kinship determination, such as insufficient primer STR resolution, incompatibility in simultaneous species identification, difficulty in quality control of STR kinship analysis, low automation, and low throughput, this invention establishes an innovative technical solution capable of simultaneously performing SNP species identification and STR kinship determination through customized primer panel development and the construction of an intelligent bioinformatics system. Specifically: 1. Development of a panel primer for SNP species identification and STR phylogenetic identification in experimental monkeys based on NGS technology. This invention independently designed and developed a primer panel based on NGS technology, compatible with both SNP species identification and STR kinship identification. First, species-specific SNP candidate sites were initially screened from the whole genome sequence of experimental monkeys, prioritizing sites with strong species specificity, clear genotyping, and high heterozygosity for accurate differentiation between cynomolgus monkeys, macaques, and their hybrids. Simultaneously, a large number of STR candidate sites were initially screened, comprehensively evaluating site polymorphism, amplification specificity, sequencing quality, and genotyping accuracy, eliminating sites with low amplification efficiency, high sequencing background noise, genotyping errors, and strong cross-interference. Multiplex PCR primers were designed for the selected SNPs and STR sites, and sequencing simulation tests were conducted for verification. Finally, a highly stable and highly specific customized primer panel suitable for cynomolgus monkeys, macaques, and their hybrids was obtained, providing a scientific basis for accurate genetic identification.
[0021] 2. Precise SNP species identification technology Based on self-selected species-specific SNP identification loci, a standardized species identification method was established. By detecting the genotype of this locus, the species of the sample can be accurately determined: a C / C genotype indicates a homozygous cynomolgus monkey, a T / T genotype indicates a homozygous rhesus monkey, and a C / T genotype indicates a hybrid individual. This technology overcomes the limitation of capillary electrophoresis STR typing in accurately identifying hybrid individuals, achieving standardized and precise species identification and providing an accurate source of species information for subsequent kinship identification.
[0022] 3. High-confidence STR kinship testing technology To address the problems of traditional capillary electrophoresis lacking quantitative indicators for STR genotyping confidence and the susceptibility to misjudgments due to low-quality STR genotyping, this invention establishes a high-confidence STR kinship identification method. By introducing a genotyping confidence level (Q-value) during the STR genotyping stage, only loci with a Q=1 high confidence level are included in the kinship calculation, automatically filtering out low-quality and unreliable genotyping, significantly reducing the probability of false positives and misjudgments.
[0023] 4. Adaptable to intelligent bioinformatics analysis systems based on the entire NGS workflow This invention independently establishes a fully intelligent bioinformatics analysis system adapted to NGS. Unlike capillary electrophoresis, which requires tedious manual interpretation of peak patterns, manual genotyping, manual counting of matching sites, and manual calculation of paternity indices, this system achieves intelligent analysis from sequencing quality control, sequence alignment, SNP species identification, STR genotyping, STR site filtering, CPI calculation, RCP calculation, to result analysis and report generation. The entire process requires no manual intervention, avoids subjective misjudgments, and provides stable, traceable, and reproducible results, significantly improving work efficiency.
[0024] 5. A high-throughput sample detection technology system capable of simultaneously identifying species and phylogenetic relationships. Based on the NGS high-throughput sequencing platform, species-specific SNP identification and STR kinship identification are integrated into a single detection workflow. A single experiment can simultaneously differentiate species and genotype individuals of cynomolgus monkeys, macaques, and hybrids, achieving multiple uses with a single test and significantly improving detection efficiency. Leveraging the high throughput and large sample size advantages of the NGS platform, it supports large-scale population tracing and kinship identification of experimental monkeys. This addresses the pain points of low sample throughput in capillary electrophoresis STR genotyping technology, which is unable to perform species identification and is difficult to scale up. It can meet the needs of large-scale population tracing and kinship identification of experimental monkeys, providing a scientific basis for the precise breeding of experimental monkeys. Detailed Implementation
[0025] The technical solution of the present invention will be further described below through specific embodiments. Those skilled in the art should understand that the embodiments described are merely illustrative of the present invention and should not be considered as specific limitations thereof.
[0026] Example 1 1. Development of a panel primer for SNP species identification and STR phylogenetic identification in experimental monkeys based on NGS technology. (1) Initial screening of candidate SNP sites from the whole genome sequence of experimental monkeys, prioritizing sites with strong species specificity, clear genotypes, and high heterozygosity, to distinguish cynomolgus monkeys, macaques and their hybrids; (2) Initially screen a large number of candidate STR genotyping sites, and prioritize sites with high polymorphism, good amplification specificity, stable sequencing quality and accurate genotyping, while eliminating sites with low amplification efficiency, high background noise, easy genotyping errors and strong cross-interference. (3) Multiplex PCR primers were designed for candidate SNPs and STR sites. A total of 30 pairs of multiplex PCR primers were designed to form a panel for identifying the relationship between SNP species and STRs (Table 1).
[0027] (4) The primer panel for identifying the kinship between this SNP species and STR was screened and eliminated after sequencing simulation verification, batch-to-batch stability testing, specificity assessment, and accuracy assessment: Reason for elimination 1: Poor quality of multiplex sequencing, a total of 6 pairs were eliminated: low amplification efficiency after primer mixing, many non-specific products, high sequencing background noise, low percentage of effective reads, and inability to stably genotype (Table 2).
[0028] Reason for elimination 2: Poor quality of bioinformatics analysis, a total of 4 pairs were eliminated: the typing results did not conform to Mendelian inheritance laws, the peak shape was ambiguous, the typing confidence was poor, and the repeatability was poor (Table 3).
[0029] A total of 10 pairs were eliminated, and 20 pairs of primers were ultimately retained to form the customized NGS primer panel of this invention, which is used for subsequent SNP species identification and STR kinship identification.
[0030] Table 1. Primer panels for experimental monkeys ;
[0031] The table shows the primers for the experimental monkey STR loci. The table lists the chromosomal location and upstream and downstream primer sequences for each primer. The "Screening Results" column indicates the final status of the primers; "Retained" indicates that the primers passed multiplexing compatibility, sequencing quality, and genotyping accuracy verification and were included in the final primer panel analysis; primers marked "Excluded" were removed due to poor multiplexing sequencing quality or poor bioinformatics analysis quality.
[0032] Table 2. Data on poor quality of panel primer multiplex sequencing in experimental monkeys
[0033] The table summarizes the data related to the removal of experimental monkey panel primers due to defects in multiplex sequencing quality. For all removed primers, the percentage of target depth sites (50% or higher) did not reach 100%, indicating that the sequencing quality did not meet the analytical requirements. Therefore, these six primer pairs were not included in subsequent data analysis.
[0034] Table 3. Data showing poor performance in panel primer bioinformatics analysis for experimental monkeys
[0035] The table summarizes primers and their data that performed poorly in bioinformatics analysis during panel multiplexing of experimental monkeys. The STR locus genotypes in the table are labeled with fragment values (Q values); the numbers represent the fragment lengths of alleles; the Q value in parentheses is the genotyping confidence level. Some primers consistently showed excessively low Q values, indicating insufficient primer genotyping accuracy; other primers exhibited disordered genotyping matching and unreliable results in STR kinship determination; these four primer pairs did not meet the kinship analysis criteria and were therefore excluded from subsequent sample kinship comparison calculations.
[0036] 2. Experimental Operation Procedure (1) Sample collection and genomic DNA extraction Peripheral blood, hair, or tissue samples were collected from experimental monkeys and stored and transported under refrigeration (2-8℃). High-quality genomic DNA was extracted using a commercially available genomic DNA extraction kit according to standard procedures. The DNA concentration and purity (OD260 / 280=1.8–2.0) were determined using Nanodrop. The distribution and integrity of DNA fragments (RIN value ≥7) were assessed using an Agilent 2100 Bioanalyzer. After passing the quality control, the next step was initiated.
[0037] (2) Construction of multiplex PCR libraries Using a self-developed panel of primers for identifying the species of experimental monkey SNPs and STR phylogenetic relationships (20 primer pairs retained in Table 1), the first round of multiplex PCR targeted amplification was carried out using qualified genomic DNA as a template (reaction conditions: first round multiplex PCR reaction system and amplification procedure) (Table 4), enriching multiple target SNPs and STR fragments at once. The amplification products were purified by magnetic beads to remove non-specific amplification fragments and primer dimers. Subsequently, the sequencing adapters and barcode tags were ligated by a second round of multiplex PCR (reaction conditions: second round multiplex PCR reaction system and amplification procedure) (Table 5), and a sequencing library was constructed.
[0038] Table 4. First-round multiplex PCR reaction system
[0039] First round of multiplex PCR amplification procedure Hot cap temperature: 105℃.
[0040] Pre-denaturation: 95℃, 3 min 30 s, 1 cycle.
[0041] Denaturation-extension cycle; 98℃, 20s; 60℃, 1min; 65℃, 1min; Number of cycles: 24 cycles.
[0042] Final extension: 72℃, 5min, 1 cycle.
[0043] Keep warm at 4℃.
[0044] Table 5. Second-round multiplex PCR reaction system
[0045] Second round of multiplex PCR amplification procedure Hot cap temperature: 105℃.
[0046] Pre-denaturation: 95℃, 3 min 30 s, 1 cycle.
[0047] Amplification cycle: 98℃, 20s; 58℃, 1min; 72℃, 30s; Number of cycles: 9.
[0048] Final extension: 72℃, 5min, 1 cycle.
[0049] Keep warm at 4℃.
[0050] (3) High-throughput sequencing Library concentration quantification was performed using a Qubit fluorometer, and library fragment quality was tested using an Agilent 2100 bioanalyzer. After passing quality control, 150 bp paired-end high-throughput sequencing was performed on the Illumina sequencing platform to obtain raw FASTQ sequencing data.
[0051] 3. Establish an intelligent bioinformatics analysis system adapted to the entire NGS workflow. (1) Sequencing data quality control, alignment and preprocessing The raw FASTQ data was processed by cuttingadapt to remove adapters, barcodes, and low-quality bases; short reads with a length of <20bp were filtered to obtain clean reads; BWA-MEM was used to align to the cynomolgus monkey reference genome; Samtools was used to sort, remove duplicates, filter low-quality alignments, and build an index to generate high-quality BAM files.
[0052] (2) Species-specific SNP identification analysis Based on the BAM file, variant detection and genotyping were performed at the chr11:56049049 locus using bcftools. To determine the bases at chr11:56049049, we uniformly compared the genomes of cynomolgus monkeys. The reference genome was C, meaning cynomolgus monkeys had C at this locus, while rhesus monkeys had T at this locus. In hybrids, both C and T were present at this locus. The software used 0 to represent consistency with the genome and 1 to represent inconsistency to classify samples: 0 / 0 for cynomolgus monkeys, 1 / 1 for rhesus monkeys, and 0 / 1 or 1 / 0 for hybrid individuals. Using this method, species identification was performed on 24 experimental monkeys, verifying an accuracy rate of 100%. Species identification results for some of the experimental monkeys are shown in Table 6.
[0053] Table 6. Species identification results of experimental monkeys
[0054] The table shows the genotyping results for species-specific SNP loci, with C as the reference base and T as the variant base. The sample genotyping format is: genotype: Phred normalized likelihood: total sequencing depth: allele depth. Genotype: 0 / 0 is the homozygous reference type (cynomolgus monkey characteristic), 0 / 1 is the heterozygous type (hybrid individual characteristic), and 1 / 1 is the homozygous variant type (macaque characteristic); Phred normalized likelihood: the confidence score for the three possible genotypes, with lower values indicating a higher probability of that genotype; total sequencing depth: the total number of reads covering this locus in the sample; allele depth: the number of reads supporting the reference base C and the variant base T. All individuals were identified as cynomolgus monkeys, macaques, or cynomolgus monkey / macaque hybrids based on their genotype at this locus.
[0055] (3) Comparison of STR genotyping with high-confidence loci filtering sites GangSTR software was used to genotype the target STR region, obtaining the allele fragment length, genotype, and genotyping confidence Q value for each locus. Before analysis, only high-confidence loci with Q=1 were retained, while low-quality loci with Q<1, ambiguous genotypes, or high background noise were automatically removed to construct a reliable STR genotype matrix. Subsequently, a customized Python script was used to automatically compare the STR loci of offspring with those of parents according to Mendelian inheritance: if any allele in the offspring overlapped with a homologous locus in the parent, it was determined to be a genetically matched locus; if neither allele in the offspring overlapped with a homologous locus in either parent, it was determined to be a genetically mismatched locus. The script automatically counted the number of matched and genetically mismatched loci and generated the final comparison results. Using this method, kinship identification was performed on 41 experimental monkeys (28 related and 13 unrelated), verifying an accuracy rate of 100%. The STR genotyping results of some experimental monkeys are shown in Tables 7 and 8. Table 7 shows the results of some related samples, and Table 8 shows the results of some unrelated samples.
[0056] Table 7. STR genotyping results of related samples
[0057] The STR locus genotypes in the table are labeled with fragment values (Q values); the numbers represent the fragment length of the allele; the Q value in parentheses is the genotyping confidence level. Only high-confidence loci with Q=1 are included in the statistics; loci with Q<1 are not included in subsequent analysis due to insufficient reliability. Interpretation rules: If any allele in the offspring overlaps with a homologous locus in the parent, it is considered a genetic match; if neither allele in the offspring overlaps with either parent, it is considered a genetic mismatch. If more than 3 valid mismatch loci are detected, no kinship is determined; if 3 or fewer mismatch loci are detected, it is considered a natural gene mutation, and a kinship is determined. In this result, fewer than 3 genetic mismatch loci were detected between all parents and offspring, therefore a kinship exists.
[0058] Table 8. STR genotyping results of unrelated samples
[0059] The STR locus genotype labeling format in the table is "fragment value (Q value); the number represents the fragment length of the allele; the Q value in parentheses is the genotyping confidence level. Only high-confidence loci with Q=1 are included in the statistics; loci with Q<1 are not included in subsequent analysis due to insufficient reliability." Interpretation rules: If any allele in the offspring overlaps with a homologous locus in the parent, it is considered a genetic match; if neither allele in the offspring overlaps with either parent, it is considered a genetic mismatch. If more than 3 valid mismatch loci are detected, no kinship is determined; if 3 or fewer mismatch loci are detected, it is considered a natural gene mutation, and a kinship is determined. In this result, more than 3 genetic mismatch loci were detected between any two individuals; therefore, kinship between any two individuals is excluded.
[0060] (4) CPI / RCP calculation and comprehensive determination of kinship A custom Python script was used to calculate the CPI for all paired samples. For each STR locus, the PI was calculated as follows: PI = 1 / p for homozygous-homozygous matches, and PI = 1 / (2p) for heterozygous matches, where p was the population allele frequency obtained from a pre-calculated frequency database. When no alleles were shared, PI was set to the mutation rate (μ = 0.001). CPI was the product of the PI values for all high-confidence loci (Q = 1). The RCP was then calculated based on the Essen-Möller formula: RCP = CPI / (CPI + 1). The kinship was determined by combining CPI and RCP values; the higher the CPI value and the closer the RCP is to 1, the stronger the kinship support. This method was used to identify the kinship of 41 experimental monkeys (28 were related and 13 were unrelated), verifying an accuracy of 100%. The CPI test results are shown in Tables 9 and 10. Table 9 shows the results of some related samples, and Table 10 shows the results of some unrelated samples.
[0061] Table 9. CPI Calculation Results for Related Samples
[0062] The table shows the CPI values between pairs of samples. Diagonal: Sample self-comparison, CPI is always 1; Kinship: CPI is a positive value (indicated by underline), indicating the strength of kinship evidence; the higher the value, the stronger the evidence supporting kinship.
[0063] Table 10. CPI Calculation Results for Unrelated Samples
[0064] The table shows the CPI values between pairs of samples. Diagonal: self-comparison, CPI is always 1; unrelated: CPI is negative, excluding related samples.
[0065] (5) Intelligent result analysis and report generation Using a customized Python script, the intelligent system integrates sample information, species identification, STR genotyping, matching / exclusion statistics, CPI / RCP values, and kinship conclusions to generate standardized official reports with a single click. The reports are formatted correctly, contain complete data, and provide clear results, and support batch export and archiving.
Claims
1. A set of primer panels for SNP species identification and STR phylogenetic identification in experimental monkeys based on NGS technology, characterized in that, Comprising the following primer pairs: F: CAAGCATGGAGATGGTCTGGTT, R: GTACATGCCTCTTTGTTGCAGTG; F: GCACTGCTAAGGCTTCTATCACA, R: GGTAGTGACATGTGCTCACTGT; F: TGAGCCTCAGAATTACCCCAGT, R: TCACTTGAACCTGTGAGGCG; F: CAGTTACGAGGAGGGTTGACATC, R: TGAGACAGTGGCATAAACCAGG; F: CCCTGGTTCTGAGGTTTTTGGA, R: TTCGGGTTCTCCAAAGAGACAG; F: CACCTGTCTCAATCCAAGACAAATC, R: ACAGGCTATCTATCTATCTATTTATTTATCATCT; F: CCATCACTTACTGGCAATGTAACC, R: TGCTGGAAACTGATAAGGGCTTTA; F: GGATCATGAAAGGGCATGAGGA, R: ACTCCCTTCTTCCCTCTCACAG; F: GTGATGGAAAAGAATCGGGACAG, R: CATCAACATCACCCCAACACCT; F: ACTTGGAAAGTATGCTGCCTCTG, R: GGATCACTTGAACCTGGGAGATG; F: AACAAAGGAGGCAGTGAGCATC, R: AGTGAGCTGAGATCACGTCACT; F: CCCAGGAGTTCCAATTTCTCCA, R: TCTGATAAGGGCTTGATATCCAGG; F: GCCAAGGATGGTGAGTTACTCA, R: TGGTAGTGATGTGGCCCTAAGT; F: CAAGTTCTAACATCACGTCCCTCT, R: GTACCTGAGGTCATCAGGACATTC; F: CTCTTCCACTGATTCTGCCCAT, R: GTGACACAGAAACAGTCTGGGA; F: GTCCATAGTGGTGCTTCTCCAT, R: GCATCTGTGTGGATTTGGGGTA; F: CCCAAGTGGGTCCAAGTGGCT, R: GGATAGGGCAACAGAGAAGAT; F: AGCCCAGATATCCCCAAGATCTC, R: CAAGTGATGGCCCAAATTTGGC; F: CCCCTATTCTTAGGAAATAAACCCTGA, R: CACAGTTGAAATCCTCTACCCAGAT; F: AGGATTCTCCAGGCAAATAGAACC, R: GACATACACCATTGGCTCCCAT; The experimental monkeys mentioned are cynomolgus monkeys, macaques, or hybrids thereof.
2. A kit for SNP species identification and STR kinship identification of experimental monkeys based on NGS technology, characterized in that, The experimental monkey contains the primer Panel as described in claim 1, and the experimental monkey is a cynomolgus monkey, a rhesus monkey, or a hybrid of the two.
3. A method for simultaneously identifying the species and kinship of experimental monkeys based on NGS technology, characterized in that, Including the following methods: (1) Collect samples from experimental monkeys and extract genomic DNA. The experimental monkeys are cynomolgus monkeys, macaques, or their hybrids. (2) Using the primer Panel described in claim 1 as the amplification primer, the first round of multiplex PCR targeted amplification is performed with genomic DNA as the template. The amplification product is purified by magnetic beads to remove non-specific amplification fragments and primer dimers. Then, the sequencing adapter and barcode tag are connected by the second round of multiplex PCR to construct a sequencing library. Finally, the original FASTQ sequencing data is obtained by high-throughput sequencing. (3) Perform quality control trimming and short fragment filtering on the original FASTQ sequencing data, compare it with the monkey reference genome using BWA-MEM, and process it with samtools to obtain a high-quality BAM file; (4) Then use SNP to perform species-specific identification analysis to determine the base at position 56049049 of chr11. The cynomolgus monkey is C and the macaque is T. If it is a hybrid species, the position is CT. (5) The kinship is then determined by comparing the STR genotype with the high-confidence filtering site; (6) Calculate the unit point paternity index PI using Python script and multiply it to obtain the cumulative paternity index CPI. Convert the relative paternity probability RCP using the Essen-Möller formula and classify the kinship determination criteria according to the CPI and RCP thresholds. (7) Customize Python scripts to integrate all test data and automatically generate standardized kinship test reports in batches.
4. The method according to claim 3, characterized in that, The two-round multiplex PCR reaction system mentioned in step (2) is: The first round of multiplex PCR reaction system consisted of: 13 μL of Nuclease-Free Water, 5 μL of primer panel with a final concentration of 100 nM, 40 ng of genomic DNA per 2 μL of reaction, and 10 μL of Amplicon Mix. First-round multiplex PCR amplification program: Heat cap temperature: 105℃; Pre-denaturation: 95℃, 3 min 30 s, 1 cycle; Denaturation-extension cycle: 98℃, 20 s; 60℃, 1 min; 65℃, 1 min; Number of cycles: 24 cycles; Final extension: 72℃, 5 min, 1 cycle. Insulation: Store at 4℃ under constant temperature. The second round of multiplex PCR reaction system consisted of: 10 μL of PCR product purified in the first round, 2 μL of 10 μM Indexed Primer, 15 μL of PCR Master Mix, and 3 μL of Nuclease-Free Water. Second-round multiplex PCR amplification program: Heat cap temperature: 105℃; Pre-denaturation: 95℃, 3 min 30 s, 1 cycle; Amplification cycles: 98℃, 20 s; 58℃, 1 min; 72℃, 30 s; Number of cycles: 9 cycles; Final extension: 72℃, 5 min, 1 cycle. Keep warm at 4℃.
5. The method according to claim 3, characterized in that, The step (4) is to use bcftools to determine the base variation at the chr11:56049049 site. The cynomolgus monkey is C, the macaque is T, and if it is a hybrid, the site is CT.
6. The method according to claim 3, characterized in that, Step (5) is as follows: Use GangSTR software to genotype the target STR region, obtain the allele fragment length, genotype, and genotyping confidence Q value for each locus. Before analysis, only high-confidence loci with Q=1 are retained, and low-quality loci with Q<1, ambiguous genotyping, and high background noise are automatically removed. Construct a reliable STR genotype matrix. According to Mendel's inheritance laws, compare the STR loci of the offspring with those of the parents: If any allele of the offspring overlaps with the homologous locus fragment of the parents, it is determined to be a genetically matched locus; if both alleles of the offspring do not overlap with the homologous locus fragments of the parents, it is determined to be a genetically mismatched locus; if more than 3 valid mismatched loci are detected, it is determined to be unrelated; if less than or equal to 3 mismatched loci are considered as natural gene mutations, and kinship is determined, including parent-child and kinship relationships.
7. The method according to claim 3, characterized in that, Step (6) is as follows: Using a customized Python script, the comprehensive paternal index (CPI) is calculated for all paired samples. For each STR locus, the formula for calculating the paternal index (CPI) is: when homozygous-homozygous matches, PI = 1 / p; when heterozygous matches, PI = 1 / (2p), where p is the population allele frequency obtained from the pre-calculated frequency database. When there is no shared allele, PI is set to the mutation rate (μ = 0.001). The cumulative paternity index (CPI) is the product of the PI values when Q = 1 for all high-confidence loci. Then, the relative paternal probability (RCP) is calculated based on the Essen-Möller formula: RCP = CPI / (CPI + 1). The kinship is determined by combining the CPI and RCP values. The larger the CPI value, the closer the RCP is to 1, and the stronger the kinship support.