Method for rapidly performing species identification and sex identification based on whole genome re-sequencing data

Through mitochondrial genome sequence alignment and Y chromosome specific mutation site screening of whole genome resequencing data, the problems of low accuracy and limited application scope of species and gender identification are solved, and efficient and accurate species and gender identification are achieved, which is suitable for rapid identification of complex samples.

CN120260682APending Publication Date: 2025-07-04GUANGDONG LABORATORY OF SOUTHERN OCEAN SCIENCE AND ENGINEERING (GUANGZHOU)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510334742.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has low accuracy and limited application scope in species and gender identification, especially when processing complex samples, it is difficult to meet the needs of modern research. The traditional methods are greatly affected by ontogenic stages, environmental factors and human judgment errors.

Method used

Whole genome resequencing data is used to identify species through mitochondrial genome sequence comparison, and Y chromosome specific mutation sites are screened and gender judging is used. Data processing is carried out in combination with NOVOPlasty, Fastp, Bwa, Samtools, Picard and GATK software to achieve efficient and accurate species and gender identification.

Benefits of technology

It has achieved efficient and accurate species and gender identification, and is suitable for a variety of biological samples, especially samples that cannot be visually identified by the naked eye. It is suitable for fields such as identification of illegal trade in wild animals, evolutionary analysis and forensic case identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005321394740000041
    Figure BDA0005321394740000041
  • Figure BDA0005321394740000051
    Figure BDA0005321394740000051
  • Figure BDA0005321394740000061
    Figure BDA0005321394740000061
Patent Text Reader

Abstract

The invention discloses a method for rapidly performing species identification and sex identification based on whole genome re-sequencing data. According to the method, whole genome DNA of tissue samples such as individual bloodstain, skin and skeleton is extracted and re-sequenced, species identification is realized by virtue of mitochondrial genome sequence alignment, and sex identification is realized by virtue of Y chromosome specific variation site screening and judgment; the method has the advantages of high efficiency, high precision and wide applicability, and can be applied to the fields of illegal trade identification of wild animals, evolution analysis, forensic case identification and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of biotechnology and molecular genetics, and particularly relates to a method for rapidly identifying species and gender based on whole-genome resequencing data. Background Art

[0002] In many fields such as wildlife conservation, evolutionary biology, and forensic medicine, rapid and accurate identification of individual species and gender has important theoretical significance and practical application value. In wildlife conservation, accurate identification of species and gender helps to evaluate population diversity, formulate conservation strategies, and combat illegal hunting; in evolutionary biology, precise identification of species and gender provides basic data for studying species origin, species evolution process, and gender determination mechanism; in forensic medicine, species identification and gender determination are key links in crime scene analysis and case detection, and can provide a scientific basis for forensic identification.

[0003] However, traditional methods for species and gender identification mainly rely on morphological feature observation or PCR amplification of specific gene fragments. Although these methods can meet the requirements to a certain extent, they have great limitations. For example, morphological identification is easily affected by individual development stage, environmental factors, or human judgment errors, while PCR amplification technology based on specific genes has low accuracy and limited applicability due to primer design limitations, making it difficult to cover a wide range of species or gender-specific gene regions. In addition, these traditional methods perform poorly when dealing with complex samples (such as mixed samples or degraded samples), and it is difficult to meet the diverse needs of modern research.

[0004] With the rapid development of molecular biology techniques, especially the wide application of high-throughput sequencing technology, whole-genome resequencing provides a new research idea and technical means for species identification and gender determination. Whole-genome data can not only comprehensively reveal the genetic information of species, but also achieve precise gender identification by analyzing sex chromosomes or gender-related genes. Compared with traditional methods, whole-genome resequencing has advantages such as high coverage, rich information content, and wide adaptability, especially showing significant potential when dealing with non-model animals or unknown samples. However, although whole-genome data shows broad application prospects in the field of species and gender identification, there is currently no research on developing a rapid species identification and gender determination method based on whole-genome data. Therefore, exploring how to use whole-genome data to achieve efficient and accurate species and gender identification has important scientific value and practical significance. Summary of the Invention

[0005] In view of the above deficiencies of the prior art, the present invention provides a method for rapidly identifying species and sex determination based on whole-genome resequencing data, which can identify species and sex for individual samples that cannot be visually identified by the naked eye (such as animal old hides, muscles, bones, forensic blood stains, bulk tissues, etc.).

[0006] The method for rapidly identifying species and sex determination based on whole-genome resequencing data of the present invention includes the following steps:

[0007] S1. Extract the whole-genome DNA of the sample, construct a second-generation sequencing library, and perform whole-genome sequencing;

[0008] S2. Based on the original second-generation whole-genome sequencing data of the sample, use the NOVOPlasty software to assemble the mitochondrial genome sequence, extract the COI and Cytb mitochondrial sequences according to the annotation file, perform an online comparison of the mitochondrial sequences with the NT database, and select the species corresponding to the sequence with the highest sequence matching degree as the species identification result of the sample to achieve species identification;

[0009] S3. Use the default parameters of the Fastp software to filter the original second-generation whole-genome sequencing data of the sample to obtain high-quality reads; select the Y chromosome sequences of the corresponding species or closely related species as the reference sequences, and use the default parameters of the index module of the Bwa software, the faidx module of the Samtools software, and the CreateSequenceDictionary module of the Picard software to construct indexes for the reference sequences; use the MEM module of the Bwa software to align the filtered reads to the reference sequences to obtain a bam file; use the sort module of the Samtools software to sort the bam file; use the Picard software to remove PCR duplicate reads from the sorted bam file, with the parameters MarkDuplicates REMOVE_DUPLICATES = true ASSUME_SORT_ORDER = coordinate; use the HaplotypeCaller module of the GATK software to process the sorted bam file with PCR duplicate reads removed, with the parameter "-ERC GVCF" and other default parameters, and screen to obtain the Y chromosome variant site information of the sample;

[0010] S4. According to the Y chromosome variant site information of the sample, use the grep software to count the number of homozygous variant sites; if the number of homozygous variant sites is large, it is determined to be male, and vice versa, it is determined to be female.

[0011] Preferably, the sample is an animal sample that needs to be identified for species and sex and cannot be visually identified by the naked eye.

[0012] Preferably, the animal sample is an old animal hide, forensic bloodstain, muscle or bone.

[0013] Preferably, the animal is a snow leopard, and the reference sequence is GCA_019924945.1.

[0014] Preferably, the animal is a takin, and the reference sequence is GCF_023091745.1. Preferably, the step S1 is as follows: extracting the whole-genome DNA of the sample, fragmenting the DNA into 200-300 bp by ultrasound or enzymatic digestion, constructing a second-generation sequencing library through end repair, adding A at the 3' end, ligation of adapters, PCR enrichment of the target fragment, and determination of fragment size and concentration, performing whole-genome sequencing on the library qualified by quality inspection using the Illumina Nova-PE150 sequencing platform, and the sequencing depth is 20×.

[0015] Preferably, the number of homozygous variant sites is one of the following (1) or (2):

[0016] (1) When the Y chromosome exists in the sample, the number of homozygous variant sites is the sum of the number of homozygous SNP sites in the homologous fragment of the X chromosome and the Y chromosome reference sequences and the number of homozygous SNP sites in the Y chromosome-specific fragment;

[0017] (2) When the Y chromosome does not exist in the sample, the number of homozygous variant sites is the number of homozygous SNP sites in the homologous fragment of the X chromosome and the Y chromosome reference sequences.

[0018] Preferably, the determination that the number of homozygous variant sites is large is for male. The batch samples are divided into two groups, a group with a lower number and a group with a higher number, according to the number of homozygous variant sites on the Y chromosome. The average value of the number of homozygous variant sites in the group with a higher number is more than 1.5 times that of the group with a lower number; if the number of homozygous variant sites of the sample conforms to the group with a higher number, it is determined to be male, and vice versa, it is determined to be female.

[0019] Compared with the prior art, the present invention has the following advantages:

[0020] (1) High efficiency: Using whole-genome resequencing data, species identification and gender determination can be completed for multiple samples simultaneously, significantly improving the detection efficiency.

[0021] (2) High precision: Through mitochondrial genome sequence alignment and screening and determination of Y chromosome-specific variant sites, species can be accurately identified and individual genders can be distinguished, avoiding misjudgment of traditional methods.

[0022] (3) Wide applicability: Applicable to a variety of biological samples, including animal old tissues, bones, forensic bloodstains, etc., especially applicable to individuals that cannot be visually identified directly, and can be applied to fields such as the identification of illegal wildlife trade, evolutionary analysis, and forensic case identification. Detailed implementation manners

[0023] The following embodiments are further descriptions of the present invention rather than limitations on the present invention.

[0024] Embodiment 1

[0025] A method for quickly identifying species and gender from whole-genome resequencing data of batch samples includes the following steps:

[0026] Step 1. Sample DNA extraction: Use DNeasy Blood&TissueKit (Qiagen) to extract whole-genome DNA from samples such as individual bloodstains, skin, or bones.

[0027] Step 2. Whole-genome resequencing of samples: First, fragment the DNA into 200 - 300 bp by ultrasound or enzymatic digestion, and then construct a second-generation sequencing library through steps such as end repair, adding A at the 3' end, ligation of adapters, PCR enrichment of target fragments, determination of fragment size and concentration, etc. Use the Illumina Nova-PE150 sequencing platform to perform whole-genome sequencing on the library qualified by quality inspection, and the sequencing depth is about 20×.

[0028] Step 3. Sample species identification: Based on the original second-generation whole-genome sequencing data of each sample, use NOVOPlasty software to assemble mitochondrial genome sequences, extract COI and Cytb mitochondrial sequences according to the annotation file, and perform an online comparison of the mitochondrial sequences with the NT database (https: / / ftp.ncbi.nlm.nih.gov / blast / db / FASTA / nr.gz), and select the species corresponding to the sequence with the highest sequence matching degree as the candidate species identification result to achieve species identification.

[0029] Step 4. Screening of Y chromosome variable sites of samples:

[0030] Step 4-1. Use the default parameters of Fastp software to filter the original second-generation whole-genome sequencing data of each sample to obtain high-quality reads;

[0031] Step 4-2. Select the Y chromosome sequences of the corresponding species or related species as reference sequences, and use the default parameters of the index module of Bwa software, the faidx module of Samtools software, and the CreateSequenceDictionary module of Picard software to construct indexes for the reference sequences;

[0032] Step 43: Align the filtered reads to the reference sequence using the MEM module of the Bwa software to obtain a bam file;

[0033] Step 44: Sort the bam file using the sort module of the Samtools software;

[0034] Step 45: Remove PCR duplicate reads from the sorted bam file using the Picard software with the parameters "MarkDuplicates REMOVE_DUPLICATES=true ASSUME_SORT_ORDER=coordinate";

[0035] Step 46: Process the sorted bam file with removed PCR duplicate reads using the HaplotypeCaller module of the GATK software with the parameter "-ERC GVCF" and other default parameters to screen for Y chromosome variant site information.

[0036] Step 5: Sample sex determination:

[0037] Step 51: According to the Y chromosome variant site information, use the grep software to count the number of homozygous variant sites in each sample (i.e., count homozygous SNP sites with genotypes 1 / 1 or 1|1, where 1 / 1 represents an unphased homozygous genotype, meaning that two alleles are known to be 1, but it is not known which parent they come from; while 1|1 represents a phased homozygous genotype, meaning that both alleles are known and which parent they come from). Since there are homologous segments between the Y chromosome and the X chromosome, in both female and male individuals, homozygous variant sites with genotypes 1 / 1 or 1|1 will appear in the variant information of all sample individuals. The number of homozygous variant sites on the Y chromosome counted above includes the number of homozygous SNP sites in the homologous segment between the X chromosome and the Y chromosome reference sequence and (if there is a Y chromosome) the number of homozygous SNP sites in the Y chromosome-specific segment (non-homologous to the X chromosome); because in addition to the homologous segment with the X chromosome, the Y chromosome also has a Y-specific segment, so theoretically, the number of homozygous variant sites in male individuals will be more than that in female individuals.

[0038] Step 52: Method for determining the sex of the sample: According to the above principle, the number of homozygous variant sites on the Y chromosome of all sample individuals usually shows two patterns. The sample with more homozygous variant sites is male, and the other is female, thus realizing sex identification.

[0039] For example, in Case 1 using the above method for rapid species identification and gender determination based on whole-genome resequencing data of batch samples (for 48 old snow leopard skin samples, where the Y chromosome sequence of the closely related species cat (GCA_019924945.1) was used as the reference sequence), the results of Y chromosome variation information for 48 snow leopard skin samples (Table 1) showed that the number of homozygous variant sites in male individuals was approximately twice that in female individuals. According to the principle that male individuals have more Y chromosome homozygous variant sites, the gender determination results for 48 snow leopard samples showed that 19 samples were identified as male and 29 samples were identified as female; among them, 14 samples had known gender information, and the gender determination results using the above method were completely consistent with the corresponding actual gender information, indicating the reliability of this method.

[0040] Table 1 Gender determination results for 48 snow leopard skin samples based on whole-genome resequencing data

[0041]

[0042]

[0043]

[0044] Note: In the table, NA under the individual gender information item indicates that the gender information of the individual sample is missing, male indicates that the gender of the individual sample is known to be male, and female indicates that the gender of the individual sample is known to be female.

[0045] In Case 2 using the above method for rapid species identification and gender determination based on whole-genome resequencing data of batch samples (for 74 old takin skin samples, where the takin Y chromosome sequence (GCF_023091745.1) was used as the reference sequence), the results of Y chromosome variation information for 74 takin skin samples (Table 2) showed that the number of homozygous variant sites in male individuals was approximately 10 times that in female individuals. According to the principle that male individuals have more Y chromosome homozygous variant sites, the gender determination results for 74 takin skin samples showed that 37 samples were identified as male and 37 samples were identified as female.

[0046] Table 2 Gender determination results for 74 takin skin samples based on whole-genome resequencing data

[0047]

[0048]

[0049]

Claims

1. A method for rapid species identification and gender determination based on whole-genome resequencing data, characterized in that, Including the following steps: S1. Extract the whole-genome DNA of the sample, construct a second-generation sequencing library, and perform whole-genome sequencing; S2. Based on the original second-generation whole-genome sequencing data of the sample, use the NOVOPlasty software to assemble the mitochondrial genome sequence, extract the COI and Cytb mitochondrial sequences according to the annotation file, perform an online alignment of the mitochondrial sequences with the NT database, and select the species corresponding to the sequence with the highest sequence matching degree as the sample species identification result to achieve species identification; S3. Use the default parameters of the Fastp software to filter the original second-generation whole-genome sequencing data of the sample to obtain high-quality reads; select the Y chromosome sequence of the corresponding species or closely related species as the reference sequence, and use the default parameters of the index module of the Bwa software, the faidx module of the Samtools software, and the CreateSequenceDictionary module of the Picard software to construct an index for the reference sequence; use the MEM module of the Bwa software to align the filtered reads to the reference sequence to obtain a bam file; use the sort module of the Samtools software to sort the bam file; use the Picard software to remove PCR duplicate reads from the sorted bam file, with the parameters MarkDuplicates REMOVE_DUPLICATES=true ASSUME_SORT_ORDER=coordinate; use the HaplotypeCaller module of the GATK software to process the sorted bam file with PCR duplicate reads removed, with the parameter "-ERC GVCF", and other default parameters, and screen to obtain the sample Y chromosome variant site information; S4. According to the sample Y chromosome variant site information, use the grep software to count the number of homozygous variant sites; if the number of homozygous variant sites is large, it is determined to be male, otherwise it is determined to be female.

2. The method according to claim 1, characterized in that, The sample is an animal sample that needs to be identified for species and sex and cannot be visually identified by the naked eye.

3. The method according to claim 1, wherein The animal sample is an old animal hide, forensic bloodstain, muscle or bone.

4. The method according to claim 3, wherein The animal is a snow leopard.

5. The method according to claim 4, characterized in that, The reference sequence is GCA_019924945.

1.

6. The method according to claim 3, characterized in that, The animal is a takin.

7. The method according to claim 6, wherein The reference sequence is GCF_023091745.

1.

8. The method according to claim 1, wherein The step S1 is: extract the whole-genome DNA of the sample, fragment the DNA into 200-300 bp by ultrasound or enzymatic digestion, construct a second-generation sequencing library through the steps of end repair, A addition at the 3' end, adapter ligation, PCR enrichment of the target fragment, fragment size and concentration determination, and perform whole-genome sequencing on the library qualified by quality inspection using the Illumina Nova-PE150 sequencing platform, with a sequencing depth of 20×.

9. The method according to claim 1, wherein The number of homozygous variant sites is one of the following (1) or (2): (1) In the case where the Y chromosome exists in the sample, the number of homozygous variant sites is the sum of the number of homozygous SNP sites in the homologous fragments of the X chromosome and the Y chromosome reference sequences and the number of homozygous SNP sites in the Y chromosome-specific fragments; (2) In the case where the Y chromosome does not exist in the sample, the number of homozygous variant sites is the number of homozygous SNP sites in the homologous fragments of the X chromosome and the Y chromosome reference sequences.

10. The method according to claim 1, characterized in that, The determination that the sample with a larger number of homozygous variant sites is male is based on dividing a batch of samples into two groups, a lower-number group and a higher-number group, according to the number of Y chromosome homozygous variant sites. The mean value of the number of homozygous variant sites in the higher-number group is more than 1.5 times that of the lower-number group. If the number of homozygous variant sites of a sample conforms to that of the higher-number group, it is determined to be male; otherwise, it is determined to be female.