A method for hybridization detection using whole genome resequencing data alone

By identifying nuclear mitochondrial DNA fragments and drawing genetic distance curves using whole-genome resequencing data, the problems of high cost and low accuracy of hybridization detection in existing technologies are solved, sensitive and accurate hybridization detection is achieved, and the scope of application of detection is expanded.

CN119541634BActive Publication Date: 2025-09-19NORTHEAST FORESTRY UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411599910.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-09-19
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

Hybridization detection methods in existing technologies are costly, have low accuracy, and require pure parental reference genome information, making it difficult to identify hybrid individuals efficiently and accurately.

Method used

Whole-genome resequencing data were used to identify nuclear mitochondrial DNA fragments (NUMTs) through mitochondrial genome assembly and alignment, and a genetic distance line graph was drawn. Hybridization detection was performed based on the genetic distance between NUMTs and mitochondrial homologous fragments.

Benefits of technology

It achieves sensitive and accurate hybridization detection, reduces costs, does not rely on pure parental reference genome information, expands the scope of application of the test, and improves the credibility of the test results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119541634B_ABST
    Figure CN119541634B_ABST
Patent Text Reader

Abstract

A method for hybridization detection using whole genome resequencing data alone belongs to the field of molecular biology technology. In order to solve the technical problems of high cost, low accuracy and the need for pure parental reference genome information in the existing hybridization detection methods, the present invention provides a method for detecting hybrid individuals based on nuclear mitochondrial DNA resequencing data. The method is based on the whole genome resequencing data of the target individual and the nuclear mitochondrial DNA fragments, and judges the hybridization detection results based on a broken line graph drawn according to the genetic distance data set between the nuclear mitochondrial DNA fragments and their mitochondrial homologous fragments. The hybrid individual detection method provided by the present invention has high accuracy, circumvents the technical difficulties of needing to design genetic markers that are effective for both parents, and does not require the detection of pure parental reference genomes of the target individual, thereby reducing detection costs and expanding the scope of application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of molecular biology, and in particular relates to a method for hybridization detection using whole genome resequencing data alone. Background Art

[0002] Hybridization is the process of mating between closely related species and is a natural phenomenon. Studies have found evidence of introgressive hybridization in many groups, including amphibians, birds, and mammals, where hybrids backcross with the parental species, thereby achieving gene flow between groups. This provides clues for a deeper understanding of the formation process of species boundaries and whether the gene exchange pattern is consistent. The introduction of non-native species is the main impact of human activities on biodiversity and also promotes hybridization. It has become the second largest threat to biodiversity after habitat loss, partly because they have the potential to hybridize with native populations, thereby affecting genetic integrity, inducing out-breeding depression, and even replacing native populations genetically and phenotypically, thereby increasing the risk of species extinction and posing a challenge to biodiversity conservation. Therefore, it is very important to propose an efficient and accurate hybridization detection method for ecological and evolutionary research and conservation practice.

[0003] Hybrids inherit traits from both parental species. Over successive generations of backcrossing, hybrids gradually lose their distinct traits and become more similar to purebred individuals, ultimately appearing native in morphology. This poses a significant challenge to morphological identification. Molecular biological identification is a sensitive detection method. Traditional molecular biological identification methods use mitochondrial DNA (mtDNA) to identify the maternal lineage and use genomic markers such as microsatellites (also known as short tandem repeats, STRs) or single nucleotide polymorphisms (SNPs) to estimate the genomic contribution of the parental species. However, the entire detection method has technical issues such as complex processes, high costs, and low accuracy. Currently, whole-genome sequencing technology allows the use of a large number of SNPs covering the entire genome, improving the accuracy of hybridization detection results. However, the use of genetic markers requires reference samples of known pure maternal and pure paternal parents. Microsatellite and SNP markers must be valid for both parental species. Otherwise, genotyping errors will also weaken the accuracy of identification. Validation of marker panels is a time-consuming and costly technical issue. For the same reason, whole-genome sequencing (WGS) methods also require pure maternal and pure paternal reference genomes. Although genome assembly is becoming more common, cheap and easy, it still has the problem of purity authentication of individuals used for genome assembly and is expensive.

[0004] Therefore, those skilled in the art are eager to develop a method for hybridization detection of target individuals that is sensitive, accurate, low-cost, and does not require pure maternal and pure paternal reference genome information. Summary of the Invention

[0005] In order to solve the technical problems of the existing hybridization detection methods such as high cost, low accuracy and the need for pure parental reference genome information, the present invention provides a method for hybridization detection using whole genome resequencing data alone.

[0006] One of the objectives of the present invention is to provide a method for hybridization detection using whole genome resequencing data alone, wherein the detection method specifically comprises the following steps:

[0007] Step 1: Collect samples from the target individual, transport them on dry ice and store them below -20°C. Use a kit to extract DNA from the samples. Construct a DNBSEQ whole-genome resequencing library for sequencing to obtain resequencing data for the target individual.

[0008] Step 2: Download the cytochrome b sequence of the species to which the target individual belongs from NCBI as the seed for de novo assembly of the mitochondrial genome. Use NOVOplasty v4.3.1 to assemble the mitochondrial genome and obtain the mitochondrial reference genome.

[0009] Step 3: Identification of nuclear mitochondrial DNA fragments;

[0010] Step 3.1: Remove the reads used to assemble the mitochondrial genome from the resequencing data reads obtained in step 1 to generate a new fastq file. Use BWA0.7.17-r1188 to align the obtained fastq file with the mitochondrial reference genome assembled in step 2 to generate a bam file.

[0011] Step 3.2: Use samtools 1.17 to filter the bam file obtained in step 3.1 to obtain potential nuclear mitochondrial DNA sequences;

[0012] Step 3.3: Use bamtobed ​​of bedtools 2.30.0 to convert the potential nuclear mitochondrial DNA sequence obtained in step 3.2 into BED format, cluster the overlapping reads within 1000 bp using Bedtools Cluster, and then convert the above sequence back to fastq format using Bedtools Bamtofastq;

[0013] Step 3.4: Perform a BLAST comparison between the mitochondrial reference genome from step 2 and the fastq obtained from step 3.3. Use the -dust no parameter to not mask low-complexity regions. The Python script extracts the highest-scoring BLAST matches. The above nuclear mitochondrial DNA sequences are compared with each other to obtain the confirmed nuclear mitochondrial DNA sequence.

[0014] Step 4: Align the nuclear mitochondrial DNA sequence determined in step 3 with the mitochondrial reference genome in step 2 separately without filtering, calculate the genetic distance between each nuclear mitochondrial DNA sequence and its mitochondrial homologous fragment; and draw a line graph, and judge the results based on the line graph.

[0015] In a preferred embodiment of the present invention, when the sample in step 1 is a blood or tissue sample, whole genome resequencing can be performed directly after DNA extraction, and the depth of whole genome resequencing is >30X.

[0016] In a preferred embodiment of the present invention, the resequencing data reads in step 1 are clean-data obtained after quality control.

[0017] In a preferred embodiment of the present invention, the genetic distance in step 4 is calculated using the Kimura2 parameter model in MEGAx.

[0018] In a preferred embodiment of the present invention, the horizontal axis of the line graph in step 4 represents the genetic distance, and the vertical axis represents the frequency of each genetic distance interval.

[0019] In a preferred embodiment of the present invention, the criterion for judging the result in step 4 is: if the frequency of the genetic distance interval decreases steadily as the X-axis increases, it indicates that the target individual is a purebred; if the frequency of the genetic distance interval shows a frequency peak as the X-axis increases, it indicates that the target individual is a hybrid.

[0020] The beneficial effects of the present invention are as follows: the present invention is based on the theory that a part of the mitochondrial genome can be transferred to the nuclear genome, namely the nuclear mitochondrial DNA fragment (NUMTs), and provides a method for hybridization detection using whole-genome resequencing data alone; the detection method assembles the mitochondrial genome using the cytochrome b fragment of the species to which the target individual belongs downloaded from NCBI as a seed after whole-genome resequencing of the target individual, and then compares the obtained resequencing data reads with the assembled mitochondrial genome using BWA, and the comparison results are screened to obtain determined NUMTs, and the genetic distance data set between the NUMTs and their mitochondrial homologous fragments is plotted into a line graph, and the target individual is identified as a hybrid individual based on the relationship between the genetic distance on the horizontal axis and the frequency of each genetic distance interval on the vertical axis in the line graph.

[0021] The present invention uses samtools 1.17 to filter bam files to remove low-quality reads (mapping quality <30), unmapped reads, and reads with secondary alignments. The clustering results are converted back to fastq format using Bedtools Bamtofastq to minimize redundant hits. Targeted reads are masked by BLAST alignment, and the "-dust no" parameter is used to not mask low-complexity regions to avoid missing information that may be useful for research, thereby improving the accuracy of the detection results of the method described in the present invention.

[0022] The present invention uses molecular means to perform hybridization detection, which has higher credibility and improves the accuracy of hybridization detection results compared with morphological hybridization detection methods. The present invention does not require the design of microsatellite and SNP markers. Compared with traditional molecular hybridization detection methods, it effectively circumvents the technical difficulty of needing to design markers that are effective for both pure parents and reduces detection costs. In addition, the detection method provided by the present invention does not require the detection of a reference genome at the chromosome level of the pure parental population of the target individual, which expands the scope of application of the detection method provided by the present invention and provides technical support for biodiversity conservation. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 The figure shows the test results of purebred sika deer in Example 1; the horizontal axis Genetic Distance is the genetic distance, and the vertical axis Frequency is the frequency;

[0024] Figure 2 This is a graph showing the test results of the sika deer and red deer hybrids in Example 1;

[0025] Figure 3 This is a graph showing the test results of the purebred Siberian tiger in Example 2;

[0026] Figure 4The figure is the test result of the hybrid of Siberian tiger and Bengal tiger in Example 2;

[0027] Figure 5 This is a graph showing the test results of the purebred Golden Pheasant in Example 3;

[0028] Figure 6 This is a diagram showing the test results of the hybrid between the red-bellied pheasant and the white-bellied pheasant in Example 3;

[0029] Figure 7 This is a graph showing the test results of the purebred green turtle in Example 4;

[0030] Figure 8 This is a graph showing the test results of the hybrid of the green turtle and the loggerhead turtle in Example 4;

[0031] Figure 9 The figure is the test result of purebred blue catfish in Example 5;

[0032] Figure 10 Graph showing the detection results of the blue catfish and channel catfish hybrids in Example 5. DETAILED DESCRIPTION

[0033] Those skilled in the art can refer to the content of this document and appropriately improve the process parameters. It is particularly important to note that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in the present invention. The methods and applications of the present invention have been described through preferred embodiments. It is obvious that relevant persons can modify or appropriately change and combine the methods and applications described herein without departing from the content and scope of the present invention to implement and apply the technology of the present invention.

[0034] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and the accompanying drawings. The experimental methods used in the following examples are all conventional methods unless otherwise specified, and the materials, reagents, methods and instruments used are all conventional materials, reagents, methods and instruments in the art unless otherwise specified, and can be obtained from commercial channels by those skilled in the art.

[0035] The English names involved in the following embodiments are:

[0036] Nuclear mitochondrial DNA fragments: NUMTs;

[0037] Cytochrome b: Cyt b.

[0038] Example 1: Detection of purebred sika deer and sika-red deer hybrids

[0039] Step 1: Ten blood samples from sika deer and eight blood samples from hybrid individuals of red deer and sika deer were obtained from a deer farm. DNA was extracted using the AxyPrep Blood Genomic DNA Miniprep Kit (Axygen Biosciences, USA). The extracted DNA was used to prepare libraries using the MGIEasy Universal DNA Library Prepared Kit (MGI, China). The libraries were then sequenced on a DNBSEQ-T1 sequencer. After quality control, clean data was obtained, which served as the resequencing reads for the target individuals.

[0040] Step 2: Download the Cyt b sequence of purebred sika deer species from NCBI as the seed for de novo assembly of the mitochondrial genome. Use NOVOplasty v4.3.1 to assemble the mitochondrial genome and obtain the mitochondrial reference genome.

[0041] Step 3: Identification of nuclear mitochondrial DNA fragments (NUMTs);

[0042] Step 3.1: Remove the reads used to assemble the mitochondrial genome from the resequencing data reads obtained in step 1 to generate a new fastq file. Use BWA0.7.17-r1188 to align the obtained fastq file with the mitochondrial reference genome assembled in step 2 to generate a bam file.

[0043] Step 3.2: Use samtools 1.17 to filter the bam files obtained in step 3.1 to remove low-quality reads (mapping quality < 30), unmapped reads, and secondary aligned reads to obtain potential NUMTs;

[0044] Step 3.3: Use bamtobed ​​of bedtools 2.30.0 to convert the potential NUMTs obtained in step 3.2 into BED format, cluster the overlapping reads within 1000 bp using Bedtools Cluster, and then convert the above sequences back to fastq format using BedtoolsBamtofastq to minimize redundant hits;

[0045] Step 3.4: Perform a BLAST comparison of the mitochondrial reference genome from step 2 with the fastq obtained in step 3.3 to mask targeted reads. Use the -dust no parameter to not mask low-complexity regions to avoid missing potentially useful information. A Python script extracts the highest-scoring BLAST matches, and the aforementioned NUMTs are aligned against each other to obtain the confirmed NUMTs.

[0046] Step 4: Align the NUMTs identified in step 3 with the mitochondrial reference genome in step 2 separately without filtering. Calculate the genetic distance between each NUMT and its mitochondrial homologous fragment using the Kimura 2-parameter model in MEGA x and plot it as a line graph. The results are then judged based on the line graph.

[0047] The horizontal axis of the line graph represents the genetic distance, and the vertical axis represents the frequency of each genetic distance interval. The genetic distance increases along the horizontal axis, such as Figure 1 As shown in , the frequency of the purebred sika deer genetic distance interval gradually decreases with the increase of the X-axis genetic distance interval, and finally tends to a stable state, indicating that the target individual (purebred sika deer) is purebred; Figure 2 As shown, the frequency of the genetic distance interval of the sika deer-red deer hybrid increases with the increase of the genetic distance interval on the X-axis. Two frequency peaks appear at the farther genetic distance interval. The presence of genomic fragments of another species or subspecies in the target individual indicates that the target individual (sika deer-red deer hybrid) is a hybrid. It can be seen that the hybridization detection method provided by the present invention using whole genome resequencing data alone is accurate for the hybridization detection results of purebred sika deer and sika deer-red deer hybrid. Example 2: Detection of purebred Siberian tiger and Siberian tiger-Bengal tiger hybrid

[0048] Step 1: Five Siberian tiger blood samples were obtained from the zoo. DNA was extracted using the AxyPrep Blood Genomic DNA Miniprep Kit (Axygen Biosciences, USA). The extracted DNA was used to prepare libraries using the MGIEasy Universal DNA Library Prepared Kit (MGI, China). The DNA was then sequenced on a DNBSEQ-T1 sequencer. After quality control, clean-data was obtained. The clean-data were the resequencing data reads of the target individuals. Then, resequencing data of seven Siberian tiger / Bengal tiger hybrid individuals were downloaded from NCBI. The NCBI numbers are NC_056660.1, NC_056661.1, NC_056662.1, NC_056663.1, NC_056664.1, NC_056665.1, and NC_056666.1, respectively.

[0049] Step 2: Download the Cyt b sequence of purebred sika deer species from NCBI as the seed for de novo assembly of the mitochondrial genome. Use NOVOplasty v4.3.1 to assemble the mitochondrial genome and obtain the mitochondrial reference genome.

[0050] Step 3: Identification of nuclear mitochondrial DNA fragments (NUMTs);

[0051] Step 3.1: Remove the reads used to assemble the mitochondrial genome from the resequencing data reads obtained in step 1 to generate a new fastq file. Use BWA0.7.17-r1188 to align the obtained fastq file with the mitochondrial reference genome assembled in step 2 to generate a bam file and obtain the aligned bam file;

[0052] Step 3.2: Use samtools 1.17 to filter the bam files obtained in step 3.1 to remove low-quality reads (mapping quality < 30), unmapped reads, and secondary aligned reads to obtain potential NUMTs;

[0053] Step 3.3: Use bamtobed ​​of bedtools 2.30.0 to convert the potential NUMTs obtained in step 3.2 into BED format, cluster the overlapping reads within 1000 bp using Bedtools Cluster, and then convert the above sequences back to fastq format using BedtoolsBamtofastq to minimize redundant hits;

[0054] Step 3.4: Perform a BLAST comparison of the mitochondrial reference genome from step 2 with the fastq obtained in step 3.3 to mask targeted reads. Use the -dust no parameter to not mask low-complexity regions to avoid missing potentially useful information. A Python script extracts the highest-scoring BLAST matches, and the aforementioned NUMTs are aligned against each other to obtain the confirmed NUMTs.

[0055] Step 4: Align the NUMTs identified in step 3 with the mitochondrial reference genome in step 2 separately without filtering. Calculate the genetic distance between each NUMT and its mitochondrial homologous fragment using the Kimura 2-parameter model in MEGA x and plot it as a line graph. The results are then judged based on the line graph.

[0056] like Figure 3 As shown in Figure 2, the frequency of the purebred Siberian tiger genetic distance interval gradually decreases with the increase of the X-axis genetic distance interval, and finally tends to a stable state, indicating that the target individual (purebred Siberian tiger) is purebred; Figure 4 As shown, the frequency of the genetic distance interval of the Siberian tiger and Bengal tiger hybrid increases with the increase of the X-axis genetic distance interval. The presence of genomic fragments of another species or subspecies in the target individual indicates that the target individual (Siberian tiger and Bengal tiger hybrid) is a hybrid. Therefore, the method for detecting hybrid individuals provided by the present invention is accurate in detecting hybridization between purebred Siberian tigers and Siberian tiger and Bengal tiger hybrids.

[0057] Example 3: Detection of purebred Golden Pheasants and hybrids of Golden Pheasants and White-bellied Pheasants

[0058] Step 1: Two purebred Golden Pheasant blood gauze samples and three Golden Pheasant / White Pheasant hybrid blood gauze samples were obtained from the farm. DNA was extracted using the AxyPrep Blood Genomic DNA Miniprep Kit (Axygen Biosciences, USA). The extracted DNA was used to prepare the library using the MGIEasy Universal DNA Library Prepared Kit (MGI, China). The library was then sequenced using the sequencing kit on the DNBSEQ-T1 sequencer. After quality control, clean-data was obtained. The above clean-data is the resequencing data reads of the target individual;

[0059] Step 2: Download the Cyt b sequence of purebred sika deer species from NCBI as the seed for de novo assembly of the mitochondrial genome. Use NOVOplasty v4.3.1 to assemble the mitochondrial genome and obtain the mitochondrial reference genome.

[0060] Step 3: Identification of nuclear mitochondrial DNA fragments (NUMTs);

[0061] Step 3.1: Remove the reads used to assemble the mitochondrial genome from the resequencing data reads obtained in step 1 to generate a new fastq file. Use BWA0.7.17-r1188 to align the obtained fastq file with the mitochondrial reference genome assembled in step 2 to generate a bam file.

[0062] Step 3.2: Use samtools 1.17 to filter the bam files obtained in step 3.1 to remove low-quality reads (mapping quality < 30), unmapped reads, and secondary aligned reads to obtain potential NUMTs;

[0063] Step 3.3: Use bamtobed ​​of bedtools 2.30.0 to convert the potential NUMTs obtained in step 3.2 into BED format, cluster the overlapping reads within 1000 bp using Bedtools Cluster, and then convert the above sequences back to fastq format using BedtoolsBamtofastq to minimize redundant hits;

[0064] Step 3.4: Perform a BLAST comparison of the mitochondrial reference genome from step 2 with the fastq obtained in step 3.3 to mask targeted reads. Use the -dust no parameter to not mask low-complexity regions to avoid missing potentially useful information. A Python script extracts the highest-scoring BLAST matches, and the aforementioned NUMTs are aligned against each other to obtain the confirmed NUMTs.

[0065] Step 4: Align the NUMTs identified in step 3 with the mitochondrial reference genome in step 2 separately without filtering. Calculate the genetic distance between each NUMT and its mitochondrial homologous fragment using the Kimura 2-parameter model in MEGA x and plot it as a line graph. The results are then judged based on the line graph.

[0066] like Figure 5 As shown in Figure 2, the frequency of the purebred Golden Pheasant genetic distance interval gradually decreases with the increase of the X-axis genetic distance interval, and finally tends to a stable state, indicating that the target individual (purebred Golden Pheasant) is purebred; Figure 6 As shown, the frequency of the genetic distance intervals of the Golden Pheasant-White Pheasant hybrid increases with the increase of the genetic distance interval on the X-axis, with four frequency peaks appearing at farther genetic distance intervals. The presence of genomic fragments of another species or subspecies in the target individual indicates that the target individual (the Golden Pheasant-White Pheasant hybrid) is a hybrid. This shows that the method for detecting hybrid individuals provided by the present invention is accurate for hybridization detection between purebred Golden Pheasants and Golden Pheasants-White Pheasants hybrids.

[0067] Example 4: Detection of purebred green turtles and green-loggerhead hybrids

[0068] Step 1: Download the resequencing data of 8 purebred green turtles and 7 green turtle-loggerhead hybrids from NCBI. The NCBI accessions are SRX21466357, SRX21466356, SRX21466348, SRX21466347, SRX8673235, SRX8673234, SRX8673232, SRX8673227, NC_057849.1, NC_057850.1, NC_057851.1, NC_057852.1, NC_051245.2, NC_051246.2, and NC_057853.1 respectively.

[0069] Step 2: Download the Cyt b sequence of purebred sika deer species from NCBI as the seed for de novo assembly of the mitochondrial genome. Use NOVOplasty v4.3.1 to assemble the mitochondrial genome and obtain the mitochondrial reference genome.

[0070] Step 3: Identification of nuclear mitochondrial DNA fragments (NUMTs);

[0071] Step 3.1: Remove the reads used to assemble the mitochondrial genome from the resequencing data reads obtained in step 1 to generate a new fastq file. Use BWA0.7.17-r1188 to align the obtained fastq file with the mitochondrial reference genome assembled in step 2 to generate a bam file.

[0072] Step 3.2: Use samtools 1.17 to filter the bam files obtained in step 3.1 to remove low-quality reads (mapping quality < 30), unmapped reads, and secondary aligned reads to obtain potential NUMTs;

[0073] Step 3.3: Use bamtobed ​​of bedtools 2.30.0 to convert the potential NUMTs obtained in step 3.2 into BED format, cluster the overlapping reads within 1000 bp using Bedtools Cluster, and then convert the above sequences back to fastq format using BedtoolsBamtofastq to minimize redundant hits;

[0074] Step 3.4: Perform a BLAST comparison of the mitochondrial reference genome from step 2 with the fastq obtained in step 3.3 to mask targeted reads. Use the -dust no parameter to not mask low-complexity regions to avoid missing potentially useful information. A Python script extracts the highest-scoring BLAST matches, and the aforementioned NUMTs are aligned against each other to obtain the confirmed NUMTs.

[0075] Step 4: Align the NUMTs identified in step 3 with the mitochondrial reference genome in step 2 separately without filtering. Calculate the genetic distance between each NUMT and its mitochondrial homologous fragment using the Kimura 2-parameter model in MEGAx and plot it as a line graph. The results are then judged based on the line graph.

[0076] like Figure 7 As shown in Figure 2, the frequency of the genetic distance interval of purebred green turtles gradually decreases with the increase of the genetic distance interval on the X axis, and finally tends to a stable state, indicating that the target individual (purebred green turtle) is purebred; Figure 8 As shown, the frequency of the genetic distance intervals of the green turtle-loggerhead hybrid increases with the increase of the genetic distance interval on the X-axis, with two frequency peaks appearing at the farther genetic distance intervals. The presence of genomic fragments of another species or subspecies in the target individual indicates that the target individual (green turtle-loggerhead hybrid) is a hybrid. Therefore, the method for detecting hybrid individuals provided by the present invention is accurate for hybridization detection of purebred green turtles and green turtle-loggerhead hybrids.

[0077] Example 5: Detection of Purebred Blue Catfish and Blue Catfish-Channel Catfish Hybrids

[0078] Step 1: Download the resequencing data of 9 purebred blue catfish and 8 blue catfish-channel catfish hybrids from NCBI. The NCBI accessions are SRX19313580, SRX19313577, SRX19313578, SRX19313576, SRX19313575, SRX19313574, SRX19313573, SRX19313572, SRX19313571, NC_071255.1, NC_071256.1, NC_071257.1, NC_071258.1, NC_071259.1, NC_071260.1, NC_071261.1, and NC_071262.1 respectively.

[0079] Step 2: Download the Cyt b sequence of purebred sika deer species from NCBI as the seed for de novo assembly of the mitochondrial genome. Use NOVOplasty v4.3.1 to assemble the mitochondrial genome and obtain the mitochondrial reference genome.

[0080] Step 3: Identification of nuclear mitochondrial DNA fragments (NUMTs);

[0081] Step 3.1: Remove the reads used to assemble the mitochondrial genome from the resequencing data reads obtained in step 1 to generate a new fastq file. Use BWA0.7.17-r1188 to align the obtained fastq file with the mitochondrial reference genome assembled in step 2 to generate a bam file.

[0082] Step 3.2: Use samtools 1.17 to filter the bam files obtained in step 3.1 to remove low-quality reads (mapping quality < 30), unmapped reads, and secondary aligned reads to obtain potential NUMTs;

[0083] Step 3.3: Use bamtobed ​​of bedtools 2.30.0 to convert the potential NUMTs obtained in step 3.2 into BED format, cluster the overlapping reads within 1000 bp using Bedtools Cluster, and then convert the above sequences back to fastq format using BedtoolsBamtofastq to minimize redundant hits;

[0084] Step 3.4: Perform a BLAST comparison of the mitochondrial reference genome from step 2 with the fastq obtained in step 3.3 to mask targeted reads. Use the -dust no parameter to not mask low-complexity regions to avoid missing potentially useful information. A Python script extracts the highest-scoring BLAST matches, and the aforementioned NUMTs are aligned against each other to obtain the confirmed NUMTs.

[0085] Step 4: Align the NUMTs identified in step 3 with the mitochondrial reference genome in step 2 separately without filtering. Calculate the genetic distance between each NUMT and its mitochondrial homologous fragment using the Kimura 2-parameter model in MEGAx and plot it as a line graph. The results are then judged based on the line graph.

[0086] like Figure 9 As shown in Figure 2, the frequency of the purebred green turtle genetic distance interval gradually decreases with the increase of the X-axis genetic distance interval, and finally tends to a stable state, indicating that the target individual (purebred blue catfish) is purebred; Figure 10 As shown, the frequency of the genetic distance interval of the green turtle and loggerhead hybrid increases with the increase of the genetic distance interval on the X-axis, with two frequency peaks appearing at the farther genetic distance interval. The presence of genomic fragments of another species or subspecies in the target individual indicates that the target individual (blue catfish and channel catfish hybrid) is a hybrid. It can be seen that the method for detecting hybrid individuals provided by the present invention is accurate for hybridization detection results of purebred blue catfish and blue catfish and channel catfish hybrids.

[0087] The contents not described in detail in the present specification are well-known technologies to those skilled in the art. Although the present invention has been disclosed above with reference to preferred embodiments, they are not intended to limit the present invention. Anyone skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be defined by the claims.

Claims

1. A method for hybridization detection using whole genome resequencing data alone, characterized in that: The detection method specifically comprises the following steps: Step 1: Collect samples from the target individual, transport them on dry ice and store them below -20°C. Use a kit to extract DNA from the samples. Construct a DNBSEQ whole-genome resequencing library for sequencing to obtain resequencing data for the target individual. Step 2: Download the cytochrome b sequence of the species to which the target individual belongs from NCBI as the seed for de novo assembly of the mitochondrial genome. Use NOVOplasty v4.3.1 to assemble the mitochondrial genome and obtain the mitochondrial reference genome. Step 3: Identification of nuclear mitochondrial DNA fragments; Step 3.1: Remove the reads used to assemble the mitochondrial genome from the resequencing data reads obtained in step 1 to generate a new fastq file. Use BWA0.7.17-r1188 to align the obtained fastq file with the mitochondrial reference genome assembled in step 2 to generate a bam file and obtain the aligned bam file; Step 3.2: Use samtools 1.17 to filter the bam file obtained in step 3.1 to obtain potential nuclear mitochondrial DNA sequences; Step 3.3: Use bamtobed ​​of bedtools 2.30.0 to convert the potential nuclear mitochondrial DNA sequence obtained in step 3.2 into BED format, cluster the overlapping reads within 1000 bp using Bedtools Cluster, and then convert the above sequence back to fastq format using Bedtools Bamtofastq; Step 3.4: Perform BLAST alignment of the mitochondrial reference genome in step 2 with the fastq obtained in step 3.

3. Use the -dust no parameter to not mask low-complexity regions. The Python script extracts the highest-scoring BLAST matches. The above nuclear mitochondrial DNA sequences are aligned with each other to obtain the confirmed nuclear mitochondrial DNA sequence. Step 4: Align the nuclear mitochondrial DNA sequence determined in step 3 with the mitochondrial reference genome in step 2 separately without filtering, calculate the genetic distance between each nuclear mitochondrial DNA sequence and its mitochondrial homologous fragment; and draw a line graph, and judge the results based on the line graph.

2. The method according to claim 1, characterized in that When the sample described in step 1 is a blood or tissue sample, whole genome resequencing can be performed directly after DNA extraction, and the depth of whole genome resequencing is >30X.

3. The method according to claim 1, characterized in that The resequencing data reads described in step 1 are clean-data obtained after quality control.

4. The method according to claim 1, wherein The genetic distance described in step 4 was calculated using the Kimura 2-parameter model in MEGAx.

5. The method according to claim 1, wherein The horizontal axis of the line graph in step 4 represents the genetic distance, and the vertical axis represents the frequency of each genetic distance interval.

6. The method according to claim 1, characterized in that The criteria for judging the results in step 4 are: if the frequency of the genetic distance interval decreases steadily as the X-axis increases, it indicates that the target individual is a purebred; if the frequency of the genetic distance interval reaches a peak as the X-axis increases, it indicates that the target individual is a hybrid.

Citation Information

Patent Citations

  • Method for identifying hybrid germplasm of actinidia based on genome heterozygosity

    CN104630382A

  • Cucumber mitochondrial genome SSR marker development and application of marker in seed purity identification

    CN105886602A