A method for drawing a giant panda genetic map

By using SNP+SV dual-type genetic marker targeted mining and an improved genetic map construction algorithm, combined with three-level sample screening and redundancy removal, the problem of insufficient accuracy and coverage of giant panda genetic maps was solved, and high-precision genetic map drawing was achieved.

CN122337355APending Publication Date: 2026-07-03SHAANXI INST OF ZOOLOGY NORTHWEST INSTOF ENDANGERED ZOOLOGICAL SPECIES
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHAANXI INST OF ZOOLOGY NORTHWEST INSTOF ENDANGERED ZOOLOGICAL SPECIES
Filing Date
2026-04-01
Publication Date
2026-07-03

Smart Images

  • Figure CN122337355A_ABST
    Figure CN122337355A_ABST
Patent Text Reader

Abstract

This invention discloses a method for constructing a giant panda genetic map, belonging to the field of genetic map construction technology. The method includes the following steps: retrieving DNA samples, performing sequencing processing to obtain a sequencing BAM file; performing SNP+SV dual-type genetic marker targeted mining and marker hierarchical filtering; using an improved partial least squares regression algorithm to correct errors in the marker filtering genotype matrix and label pedigree relationships; using an improved minimum spanning tree-based genetic map construction algorithm to construct linkage groups and calculate preliminary genetic distances between markers; correcting the preliminary genetic distances between markers to construct the giant panda genetic map; and performing genome anchoring and functional annotation to obtain the final giant panda genetic map. This application improves the accuracy, coverage, and species suitability of the giant panda genetic map by combining SNP+SV dual-type genetic marker targeted mining, an improved error correction algorithm, and an improved genetic map construction algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of genetic mapping technology, and in particular to a method for mapping the genetic map of a giant panda. Background Technology

[0002] In the fields of animal molecular genetics and conservation genomics, the giant panda, as a flagship species, relies on its unique genetic map as a core foundational tool for functional gene research and population genetic management. General mammalian genetic mapping methods are ill-suited to the characteristics of the giant panda genome, including low heterozygosity, a high proportion of repetitive sequences, and high inbreeding rates. These methods suffer from technical shortcomings such as low marker mining efficiency, insufficient genotyping accuracy, and large deviations in genetic distance estimation, failing to meet the technical requirements of giant panda genomics research. Therefore, research is needed on methods for mapping the giant panda's genetic map.

[0003] In the prior art, Chinese patent CN117409868A discloses a method and system for drawing a giant panda genetic map. The method involves drawing a disease genetic map configuration scenario, which includes a genetic map data stream configuration subgraph and a genetic map node configuration subgraph. Based on the data stream configuration operations within the genetic map data stream configuration subgraph, a custom genetic map data stream is determined. Based on the node input operations within the genetic map node configuration subgraph, input genetic map nodes are obtained. Based on the received genetic map sending operation and the fact that the giant panda to be tested has a disease genetic map potential value, genetic map nodes are drawn in the family tree scenario according to the custom genetic map data stream. The disease genetic map potential value is obtained by removing a preset number of non-genetic diseases from the giant panda to be tested.

[0004] However, the aforementioned existing technologies only focus on the scenario-based configuration and visualization of giant panda disease genetic maps, and do not optimize the core algorithms for marker mining and map construction for the unique genomic characteristics of giant pandas. The accuracy, coverage and species adaptability of giant panda genetic maps need to be improved. Summary of the Invention

[0005] This application provides a method for drawing a giant panda genetic map, which addresses the problem that the accuracy, coverage, and species adaptability of existing giant panda genetic maps need to be improved.

[0006] On the one hand, this application provides a method for drawing a genetic map of giant pandas, including the following steps: Step 1: Directly retrieve DNA samples from the giant panda genetic sample bank, perform sequencing processing, and obtain sequencing BAM files.

[0007] Step 2: Perform SNP+SV dual-type genetic marker targeting mining on the sequencing BAM files of all samples to obtain an initial genotyping matrix. Then, perform marker-based hierarchical filtering on the initial genotyping matrix to obtain a marker-filtered genotyping matrix.

[0008] Step 3: The improved partial least squares regression algorithm is used to correct the label filtering fractal matrix to obtain the error-corrected fractal matrix. The kinship relationship is labeled on the error-corrected fractal matrix to obtain the fractal matrix with kinship relationship.

[0009] Step four: Using an improved genetic map construction algorithm based on minimum spanning tree, linkage groups adapted to dual-type markers of giant pandas are constructed according to the kinship genotyping matrix, and the preliminary genetic distance between markers within each linkage group is calculated.

[0010] Step 5: Correct the preliminary genetic distance between markers to obtain the corrected genetic distance between markers, and construct a giant panda genetic map based on the corrected genetic distance between markers for each linkage group.

[0011] Step six: Perform genome anchoring and functional annotation on the giant panda genetic map to obtain the final giant panda genetic map.

[0012] In one possible implementation, in step one, a three-level targeted screening strategy is used to selectively retrieve DNA samples from a giant panda genetic sample bank. The three-level targeted screening strategy includes: DNA samples were retrieved from the Giant Panda Genetic Sample Bank, consisting of three levels: core family samples, half-sib supplementary samples, and wild distant population samples.

[0013] In one possible implementation, in step one, a k-mer weighted redundancy removal algorithm is used to remove redundancy from the sequencing data obtained from the sequencing process, resulting in a sequencing BAM file.

[0014] In one possible implementation, step two, the SNP+SV dual-type genetic marker targeting mining, includes: SNP marker mining and SV marker mining were performed on the sequencing BAM files of all samples to obtain SNP marker genotyping files and SV marker genotyping files.

[0015] The SNP marker genotyping file and the SV marker genotyping file are merged to construct an initial genotyping matrix. The rows in the initial genotyping matrix represent all detected markers, the columns represent all samples, and the matrix elements are the genotypes of the corresponding samples at the marker loci.

[0016] In one possible implementation, step two, the tag-based hierarchical filtering, includes: The markers in the initial genotyping matrix are initially filtered based on the genotyping loss rate and Mendelian genetic error rate.

[0017] A comprehensive score for each marker is calculated by weighting the minimum allele frequency, observed heterozygosity, genotyping loss rate, and Mendelian inheritance error rate. The markers after the initial filtering are then filtered a second time based on the comprehensive score.

[0018] Chained unbalanced filtering is performed on the tags after the second filtering to generate a tag filtering fractal matrix.

[0019] In one possible implementation, in step three, Mendelian genetic constraints and inbreeding coefficient correction terms are added to the objective function of the partial least squares regression algorithm to improve the algorithm.

[0020] In one possible implementation, in step four, a type weighting factor is added to the recombination rate calculation of the genetic map construction algorithm based on minimum spanning tree, the corrected recombination rate between all pairs of markers is calculated, and the corrected recombination rate full matrix is ​​constructed.

[0021] Based on the corrected recombination rate matrix, a minimum spanning tree of all tags is generated using the minimum spanning tree algorithm. Linkage groups are partitioned using a dynamic threshold, and a linkage group adapted to the dual-type tags of giant pandas is constructed.

[0022] After linearly orienting the markers for each linkage group, the preliminary genetic distances between markers within each linkage group are calculated.

[0023] In one possible implementation, in step five, an inbreeding coefficient correction term and a marker type bias correction term are introduced to correct the preliminary genetic distance between the markers, resulting in the corrected genetic distance between the markers. Based on the corrected genetic distance between the markers for each linkage group, a giant panda genetic map is constructed.

[0024] In one possible implementation, in step six, the giant panda genetic map is subjected to Mendelian genetic consistency verification, collinearity verification of genetic distance and physical distance, and map resolution and coverage verification. The giant panda genetic map that has passed the verification is then subjected to genome anchoring and functional annotation to obtain the final giant panda genetic map.

[0025] The method for drawing a genetic map of giant pandas in this application has the following advantages: By combining SNP+SV dual-type genetic marker targeted mining, improved error correction algorithms, and improved genetic map construction algorithms, the accuracy, coverage, and species suitability of the giant panda genetic map were improved.

[0026] By employing a three-tiered targeted screening strategy involving core family samples, half-sib supplementary samples, and wild distant relatives, the genetic representativeness of the samples was improved, and the problem of insufficient marker polymorphism in giant pandas was solved.

[0027] By employing a k-mer weighted redundancy removal algorithm to remove redundancy from the sequencing data, the false deletion rate of effective data in highly repetitive regions of giant pandas was reduced, thereby improving the utilization rate of sequencing data and the accuracy of genotyping.

[0028] By performing SNP marker mining and SV marker mining on the sequencing BAM files of all samples, an initial genotyping matrix was constructed, which improved the genomic coverage of the markers and filled the coverage gap of single-type markers.

[0029] By using three levels of marker-based filtering, the polymorphism and typing accuracy of genetic markers are improved, and low-quality, false-positive markers are effectively eliminated.

[0030] By adding Mendelian genetic constraints and inbreeding coefficient correction terms to the objective function of the partial least squares regression algorithm, the accuracy of giant panda typing error correction was improved, and the typing error rate and missing rate were significantly reduced.

[0031] By incorporating a type weight factor into the recombination rate calculation of the minimum spanning tree-based genetic map construction algorithm and using a dynamic threshold for linkage group division, the accuracy of linkage grouping and marker sorting precision of giant pandas were improved.

[0032] By introducing inbreeding coefficient correction terms and marker type bias correction terms to correct the initial genetic distance between markers, the distance compression problem caused by inbreeding is solved, and the accuracy of the genetic map is improved.

[0033] By combining triple precision validation with genome anchoring and functional annotation, the reliability and practicality of the giant panda genetic map have been improved, providing precise targets for subsequent research. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a flowchart illustrating a method for drawing a genetic map of a giant panda, as provided in an embodiment of this application. Detailed Implementation

[0036] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0037] like Figure 1 As shown in the embodiment of this application, a method for drawing a genetic map of a giant panda is provided, including the following steps: Step 1: Directly retrieve DNA samples from the giant panda genetic sample bank, perform sequencing processing, and obtain sequencing BAM files.

[0038] Step 2: Perform SNP+SV dual-type genetic marker targeting mining on the sequencing BAM files of all samples to obtain an initial genotyping matrix. Then, perform marker-based hierarchical filtering on the initial genotyping matrix to obtain a marker-filtered genotyping matrix.

[0039] Step 3: The improved partial least squares regression algorithm is used to correct the label filtering fractal matrix to obtain the error-corrected fractal matrix. The kinship relationship is labeled on the error-corrected fractal matrix to obtain the fractal matrix with kinship relationship.

[0040] Step four: Using an improved genetic map construction algorithm based on minimum spanning tree, linkage groups adapted to dual-type markers of giant pandas are constructed according to the kinship genotyping matrix, and the preliminary genetic distance between markers within each linkage group is calculated.

[0041] Step 5: Correct the preliminary genetic distance between markers to obtain the corrected genetic distance between markers, and construct a giant panda genetic map based on the corrected genetic distance between markers for each linkage group.

[0042] Step six: Perform genome anchoring and functional annotation on the giant panda genetic map to obtain the final giant panda genetic map.

[0043] For example, in step one, a three-level targeted screening strategy is used to selectively retrieve DNA samples from the giant panda genetic sample bank. The three-level targeted screening strategy includes: DNA samples were retrieved from the Giant Panda Genetic Sample Bank, consisting of three levels: core family samples, half-sib supplementary samples, and wild distant population samples.

[0044] Specifically, in this embodiment, 5 non-related captive giant panda F1 core family samples were retrieved定向调取 from a standardized and traceable genetic sample library dedicated to giant pandas. Each core family contains complete genomic DNA samples of the sire, dam, and ≥8 offspring; 7 half-sibling supplementary samples (half-siblings with different mothers / fathers) that have no genetic relationship with the core family were retrieved继续调取, and each half-sibling supplementary sample contains ≥5 offspring DNA samples; 10 wild individual DNA samples (i.e., wild outbred population samples) with no inbreeding records and no genetic associations were retrieved继续调取.

[0045] Exemplarily, in step one, a k-mer distribution weighted redundancy removal algorithm was used to perform redundancy removal on the sequencing data obtained from the sequencing process, and a sequencing BAM file was obtained.

[0046] Specifically, in this embodiment, after the retrieved DNA samples passed the quality inspection, the qualified DNA samples were uniformly subjected to PE150 high-throughput sequencing using the Illumina NovaSeq 6000 sequencing platform. The sequencing depth of each sample was not less than 10×, and the sequencing depth of the core family parental samples was not less than 20× to ensure the sequencing coverage of low heterozygosity regions.

[0047] The Trimmomatic software was used to perform quality control on the original off-machine data, removing sequencing adapters, bases with terminal quality value Q<20, and short sequences with a length less than 50bp, to obtain the purified data after preliminary quality control.

[0048] The purified data after preliminary quality control was aligned to the latest reference genome of the giant panda using the BWA-MEM algorithm, with the alignment similarity set to ≥95% and the maximum number of mismatches set to ≤3, to obtain the initial SAM alignment file, which was converted to a sorted BAM file by Samtools.

[0049] Redundancy removal was performed on the sorted BAM file (i.e., the sequencing data): the 17-mer frequency of all alignment sequences in the sorted BAM file was counted to obtain the whole-genome k-mer distribution frequency matrix F, where F(i) is the 17-mer occurrence frequency of the i-th locus in the genome; for each alignment sequence, its weighted repeat score W was calculated, and the formula is: W = L×(1 - 1 / mean(F(s))), where L is the sequence alignment length and mean(F(s)) is the average k-mer frequency of all genomic loci covered by this sequence; the weighted deduplication threshold T = ---0.2 (preliminary experiments verified that this threshold can balance the deduplication effect and the retention rate of valid data) was set. When the alignment positions and insertion fragment lengths of two alignment sequences are exactly the same, and W<T, they are determined to be PCR repeat sequences and removed; if W≥T, they are determined to be real alignment sequences in high-repeat regions and retained. The sequencing BAM file was obtained.

[0050] For example, in step two, the SNP+SV dual-type genetic marker targeting mining includes: SNP marker mining and SV marker mining were performed on the sequencing BAM files of all samples to obtain SNP marker genotyping files and SV marker genotyping files.

[0051] The SNP marker genotyping file and the SV marker genotyping file are merged to construct an initial genotyping matrix. The rows in the initial genotyping matrix represent all detected markers, the columns represent all samples, and the matrix elements are the genotypes of the corresponding samples at the marker loci.

[0052] Specifically, SNP marker mining includes: using the GATK HaplotypeCaller tool to perform population-level SNP mining on the BAM files of all samples, with mining parameters set as follows: minimum allele depth ≥ 3, minimum genotype quality value (GQ) ≥ 20, and minimum alignment quality (MQ) ≥ 30. This yields SNP marker genotyping files.

[0053] SV marker mining includes: using two tools, Delly and Lumpy, to perform SV mining, covering four types: missing (DEL), duplicate (DUP), inversion (INV), and translocation (TRA). The SV length range is set to 50bp-1Mb. SVs with lengths exceeding the range are filtered out. SV markers detected by both tools are retained to obtain the SV marker classification file.

[0054] The SNP marker genotyping file and the SV marker genotyping file are merged to construct an initial genotyping matrix. The rows in the initial genotyping matrix represent all detected markers, the columns represent all samples, and the matrix elements are the genotypes of the corresponding samples at the marker loci (0 / 0 homozygous reference, 0 / 1 heterozygous, 1 / 1 homozygous variant).

[0055] For example, in step two, the tag-based hierarchical filtering includes: The markers in the initial genotyping matrix are initially filtered based on the genotyping loss rate and Mendelian genetic error rate.

[0056] A comprehensive score for each marker is calculated by weighting the minimum allele frequency, observed heterozygosity, genotyping loss rate, and Mendelian inheritance error rate. The markers after the initial filtering are then filtered a second time based on the comprehensive score.

[0057] Chained unbalanced filtering is performed on the tags after the second filtering to generate a tag filtering fractal matrix.

[0058] Specifically, in this embodiment, the preliminary filtering includes filtering out markers with a typing loss rate greater than 20% and a Mendelian genetic error rate greater than 5%.

[0059] The second filtering involves designing a comprehensive scoring formula based on minimum allele frequency, observed heterozygosity, genotyping loss rate, and Mendelian error rate, as follows: S = α × MAF + β × H o +γ×(1-F miss )-δ×F me .

[0060] Where S represents the overall score, MAF represents the minimum allele frequency of the marker, and H o F represents the observed heterozygosity of the label. miss F represents the typing loss rate of the markers. me The Mendelian error rate of the marker is represented by α, β, γ, and δ, which represent the weight coefficients of the corresponding terms. In this embodiment, α=0.3, β=0.3, γ=0.2, and δ=0.2. A comprehensive score is calculated for each marker based on the comprehensive scoring formula, and markers with a comprehensive score less than 0.15 are filtered out.

[0061] After the second filtering, the labels were subjected to chain imbalance filtering using PLINK software, including: for the coefficient of determination r 2 For two tags with a score greater than 0.8, retain the tag with the higher overall score. The final result is a tag filtering fractal matrix without redundancy.

[0062] For example, in step three, Mendelian genetic constraint terms and inbreeding coefficient correction terms are added to the objective function of the partial least squares regression algorithm to improve the partial least squares regression algorithm.

[0063] Specifically, in this embodiment, constructing and training an error correction model based on an improved partial least squares regression algorithm includes: Set the model input matrix: the independent variable matrix X is the label filtering fractal matrix obtained in step two, with dimensions n×m, where n is the number of labels and m is the number of samples. Matrix elements X ij Let Y be the genotype encoding of the i-th marker in the j-th sample; the dependent variable matrix Y is the theoretical genotype matrix of the core family, with dimensions n×m0, where m0 is the number of samples in the core family, and the matrix elements Y ij The theoretical genotype encoding of the i-th marker in the j-th core family sample, derived based on Mendelian inheritance laws.

[0064] Set the objective function to include Mendelian genetic constraints and inbreeding coefficient correction terms: J = max[cov(T,U)] - λ × ||X × WY|| 2 -μ×F inbreed .

[0065] Where J represents the objective function value of the error correction model, and the training objective of the error correction model is to maximize J; cov(T,U) represents the covariance between the score matrix T of the independent variable matrix X and the score matrix U of the dependent variable matrix Y; W represents the projected weight matrix of the independent variable matrix X; ||X×WY|| 2 F represents the sum of squared residuals between the model's predicted genotypes and the theoretical genotypes of the core family, i.e., the Mendelian genetic constraint term, where λ represents the weighting coefficient of the constraint term; inbreed The average insection coefficient of the sample is represented by μ, which is the insection coefficient correction term. μ represents the weight coefficient of the correction term.

[0066] The improved partial least squares regression algorithm model was trained using 10-fold cross-validation. The optimal projection weight matrix W was iteratively solved using the core pedigree theory genotype matrix as the gold standard. The iteration stopped when the model prediction error rate was lower than 0.1%, and the trained error correction model was obtained.

[0067] Input the label-filter fractal matrix obtained in step two into the trained error correction model, and project the label-filter fractal matrix onto the projection weight matrix W to obtain the error correction fractal matrix X. correct .

[0068] Based on the error-correcting classification matrix, Plink software was used to calculate the kinship coefficients among all samples. The calculation results were compared with the pedigree information recorded in the sample records to verify the accuracy of the kinship relationships: the kinship coefficients of the core family's parent-child generations should be between 0.45 and 0.55, the kinship coefficients of half-siblings should be between 0.2 and 0.3, and the kinship coefficients of full-siblings should be between 0.45 and 0.55. Samples whose calculation results did not match the pedigree information recorded in the sample records were removed. The kinship relationships were then labeled on the error-correcting classification matrix to obtain the kinship classification matrix.

[0069] For example, in step four, a type weighting factor is added to the recombination rate calculation of the genetic map construction algorithm based on minimum spanning tree, the corrected recombination rate between all pairs of markers is calculated, and the corrected recombination rate full matrix is ​​constructed.

[0070] Based on the corrected recombination rate matrix, a minimum spanning tree of all tags is generated using the minimum spanning tree algorithm. Linkage groups are partitioned using a dynamic threshold, and a linkage group adapted to the dual-type tags of giant pandas is constructed.

[0071] After linearly orienting the markers for each linkage group, the preliminary genetic distances between markers within each linkage group are calculated.

[0072] Specifically, in this embodiment, the formula for calculating the corrected recombination rate is set as follows: r ij =w i ×w j ×(Nrecom_ij / N total_ij ).

[0073] Where, r ij w represents the corrected recombination rate between marker i and marker j. i and w j Let N represent the type weight factors for marker i and marker j, respectively. In this embodiment, let the type weight factor corresponding to the SNP marker be 1, and the type weight factor corresponding to the SV marker be 0.95. recom_ij N represents the total number of recombinant gametes in the entire population between marker i and marker j (obtained by counting the number of gametes that recombine between two markers in the phylogenetic matrix with kinship), total_ij This represents the total number of effective gametes in the entire population between marker i and marker j (the total number of effective gametes that can be used for statistics between two markers in the phylogenetic matrix with kinship).

[0074] Calculate the corrected recombination rate between all pairs of tags, and construct the corrected recombination rate matrix D, where the matrix elements are the corrected recombination rates r between tags i and j. ij .

[0075] Using all tags as nodes, and the corrected recombination rate r between tags... ij Using edge weights, construct a fully labeled undirected weighted graph. Based on the minimum spanning tree algorithm, generate a fully labeled minimum spanning tree from the undirected weighted graph, cutting edges with weights greater than the initial dynamic threshold to obtain initial linkage groups. The initial dynamic threshold is set to 0.3, and the first round of linkage group partitioning is performed: labels with a corrected recombination rate less than 0.3 are determined to have significant linkage relationships and are assigned to the same linkage group.

[0076] If the initial number of linkage groups is greater than 21, the dynamic threshold is gradually increased (in this embodiment, it is increased by 0.01 each time) to reduce the number of linkage groups.

[0077] If the initial number of multiple linkage groups is less than 21, the dynamic threshold is gradually reduced (in this embodiment, it is reduced by 0.01 each time) to split up the excessively merged linkage groups.

[0078] When the number of linkage groups stabilizes at 21, and the distribution of marker counts in each linkage group is positively correlated with the physical length of the corresponding chromosome in the giant panda reference genome, the partitioning is terminated. This ensures that the final 21 linkage groups correspond to the 21 pairs of chromosomes of the giant panda, thus obtaining linkage groups adapted to the dual-type markers of the giant panda.

[0079] For a single linkage group, a minimum spanning tree of the labels within the group is constructed based on the undirected weighted graph specific to the group. Based on the path relationships of the minimum spanning tree, the initial relative linear order between the labels is determined, resulting in the initial sorting graph of the labels for the linkage group.

[0080] The marker order of the initial marker ordination map is iteratively optimized based on pedigree genotype data in the kinship typology matrix. A likelihood optimization strategy is used, where the position of one marker is randomly adjusted in each iteration, and the total likelihood value of the adjusted marker order ordination map is calculated. The marker order with the highest total likelihood value is retained. When the change in the total likelihood value of the map is less than 1e-6, the resulting marker order is the optimal linear ordination for that linkage group.

[0081] For linkage groups after marker linear ordering, the Kosambi plotting function is used to convert the corrected recombination rate between adjacent markers into the preliminary genetic distance L. raw ,as follows: L raw =0.25×ln[(1+2r ij ) / (1-2r ij )).

[0082] For example, in step five, an inbreeding coefficient correction term and a marker type deviation correction term are introduced to correct the preliminary genetic distance between the markers, and the corrected genetic distance between the markers is obtained. Based on the corrected genetic distance between the markers of each linkage group, a giant panda genetic map is constructed.

[0083] Specifically, in this embodiment, a genetic distance correction formula is designed that incorporates an inbreeding coefficient correction term and a marker type bias correction term, as follows: L correct =L raw ×(1+k×F inbreed )×(1+b×T marker ).

[0084] Among them, L correct Indicates the genetic distance between corrected markers; F inbreed T represents the average concordance coefficient of the samples, i.e., the concordance coefficient correction term, where k represents the concordance correction coefficient; marker This indicates the label type deviation correction item (in this embodiment, if both labels are SNP labels, then T). marker =0, if one of the tags is an SV tag, then T marker =0.03, if both tags are SV tags, then T marker =0.06), b represents the deviation correction coefficient.

[0085] For each linkage group, the corrected genetic distances between all adjacent markers are summed to obtain the cumulative genetic length of each linkage group.

[0086] For the 21 linkage groups, they were sorted from largest to smallest according to their cumulative genetic length and matched one-to-one with the chromosome number of the giant panda reference genome. A visual linear genetic map was constructed for each linkage group, with the horizontal axis representing the cumulative genetic length and the vertical axis representing the linkage group number. The position of the marker on the map corresponds one-to-one with the genetic distance between the corrected markers, thus obtaining the giant panda genetic map.

[0087] For example, in step six, the giant panda genetic map is verified for Mendelian genetic consistency, collinearity of genetic distance and physical distance, and map resolution and coverage. The giant panda genetic map that has passed the verification is then subjected to genome anchoring and functional annotation to obtain the final giant panda genetic map.

[0088] Specifically, in this embodiment, the Mendelian genetic consistency verification includes: calculating the genetic error rate of each marker in the giant panda genetic map. The verification criteria are: the average Mendelian genetic error rate of all markers in the map is ≤0.01%, and the error rate of a single marker is ≤0.1%. If the criteria are not met, return to step two to filter the corresponding markers.

[0089] The collinearity verification of genetic and physical distances includes: for each linkage group in the giant panda genetic map, comparing the marked genetic location in the map with the physical location of the corresponding reference genome, and drawing a collinearity scatter plot; using the Pearson correlation coefficient to calculate the correlation between the genetic location and the physical location of each linkage group, the verification criterion is: the Pearson correlation coefficient of all linkage groups is greater than 0.95. If the criterion is not met, return to step four to optimize the linear ordination of markers.

[0090] Map resolution and coverage validation includes: calculating the average marker spacing of the entire map, with validation criteria of average marker spacing ≤ 0.5 cM and maximum marker spacing ≤ 5 cM, ensuring high density and high resolution of the map; calculating the map's coverage of the giant panda reference genome, with validation criteria of genome coverage ≥ 98%, of which euchromatin region coverage ≥ 99.5%; and calculating the coverage of SV markers, with validation criteria of SV markers covering more than 90% of structural variation hotspots in the genome, filling the coverage gaps of conventional SNP maps. If the criteria are not met, return to step two to target and mine SNP+SV markers for the marker gap regions, adjust the marker selection threshold to supplement effective markers, optimize the marker set, and then re-execute the subsequent full process.

[0091] Genome anchoring involves: anchoring each of the 21 linkage groups of the giant panda genetic map to one of the 21 pairs of chromosomes of the giant panda based on the physical location of the markers on the reference genome, and determining the chromosome number corresponding to each linkage group.

[0092] Functional annotation includes: performing functional annotation on all markers in the map, using ANNOVAR software to align the markers with the giant panda functional gene set, annotating the gene regions (exons, introns, promoter regions, intergenic regions) where the markers are located, as well as the gene names and gene functions corresponding to the markers.

[0093] This application's embodiments improve the accuracy, coverage, and species suitability of the giant panda genetic map by combining SNP+SV dual-type genetic marker targeted mining, an improved error correction algorithm, and an improved genetic map construction algorithm.

[0094] By employing a three-tiered targeted screening strategy involving core family samples, half-sib supplementary samples, and wild distant relatives, the genetic representativeness of the samples was improved, and the problem of insufficient marker polymorphism in giant pandas was solved.

[0095] By employing a k-mer weighted redundancy removal algorithm to remove redundancy from the sequencing data, the false deletion rate of effective data in highly repetitive regions of giant pandas was reduced, thereby improving the utilization rate of sequencing data and the accuracy of genotyping.

[0096] By performing SNP marker mining and SV marker mining on the sequencing BAM files of all samples, an initial genotyping matrix was constructed, which improved the genomic coverage of the markers and filled the coverage gap of single-type markers.

[0097] By using three levels of marker-based filtering, the polymorphism and typing accuracy of genetic markers are improved, and low-quality, false-positive markers are effectively eliminated.

[0098] By adding Mendelian genetic constraints and inbreeding coefficient correction terms to the objective function of the partial least squares regression algorithm, the accuracy of giant panda typing error correction was improved, and the typing error rate and missing rate were significantly reduced.

[0099] By incorporating a type weight factor into the recombination rate calculation of the minimum spanning tree-based genetic map construction algorithm and using a dynamic threshold for linkage group division, the accuracy of linkage grouping and marker sorting precision of giant pandas were improved.

[0100] By introducing inbreeding coefficient correction terms and marker type bias correction terms to correct the initial genetic distance between markers, the distance compression problem caused by inbreeding is solved, and the accuracy of the genetic map is improved.

[0101] By combining triple precision validation with genome anchoring and functional annotation, the reliability and practicality of the giant panda genetic map have been improved, providing precise targets for subsequent research.

[0102] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0103] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for drawing a genetic map of giant pandas, characterized in that, Includes the following steps: Step 1: Directly retrieve DNA samples from the giant panda genetic sample bank, perform sequencing processing, and obtain sequencing BAM files; Step 2: Perform SNP+SV dual-type genetic marker targeting mining on the sequencing BAM files of all samples to obtain an initial genotyping matrix. Then, perform marker-based hierarchical filtering on the initial genotyping matrix to obtain a marker-filtered genotyping matrix. Step 3: The improved partial least squares regression algorithm is used to correct the label filtering fractal matrix to obtain the error-corrected fractal matrix. The kinship relationship is labeled on the error-corrected fractal matrix to obtain the fractal matrix with kinship relationship. Step 4: Using an improved genetic map construction algorithm based on minimum spanning tree, linkage groups adapted to dual-type markers of giant pandas are constructed according to the kinship genotyping matrix, and the preliminary genetic distance between markers within each linkage group is calculated. Step 5: Correct the preliminary genetic distance between markers to obtain the corrected genetic distance between markers, and construct a giant panda genetic map based on the corrected genetic distance between markers for each linkage group; Step six: Perform genome anchoring and functional annotation on the giant panda genetic map to obtain the final giant panda genetic map.

2. The method for drawing a genetic map of a giant panda according to claim 1, characterized in that, In step one, a three-level targeted screening strategy is used to selectively retrieve DNA samples from the giant panda genetic sample bank. This three-level targeted screening strategy includes: DNA samples were retrieved from the Giant Panda Genetic Sample Bank, consisting of three levels: core family samples, half-sib supplementary samples, and wild distant population samples.

3. The method for drawing a genetic map of a giant panda according to claim 1, characterized in that, In step one, the k-mer weighted redundancy removal algorithm is used to remove redundancy from the sequencing data obtained from the sequencing process, resulting in a sequencing BAM file.

4. The method for drawing a genetic map of a giant panda according to claim 1, characterized in that, Step two, the SNP+SV dual-type genetic marker targeting mining includes: SNP marker mining and SV marker mining were performed on the sequencing BAM files of all samples to obtain SNP marker genotyping files and SV marker genotyping files; The SNP marker genotyping file and the SV marker genotyping file are merged to construct an initial genotyping matrix. The rows in the initial genotyping matrix represent all detected markers, the columns represent all samples, and the matrix elements are the genotypes of the corresponding samples at the marker loci.

5. The method for drawing a genetic map of a giant panda according to claim 1, characterized in that, In step two, the tag-based hierarchical filtering includes: Initial filtering of markers in the initial genotyping matrix is ​​performed based on genotyping loss rate and Mendelian genetic error rate; A comprehensive score for each marker is calculated by weighting the minimum allele frequency, observed heterozygosity, genotyping loss rate, and Mendelian inheritance error rate. The markers after the initial filtering are then filtered a second time based on the comprehensive score. Chained unbalanced filtering is performed on the tags after the second filtering to generate a tag filtering fractal matrix.

6. The method for drawing a genetic map of a giant panda according to claim 1, characterized in that, In step three, Mendelian genetic constraints and inbreeding coefficient correction terms are added to the objective function of the partial least squares regression algorithm to improve the algorithm.

7. The method for drawing a genetic map of a giant panda according to claim 1, characterized in that, In step four, a type weighting factor is added to the recombination rate calculation of the genetic map construction algorithm based on the minimum spanning tree, the corrected recombination rate between all pairs of markers is calculated, and the corrected recombination rate full matrix is ​​constructed. Based on the corrected recombination rate matrix, a minimum spanning tree of all tags is generated using the minimum spanning tree algorithm. Linkage groups are partitioned using a dynamic threshold, and a linkage group adapted to the dual-type tags of giant pandas is constructed. After linearly orienting the markers for each linkage group, the preliminary genetic distances between markers within each linkage group are calculated.

8. The method for drawing a genetic map of a giant panda according to claim 1, characterized in that, In step five, an inbreeding coefficient correction term and a marker type bias correction term are introduced to correct the preliminary genetic distance between the markers, resulting in the corrected genetic distance between the markers. Based on the corrected genetic distance between the markers for each linkage group, a giant panda genetic map is constructed.

9. The method for drawing a genetic map of a giant panda according to claim 1, characterized in that, In step six, the giant panda genetic map is subjected to Mendelian genetic consistency verification, collinearity verification of genetic distance and physical distance, and map resolution and coverage verification. The giant panda genetic map that has passed the verification is subjected to genome anchoring and functional annotation to obtain the final giant panda genetic map.

Citation Information

Patent Citations

  • Panda genetic map drawing method and system

    CN117409868A