Device and method for genetic map construction or gene localization based on low-depth sequencing data and application

By using low-depth resequencing data and hidden Markov models, the problem of low-cost and high-precision gene localization and genetic map construction for quantitative traits in species with large genomes was solved, and efficient genetic map construction and gene localization for species such as wheat were achieved.

CN120913646APending Publication Date: 2025-11-07CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410548440.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-06
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In species with large genomes, existing technologies struggle to achieve low-cost, rapid, and high-precision quantitative trait gene localization and genetic map construction. This is especially true in species like wheat, where SNP microarrays have limited characterization capabilities, making it difficult to comprehensively characterize genomic structural variations. Furthermore, pooled sequencing struggles to obtain individual genotype information.

Method used

By receiving low-depth resequencing data from the mapping population and parents, SNP variant detection and structural variant identification are performed. Combined with a hidden Markov model, genomic windows are divided to identify parental haplotypes, and a genetic map is constructed. Gene localization and genetic map construction are then performed using the low-depth sequencing data.

Benefits of technology

It achieves low-cost, high-precision quantitative trait gene localization and genetic map construction, accurately identifies heterozygous regions and recombination sites in each strain, reduces sequencing costs, and improves the resolution and continuity of haplotype information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0004823477510000011
    Figure HDA0004823477510000011
  • Figure HDA0004823477510000012
    Figure HDA0004823477510000012
  • Figure HDA0004823477510000021
    Figure HDA0004823477510000021
Patent Text Reader

Abstract

The invention discloses a device and method for genetic map construction or gene localization based on low-depth sequencing data and application. According to the method, about 0.1 * ultra-low deep sequencing is carried out on each line of the mapping group by utilizing a re-sequencing technology, relatively high deep sequencing is carried out on parents, and then type information of amphiphilic haplotypes (including SNP and structural variation) in a whole genome window of each line of the mapping group is detected and traced based on sequencing data; a genetic map is constructed on the basis of haplotype information, phenotypic characters are combined, phenotypic-related quantitative character haplotypes are further positioned through linkage analysis, and gene positioning is achieved. The device or the method provided by the invention can be used for rapidly representing the genetic source of the offspring strain at lower cost and positioning quantitative character genes at the haplotype level, and can be applied to plant breeding and strain screening.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of bioinformatics, in particular to a device and method for constructing a genetic map or locating a gene based on low-depth sequencing data and application. BACKGROUND

[0002] Gene cloning plays an important role in crop improvement. However, in species with a relatively large genome, gene cloning work is often slow due to high time and financial costs. Quantitative trait locus positioning is the first step of gene cloning. Map-based cloning is one of the important methods of quantitative trait locus positioning, which uses molecular markers such as SNPs and phenotype data of mapping populations to carry out linkage analysis and locate quantitative trait loci to a specific interval of the genome. SNP chips are an important means of genotyping and identifying SNP markers for mapping populations. However, due to cost factors, in species with a relatively large genome such as wheat, SNP chips can only represent the genotypes of a small part of the genome. At the same time, SNP chips often cannot effectively represent structural variations in the genome, which are very common in plant genomes and can affect important agronomic traits. This further limits the ability to locate quantitative trait loci in species with a relatively large genome. The scheme based on pooled sequencing (such as BSA-seq) is another commonly used method for locating quantitative trait loci in mapping populations. It divides the mapping population into two extreme phenotype groups according to the phenotype value, and mixes and sequences the DNA or RNA of the lines belonging to the same group in equal amounts to identify quantitative trait loci. However, although this method is low-cost (only two groups need to be sequenced, not each line), it cannot obtain the genotype information of each line. The lack of this information is not conducive to the subsequent screening of remaining heterozygous lines and the construction of near-isogenic lines. In summary, a technology that can obtain high marker density at low cost and represent the genetic background of each line is very important for gene cloning in species with a relatively large genome. SUMMARY

[0003] The technical problem to be solved by the present application is how to realize low-cost, rapid and high-precision positioning of quantitative trait genes at the haplotype level and / or how to construct a genetic map at low cost and / or how to obtain high marker density representation of each line in a gene cloning mapping population at low cost and / or how to effectively map-based cloning or gene cloning in species with a large genome at low cost and / or how to genotype and identify SNP markers for crop inbred lines at low cost.

[0004] To solve the above technical problems, the present application first provides a computer device comprising a memory, a processor and a computer program stored on the memory, wherein the processor can execute the computer program to implement the following steps:

[0005] S1) Receiving resequencing data: receiving DNA resequencing data of each strain in the mapping population and DNA resequencing data of two parents in the mapping population;

[0006] S2) Mining of variant sites: aligning the low-depth resequencing data to the reference genome of the corresponding species of the mapping population to obtain the sequencing alignment result of each strain in the mapping population; aligning the resequencing data of the parents to the reference genome of the corresponding species of the mapping population to obtain the sequencing alignment result of the parents; using a variant detection software to detect the SNP variations in the sequencing alignment result of each strain and the sequencing alignment result of the parents to obtain the SNP variation detection result of each strain and the parents in the mapping population;

[0007] S3) Identification of biparental haplotypes: obtaining the haplotype identification result in the parent genome based on the sequencing alignment result of the parents and the SNP variation detection result;

[0008] S4) Identification of structural variations: obtaining the structural variation interval of the parents based on the analysis of the sequencing alignment result of the parents; obtaining the structural variation interval of each strain based on the sequencing alignment result of each strain in the mapping population;

[0009] S5) Identifying the effective differentiation interval of the mapping population and the genetic source of the effective differentiation interval: obtaining the genomic interval covered by the sequencing reads in the low-depth resequencing data based on the sequencing alignment result of each strain in the mapping population; comparing the SNP variation sites of each strain and the SNP variation sites of the parents in the genomic interval based on the SNP variation detection result of each strain and the parents in the mapping population to obtain the effective differentiation interval; and determining the genetic source of each strain in the effective differentiation interval according to the genotype of each strain at the differential SNP sites;

[0010] The effective differentiation interval contains a differential SNP site with different genotypes in the two parent individuals of the parents, and the genotype of the strain at the differential SNP site is consistent with that of one of the two parent individuals;

[0011] S6) Biparental haplotype identification of genomic windows: dividing the reference genome into (continuous) genomic windows, determining the biparental haplotype identification result of the genomic windows of each strain based on the genetic source of each strain in the effective differentiation interval obtained in S5) and the structural variation interval of each strain obtained in S4).

[0012] S7) filtering and binning of haplotype data: based on the haplotype identification results in the parental genomes described in S3), the parental haplotype identification results of the genomic windows of each line described in S6) are filtered, and the (fully) linked genomic windows are merged, obtaining the filtered and binned haplotype data of each line;

[0013] The sequencing depth of the DNA resequencing data of each line described above can be greater than or equal to 0.1x; the sequencing depth of the DNA resequencing data of the two parents described above can be greater than or equal to 4x.

[0014] The sequencing depth of the DNA resequencing data of each line described above can be 0.13x. The sequencing depth of the DNA resequencing data of the two parents described above can be 6x or 8x.

[0015] In the above computer device, the structural variation intervals of the parents in S4) can be obtained by a method comprising the following steps:

[0016] dividing the reference genome into genomic windows, determining the number of sequencing reads aligned to the reference genome in each genomic window in the genome of the parent based on the sequencing alignment results of the parent Count bin -P, calculating the non-zero mode of the number of sequencing reads in all genomic windows of the parent Count mode -P, using the Count mode -P, obtaining the non-zero mode of the number of sequencing reads in each genomic window of the parent Count bin -P, obtaining the non-zero mode of the number of sequencing reads in each genomic window of the parent Count bin -P / Count mode -P, normalizing the value of Count bin -P-N, obtaining the structural variation intervals of the parent based on the normalized value; the structural variation intervals of each line in S4) are obtained by a method comprising the following steps: dividing the reference genome into genomic windows, determining the number of sequencing reads aligned to the reference genome in each genomic window in the genome of each line based on the sequencing alignment results of each line Count bin -S, calculating the non-zero mode of the number of sequencing reads in all genomic windows of each line Count mode -S, using the Count mode -S, obtaining the non-zero mode of the number of sequencing reads in each genomic window of each line Count bin -S, obtaining the non-zero mode of the number of sequencing reads in each genomic window of each line Count bin -S / Count modeS is normalized to obtain a normalized value Count of the number of sequencing reads in each of the genomic windows bin S-N, based on the normalized value, obtaining a preliminary result of the structural variation interval of the strain; correcting the preliminary result of the structural variation type of the strain in each of the genomic windows using the structural variation interval of the parent to obtain the structural variation interval of each of the strains.

[0017] In the above computer device, S6) the parental haplotype identification of the genomic window can comprise the following steps:

[0018] S6-1) Genomic window division: dividing the reference genome into genomic windows;

[0019] S6-2) SNP variation parental haplotype identification: based on the genetic source of each strain in the effective interval obtained in S5), the proportion of the effective interval from which each strain in each genomic window is respectively derived from two parents is counted, and the parental SNP haplotype identification result of each strain in the genomic window is determined based on the proportion;

[0020] S6-3) Structural variation parental haplotype identification: based on the structural variation interval of each strain obtained in S4) and the genomic window where the structural variation interval is located, the parental CNV haplotype identification result of each strain in the genomic window is obtained;

[0021] S6-4) Parental haplotype result output: based on the parental SNP haplotype identification result of the genomic window and the parental CNV haplotype identification result of the genomic window, the parental haplotype of the whole genomic window of each strain is predicted using a hidden Markov model, and the parental haplotype result of the whole genomic window of each strain is output.

[0022] In the above computer device, the size of the genomic window in S6-1) can be 50Kbp and / or 1Mbp. The size of the genomic window in S4) can be 1Mbp.

[0023] In order to solve the above technical problems, the present application also provides a device for constructing a genetic map based on low-depth sequencing data, which can comprise the following modules:

[0024] A1) Resequencing data receiving module: for receiving DNA resequencing data of each strain in the mapping population and DNA resequencing data of two parents in the mapping population;

[0025] A2) Variant site mining module: for aligning the low-depth re-sequencing data to the reference genome of the mapping population corresponding species to obtain the sequencing alignment results of each line in the mapping population; for aligning the re-sequencing data of the parents to the reference genome of the mapping population corresponding species to obtain the sequencing alignment results of the parents; for detecting the SNP variations in the sequencing alignment results of each line and the sequencing alignment results of the parents using a variation detection software to obtain the SNP variation detection results of each line and the parents in the mapping population;

[0026] A3) Biparental haplotype identification module: for obtaining the haplotype identification results in the genomes of the parents based on the sequencing alignment results of the parents and the SNP variation detection results;

[0027] A4) Structural variation identification module: for obtaining the structural variation intervals of the parents based on the analysis of the sequencing alignment results of the parents; for obtaining the structural variation intervals of each line based on the sequencing alignment results of each line in the mapping population;

[0028] A5) Module for identifying the effective differentiation intervals of the mapping population and the genetic sources of the effective differentiation intervals: for obtaining the genomic intervals covered by the sequencing reads in the low-depth re-sequencing data based on the sequencing alignment results of each line in the mapping population; for comparing the SNP variation sites of each line and the SNP variation sites of the parents in the genomic intervals based on the SNP variation detection results of each line and the parents in the mapping population to obtain the effective differentiation intervals; and for determining the genetic sources of each line in the effective differentiation intervals according to the genotypes of each line at the differential SNP sites;

[0029] The effective differentiation intervals contain differential SNP sites with different genotypes in the two parental individuals of the parents, and the genotype of the line at the differential SNP sites is consistent with the genotype of one of the two parental individuals;

[0030] A6) Biparental haplotype identification module for genomic windows: for dividing the reference genome into (continuous) genomic windows, determining the biparental haplotype identification results of the genomic windows of each line based on the genetic sources of each line in the effective differentiation intervals obtained in A5) and the structural variation intervals of each line obtained in S4);

[0031] A7) Haplotype data filtering and binning module: for filtering the biparental haplotype identification results of the genomic windows of each line in A6) based on the haplotype identification results in the genomes of the parents in A3), and merging the (completely) linked genomic windows to obtain the filtered and binned haplotype data of each line;

[0032] A8) a genetic map construction module for constructing a genetic map of the mapping population based on the filtered and binned haplotype data of each line of A7);

[0033] The sequencing depth of the low-depth resequencing data can be greater than or equal to 0.1x; the sequencing depth of the DNA resequencing data of the two parents can be greater than or equal to 4x.

[0034] In the above device, the sequencing depth of the DNA resequencing data of each line can be 0.13x. The sequencing depth of the DNA resequencing data of the two parents described above can be 6x or 8x.

[0035] In the above device, the structural variation interval of the parent in A4) can be obtained by a method comprising the following steps:

[0036] dividing the reference genome into genomic windows, determining the number of sequencing reads aligned to the reference genome in each genomic window in the genome of the parent based on the sequencing alignment result of the parent Count bin -P, calculating the non-zero mode of the number of sequencing reads in all genomic windows of the parent Count mode -P, using the Count mode -P to each of the Count bin -P by calculating Count bin -P / Count mode -P, normalizing the value of each genomic window Count bin -P-N, obtaining the structural variation interval of the parent based on the normalized value.

[0037] The size of the genomic window in A4) can be 1Mbp.

[0038] In the above device, the structural variation interval of each line in A4) can be obtained by a method comprising the following steps: dividing the reference genome into genomic windows, determining the number of sequencing reads aligned to the reference genome in each genomic window in the genome of each line based on the sequencing alignment result of each line Count bin -S, calculating the non-zero mode of the number of sequencing reads in all genomic windows of each line Count mode -S, using the Count mode -S to each of the Count bin -S by calculating Count bin -S / Count mode-S is normalized to obtain a normalized value Count of the number of sequencing reads in each of the genomic windows bin -S-N, based on the normalized value, obtaining a preliminary result of the structural variation interval of the strain; and correcting the preliminary result of the structural variation type of the strain in each of the genomic windows using the structural variation interval of the parent to obtain the structural variation interval of each of the strains.

[0039] In the above device, A6) the haplotype identification of the genomic window can include the following modules:

[0040] A6-1) Genomic window division module: for dividing the reference genome into genomic windows;

[0041] A6-2) SNP variation haplotype identification module: for determining the proportion of the effective interval in each genomic window that is respectively derived from the two parents of each strain based on the genetic source of each strain in the effective interval obtained in A5), and determining the parental SNP haplotype identification result of the genomic window of each strain based on the proportion;

[0042] A6-3) Structural variation haplotype identification module: for obtaining the parental CNV haplotype identification result of the genomic window of each strain based on the structural variation interval of each strain obtained in A4) and the genomic window where the structural variation interval is located;

[0043] A6-4) Parental haplotype result output module: for obtaining the parental haplotype of the whole genomic window of each strain using a hidden Markov model based on the parental SNP haplotype identification result of the genomic window and the parental CNV haplotype identification result of the genomic window, and outputting the parental haplotype result of the whole genomic window of each strain.

[0044] The size of the genomic window in A6-1) can be 50Kbp and / or 1Mbp.

[0045] To solve the above technical problems, the present application also provides a method for genetic mapping based on low-depth sequencing data, which can include the following steps:

[0046] B1) Re-sequencing data receiving: receiving DNA low-depth re-sequencing data of each strain in the mapping population and DNA re-sequencing data of two parents in the mapping population;

[0047] B2) Variant site mining: aligning the low-depth resequencing data to the reference genome of the corresponding species of the mapping population to obtain the sequencing alignment results of each line in the mapping population; aligning the resequencing data of the parents to the reference genome of the corresponding species of the mapping population to obtain the sequencing alignment results of the parents; detecting the SNP variations in the sequencing alignment results of each line and the sequencing alignment results of the parents using a variant detection software to obtain the SNP variation detection results of each line and the parents in the mapping population;

[0048] B3) Biparental haplotype identification: obtaining the haplotype identification results in the genomes of the parents based on the sequencing alignment results of the parents and the SNP variation detection results;

[0049] B4) Structural variation identification: obtaining the structural variation intervals of the parents based on the analysis of the sequencing alignment results of the parents; obtaining the structural variation intervals of each line based on the sequencing alignment results of each line in the mapping population;

[0050] B5) Identifying the effective differentiation intervals of the mapping population and the genetic sources of the effective differentiation intervals: obtaining the genomic intervals covered by the sequencing reads in the low-depth resequencing data based on the sequencing alignment results of each line in the mapping population; comparing the SNP variation sites of each line and the SNP variation sites of the parents in the genomic intervals based on the SNP variation detection results of each line and the parents in the mapping population to obtain the effective differentiation intervals; and determining the genetic sources of each line in the effective differentiation intervals according to the genotypes of each line at the differential SNP sites;

[0051] The effective differentiation intervals contain differential SNP sites with different genotypes in the two parental individuals of the parents, and the genotype of the line at the differential SNP site is consistent with that of one of the two parental individuals;

[0052] B6) Biparental haplotype identification of genomic windows: dividing the reference genome into (continuous) genomic windows, determining the biparental haplotype identification results of the genomic windows of each line based on the genetic sources of each line in the effective differentiation intervals obtained in B5) and the structural variation intervals of each line obtained in B4);

[0053] B7) Filtering and binning of haplotype data: filtering the biparental haplotype identification results of the genomic windows of each line in B6) based on the haplotype identification results in the genomes of the parents in B3), and merging the (completely) linked genomic windows to obtain the filtered and binned haplotype data of each line;

[0054] B8) Genetic map construction: constructing a genetic map of the mapping population based on the filtered and binned haplotype data of each line described in B7);

[0055] B9) Gene localization: obtaining genomic intervals associated with the phenotypic trait of interest in the mapping population based on the filtered and binned haplotype data of each line described in B7) and the genetic map described in B8), in combination with the phenotypic trait data of each line in the mapping population;

[0056] In the above method, the sequencing depth of the low-depth resequencing data can be greater than or equal to 0.1x; and the sequencing depth of the DNA resequencing data of the two parents can be greater than or equal to 4x.

[0057] The sequencing depth of the DNA resequencing data of each line described above can be 0.13x. The sequencing depth of the DNA resequencing data of the two parents described above can be 6x or 8x.

[0058] In the above method, the structural variation interval of the parent in B4) can be obtained by a method comprising the following steps:

[0059] dividing the reference genome into genomic windows, determining the number of sequencing reads aligned to the reference genome within each genomic window in the genome of the parent based on the sequencing alignment result of the parent Count bin -P, calculating the non-zero mode of the number of sequencing reads within all the genomic windows of the parent Count mode -P, using the Count mode -P, obtaining the non-zero mode of the number of sequencing reads within each genomic window Count bin -P, obtaining the non-zero mode of the number of sequencing reads within each genomic window Count bin -P / Count mode -P, normalizing the value of Count bin -P-N, obtaining the structural variation interval of the parent based on the normalized value.

[0060] The size of the genomic window in B4) can be 1Mbp.

[0061] In the above method, the structural variation interval of each line in B4) can be obtained by a method comprising the following steps: dividing the reference genome into genomic windows, determining the number of sequencing reads aligned to the reference genome within each genomic window in the genome of each line based on the sequencing alignment result of each line Count bin -S, calculating the non-zero mode of the number of sequencing reads within all the genomic windows of each line Count mode -S, using the Countmode -S is the number of sequencing reads in each of the genomic windows bin -S is the number of sequencing reads in each of the genomic windows bin -S is the number of sequencing reads in each of the genomic windows mode -S is the number of sequencing reads in each of the genomic windows bin -S is the number of sequencing reads in each of the genomic windows

[0062] In the above method, B6) the parental haplotype identification of the genomic window can comprise the following steps:

[0063] B6-1) Genomic window division: dividing the reference genome into genomic windows;

[0064] B6-2) SNP variation parental haplotype identification: based on the genetic source of each strain in the effective interval obtained in B5), the proportion of the effective interval from two parents in each genomic window of each strain is counted, and the parental SNP haplotype identification result of the genomic window of each strain is determined based on the proportion;

[0065] B6-3) Structural variation parental haplotype identification: based on the structural variation interval of each strain obtained in B4) and the genomic window where the structural variation interval is located, the parental CNV haplotype identification result of the genomic window of each strain is obtained;

[0066] B6-4) Parental haplotype result output: based on the parental SNP haplotype identification result of the genomic window and the parental CNV haplotype identification result of the genomic window, the parental haplotype of the whole genomic window of each strain is predicted using a hidden Markov model, and the parental haplotype result of the whole genomic window of each strain is output.

[0067] In B6-1), the size of the genomic window can be 50Kbp and / or 1Mbp.

[0068] The above method can further comprise the following steps: selecting two strains M and N with the largest difference in phenotypic traits in the mapping population, based on the parental haplotype result of the whole genomic window of the strain M and the parental haplotype result of the whole genomic window of the strain N, obtaining haplotype localization of the gene related to the phenotypic trait to be tested, and based on the haplotype localization, obtaining localization of the phenotype-related gene.

[0069] To solve the above technical problems, the application further provides a computer readable storage medium, which stores a computer program / instruction, and the computer program / instruction is executed by a processor to realize the steps of the above method.

[0070] To solve the above technical problems, the application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to realize the steps of the above method.

[0071] Any of the following applications of the above computer device also belongs to the protection scope of the application:

[0072] P1, in the application of constructing a genetic map of a to-be-tested population;

[0073] P2, in the application of phenotype-related gene positioning or quantitative trait locus positioning;

[0074] P3, in the application of marker-assisted selection;

[0075] P4, in the application of genetic breeding and / or quality improvement;

[0076] P5, in the application of assisting in strain screening.

[0077] Any of the following applications of the above device and / or the above computer readable storage medium and / or the above computer product also belongs to the protection scope of the application:

[0078] P1, in the application of marker-assisted selection;

[0079] P2, in the application of genetic breeding and / or quality improvement;

[0080] P3, in the application of assisting in strain screening.

[0081] In the above applications, the genetic breeding can be crop genetic breeding; and the quality improvement can be crop quality improvement.

[0082] The above mapping population can be a plant mapping population, and the plant can be wheat. The above mapping population can be a recombinant inbred line. The above mapping population can also be a doubled haploid population or a backcross introgression line population, etc.

[0083] The linkage mentioned above can be complete linkage, and the genome window of the complete linkage can be defined as follows: if a line in a previous window is a "p1" haplotype, a line in a next window is also a "p1", and a sample in the previous window haplotype is a "p2" haplotype, the sample in the next window is also a "p2", and the adjacent genome window is considered as a completely linked genome window.

[0084] The hidden states of the hidden Markov model described above can be three, representing three haplotype types respectively: two types represented by the parents (denoted as "p1" and "p2") and a hybrid type (denoted as "h").

[0085] The purpose of the present application is to provide a low-cost, fast and high-precision method for positioning quantitative traits at the haplotype level.

[0086] The main flowchart of the present application is as follows Figure 1 .

[0087] Compared with the prior art, the present application has the following beneficial effects:

[0088] 1) The cost is relatively low. Taking the field of wheat as an example, the cost required for ultra-low-depth sequencing of the mapping population is about half of that of genotyping using a common chip.

[0089] 2) The hybrid interval and recombination site of each strain of the mapping population can be accurately identified. Based on the SNP chip technology, it is difficult to fully characterize the genome due to the limited number of SNPs designed; based on the mixed pool sequencing technology, it is difficult to characterize the genome of an individual due to the need to mix the sample DNA for sequencing. The present application combines low-depth and statistical models, and the haplotype identification is accurate.

[0090] 3) The resolution of the obtained haplotype information is high and the distribution is continuous and uniform. Since the reads obtained by second-generation sequencing are uniformly distributed, and the sliding window method is used when dividing haplotypes, the obtained haplotype information is also continuously and uniformly distributed on the genome.

[0091] The present application provides a quantitative trait haplotype positioning method based on ultra-low-depth sequencing, relating to the field of agricultural technology. For genetic mapping populations such as recombinant inbred line populations, doubled haploid populations, backcross introgression line populations, etc., the present application first uses whole-genome resequencing technology to perform ultra-low-depth sequencing (sequencing depth of about 0.1x) on each strain, and performs higher-depth (about 5x) sequencing on the parents. Then, the hidden Markov model is used to trace the parental haplotype types on the genome of each strain, to reconstruct the recombination event site of each strain during construction and to identify the remaining hybrid interval. Based on this haplotype information, combined with field phenotypes, linkage analysis is used to further position quantitative trait haplotypes. The present application can quickly characterize the genetic origin of the offspring strains at a relatively low cost, and position quantitative trait genes at the haplotype level. BRIEF DESCRIPTION OF DRAWINGS

[0092] Figure 1 The flowchart of the method or device of the present application is shown.

[0093] Figure 2 The parameters of the hidden Markov model used in the method of the present application are shown.

[0094] Figure 3 The figure is the LOD curve obtained by using the method of the present application to locate the quantitative trait loci. The upper figure is the LOD curve on the A subgenome; the middle figure is the LOD curve on the B subgenome; the lower figure is the LOD curve on the D subgenome; the vertical coordinate is the LOD value; and the horizontal coordinate is the genetic distance.

[0095] Figure 4 The figure is the ear length phenotype value of 11 recombinant inbred lines obtained in the embodiment of the present application, and the haplotype thereof in the QTL interval on the short arm of chromosome 4B.

[0096] Figure 5 The figure is the ear length value and the difference significance test of the recombinant inbred lines carrying Rht-B1b allele, r-e-z allele and hybrid type on the Rht-B1 gene respectively in the embodiment of the present application. DETAILED DESCRIPTION

[0097] The present application will be further described in detail below with specific embodiments. The embodiments provided below are only for illustrating the present application, and are not intended to limit the scope of the present application. The embodiments provided below can serve as a guide for further improvement by those skilled in the art, and do not constitute any limitation on the present application in any way.

[0098] In the following examples, the experimental methods are all conventional methods, and are performed according to the techniques or conditions described in the literature in the art or according to the product instructions, unless otherwise specified. The materials, reagents and the like used in the following examples can be obtained from commercial channels, unless otherwise specified.

[0099] Example 1, establishment of a method for rapid gene location based on low-depth sequencing

[0100] The haplotype of the ear length trait gene of the recombinant inbred lines (including 266 F7 generation lines) constructed by the commercially available and commonly known wheat varieties Nongda 3097 and Luanxuan 987 was located. The recombinant inbred lines were constructed by the laboratory in the previous stage, and the construction method followed the common recombinant inbred line construction method, that is, the parents were artificially crossed (Luanxuan 987 as the father and Nongda 3097 as the mother), and the F1 was continuously self-crossed until the 7th generation.

[0101] 1. Resequencing data acquisition

[0102] 1.1 Ultra-low-depth sequencing of the mapping population

[0103] For each line of recombinant inbred lines, young leaves were used to extract DNA by CTAB method, and low-depth whole-genome resequencing with a sequencing amount of 2Gbp (about 0.13x) was performed to obtain the raw resequencing data of each line of recombinant inbred lines.

[0104] 1.2 Higher-depth sequencing of mapping population parents

[0105] The resequencing data of Nongda 3097 and Luanxuan 987 were downloaded from published literature (relevant literature: YANG Z, WANG Z, WANG W, et al. ggComp enables dissection of germplasm resources and construction of a multiscale germplasm network in wheat [J]. Plant Physiology, 2022, 188(4): 1950-1965) as the raw resequencing data of parents. The sequencing depth of the resequencing data of the two materials was 6x and 8x, respectively.

[0106] 2, Variation site mining

[0107] 2.1 Quality control of raw data

[0108] For the raw resequencing data obtained in step 1 (including the raw resequencing data of the recombinant inbred line lines and the raw resequencing data of the parents), Trimmomatic software was used for filtering to remove low-quality bases in the raw resequencing data, and effective resequencing data was obtained.

[0109] 2.2 Aligning sequencing reads to reference genome

[0110] The BWA-MEM tool of BWA software was used to align the filtered sequencing reads (effective resequencing data) to the reference genome of wheat IWGSC RefSeq v1.0, and the sequencing alignment result was obtained.

[0111] The samtools software is used for quality control of the sequencing alignment results, PCR repeated reads are removed, and the effective alignment results of the sequencing data are obtained, including the effective alignment results of the two parents of the recombinant inbred lines and the effective alignment results of each line in the recombinant inbred lines. Using the HaplotypeCaller function of the GATK software package, based on the effective alignment results, the single nucleotide polymorphism (SNP) variation information of the recombinant inbred lines and the parents is detected and identified, and the GVCF format variation detection file containing the SNP detection results of each line in the recombinant inbred lines and the GVCF format variation detection file containing the SNP detection results of the two parents of the recombinant inbred lines are obtained. All the above GVCF format files are merged using Glnexus software to obtain a VCF format file containing the SNP variation detection results of the two parents of the recombinant inbred lines and all lines in the recombinant inbred lines.

[0112] 3. Parental haplotype identification

[0113] Using the effective alignment results and VCF format files of the two parents of the recombinant inbred lines obtained in step 2, the ggComp software (related website: https: / / zack-young.github.io / ggComp / ) is used for parental haplotype identification to obtain the haplotype identification results of the parents. During identification, the ggComp software will identify whether the two parents are the same haplotype in each window of the genome in the form of sliding window with a window size of 1 Mbp. For the genomic windows with the same genetic background of the parents, mark as no haplotype diversity; for the genomic windows with different genetic backgrounds of the parents, mark as the genomic interval (window) exists haplotype diversity.

[0114] 4. Structural variation detection

[0115] 4.1 Parental structural variation identification

[0116] Using the effective alignment results of the two parents of the recombinant inbred lines obtained in step 2, the interval of the two parents with structural variation on the genome is identified. In the process of judging the existence of structural variation, the copy number variation (CNV) signal is used to represent the structural variation.

[0117] Firstly, the bedtools software is used to divide the whole wheat reference genome into continuous genomic windows, and the window size is specifically 1 Mbp, and the number of windows obtained is specifically 14075.

[0118] For each parent, the number of sequencing reads (Count bin -P) that successfully mapped to the reference genome in each genomic window (1Mbp long) was counted based on the valid mapping results of the parent re-sequencing data obtained in step 3, and the non-zero mode (Count mode -P) of the number of sequencing reads in all genomic windows was calculated.

[0119] The non-zero mode was calculated as follows: first, Count bin -P) of all intervals of the whole genome of the parent were taken as a set, then the items with value 0 were removed from the set, and finally the mode function of the python language was used to calculate the mode, thereby obtaining Count mode -P.

[0119] The sequencing read count in each window was normalized (Count mode -P / Count bin -P) using the mode Count mode -P, and the normalized value Count bin -P-N of the sequencing read count in each window (1Mbp long) was obtained. The window with a normalized value Count bin -P-N lower than 0.5 was recorded as an interval with a "deletion" variation, the window with a normalized value Count bin -P-N higher than 1.5 was recorded as an interval with a "duplication" variation, and the window with a normalized value Count bin -P-N between 0.5 and 1.5 was recorded as an interval without structural variation. The intervals with "deletion" variation and the intervals with "duplication" variation together constituted the structural variation intervals of the recombinant inbred lines.

[0120] 4.2 Mapping population structural variation scanning

[0121] The structural variations of each line in the recombinant inbred lines were scanned using the valid mapping results of each line in the recombinant inbred lines obtained in step 2 and the structural variation intervals of the parents obtained in step 4.1. First, the read count normalization value Count bin -S-N of the continuous genomic windows of the whole genome of each line of the recombinant inbred lines was calculated according to the method of step 4.1, and the preliminary results of the structural variation intervals of each line were obtained:

[0122] The reference genome was divided into continuous genomic windows with a length of 1Mbp, and the number of sequencing reads (Count bin-S, calculate the non-zero mode Count of the number of sequencing reads in all the genomic windows of each line mode -S, using Count mode -S, for each Count bin -S, by calculating Count bin -S / Count mode -S, normalize the value of each genomic window by the normalized value Count of the number of sequencing reads bin -S-N. The normalized value Count bin -S-N, the windows with a value lower than 0.5 are recorded as intervals with "deletion" variation, Count bin -S-N, the windows with a value higher than 1.5 are recorded as intervals with "duplication" variation, and the others (Count bin -S-N, the windows with a value between 0.5 and 1.5 are recorded as intervals without structural variation.

[0123] Subsequently, the preliminary results of the structural variation intervals of each line of the recombinant inbred lines, i.e. the copy number variation state of each genomic window in each line (including deletion, duplication and no structural variation), are compared with the parents. For the genomic windows with structural variation in the parents, if the line in the recombinant inbred lines has the same type of variation as either parent (both "deletion" variation or both "duplication" variation) in the same genomic window, it is recorded as the line having the same structural variation as the parent.

[0124] 5, the read coverage interval of the mapping population is used to identify the genetic origin of the SNP variation:

[0125] Since the recombinant inbred lines are sequenced at a depth of about 0.13x, the sequencing reads cover at most 13% of the reference genome after alignment to the reference genome. Due to the randomness of second-generation sequencing, these 13% intervals are randomly distributed on the genome. From the distribution of the length of the continuous genomic intervals covered by the sequencing reads, these intervals with a length of less than or equal to 300bp account for 96.3% of all fragments, and those with a length of 150bp account for 66%. This length distribution is consistent with the characteristics of double-end 150bp resequencing. Therefore, each continuous genomic interval covered by a sequencing read is approximately considered to be covered by a single sequencing read. Further, by analyzing these intervals, the parent origin of each sequencing read can be directly determined, and the inheritance source of each window on the genome of the recombinant inbred lines can be determined based on sequencing reads.

[0126] Using the effective alignment results of each line of the recombinant inbred lines obtained in step 2, for each line, the continuous genomic intervals covered by sequencing reads on the genome (genomic intervals covered by greater than or equal to 1 sequencing read) were extracted using bedtools software. Then, based on the VCF format file obtained in step 2.2, the SNP variation information of the parents and the line in each genomic interval was extracted.

[0127] For each genomic interval of each line where sequencing reads exist, if the parents have at least one SNP site with genotype difference (denoted as a difference SNP site) in the interval, and the genotype of the line at the difference SNP site is consistent with one of the parents, the site is considered to be an effective distinguishing variation site, the interval has effective variation information for distinguishing genetic sources, and the interval is retained as an effective distinguishing interval. For the effective distinguishing interval, if the genotype of the line at all effective distinguishing variation sites in the interval is consistent with one of the parents, the effective distinguishing interval is recorded as being inherited from the parent. If the genotypes of the line at multiple effective distinguishing variation sites in the interval are respectively consistent with different parents, the interval is discarded.

[0128] Finally, all effective distinguishing intervals of each line of the wheat recombinant inbred line population constructed by Nongda 3097 and Lunxuan 987, and the genetic source identification results in the effective distinguishing intervals of the genome of each line of the recombinant inbred line were obtained.

[0129] 6. Parental haplotype identification of the mapping population genome:

[0130] The entire wheat reference genome was divided into continuous genomic windows using bedtools software, and the window size was the resolution required for identifying parental haplotypes. In this step, the window size was set to 50Kbp (obtaining 281500 windows) and 1Mbp (obtaining 14075 windows), respectively, and the parental haplotypes of the mapping population genome at 50Kbp and 1Mbp haplotype resolution were identified.

[0131] 6.1 Determine the haplotype of each line genomic window based on the genetic source identification results based on SNP variation

[0132] Using the genetic source identification results in the effective distinguishing intervals of the genome of each line of the recombinant inbred line obtained in step 5, the number of sequencing read coverage intervals inherited from the two parents in each genomic window of each line was counted, respectively, to determine the proportion of effective distinguishing intervals of different inheritance sources of each line genomic window.

[0133] Subsequently, the proportions of effective distinguishing intervals from different inheritance sources were discretized to obtain the preliminary inheritance source for each genomic window of each strain. Specifically, during discretization, a proportion of effective distinguishing intervals consistent with parent 1 (Nongda 3097) greater than or equal to 0.9 was denoted as "P1", and a proportion greater than 0.75 but less than 0.9 was denoted as "P1H". Similarly, a proportion of effective distinguishing intervals consistent with parent 2 (Lunxuan 987) greater than or equal to 0.9 was denoted as "P2", and a proportion greater than 0.75 but less than 0.9 was denoted as "P2H". All other cases were denoted as "H".

[0134] 6.2 Correcting the preliminary inheritance source of the genomic window using CNV variant results

[0135] Based on the structural variation scanning results of each strain obtained in step 4, if a structural variation (CNV) exists between the parents in the genomic window, the preliminary inheritance source of the genomic window is directly derived from the structural variation and covers the preliminary inheritance source obtained in step 6.1. That is, if the strain's structural variation type is consistent with parent 1, it is denoted as "P1"; if it is consistent with parent 2, it is denoted as "P2".

[0136] 6.3 Using Hidden Markov Models to Determine the Parental Inheritance Source of the Genomic Window

[0137] The preliminary inheritance source of the genomic window obtained from steps 6.1 and 6.2 is used as the observed sequence (i.e., model input data) and input into the Hidden Markov Model (HMM). The Viterbi algorithm is then used to infer the haplotype type of each interval. Therefore, in the model design, three hidden states are set, each representing one of the three haplotype types: the two types represented by the parents (denoted as "p1" and "p2"), and the heterozygous type (denoted as "h"). The model parameters of the HMM are as follows: Figure 2 The Hidden Markov Model observes 5 states. Figure 2 As shown in the figure, p1, p2, p1h, p2h, and h are respectively. The creation of the Hidden Markov Model utilizes the hmmlearn package in Python.

[0138] By combining the SNP and CNV identification results of each line through steps 6.1 and 6.2, the genomic window of the whole genome of each recombinant inbred line was finally obtained as the parental haplotype identification results.

[0139] 7. Filtering and Binning of Haplotype Data for Plotting Populations

[0140] Using the parental haplotype identification results of both 50Kbp and 1Mbp haplotype resolution obtained in step 6, the following operations were performed. First, the genomic windows in which the parents carried the same haplotype were screened, and based on this, the parental haplotype identification results of the genomic windows of each line of the recombinant inbred lines obtained in step 6 were filtered, and the haplotype data of the genomic windows located in the genomic intervals with the same genetic background of the parents were removed. Subsequently, adjacent genomic windows that were completely linked (i.e., if the previous window was a line with a “p1” haplotype, all the next windows were also “p1”, and the previous window haplotype was a “p2” haplotype, all the next windows were also “p2”, then the adjacent genomic windows were considered to be completely linked adjacent genomic windows) were combined to refine the haplotype information. Finally, the haplotype data of each line after filtering and binning at 50Kbp and 1Mbp haplotype resolution were obtained.

[0141] 8. Genetic map construction

[0142] The lines of the recombinant inbred lines were used as a genetic mapping population, and the filtered and binned haplotype data of the mapping population obtained in step 7 were used to construct a genetic map at 1Mbp haplotype resolution using the “est.map()” function of the qtl package of R language with the “error.prob=0.001, map.function=“kosambi”, verbose=FALSE” parameters.

[0143] 9. Quantitative trait locus genomic interval positioning

[0144] Taking the ear length trait as an example, the quantitative trait locus interval related to the ear length phenotype was positioned.

[0145] Using the filtered and binned haplotype data of the mapping population at 1Mbp haplotype resolution obtained in step 7 and the genetic map obtained in step 8, combined with the ear length data of each line of the recombinant inbred line population, the “scanone()” function of the qtl package of R language was used for QTL preliminary positioning, and the “method=“em”” parameter was used to obtain the genomic interval with significant correlation (LOD>3) with the phenotype variation of the mapping population as 2D:281-356cM, 3A:406-412cM, 3B:253-381cM, 4A:279-325cM, 4B:22-103cM( Figure 3 ). Among them, the LOD value in the 4B:45-78cM interval was greater than 9.

[0146] 10. Quantitative trait candidate haplotype screening and quantitative trait locus major gene positioning

[0147] Using the significant association interval (4B:45-78 cM) identified in step 9, the quantitative trait candidate haplotype screening and major gene mapping related to the ear length trait were performed using the parental haplotype identification results of the genomic windows of the recombinant inbred lines of each strain obtained in step 7 at a resolution of 50Kbp.

[0148] First, strains with only one recombination in the significant association interval (4B:45-78 cM) were screened, and strains RIL128, RIL021, RIL271, RIL184, RIL296, RIL186, RIL242 and RIL054 were obtained. By sorting the phenotypic values of these strains from small to large, the haplotype difference interval between the two strains with the largest adjacent phenotypic value change was recorded as the quantitative trait candidate haplotype interval. Then, by comparing the haplotype information of other materials in the recombinant inbred lines, the interval was further narrowed to 1.7Mbp (4B: 66.7-68.4 cM) by RIL105, RIL049 and RIL255. Figure 4 ).

[0149] Specifically, it can be seen from Figure 4 that the ear length phenotypic value is most different between strain RIL255 (ear length value of 8.9 cm) and RIL021 (ear length value of 9.8 cm). Based on this result, the 11 strains can be divided into two groups, one group consisting of strains RIL105, RIL049, RIL128 and RIL255 (ear length range 8.4-8.9 cm) and the other group consisting of RIL021, RIL271, RIL184, RIL296, RIL186, RIL242 and RIL054 (ear length range 9.8-10.9 cm). Using the rule that the haplotypes of the short ear length group are consistent, the haplotypes of the long ear length group are consistent, and the haplotypes of the short ear length group and the long ear length group are inconsistent, a genomic interval of 1.7Mbp was screened.

[0150] Comparison of genetic differences between the parents within this interval revealed a genomic segment deletion diversity of approximately 500 kbp between them. This deletion is associated with the rez allele of the dwarf gene Rht-B1, and previous literature has demonstrated through transgenic experiments that the rez allele increases spike length (related literature: SONG L, LIU J, CAO B, et al. Reducing brassinosteroid signalling enhances grain yield in semi-dwarfwheat[J]. Nature, 2023, 617(7959): 118-124). This effect is consistent with that observed in this recombinant inbred line, namely the line carrying the rez allele inherited from the parent Nongda 3097 (long spike long group line). Figure 5 The "rez" strain and the line carrying the Rht-B1b allele inherited from the parental selection 987 (short spike long group strain) Figure 5 Compared to Rht-B1b (represented by Rht-B1b), it has a longer spike length ( Figure 5 (P < 0.01); meanwhile, the spike length of the heterozygous line was between that of the lines carrying the rez allele and the Rht-B1b allele.

[0151] Through the above steps, the method of the present invention successfully located the major genes related to ear length in recombinant inbred lines at the haplotype level, proving the effectiveness of the method of the present invention.

[0152] 11. Strain selection

[0153] For specific breeding objectives, based on the method of this invention, the parental genomic variation information obtained in steps 2 and 3 of this invention, along with the allelic genotypes and effects of cloned functional genes, can be combined to design the ideal haplotype composition of the offspring genome. Furthermore, the mapping population parental haplotype data obtained in steps 8 and 9 of this invention can be used to screen offspring lines for specific breeding objectives, thereby accelerating the variety improvement process. For example, when it is necessary to screen for remaining heterozygous lines, lines with heterozygous haplotypes within the highly significant association interval (4B: 45-78 cM) can be directly screened.

[0154] The application has been described in detail. For those skilled in the art, the application can be implemented in a wider range under the same parameters, concentrations and conditions without departing from the spirit and scope of the application and without unnecessary experiments. Although the application gives a special example, it should be understood that the application can be further improved. In summary, according to the principle of the application, the application intends to include any change, use or improvement of the application, including the change made by the conventional technology known in the art, which is out of the range disclosed in the application.

Claims

1. Computer apparatus comprising a memory, a processor and a computer program stored on the memory, characterised in that, The processor executes the computer program to implement the following steps: S1) receiving resequencing data: receiving DNA resequencing data of each strain in the mapping population and DNA resequencing data of two parents in the mapping population; S2) mining of variant sites: aligning the low-depth resequencing data to the reference genome of the corresponding species of the mapping population to obtain the sequencing alignment result of each strain in the mapping population; aligning the resequencing data of the parents to the reference genome of the corresponding species of the mapping population to obtain the sequencing alignment result of the parents; using a variant detection software to detect the SNP variations in the sequencing alignment result of each strain and the sequencing alignment result of the parents to obtain the SNP variation detection result of each strain and the parents in the mapping population; S3) identification of biparental haplotypes: obtaining the haplotype identification result in the parent genome based on the sequencing alignment result of the parents and the SNP variation detection result; S4) identification of structural variations: obtaining the structural variation interval of the parents based on the analysis of the sequencing alignment result of the parents; obtaining the structural variation interval of each strain based on the sequencing alignment result of each strain in the mapping population; S5) identifying the effective distinguishing interval of the mapping population and the genetic source of the effective distinguishing interval: obtaining the genomic interval covered by the sequencing reads in the low-depth resequencing data based on the sequencing alignment result of each strain in the mapping population; comparing the SNP variation sites of each strain and the SNP variation sites of the parents in the genomic interval based on the SNP variation detection result of each strain and the parents in the mapping population to obtain the effective distinguishing interval; and determining the genetic source of each strain in the effective distinguishing interval according to the genotype of each strain at the differential SNP sites; The effective distinguishing interval contains a differential SNP site with a differential genotype in the two parent individuals of the parents, and the genotype of the strain at the differential SNP site is consistent with that of one of the two parent individuals; S6) identification of biparental haplotypes of genomic windows: dividing the reference genome into genomic windows, determining the biparental haplotype identification result of the genomic windows of each strain based on the genetic source of each strain in the effective distinguishing interval obtained in S5) and the structural variation interval of each strain obtained in S4); S7) filtering and binning of haplotype data: filtering the biparental haplotype identification result of the genomic windows of each strain in S6) based on the haplotype identification result in the parent genome in S3), and merging the linked genomic windows to obtain the filtered and binned haplotype data of each strain; The sequencing depth of the DNA resequencing data of each strain is greater than or equal to 0.1x; and the sequencing depth of the DNA resequencing data of the two parents is greater than or equal to 4x.

2. The computer apparatus of claim 1, wherein: The structural variation interval of the parents in S4) is obtained by a method comprising the following steps: dividing the reference genome into genomic windows, determining the number of sequencing reads aligned to the reference genome within each of the genomic windows in the genome of the parent based on the sequencing alignment results of the parent Count bin -P, determining the non-zero mode of the number of sequencing reads within all of the genomic windows of the parent Count mode -P, using the Count mode -P, obtaining the number of sequencing reads within each of the genomic windows of the parent Count bin -P, obtaining the number of sequencing reads within each of the genomic windows of the parent Count bin -P / Count mode -P, normalizing the value of each of the genomic windows by Count bin -P-N, obtaining the structural variant interval of the parent based on the normalized value The structural variation intervals of each strain mentioned in S4) are obtained by a method including the following steps: dividing the reference genome into genomic windows, and determining the number of sequencing reads aligned to the reference genome within each genomic window of each strain based on the sequencing alignment results of each strain. bin -S, calculates the non-zero mode (Count) of the number of sequencing reads within all said genomic windows for each strain. mode -S, using the Count mode -S for each of the Counts bin -S calculates Count bin -S / Count mode The value of -S is normalized to obtain the normalized value Count of the number of sequencing reads within each genome window. bin -SN, based on the normalized value, obtain the preliminary results of the structural variation interval of the strain; correct the preliminary results of the structural variation type of the strain within each genomic window using the structural variation interval of the parent to obtain the structural variation interval of each strain.

3. The computer apparatus according to claim 1 or 2, wherein: S6) the identification of biparental haplotypes of the genomic windows comprises the following steps: S6-1) Genomic window division: dividing the reference genome into genomic windows; S6-2) SNP variant parent haplotype identification: based on the genetic origin of each strain in the effective differential interval obtained in S5), the proportion of the effective differential interval from which each strain in each genomic window is respectively derived from two parents is counted, and the parent SNP haplotype identification result of each strain in the genomic window is determined based on the proportion; S6-3) Structural variant parent haplotype identification: based on the structural variant interval of each strain obtained in S4) and the genomic window where the structural variant interval is located, the parent CNV haplotype identification result of each strain in the genomic window is obtained; S6-4) Parent haplotype result output: based on the parent SNP haplotype identification result of the genomic window and the parent CNV haplotype identification result of the genomic window, the parent haplotype of the whole genomic window of each strain is predicted and obtained using a hidden Markov model, and the parent haplotype result of the whole genomic window of each strain is output.

4. Apparatus for constructing a genetic map based on low-depth sequencing data, characterized in that: The device comprises the following modules: A1) Resequencing data receiving module: for receiving DNA resequencing data of each strain in the mapping population and DNA resequencing data of two parents in the mapping population; A2) Variant site mining module: for aligning the low-depth resequencing data to the reference genome of the corresponding species of the mapping population to obtain the sequencing alignment result of each strain in the mapping population; aligning the resequencing data of the parents to the reference genome of the corresponding species of the mapping population to obtain the sequencing alignment result of the parents; using a variant detection software to detect the SNP variants in the sequencing alignment result of each strain and the sequencing alignment result of the parents to obtain the SNP variant detection result of each strain and the parents in the mapping population; A3) Parent haplotype identification module: for obtaining haplotype identification results in the genomes of the parents based on the sequencing alignment results of the parents and the SNP variant detection results; A4) Structural variant identification module: for obtaining the structural variant interval of the parents based on the sequencing alignment results of the parents; obtaining the structural variant interval of each strain based on the sequencing alignment results of each strain in the mapping population; A5) Effective differential interval identification module: for obtaining the genomic interval covered by the sequencing reads in the low-depth resequencing data based on the sequencing alignment results of each strain in the mapping population; comparing the SNP variant sites in the genomic interval of each strain and the SNP variant sites of the parents based on the SNP variant detection results of each strain and the parents in the mapping population to obtain the effective differential interval; and determining the genetic origin of each strain in the effective differential interval according to the genotype of each strain at the differential SNP site; the effective distinguishing interval contains a difference SNP site with different genotypes in the two parent individuals of the parents, and the strain is consistent with the genotype of one of the two parent individuals at the difference SNP site; A6) A module for identifying the parental haplotype of a genomic window: for dividing the reference genome into genomic windows, determining the parental haplotype identification result of the genomic window of each strain based on the genetic origin of each strain in the effective distinguishing interval obtained in A5) and the structural variation interval of each strain obtained in S4); A7) A module for filtering and binning haplotype data: for filtering the parental haplotype identification result of the genomic window of each strain in A6) based on the haplotype identification result in the parent genome in A3), and combining the linked genomic windows to obtain the filtered and binned haplotype data of each strain; A8) A genetic map construction module: for constructing the genetic map of the mapping population based on the filtered and binned haplotype data of each strain in A7); The sequencing depth of the low-depth resequencing data is greater than or equal to 0.1x, and the sequencing depth of the DNA resequencing data of the two parents is greater than or equal to 4x.

5. A method for genetic mapping based on low-depth sequencing data, the method comprising the following steps: B1) Receiving resequencing data: receiving DNA low-depth resequencing data of each strain in a mapping population and DNA resequencing data of two parents in the mapping population; B2) Mining variation sites: aligning the low-depth resequencing data to the reference genome of the corresponding species of the mapping population to obtain the sequencing alignment result of each strain in the mapping population; aligning the resequencing data of the parents to the reference genome of the corresponding species of the mapping population to obtain the sequencing alignment result of the parents; detecting SNP variations in the sequencing alignment result of each strain and the sequencing alignment result of the parents using a variation detection software to obtain SNP variation detection results of each strain and the parents in the mapping population; B3) Parental haplotype identification: obtaining haplotype identification results in the parent genome based on the sequencing alignment result of the parents and the SNP variation detection results; B4) Structural variation identification: analyzing the sequencing alignment result of the parents to obtain the structural variation interval of the parents; obtaining the structural variation interval of each strain based on the sequencing alignment result of each strain in the mapping population; B5) identifying effective differentiating intervals of the mapping population and the genetic origin of the effective differentiating intervals: obtaining, based on the sequencing alignment results for each line of the mapping population, the genomic intervals covered by the sequencing reads in the low-depth resequencing data; Comparing the SNP variation sites of each strain and the SNP variation sites of the parents in the genomic interval based on the SNP variation detection results of each strain and the parents in the mapping population to obtain an effective distinguishing interval; and determining the genetic origin of each strain in the effective distinguishing interval according to the genotype of each strain at the difference SNP site; the effective distinguishing interval contains a difference SNP site with different genotypes in the two parent individuals of the parents, and the strain is consistent with the genotype of one of the two parent individuals at the difference SNP site; the effective distinguishing interval contains a difference SNP site with different genotypes in the two parent individuals of the parents, and the strain is consistent with the genotype of one of the two parent individuals at the difference SNP site; B6) Genome window-based parental haplotype identification: dividing the reference genome into genome windows, determining the parental haplotype identification results of the genome windows of each line based on the genetic origin of each line in the effective division interval obtained in B5) and the structural variation interval of each line obtained in B4); B7) Filtering and binning of haplotype data: filtering the parental haplotype identification results of the genome windows of each line in B6) based on the haplotype identification results in the parental genomes in B3), and merging the linked genome windows to obtain the filtered and binned haplotype data of each line; B8) Genetic map construction: constructing a genetic map of the mapping population based on the filtered and binned haplotype data of each line in B7); B9) Gene mapping: obtaining a genomic interval related to the phenotypic trait based on the filtered and binned haplotype data of each line in B7) and the genetic map in B8), combined with the phenotypic trait data of each line in the mapping population. The sequencing depth of the low-depth resequencing data is greater than or equal to 0.1x; the sequencing depth of the DNA resequencing data of the two parents is greater than or equal to 4x.

6. The method of claim 5, wherein: The method further comprises the following steps: selecting two lines M and N with the largest difference in phenotypic traits in the mapping population, obtaining haplotype mapping of the phenotype-related gene to be tested based on the parental haplotype results of the whole genome windows of the line M and the parental haplotype results of the whole genome windows of the line N, and obtaining the mapping of the phenotype-related gene based on the haplotype mapping.

7. A computer readable storage medium having stored thereon computer programs / instructions, characterized in that, The computer program / instructions are executed by the processor to implement the steps of the method of claim 5 or 6.

8. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of claim 5 or 6.

9. The following application of the computer device of any one of claims 1-3: P1, in the construction of the genetic map of the test population; P2, in the mapping of phenotype-related genes or quantitative trait loci; P3, in marker-assisted selection; P4, in genetic breeding and / or quality improvement; P5, in the screening of lines.

10. The following application of the device of claim 4 and / or the computer readable storage medium of claim 7 and / or the computer product of claim 8: P1, in marker-assisted selection; P2, in genetic breeding and / or quality improvement; P3, in the screening of lines.